Search NASA⌕ Search

SEARCH · Search NASA

Results for “knowledge networks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31

Utilization of Data Augmentation Techniques in Automated Inspection Systems for Defect Detection in Metals With Limited Data

Accurate identification of defects on metal surfaces is of great interest to many industry sectors, such as the automotive and aerospace industries. In contrast to conventional manual inspection techniques, recent automated inspection systems employ deep learning models trained to detect defects rapidly and precisely. The development of these models often requires a substantial image dataset to acquire adequate knowledge of defect features and enhance their predictive accuracy. When data is limited, augmentation techniques are often used to improve the precision and accuracy of defect detection systems. This study examined the prediction performance of two object detection models, namely Faster Region‐based Convolutional Neural Network (Faster R‐CNN) and You Only Look Once version 8 (YOLOv8), to identify dent defects in limited images of cast iron cylinder head surfaces. The original image set contains 46 images with 563 dents. To overcome limited data availability, common image augmentation techniques along with a copy‐paste method were applied. Results show that standard augmentation improved YOLOv8 accuracy by 8.00% and average precision (AP) by 3.00%. On the other hand, the copy‐paste technique achieved a 20.00% increase in accuracy and a 1% increase in AP with just 200 synthetic dents. Furthermore, these results provide support for using the copy‐paste augmentation strategy to enhance defect detection performance, with a limited dataset, contributing to more accurate defect identification in remanufacturing processes.

36 MATERIALS SCIENCE↗

ChemGraph as an agentic framework for computational chemistry workflows

Atomistic simulations are essential in chemistry and materials science but remain challenging to run due to the expert knowledge required for the setup, execution, and validation stages of these calculations. We present ChemGraph, an agentic framework powered by artificial intelligence and state-of-the-art simulation tools to streamline and automate computational chemistry and materials science workflows. ChemGraph leverages graph neural network-based foundation models for accurate yet computationally efficient calculations and large language models (LLMs) for natural language understanding, task planning, and scientific reasoning to provide an intuitive and interactive interface. We evaluate ChemGraph across 13 benchmark tasks and demonstrate that smaller LLMs (GPT-4o-mini, Claude-3.5-haiku, Qwen-2.5-14B) perform well on simple workflows, while more complex tasks benefit from using larger models. Importantly, we show that decomposing complex tasks into smaller subtasks through a multi-agent framework enables GPT-4o to reach perfect accuracy and smaller LLMs to match or exceed single-agent GPT-4o's performance in these benchmarks.

Computational chemistry↗

Rdesign: A data dictionary with relational database design capabilities in Ada

Data Dictionary is defined to be the set of all data attributes, which describe data objects in terms of their intrinsic attributes, such as name, type, size, format and definition. It is recognized as the data base for the Information Resource Management, to facilitate understanding and communication about the relationship between systems applications and systems data usage and to help assist in achieving data independence by permitting systems applications to access data knowledge of the location or storage characteristics of the data in the system. A research and development effort to use Ada has produced a data dictionary with data base design capabilities. This project supports data specification and analysis and offers a choice of the relational, network, and hierarchical model for logical data based design. It provides a highly integrated set of analysis and design transformation tools which range from templates for data element definition, spreadsheet for defining functional dependencies, normalization, to logical design generator.

Lekkos, Anthony A.↗

The Impact of Information Technology on the Design, Development, and Implementation of a Lunar Exploration Mission

From the beginning to the present expeditions to the Moon have involved a large investment of human labor. This has been true for all aspects of the process, from the initial design of the mission, whether scientific or technological, through the development of the instruments and the spacecraft, to the flight and operational phases. In addition to the time constraints that this situation imposes, there is also a significant cost associated with the large labor costs. As a result lunar expeditions have been limited to a few robotic missions and the manned Apollo program missions of the 1970s. With the rapid rise of the new information technologies, new paradigms are emerging that promise to greatly reduce both the time and cost of such missions. With the rapidly increasing capabilities of computer hardware and software systems, as well as networks and communication systems, a new balance of work is being developed between the human and the machine system. This new balance holds the promise of greatly increased exploration capability, along with dramatically reduced design, development, and operating costs. These new information technologies, utilizing knowledge-based software and very highspeed computer systems, will provide new design and development tools, scheduling mechanisms, and vehicle and system health monitoring capabilities that have hitherto been unavailable to the mission and spacecraft designer and the system operator. This paper will utilize typical lunar missions, both robotic and crewed, as a basis to describe and illustrate how these new information system technologies could be applied to all aspects such missions. In particular, new system design tradeoff tools will be described along with technologies that will allow a very much greater degree of autonomy of exploration vehicles than has heretofore been possible. In addition, new information technologies that will significantly reduce the human operational requirements will be discussed.

Gross, Anthony R.↗

Optimizing Cryo-Focused Pyrolysis GC/MS for Tracing Soil Organic Matter Across Diverse Ecosystems

The cycling of organic matter in terrestrial soils and sediments is central to a range of biogeochemical processes that regulate nutrient cycling, crop productivity, trace gas emissions, and contaminant transport. Pyrolysis-gas chromatography/mass spectrometry (py-GC/MS) is a powerful tool for characterizing bulk soil organic matter (SOM) at the molecular level. In this study, we used a cryo-focused py-GC/MS system to analyze soil samples from seven diverse ecosystems: vernal pool, prairie pothole, temperate forest, tropical forest, tundra, wildfire-affected boreal forest, and grassland. We addressed a key bottleneck in molecular-level SOM characterization by developing an automated data analysis pipeline to optimize py-GC/MS and complementary evolved gas analysis/mass spectrometry (EGA/MS) methods, incorporating advanced tools for peak deconvolution, developing a custom compound class library, and implementing fragmentation spectrum-based molecular networking for the first time. This improved workflow was applied to soil samples from all seven ecosystems, including multiple depths and density fractions. Our findings demonstrate that ecosystem type plays a dominant role in shaping compositional differences in SOM. We also identified trends in the source of SOM compounds (e.g., microbial vs plantderived) across soil depth and density fractions, which are critical for understanding persistence and turnover of SOM. Our molecular networking analysis indicated that although many compounds are widespread across ecosystems, others are restricted to specific environments, such as wetlands. This underscores the utility of molecular-level data in elucidating the complexity of SOM composition and the environmental drivers that shape it. Such molecular-level insights can deepen our knowledge of biogeochemical SOM cycles.

54 ENVIRONMENTAL SCIENCES↗

AI-Powered Knowledge Graphs for Neuromorphic and Energy-Efficient Computing

The surge in scientific literature obscures breakthroughs and hinders the discovery of new research paths. We propose an artificial intelligence (AI) powered framework using large language models (LLMs) and knowledge graphs (KGs) to automate parts of scientific discovery, focusing on energy-efficient AI circuits. Our hybrid approach combines LLMs, structured data, and ontology-based reasoning to construct a comprehensive knowledge graph that integrates insights across computational neuroscience, spiking neuron models, learning rules, architectural motifs, and neuromorphic device technologies. This multi-domain representation enables the generation of hypotheses that connect biological function with implementable, energy-efficient hardware architectures. Using KG embeddings and graph neural networks, the framework generates hypotheses for novel circuits, validates them through optimization on exascale HPC systems, and with tools like SuperNeuro and Fugu, the most promising designs will be prototyped in hardware. This open-source system aims to accelerate discoveries and bridging neuroscience with hardware innovation, drive collaboration, and unlock new opportunities in low-power AI computing.

Gautam, Ashish [ORNL]↗

Cloud Thermodynamic Phase Detection with Polarimetrically Sensitive Passive Sky Radiometers

The primary goal of this project has been to investigate if ground-based visible and near-infrared passive radiometers that have polarization sensitivity can determine the thermodynamic phase of overlying clouds, i.e. if they are comprised of liquid droplets or ice particles. While this knowledge is important by itself for our understanding of the global climate, it can also help improve cloud property retrieval algorithms that use total (unpolarized) radiance to determine Cloud Optical Depth (COD). This is a potentially unexploited capability of some instruments in the NASA Aerosol Robotic Network (AERONET), which, if practical, could expand the products of that global instrument network at minimal additional cost. We performed simulations that found, for zenith observations, cloud thermodynamic phase is often expressed in the sign of the Q component of the Stokes polarization vector. We chose our reference frame as the plane containing solar and observation vectors, so the sign of Q indicates the polarization direction, parallel (negative) or perpendicular (positive) to that plane. Since the quantity of polarization is inversely proportional to COD, optically thin clouds are most likely to create a signal greater than instrument noise. Besides COD and instrument accuracy, other important factors for the determination of cloud thermodynamic phase are the solar and observation geometry (scattering angles between 40 and 60 degrees are best), and the properties of ice particles (pristine particles may have halos or other features that make them difficult to distinguish from water droplets at specific scattering angles, while extreme ice crystal aspect ratios polarize more than compact particles). We tested the conclusions of our simulations using data from polarimetrically sensitive versions of the Cimel 318 sun photometerradiometer that comprise AERONET. Most algorithms that exploit Cimel polarized observations use the Degree of Linear Polarization (DoLP), not the individual Stokes vector elements (such as Q). For this reason, we had no information about the accuracy of Cimel observed Q and the potential for cloud phase determination. Indeed, comparisons to ceilometer observations with a single polarized spectral channel version of the Cimel at a site in the Netherlands showed little correlation. Comparisons to Lidar observations with a more recently developed, multi-wavelength polarized Cimel in Maryland, USA, show more promise. This divergence between simulations and observations has prompted us to begin the development of a small test instrument called the Sky Polarization Radiometric Instrument for Test and Evaluation (SPRITE). This instrument is specifically devoted to the accurate observation of Q, and the testing of calibration and uncertainty assessment techniques, with the ultimate goal of understanding the practical feasibility of these measurements.

Knobelspiesse, Kirk D.↗

3D-CHESS: Decentralized, Distributed, Dynamic, and Context-aware Heterogeneous Sensor Systems

This paper describes the objectives and current status of the 3D-CHESS project which aims to demonstrate a new Earth observing strategy based on a context-aware Earth observing sensor web. This sensor web consists of a set of nodes with a knowledge base, heterogeneous sensors, edge computing, and autonomous decision-making capabilities. Context awareness is defined as the ability for the nodes to gather, exchange, and leverage contextual information (e.g., state of the Earth system, state and capabilities of itself and of other nodes in the network, and how those states relate to the dy- namic mission objectives) to improve decision making and planning. The current goal of the project is to demonstrate proof of concept by comparing the performance of a 3D- CHESS sensor web with that of status quo architectures in the context of a multi-sensor inland hydrologic and ecologic monitoring system.

David, Cedric H.↗

Baseflow Identification via Explainable AI With Kolmogorov‐Arnold Networks

Abstract Hydrological models often involve constitutive laws that may not be optimal in every application. We propose to replace such laws with the Kolmogorov‐Arnold networks (KANs), a class of neural networks designed to identify symbolic expressions. We demonstrate KAN's potential on the problem of baseflow identification, a notoriously challenging task plagued by significant uncertainty. KAN‐derived functional dependencies of the baseflow components on the aridity index outperform their original counterparts; they demonstrate that water availability, rather than potential evapotranspiration, drives baseflow by constraining actual evapotranspiration under arid conditions. On a test set, they increase the Nash‐Sutcliffe efficiency (NSE) by 65%, decrease the root mean squared error by 29%, and increase the Kling‐Gupta efficiency by 34%. This superior performance is achieved while reducing the number of fitting parameters from three to two. Next, we use data from 378 catchments across the continental United States to refine the water‐balance equation at the mean‐annual scale. The KAN‐derived equations based on the refined water balance outperform both the current aridity index model, with up to a 105% increase in NSE, and the KAN‐derived equations based on the original water balance. While the performance of our model and tree‐based machine learning methods is similar, KANs offer the advantage of simplicity and transparency and require no specific software or computational tools. This case study focuses on the aridity index formulation, but the approach is flexible and transferable to other hydrological processes. Plain Language Summary Equations used in hydrologic model are often suboptimal, resulting in reduced prediction accuracy and efficiency. We implemented Kolmogorov‐Arnold networks (KAN), a machine learning algorithm for deriving symbolic formulations, to estimate groundwater recharge and showed that it outperforms an existing state‐of‐the‐art semi‐empirical formulation. In hydrology, Nash‐Sutcliffe efficiency (NSE), root mean squared error (RMSE), and Kling‐Gupta efficiency (KGE) are commonly used to evaluate model performance. Higher NSE and KGE values indicate better performance, while lower RMSE values are preferable. Our results show that NSE increased by 71%, RMSE decreased by 32%, and KGE improved by 25%. In addition, KAN identifies an optimal functional form and can be used to derive new analytical formulas using the prior knowledge. The KAN‐inspired equation outperformed the original formulation and reduced the fitting parameters. Furthermore, we refined the water‐balance equation at the mean‐annual scale and showed that, based on the new water‐balance equation, KAN can derive new formulations that are superior to the original aridity index formulations (up to 105% increase in NSE) and KAN‐derived equations based on the original water balance. These findings highlight the significant potential of KAN to advance the scientific understanding of a wide range of hydrologic processes. Key Points Kolmogorov‐Arnold networks (KANs) enhance interpretability of machine‐learned hydrological models KAN‐derived symbolic formulations outperform state‐of‐the‐art semi‐empirical aridity indices KAN‐identified functional form yields an analytical index with fewer fitting parameters and improved performance

baseflow↗

Voucher Opportunity 5-15: Independent Assessment of Monitoring, Reporting, and Verification (MRV) Technologies and Practices for Enhanced Rock Weathering (CRADA 718) Abstract

Development of robust, transparent, and precise monitoring, reporting, and verification (MRV) technologies and practices is critical for carbon dioxide removal (CDR) project developers to comply with regulatory and permitting requirements, voluntary carbon market (VCM) protocols, and to ensure safety while reducing environmental impacts. Enhanced rock weathering (ERW)-based CDR technologies focus on removing atmospheric carbon through conversion into thermodynamically stable solid or aqueous carbonate forms for permanent storage (i.e., mineralization). This highly durable form of CDR enhances naturally occurring silicate rock weathering cycles by optimizing application of finely-ground silicate rock particles (i.e., from basalt) on terrestrial agricultural lands to accelerate natural silicate rock weathering and mineralization. Enhanced rock weathering may also provide improved crop yields and enhance soil health. A critical aspect for commercialization of these technologies is the development of MRV to quantify the net removal and durable storage of atmospheric CO 2 . For ERW systems, it is essential to accurately characterize the mineral feedstock selected for application to establish the baseline geochemical composition, mineral dissolution rates, and carbon removal potential of the feedstocks to estimate overall net removal. Given the difficulty with conducting MRV for ERW in diverse soil/environment types, over large application areas, and due to complex chemical reaction networks, this project will accelerate understanding towards consensus on best practices for MRV. The overall objectives of the proposed voucher project are to: 1) Characterize and analyze feedstock(s) intended for ERW field application by Lithos Carbon (“Voucher Recipient”/ “CRADA Participant”) to determine overall mineralization potential; 2) Facilitate knowledge transfer and documentation of experimental protocols, instrumentation, and other relevant best practices; and 3) Support the Voucher Recipient’s broader technology commercialization and ERW Research Facility development plans. This work will align with the Voucher Recipient’s MRV plans for field sites and build upon complementary efforts conducted by PNNL on mineralization MRV.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Using Transparent Informed Prefetching (TIP) to reduce file read latency

As processor performance gains continue to outstrip Input/Output gains, I/O performance is becoming critical to overall system performance. File read latency is the most significant bottleneck for high performance I/O. Other aspects of I/O performance benefit from recent advances in disk bandwidth and throughput resulting from disk arrays, and in write performance derived from buffered write behind and the Log-structured File System. The access gap problem limiting improvements in read latency is exacerbated by distributed file systems operating over networks with diverse bandwidth. Focus is on extending the power of caching and prefetching to reduce file read latencies by exploiting hints from high-levels of a system. Such Transparent Informed Prefetching, TIP, and its benefits are described. It is argued that hints that disclose high level knowledge are a means for transferring optimization information across, without violating, module boundaries. How TIP can be used to convert the high throughput of new technologies such as disk arrays and log-structured file systems into low latency for applications is discussed. Our preliminary experiments show reductions in wall - clock execution time of 13 percent and 20 percent for a multiple module compilation tool (make) accessing data on a local disk and remote Coda file server, respectively, and a reduction of 30 percent for a text search (grep) remotely accessing many small files.

Patterson, R. H.↗

Forward and Backscattering Measurements of Rainfall using the NASA Microwave Link

This paper will present results from a study on the feasibility of making backscatter measurements from rainfall with the NASA/Microwave Link system at Wallops Island, VA. The study entails the implementation of an FMCW radar at the Link frequencies to enable simultaneous forward and backscatter measurements from rain. The Microwave Link system has been successfully employed in the development and testing of rainfall retrieval techniques. As presently configured, the Microwave Link measures attenuation and phase-shift due to rain over a 2.3 km path between the transmitting and receiving antennas. By their very nature, these measured quantities are averaged over the propagation path. As a result, the rainfall estimates obtained from the Link data are also averaged over the propagation path. However, rainfall is a highly variable process in space (as well as time). In order to gain a more detailed knowledge of its microphysics finer spatial resolutions are required. The backscatter measurements would enable range profiling over the Link path permitting a detailed study of the rainfall process. The backscatter measurements will be used in conjunction with the forward measurements and the measurements from a ground-based network of disdrometers and rain gauges located under the propagation path to develop new microwave retrieval techniques, and to test established single-frequency and dual-frequency radar retrieval algorithms relevant to the ongoing TRMM and up coming GPM missions.

Rincon, Rafael F.↗

Feature Acquisition with Imbalanced Training Data

This work considers cost-sensitive feature acquisition that attempts to classify a candidate datapoint from incomplete information. In this task, an agent acquires features of the datapoint using one or more costly diagnostic tests, and eventually ascribes a classification label. A cost function describes both the penalties for feature acquisition, as well as misclassification errors. A common solution is a Cost Sensitive Decision Tree (CSDT), a branching sequence of tests with features acquired at interior decision points and class assignment at the leaves. CSDT's can incorporate a wide range of diagnostic tests and can reflect arbitrary cost structures. They are particularly useful for online applications due to their low computational overhead. In this innovation, CSDT's are applied to cost-sensitive feature acquisition where the goal is to recognize very rare or unique phenomena in real time. Example applications from this domain include four areas. In stream processing, one seeks unique events in a real time data stream that is too large to store. In fault protection, a system must adapt quickly to react to anticipated errors by triggering repair activities or follow- up diagnostics. With real-time sensor networks, one seeks to classify unique, new events as they occur. With observational sciences, a new generation of instrumentation seeks unique events through online analysis of large observational datasets. This work presents a solution based on transfer learning principles that permits principled CSDT learning while exploiting any prior knowledge of the designer to correct both between-class and withinclass imbalance. Training examples are adaptively reweighted based on a decomposition of the data attributes. The result is a new, nonparametric representation that matches the anticipated attribute distribution for the target events.

Thompson, David R.↗

The Geostationary Operational Satellite R Series SpaceWire Based Data System

The Geostationary Operational Environmental Satellite R-Series Program (GOES-R, S, T, and U) mission is a joint program between National Oceanic & Atmospheric Administration (NOAA) and National Aeronautics & Space Administration (NASA) Goddard Space Flight Center (GSFC). SpaceWire was selected as the science data bus as well as command and telemetry for the GOES instruments. GOES-R, S, T, and U spacecraft have a mission data loss requirement for all data transfers between the instruments and spacecraft requiring error detection and correction at the packet level. The GOES-R Reliable Data Delivery Protocol (GRDDP) [1] was developed in house to provide a means of reliably delivering data among various on board sources and sinks. The GRDDP was presented to and accepted by the European Cooperation for Space Standardization (ECSS) and is part of the ECSS Protocol Identification Standard [2]. GOES-R development and integration is complete and the observatory is scheduled for launch November 2016. Now that instrument to spacecraft integration is complete, GOES-R Project reviewed lessons learned to determine how the GRDDP could be revised to improve the integration process. Based on knowledge gained during the instrument to spacecraft integration process the following is presented to help potential GRDDP users improve their system designs and implementation.

Networks↗

SafeAeroBERT: Towards a Safety-Informed Aerospace-Specific Language Model

As aviation systems continue to operate with high traffic, large amounts of documents containing safety-relevant data continue to be generated via reporting systems such as the ASRS. Advanced natural language processing techniques, specifically pre-trained language models, have shown great success in domain-specific applications; however, the text in aviation safety reports is inundated with jargon and thus not fully utilized by general pre-trained models. In this research, we work towards developing a safety-informed aerospace-specific language model by pre-training a Bidirectional Encoder Representations from Transformer (BERT) model on reports from the Aviation Safety Reporting System and the National Transportation Safety Board. The resulting model, called SafeAeroBERT, is fine-tuned for the specific task of document classification, and can be further tuned for named-entity recognition, relation detection, information retrieval, and summarization. Results from the classification task are compared between SafeAeroBERT, the base BERT, and SciBERT models and show SafeAeroBERT outperforms the general BERT and SciBERT on classifying reports about human factors, aircraft, and procedure. SafeAeroBERT can be used on custom tasks, not limited to document classification, and is intended to aid an intelligent knowledge manager for safety report repositories.

Aviation↗

Intelligent Systems Technologies and Utilization of Earth Observation Data

The addition of raw data and derived geophysical parameters from several Earth observing satellites over the last decade to the data held by NASA data centers has created a data rich environment for the Earth science research and applications communities. The data products are being distributed to a large and diverse community of users. Due to advances in computational hardware, networks and communications, information management and software technologies, significant progress has been made in the last decade in archiving and providing data to users. However, to realize the full potential of the growing data archives, further progress is necessary in the transformation of data into information, and information into knowledge that can be used in particular applications. Sponsored by NASA s Intelligent Systems Project within the Computing, Information and Communication Technology (CICT) Program, a conceptual architecture study has been conducted to examine ideas to improve data utilization through the addition of intelligence into the archives in the context of an overall knowledge building system (KBS). Potential Intelligent Archive concepts include: 1) Mining archived data holdings to improve metadata to facilitate data access and usability; 2) Building intelligence about transformations on data, information, knowledge, and accompanying services; 3) Recognizing the value of results, indexing and formatting them for easy access; 4) Interacting as a cooperative node in a web of distributed systems to perform knowledge building; and 5) Being aware of other nodes in the KBS, participating in open systems interfaces and protocols for virtualization, and achieving collaborative interoperability.

Ramapriyan, H. K.↗

Rewiring the unfolded protein response for plant growth recovery after stress

The unfolded protein response (UPR) is a highly coordinated signaling network that alleviates endoplasmic reticulum (ER) stress, a condition induced by diverse environmental challenges in plants. Over the past two decades, substantial progress has been made in elucidating the genetic and molecular mechanisms of ER stress sensing and signal transduction in plants, largely through studies in the model plant Arabidopsis thaliana . These advances have established the UPR as a central regulator of proteostasis and underscored its broader relevance to plant growth and development and crop productivity under stress conditions. Despite this progress, critical knowledge gaps remain, particularly concerning the downstream biological processes required for growth recovery once ER stress has subsided and how these processes are coordinated by UPR regulators. Recent systems-level and integrative studies have begun to reveal critical roles of UPR signaling in pathways governing growth re-establishment and homeostasis of nutrient allocation and energy metabolism. In this review, we highlight recent findings on the functional roles of the plant UPR in recovery from ER stress, with a focus on mechanisms mediated by UPR regulators and downstream biological pathways that enable the transition from stress mitigation to growth restoration. Although this research area is still emerging, accumulating evidence supports a model in which the UPR functions as a dynamic regulatory network that actively coordinates post-stress physiological recovery to support plant fitness.

ER stress↗

Semantic Search with Sentence-BERT for Design Information Retrieval

Managing and referencing design knowledge is a critical activity in the design process. However, reliably retrieving useful knowledge can be a frustrating experience for users of knowledge management systems due to inherent limitations of standard keyword-based searches. In this research, we consider the task of retrieving relevant lessons learned from the NASA Lessons Learned Information System (LLIS). To this end, we apply a state-of-the-art natural language processing (NLP) technique for information retrieval (IR): semantic search with sentence-BERT, which is a modification of a Bidirectional Encoder Representations from Transformers (BERT) model that uses siamese and triplet network architectures to obtain semantically meaningful sentence embeddings. While the pre-trained sBERT model performs well out-of-the-box, we further fine-tune the model on data from the LLIS so that it learns on design engineering-relevant vocabulary. We quantify the improvement in query results using both standard sBERT and fine-tuned sBERT over a keyword search. Our use case throughout the paper is to use queries related to specific requirements from a NASA project. Fine tuning the sBERT model on LLIS data yields a mean average precision (MAP) of 0.807 on queries based on information needs from a real NASA project. Results indicate that applying state-of-the-art natural language processing techniques, especially when finetuned using engineering data, to design information retrieval tasks shows significant promise in modernizing design knowledge management systems.

Hannah S. Walsh↗