Search NASA⌕ Search

SEARCH · Search NASA

Results for “Domain knowledge”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Constructing A New CHF Look-Up Table Based on the Domain Knowledge Informed Machine Learning Methodology

Accurate prediction of CHF under various fluid flow conditions continues to be required for design, operation and safety analysis of light water reactor rod bundles. Due to the lack of in-depth physical understanding as well as limited high-resolution data in the micro-scale flow and heat transfer, the existing models feature a sub-optimal uncertainty band. In this study, driven by the prior domain knowledge information obtained, an improved CHF look-up table is developed through unified machine learning algorithms for the vertical flow conditions within tube and annulus geometry. The Groeneveld 2006 look-up table is used as the domain knowledge to train machine learning process against tube and annulus CHF data for both DNB and DO type. The new look-up table shows improved accuracy for conditions relevant to PWRs and BWRs. In addition, its domain knowledge informed nature ensures that a rationale prediction can be made, thus accounting for previous valuable information in the machine learning model training process.

Jin, Yue↗

Di-CNN: Domain-Knowledge-Informed Convolutional Neural Network for Manufacturing Quality Prediction

In manufacturing, convolutional neural networks (CNNs) are widely used on image sensor data for data-driven process monitoring and quality prediction. However, as purely data-driven models, CNNs do not integrate physical measures or practical considerations into the model structure or training procedure. Consequently, CNNs’ prediction accuracy can be limited, and model outputs may be hard to interpret practically. This study aims to leverage manufacturing domain knowledge to improve the accuracy and interpretability of CNNs in quality prediction. A novel CNN model, named Di-CNN, was developed that learns from both design-stage information (such as working condition and operational mode) and real-time sensor data, and adaptively weighs these data sources during model training. It exploits domain knowledge to guide model training, thus improving prediction accuracy and model interpretability. A case study on resistance spot welding, a popular lightweight metal-joining process for automotive manufacturing, compared the performance of (1) a Di-CNN with adaptive weights (the proposed model), (2) a Di-CNN without adaptive weights, and (3) a conventional CNN. The quality prediction results were measured with the mean squared error (MSE) over sixfold cross-validation. Model (1) achieved a mean MSE of 6.8866 and a median MSE of 6.1916, Model (2) achieved 13.6171 and 13.1343, and Model (3) achieved 27.2935 and 25.6117, demonstrating the superior performance of the proposed model.

47 OTHER INSTRUMENTATION↗

Domain Knowledge Guided Bayesian Optimization For Autonomous Alignment Of Complex Scientific Instruments

Bayesian Optimization (BO) is a powerful tool for optimizing complex non-linear systems. However, its performance degrades in high-dimensional problems with tightly coupled parameters and highly asymmetric objective landscapes, where rewards are sparse. In such needle-in-a-haystack scenarios, even advanced methods like trust-region BO (TurBO) often lead to unsatisfactory results. We propose a domain knowledge guided Bayesian Optimization approach, which leverages physical insight to fundamentally simplify the search problem by transforming coordinates to decouple input features and align the active subspaces with the primary search axes. We demonstrate this approach's efficacy on a challenging 12-dimensional, 6-crystal Split-and-Delay optical system, where conventional approaches, including standard BO, TuRBO and multi-objective BO, consistently led to unsatisfactory results. When combined with an reverse annealing exploration strategy, this approach reliably converges to the global optimum. The coordinate transformation itself is the key to this success, significantly accelerating the search by aligning input co-ordinate axes with the problem's active subspaces. As increasingly complex scientific instruments, from large telescopes to new spectrometers at X-ray Free Electron Lasers are deployed, the demand for robust high-dimensional optimization grows. Our results demonstrate a generalizable paradigm: leveraging physical insight to transform high-dimensional, coupled optimization problems into simpler representations can enable rapid and robust automated tuning for consistent high performance while still retaining current optimization algorithms.

FOS: Computer and information sciences↗

Impact of Domain Knowledge on the Property Prediction of Specialized Machine Learning Models

Developing transferable machine learning models is trending in data-driven materials research. However, how to apply such models to a specific research domain remains unclear. Here, in this work, we choose high-entropy materials as a platform with a specialized data set containing 145,323 DFT-relaxed materials. This data set is used to explore the role of domain-specific knowledge in training effective models. Our tests with three representative graph neural network architectures indicate the model complexity has much smaller influence on performance than the data itself. Specifically, the consideration of low-energy atomic ordering, structures with diverse elemental coverage, and high-order interactions significantly influences the model performance. We also find that domain knowledge-driven sampling can greatly enhance unsupervised learning techniques. This research highlights that developing specialized data sets is more beneficial than further complicating deep learning architectures. Additionally, physics-inspired sampling algorithms are crucially needed for better machine learning models for a specific materials research domain.

36 MATERIALS SCIENCE↗

Physics-guided neural networks with engineering domain knowledge for hybrid process modeling

As neural networks are more frequently used to solve problems in science and engineering, the methods used to incorporate scientific knowledge into these networks are becoming increasingly complex. Here, this work breaks down these complicated techniques into a set of basic strategies which can easily be applied to diverse situations. Several novel neural networks are built using the categories laid out in this work. These networks are tested on simulated data from a continuous stirred tank reactor (CSTR) model to evaluate the advantages provided by each network. The three points demonstrated in this work are: (1) architectural hybrid models can speed up convergence and reduce the amount of data necessary to train a model; (2) adding a physics-guided loss function can improve model generalization and make models more physically consistent; (3) using physics-guided initialization and transfer learning improves accuracy and speeds up convergence, but can harm generalizability if used incorrectly.

42 ENGINEERING↗

Materials Data Science Ontology(MDS-Onto): Unifying Domain Knowledge in Materials and Applied Data Science

Ontologies have gained popularity in the scientific community as a way to standardize terminologies in organizations’ data. Although certain cohorts have created frameworks with rules and guidelines on creating ontologies, there exist significant variations in how Materials Science ontologies are currently developed. We seek to provide guidance in the form of a unified automated framework for developing interoperable and modular ontologies for Materials Data Science that simplifies the ontology terms matching by establishing a semantic bridge up to the Basic Formal Ontology(BFO). This framework provides key recommendations on how ontologies should be positioned within the semantic web, what knowledge representation language is recommended, and where ontologies should be published online to boost their findability and interoperability. Two fundamental components of the MDS-Onto framework are the bilingual package called FAIRmaterials for ontology creation and FAIRLinked, for FAIR data creation. To showcase the practical capabilities of FAIRmaterials, we present two exemplar domain ontologies of MDS-Onto: Synchrotron X-Ray Diffraction and Photovoltaics.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Multimodal Approaches for Leveraging Domain Knowledge with State-of-the-Art Machine Learning to Engineer Biocatalysts

This grant aimed to accelerate the development of specialized enzymes—biological catalysts essential for sustainable manufacturing and medicine—by integrating traditional laboratory evolution with cutting-edge artificial intelligence. To achieve this, we developed a suite of high-throughput sequencing tools and a centralized database to bridge the gap between a protein’s genetic "code" and its physical function. By training machine learning models on large datasets, we also demonstrated the ability to move beyond slow, trial-and-error testing to a "generative" approach, where AI can independently design new, versatile enzymes like tryptophan synthases. Ultimately, these findings demonstrate that combining laboratory data with computer-guided design enables the engineering of highly efficient biological tools with unprecedented speed and precision.

59 BASIC BIOLOGICAL SCIENCES↗

Performance Results for Sensor Assignment Problem as Solved on a Multi-Node Cluster

An earlier report described a procedure for optimal sensor set selection and its implementation on a computational cluster. This new and innovative capability was developed to facilitate a reduction in operations staffing levels to improve plant economics. By automating surveillance and maintenance tasks through early detection of degrading sensors and equipment, staff can be more efficiently deployed. The method uses automated reasoning and domain knowledge in the form of the conservation equations to infer from plant measurements the state of equipment health. Inclusion of domain knowledge addresses the problem that exists with pure data-driven methods that there are no rigorous guidelines for determining what constitutes an adequate sensor set. Formalizing the procedure for sensor set selection as we have done results in a more reliable and explainable diagnosis of plant equipment health. Importantly, from the standpoint of the plant owner, personnel are provided with an early and explicit diagnosis of an equipment problem. That in principle automates the process and eliminates having to send personnel into the plant to find the cause as typically occurs when a data-driven method detects an anomaly. In this report we describe first results obtained using a computational cluster to solve the sensor set selection problem as framed above. The case described addresses the problem of equipment health monitoring in the high-pressure (HP) feedwater system of a pressurized light water reactor as seen through the eyes of our collaborating utility partner. Maintenance of this system can amount to millions of dollars per year if equipment health issues go undiagnosed and lead to loss of function. On examining the potential that is inherent in the installed sensor set for diagnosing equipment health degradation, it was found that greater fault resolution capability can be achieved using a sensor set that is 20 percent fewer in number. The take-away is that compared to the installed sensor set there exists a more strategic assignment of sensors that will furnish better health monitoring capability and with fewer sensors. Where the problem defies solution by manual inspection, as is the case here, one can be found by an algorithm. The solution was obtained in four hours using 30 computational cores. The HP feedwater problem as posed above illustrates the added value of approaching the sensor selection problem as one amenable to algorithmic solution. This problem is of interest to advanced reactor designers and to utilities that are setting up remote monitoring and diagnostic centers.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Empirically Categorizing the Built Environment in Relation to Height

Buildings are a core component of the urban environment and affect human populations, energy usage, city development, city planning, and urban heat islands. Buildings span an enormous range of sizes, from a 2m tall shelter to the Burj Khalifa; and at the same time there are widely recognized categories of similar buildings, with homes, office buildings, or skyscrapers as some examples. Currently, there is no consistent method to quantitatively determine how a building should be categorized by its height, or how many categories there should be within the built environment. Additionally, these categories vary spatially, leading to multiple definitions at local scales of what it means to be a tall, medium, or short building. Here, we find across 17.59 million buildings in the United States, Germany, and Japan, that applying a K-nearest neighbor approach to quantitatively bin the built environment outperforms the current state-of-the-art, subjective domain knowledge. This was evidenced as our method of leveraging a K-nearest neighbor improved upon the existing approach of using domain knowledge by 10% with respect to precision, recall, F1-score and accuracy. Our results showcase the finding that it is possible to generate a global and consistent approach to categorizing the built environment in relation to height. This is significant in that there is now a quantitative way to categorize the built environment based on building height at a global scale, allowing researchers a consistent platform for comparison and collaboration across various applications.

Stipek, Clinton↗

Chemical reaction enhanced graph learning for molecule representation

Abstract Motivation Molecular representation learning (MRL) models molecules with low-dimensional vectors to support biological and chemical applications. Current methods primarily rely on intrinsic molecular information to learn molecular representations, but they often overlook effectively integrating domain knowledge into MRL. Results In this article, we develop a reaction-enhanced graph learning (RXGL) framework for MRL, utilizing chemical reactions as domain knowledge. RXGL introduces dual graph learning modules to model molecule representation. One module employs graph convolutions on molecular graphs to capture molecule structures. The other module constructs a reaction-aware graph from chemical reactions and designs a novel graph attention network on this graph to integrate reaction-level relations into molecular modeling. To refine molecule representations, we design a reaction-based relation learning task, which considers the relations between the reactant and product sides in reactions. In addition, we introduce a cross-view contrastive task to strengthen the cooperative associations between molecular and reaction-aware graph learning. Experiment results show that our RXGL achieves strong performance in various downstream tasks, including product prediction, reaction classification, and molecular property prediction. Availability and implementation The code is publicly available at https://github.com/coder-ACAC/RLM.

Biochemistry & Molecular Biology↗

Deep Learning-Based Failure Prognostic Model for PV Inverter Using Field Measurements

Here, this study presents a novel approach for the precise monitoring and prognosis of photovoltaic (PV) inverter status, which is crucial for the proactive maintenance of PV systems. It addresses the gaps in traditional model-based methods, which tend to neglect the overall reliability of inverters, and the limitations of data-driven approaches that largely depend on simulated data. This research presents a robust solution applicable to real-world scenarios. The proposed data-driven model for PV inverter failure prognosis employs actual inverter measurements, integrating various operational and weather-related factors based on domain knowledge. This approach effectively represents inverter stressors and operational status. Utilizing an Enhanced Siamese Convolutional Neural Network (ESCNN), the model merges operational data with domain knowledge features, redefining the prognosis challenge as a classification task. Furthermore, the paper discusses an ESCNN-based real-time inverter failure monitoring method developed on the well-trained model. The proposed models are rigorously trained and tested with real inverter data and a novel filtering method is included to address accidental failures in practical scenarios. The results validate the model's efficacy, and the directions for future research are also outlined.

42 ENGINEERING↗

Spatio-Temporal Denoising Graph Autoencoders with Data Augmentation for Missing Photovoltaic Data Imputation

The integration of the global Photovoltaic (PV) market with real time data-loggers has enabled large scale PV data analytical pipelines for power forecasting and long-term reliability assessment of PV fleets. Nevertheless, the performance of PV data analysis heavily depends on the quality of PV timeseries data. This paper proposes a novel Spatio-Temporal Denoising Graph Autoencoder (STD-GAE) framework to impute missing PV Power Data. STDGAE exploits temporal correlation, spatial coherence, and value dependencies from domain knowledge to recover missing data. It is empowered by two modules. (1) To cope with sparse yet various scenarios of missing data, STD-GAE incorporates a domain-knowledge aware data augmentation module that creates plausible variations of missing data patterns. This generalizes STD-GAE to robust imputation over different seasons and environment. (2) STD-GAE nontrivially integrates spatiotemporal graph convolution layers (to recover local missing data by observed “neighboring” PV plants) and denoising autoencoder (to recover corrupted data from augmented counterpart) to improve the accuracy of imputation accuracy at PV fleet level. We have evaluated our proposed model on two realworld PV datasets. Experimental results show that STD-GAE can achieve a gain of 43.14% in imputation accuracy and remains less sensitive to missing rate, different seasons, and missing scenarios, compared with state-of-the-art data imputation methods such as MIDA and LRTC-TNN.

Fan, Yangxin↗

Deep-freeze graph training for latent learning

Scientific and engineering advances are primarily driven by multi-tier conceptual constructs and conditional theoretical frameworks. The theories allow predictions of hypothetical system responses, given a set of approximate conditions (ranges of applicability) imposed on latent parameters that cannot be measured directly. Learning to estimate the latent variables (Latent Learning) helps to pinpoint the anticipated range-edge anomalies and improves the confidence in interpretation, interpolation and extrapolation of limited experimental data. Due to high dimensionality and extreme non-linearity of the materials science problems, very large datasets are typically required for conventional data-driven model development. The vital experimental data collection, particularly on microstructural phases, is very challenging, which makes it difficult to compile a high-quality database. Incorporation of the domain knowledge into the computational graph structure, initialization and optimization processes presents a viable mechanism for developing accurate models, with limited datasets. Furthermore, this study successfully utilized the approach to build the Deep Freeze Graph (DeepFreG) by mapping known causality relationships and by digitizing empirical domain knowledge for Latent Learning (LL), with specific applications in materials science.

36 MATERIALS SCIENCE↗

Real-time Event Detection Using Rank Signatures of Real-world PMU Data

Timely detection of power system events is a crucial task, which can facilitate the implementation of remedial actions to improve reliability, resiliency, and security of the system. Meanwhile, the widespread deployment of phasor measurement units (PMUs) makes it possible to develop data-driven event detection techniques. However, relying purely on data without incorporating domain knowledge for the event detection task in power systems poses substantial security and stability risks due to issues associated with data misinterpretation and model accuracy. In this regard, we propose a real-time event detection method using real-world PMU data by incorporating domain knowledge to adequately capture the event signatures. Specifically, we track the change in rank signatures of PMU data to accurately localize the events. To optimize the detection process, we incorporate an offline Bayesian optimization algorithm to tune the parameters by efficiently searching for the best values. The experiments using the real-world PMU dataset from a U.S. interconnection show that the proposed event detection approach can efficiently detect the events from PMU data streams with high accuracy.

Ghasemkhani, Amir↗

Machine learning–assisted prediction of heat fluxes through thermally anisotropic building envelopes

Thermally anisotropic building envelope (TABE) is a novel active building envelope that can save energy use to maintain thermal comfort in buildings by redirecting heat and coolness from building envelopes to thermal loops. Finite element models (FEMs) can be used to compute the heat fluxes through TABEs, but the high computational cost of finite element simulations has prevented parametric studies and design optimizations. This paper proposes a domain knowledge–informed, finite element–based machine learning framework to reduce the computation cost for the energy management of buildings installed with TABE that uses a ground thermal loop. First, the training heat flux data set was generated by FEM simulations with different thermal loop schedules. Then, both shallow learning models (i.e., multivariate linear regression and eXtreme Gradient Boost, or XGBoost) and a deep learning model (i.e., deep neural network, or DNN) were trained to predict the heat fluxes. Domain knowledge was used for data preprocessing and feature selection. Finally, the suitability of the selected machine learning model was tested under different thermal loop schedules. Herein, the case study results showed that: (1) XGBoost can be as accurate as DNN (coefficient of determination equal to 0.81) with much less training time; (2) the annual energy cost savings for different thermal loop schedules obtained by the XGBoost-predicted and FEM-calculated heat fluxes are consistent, having a difference of only 4%; and (3) XGBoost can reduce the computation time for the annual energy analysis of the case study building with a given thermal loop schedule from around 12 h by using FEM to less than 1 min.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗