Search NASASearch

SEARCH · Search NASA

Results for “Feature engineering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Recent Progress on Surface Water Quality Models Utilizing Machine Learning Techniques

Surface waterbodies are heavily exposed to pollutants caused by natural disasters and human activities. Empowering sensor technologies in water quality monitoring, sufficient measurements have become available to develop machine learning (ML) models. Numerous ML models have quickly been adopted to predict water quality indicators in various surface waterbodies. This paper reviews 78 recent articles from 2022 to October 2024, categorizing water quality models utilizing ML into three groups: Point-to-Point (P2P), which estimates the current target value based on other measurements at the same time point; Sequence-to-Point (S2P), which utilizes previous time series data to predict the target value at one time point ahead; and Sequence-to-Sequence (S2S), which uses previous time series data to forecast sequential target values in the future. The ML models used in each group are classified and compared according to water quality indicators, data availability, and model performance. Widely used strategies for improving performance, including feature engineering, hyperparameter tuning, and transfer learning, are recognized and described to enhance model effectiveness. The interpretability limitations of ML applications are discussed. This review provides a perspective on emerging ML for surface water quality models.

machine learning (ML)

Tandem Predictions for HPC Jobs: Preprint

At the core of the predictive analytics applied to High Performance Computing (HPC), the most prominent tasks are the prediction of job runtimes and the prediction of job queue times, both of which have the potential for informing HPC users during their every-day decision making. Accurate runtime predictions can help users better choose so-called wallclock times at job submission, decreasing the odds of their jobs waiting in queues longer than necessary. The accurate and timely queue time predictions offered for the available partitions can inform the favorable selection of partitions for running jobs. This potential is well understood as we see in the abundance of research studies that propose solutions for these tasks, including the work published in the last several years. These tasks are seemingly receptive to the Machine Learning (ML) solutions, considering that there is no shortage of training data where HPC centers over time run millions and millions of jobs. However, we study the existing research literature, as well as look for examples in the toolchains supported on the exemplar HPC facilities, and, surprisingly, do not find any practical solutions that are ready to be adopted. We interpret this as a manifestation of the shortage of UX/UI efforts that support HPC analytics and also as a sign that the research has not come to the consensus on solving these tasks. In this study, we aim to shed new light on the long-running task of job queue time prediction by exploring the utility of runtime predictions in improving prediction accuracy and, actually, predicting these two metrics together, in tandem. In other words, we show how runtime predictions become valuable input in the queue time modeling. We challenge the existing approaches to feature engineering for the queue time prediction and describe promising results we obtained for a large dataset of HPC jobs from a supercomputer at the National Renewable Energy Laboratory.

97 MATHEMATICS AND COMPUTING

CFD modeling of non-catalytic, partial-oxidation engine reformer for flare mitigation

Flaring associated natural gas is commonly employed in the oil and gas industry to reduce methane (CH 4 ) emissions but generates carbon dioxide (CO 2 ) and harmful pollutants, significantly contributing to air pollution and posing risks to public health. To mitigate this impact, M2X Energy Inc. has developed a small-scale, modular gas-to-methanol system. This system features an engine reformer that performs fuel-rich partial oxidation of wellhead gas to produce syngas—a mixture of carbon monoxide (CO) and hydrogen (H 2 )—followed by a downstream reactor for methanol synthesis. This study focused on computational fluid dynamics (CFD) modeling of the engine reformer to simulate partial oxidation chemistry, predict the rich-burn operating limit, and assess syngas quality, ultimately aiding in design and operational optimization. The CFD model, developed within a Reynolds-Averaged Navier-Stokes (RANS) turbulence framework, incorporated sub-models for turbulent combustion, a chemical mechanism with polycyclic aromatic hydrocarbon (PAH) pathways, and soot emissions to accurately capture the fuel-rich, turbulent jet ignition and combustion processes. Model validation against experimental data showed good agreement across pre- and main-chamber pressures, apparent heat release rates, and exhaust gas concentrations of key species (H 2 , CO, CO 2 , CH 4 ) for varying intake equivalence ratios. Here, the model identified a rich-burn operating limit near a fuel-air equivalence ratio of 2.35, consistent with experimental observations. Furthermore, syngas quality analysis revealed that extending the rich-burn limit through engine reformer optimization could enhance syngas production, contributing to higher methanol synthesis efficiency.

Computational Fluid Dynamics

Feature issue introduction: laser driven inertial confinement fusion and bridging the gaps to inertial fusion energy systems

Major fusion research milestones have been achieved using laser driven inertial confinement fusion (ICF) in recent years, and these successes have ignited tremendous enthusiasm for inertial fusion energy (IFE). However, the complexity and difficulty of obtaining fusion ignition with a laser driver in a research setting are often underappreciated, as are the gaps to high driver efficiency, high repetition rates, and laser and target durability requirements needs for IFE. On the academic side, several new research laser systems have been constructed over the past few years, enabling researchers to probe the limits of ICF physics and engineering. This feature issue highlights the challenges and capabilities of laser research and development targeted towards advancing IFE.

Physics - Plasma physics

A systematic review of machine learning in groundwater monitoring

With increasing concerns about water scarcity, groundwater has become crucial since this resource provides most of the freshwater needs. However, various human and natural activities often contaminate the groundwater, making it unsuitable for use. Over the years, scientists and engineers have used many methods to predict and track groundwater contamination as part of environmental monitoring. Consequently, there is an urgent need for improved methods, particularly in the face of increasing contamination. Machine learning has sometimes been used to monitor groundwater, air quality, and climate. Traditional methods must be improved due to the complexity and large amount of environmental data. This includes using hybrid models that combine traditional and new techniques. Despite the use of machine learning in many scientific areas, there is a lack of comprehensive reviews focusing on its use in environmental monitoring, especially groundwater monitoring. We aim to fill this gap by exploring machine-learning applications in groundwater monitoring. We discuss relevant methods, their limitations, and future potential. We summarize research on automating data processing and model training using groundwater sensor data. Our research underscores the transformative potential of machine learning to revolutionize long-term groundwater monitoring and contamination detection, providing valuable insights for future research and practical applications.

AI/ML

Systematic feature design for cycle life prediction of lithium-ion batteries during formation

Optimization of the formation step in lithium-ion battery manufacturing is challenging due to limited physical understanding of solid-electrolyte interphase formation and the long testing time (∼100 days) for cells to reach the end of life. We propose a systematic feature-design framework that requires minimal domain knowledge for accurate cycle life prediction during formation. By only using two simple Q (V) features designed from our framework, extracted from formation data without any additional diagnostic cycles, we achieved an average of 9.87% error for cycle life prediction. Here, the physics-based investigation guided by the two designed features shows that the voltage ranges identified by our framework capture the effects of formation temperature and microscopic-particle resistance heterogeneity. By designing highly predictive, robust, and interpretable features, our approach can accelerate industrial battery formation research, leveraging the interplay between data-driven feature design and mechanistic understanding.

25 ENERGY STORAGE

Comparison of Expert Vocabulary Usage Patterns Between Mental Health and Nonmental Health Clinicians When Diagnosing Pediatric Anxiety Disorders

Objective: To compare the utilization patterns of expert vocabulary (EVo) in diagnosing pediatric anxiety between mental health and non-mental health clinical notes from electronic health records to understand the role of Evo in informing classification and decision-making in anxiety diagnoses. Study design: We conducted a retrospective study using a cohort less than age 25 from Cincinnati Children's Hospital including 897 685 patients with 61 586 446 notes. We analyzed EVo, collected from mental health clinicians, in both mental and nonmental health notes. We compared classification accuracy using EVo-based patient-level embedding from all clinical notes, mental-health notes, and nonmental health notes for 2 tasks: 1) pre-vs postdiagnosis anxiety patients, and 2) prediagnosis anxiety vs nonanxiety patients. Results: EVo usage was highest in prediagnosis anxiety, lower in nonanxiety, and lowest in post-diagnosis. Classification models using EVo features from all, mental-health, and non-mental health notes showed similar F1 scores for prediagnosis anxiety (0.70 ± 0.2 for 2 categories). For anxiety vs nonanxiety classification, all clinical and nonmental health notes had better F1 scores than mental-health notes (above 0.90 for 3 categories). There was a notable difference in class-wise performance across both tasks. Conclusions: There are significant differences in anxiety EVo use between mental health and nonmental health clinicians. Despite less anxiety-specific terminology, non-mental health notes still captured key aspects of patient presentations, emphasizing the importance of including all clinicians' notes in analysis. EVo's utility for anxiety classification is most effective in prediagnostic phases, suggesting the need for a dedicated diagnostic lexicon and further study before incorporating EVo into classification models.

feature engineering

Monitoring Plan for the Idaho National Laboratory Remote Handled Low Level Waste Disposal Facility

This monitoring plan for Idaho National Laboratory’s Remote-Handled Low Level Waste Disposal Facility was developed to meet the requirements for monitoring low-level waste disposal facilities according to the U.S. Department of Energy (DOE) Order 435.1, “Radioactive Waste Management,” and the guidance provided in the associated technical standard “Disposal Authorization Statement and Tank Closure Documentation” (DOE-STD-5002-2017). The purpose of this monitoring plan is to document a monitoring strategy that includes (1) compliance monitoring activities to demonstrate compliance with regulatory standards/limits and (2) performance monitoring to build confidence the facility is performing as demonstrated in the facility performance assessment (PA) (DOE-ID 2018a), composite analysis (CA) (DOE ID 2012), and CA addendum (DOE-ID 2018b). The de minimus impact to the aquifer predicted by the PA suggests that aquifer compliance monitoring should be augmented with performance monitoring of the drainage course materials and sedimentary interbeds in the vadose zone beneath the facility to provide a more effective means of identifying performance deviations. The monitoring approach delineated in this document was informed by the systems evaluation of natural and engineered facility features presented in the PA, an assessment of aquifer baseline conditions (INL 2017d), the dose analysis conducted in support of the PA and CA, and monitoring data collected during the first four years of facility operations (baseline monitoring phase) (INL 2023b). This plan provides monitoring locations, sampling frequencies, and sampling methods; recommendations for data evaluation; and a description of the monitoring plan implementation. Collected data will be used to demonstrate facility compliance and to identify conditions that are not consistent with the key assumptions made by the PA and CA.

12 - MGMT OF RADIOACTIVE AND NON-RADIOACTIVE WASTE

Hydropower Infrastructure - LAkes, Reservoirs, and RIvers (HILARRI), v4

HILARRI is a database of links between major datasets of operational hydropower dams and powerplants, and inland water bodies. These connections are critical for conducting large-scale analysis of hydropower infrastructure and their associated natural and engineered water systems. Features include: – Dams from the National Inventory of Dams (2025) and the Global Reservoir and Dam Database (GRanD v1.3) – Hydropower plants from the Existing Hydropower Assets dataset (EHA 2025) – Power plants that are listed in the 2025 U.S. Hydropower Development Pipeline Data or were listed in previous versions of the dataset These hydropower infrastructure features are linked to several major datasets that provide hydrologic and hydraulic information relevant for analysis of hydropower systems that includes the integral water resources. That information comes from: – Products from the National Hydrography Dataset (NHD) – NHDPlusV2 Medium Resolution river network flowlines, – NHD waterbodies (limited to lakes and reservoirs), – NHD Watershed Boundary Dataset (HUC12-level for the Conterminous United States (CONUS)) – NHD High Resolution waterbodies – HydroLAKES water bodies (lakes and reservoirs) – LAGOS-US lakes and reservoirs – EPA National Lakes Assessment (2007, 2012, 2017, and 2022) – The Reservoir Sedimentation Database (RESSED) – EPA SuRGE sampling locations Unique identifiers are used to facilitate joining to the original full datasets. For example, characteristics of NHD flowlines such as estimated average flow rate can be joined from the NHDPlusV2 dataset to a dam or power plant listed in HILARRI based on the ID field, “COMID”, that is common to both datasets. HILARRI only includes basic information about identifiers, location, and data quality or usage notes. It does not contain the attributes or time series data associated with these sites. The HILARRI dataset incorporates information from several datasets to facilitate more effective and accurate analysis of hydropower infrastructure and their associated waterbodies. For example, dams were checked against the most recent American Rivers Dam Removal Database to identify and flag facilities that may no longer exist. Additionally, dams that are listed multiple times in the NID are identified and flagged to avoid double-counting when analyzing and summarizing information. Other quality flags include certainty of operational hydropower (i.e., if one or more datasets indicates hydropower at a particular location), whether an associated water body is accurate or composed of multiple polygons, or whether there is a known issue with reported characteristics in one of the underlying datasets. These additional data flags are designed to increase confidence in data usage for individual to large-scale analyses.

Hansen, Carly [ORNL] (ORCID:0000000193280838)

Using Explainable Artificial Intelligence to Predict Perovskite Solar Cell Electrical Metastability from Operando Photoluminescence Images in Accelerated Stress Testing

Metal halide perovskite (MHP) solar cells exhibit a metastable response to bias governed by coupled ionic–electronic processes, complicating the conventional reciprocity relation between luminescence intensity and device open-circuit voltage (V oc ). This limits the use of luminescence as a diagnostic for device screening or accelerated stress testing, motivating new approaches that can interpret photoluminescence (PL) signals under nonequilibrium conditions. From the artificial intelligence perspective, we develop an explainable deep learning framework that integrates convolutional neural networks (CNN), long short-term memory (LSTM) layers, and an attention mechanism to learn spatiotemporal features from operando photoluminescence PL image sequences. The model achieves a mean absolute error of ±0.027 V in predicting open-circuit voltage transients and reduces extreme-tail errors by up to 78% compared to physics-based reciprocity calculations. Gradient-weighted Class Activation Mapping (Grad-CAM) provides interpretability by highlighting physically meaningful regions such as electrode edges and emergent defect features. From the engineering application perspective, this framework enables accurate, contactless prediction of device V oc and identification of degradation-relevant features during accelerated aging of perovskite solar cells. This approach demonstrates how explainable AI can enhance operando diagnostics and reliability analysis in photovoltaic devices under nonequilibrium conditions.

14 SOLAR ENERGY

RMCProfile7 : reverse Monte Carlo for multiphase systems

This work introduces a completely rewritten version of the programRMCProfile(version 7), big-box, reverse Monte Carlo modelling software for analysis of total scattering data. The major new feature ofRMCProfile7is the ability to refine multiple phases simultaneously, which is relevant for many current research areas such as energy materials, catalysis and engineering. Other new features include improved support for molecular potentials and rigid-body refinements, as well as multiple different data sets. An empirical resolution correction and calculation of the pair distribution function as a back-Fourier transform are now also available.RMCProfile7is freely available for download at https://rmcprofile.ornl.gov/.

Chemistry

An Innovative High Throughput Genome Releaser for Rapid and Efficient PCR Screening

High-throughput PCR screening is vital in synthetic biology and metabolic engineering as it allows researchers to rapidly analyze and detect numerous targeted genetic mutation in the genome. Current challenges for high-throughput PCR screening in synthetic biology include efficiently preparing genomic DNA, optimizing protocols for diverse sample types, managing contamination risks, and effectively analyzing the large volumes of data generated while ensuring consistent and accurate results. In this study, we present the development of a High Throughput Genome Releaser (HTGR), an innovative device addressing common challenges in screening PCR. This genome DNA releaser is designed based on a squash method for rapid, cost-effective, and efficient DNA release, optimized for subsequent PCR reactions. After experimenting with various synthetic materials, we selected a plastic that closely replicates the smooth surface and compression properties of microscope slides, ensuring reliable performance. We engineered a device featuring a 96-Well Plate and a shear applicator, operable both manually and automatically, and compatible with standard liquid-handling robot platform. This compatibility enhances ease of use in high-throughput PCR workflows. Additionally, we developed software to support its automatic functions. Our results demonstrated that the specially engineered 96-Well Plate and HTGR can effectively squash fungal spores , which release enough genome DNA for PCR screening. The genome releaser facilitates the preparation of PCR-amplifiable genomic DNA substrate from 96 samples within minutes, eliminates the need for extraction buffers, and is adaptable to a wide range of microorganisms and cells, which could significantly advance biomanufacturing processes.

Yuan, Guoliang [BATTELLE (PACIFIC NW LAB)]

The spherical tokamak advanced reactor (STAR) fusion power plant design

Scientific and technical advancements have been made that improve fusion’s prospects to provide a new energy source, showing enhanced plasma confinement conditions with plasma temperatures reaching or exceeding 100 million degrees. Overshadowing this progress is the challenge involved in developing an economically viable fusion power plant design. Many proposed next-step DEMO and pilot plant designs are extensions of existing physics-focused experimental devices defined to understand and control plasma operations to achieve and sustain a fusion reaction. Transitioning scientific and technical advancements into a functional power plant requires a dedicated focus on architectural designs that integrate diverse technologies, while optimizing physics conditions, with a focus on economic viability. This holistic approach is essential in turning the promise of fusion energy into a reality. The Spherical Tokamak Advanced Reactor (STAR) is a fusion power plant conceptual design with the architectural focus that strives to balance physics, engineering, and cost considerations. In conclusion, it has been set up to introduce relevant physics, engineering and concept features that an intermediate pilot plant might follow, with the goal of meeting system performances and economic requirements that lead to a commercially competitive fusion power plant.

Blanket segmentation

Distribution Substation Planning Toolkit (dsp-toolkit) v1.0

The Distribution Substation Planning Toolkit (DSP Toolkit) is a software suite designed to streamline the planning and optimization of distribution substations. This toolkit offers a comprehensive set of tools and APIs for data curation, short-term electric load forecasting, and weather-sensitive load adjustment, making it an essential resource for utility companies, engineers, and researchers. Features • Data Preprocessing and Curation: Efficiently manage and preprocess large datasets to ensure high-quality input for analysis. • Short-Term Load Forecasting: Utilize data-driven models to predict short-term electric loads accurately. • Weather-Sensitive Modeling: Automatically adjust load forecasts based on weather data to predict future peak demands more precisely. Uses The DSP Toolkit is ideal for planning and optimizing distribution substations, providing a user-friendly interface and comprehensive documentation. It is suitable for both novice and experienced users, facilitating efficient and accurate planning processes. Advantages • Efficiency: Automates complex planning tasks, reducing manual effort and minimizing errors. • Scalability: Handles large datasets and complex models, making it suitable for large-scale projects. • Community and Support: Open-source with active community contributions, ensuring continuous improvement and support. • Extensibility: Easily extendable with custom modules and plugins, allowing users to tailor the toolkit to their specific needs. The DSP Toolkit stands out by offering a robust, flexible, and user-friendly solution for distribution substation planning. Public Abstract

Li, Han [Lawrence Berkeley National Laboratory (LB

Hydropower Infrastructure – LAkes, Reservoirs, and RIvers (HILARRI)

HILARRI is a database of links between major datasets of operational hydropower dams and powerplants, and inland water bodies. These connections are critical for conducting large-scale analysis of hydropower infrastructure and their associated natural and engineered water systems. Features include: – Dams from the National Inventory of Dams (2024) and the Global Reservoir and Dam Database (GRanD v1.3) – Hydropower plants from the Existing Hydropower Assets dataset (EHA 2024) These hydropower infrastructure features are linked to several major datasets that provide hydrologic and hydraulic information relevant for analysis of hydropower systems that includes the integral water resources. That information comes from: – Products from the National Hydrography Dataset (NHD) – NHDPlusV2 Medium Resolution river network flowlines, – NHD waterbodies (limited to lakes and reservoirs), – NHD Watershed Boundary Dataset (HUC12-level for the Conterminous United States (CONUS)) – NHD High Resolution waterbodies – HydroLAKES water bodies (lakes and reservoirs) – LAGOS-US lakes and reservoirs – EPA National Lakes Assessment (2007, 2012, 2017, and 2022) – The Reservoir Sedimentation Database (RESSED) Unique identifiers are used to facilitate joining to the original full datasets. For example, characteristics of NHD flowlines such as estimated average flow rate can be joined from the NHDPlusV2 dataset to a dam or power plant listed in HILARRI based on the ID field, “COMID”, that is common to both datasets. HILARRI only includes basic information about identifiers, location, and data quality or usage notes. It does not contain the attributes or time series data associated with these sites. The HILARRI dataset incorporates information from several datasets to facilitate more effective and accurate analysis of hydropower infrastructure and their associated waterbodies. For example, dams were checked against the most recent American Rivers Dam Removal Database to identify and flag facilities that may no longer exist. Additionally, dams that are listed multiple times in the NID are identified and flagged to avoid double-counting when analyzing and summarizing information. Other quality flags include certainty of operational hydropower (i.e., if one or more datasets indicates hydropower at a particular location), whether an associated water body is accurate or composed of multiple polygons, or whether there is a known issue with reported characteristics in one of the underlying datasets. These additional data flags are designed to increase confidence in data usage for individual to large-scale analyses.

13 HYDRO ENERGY

Reduced Order Modeling conditioned on monitored features for response and error bounds estimation in engineered systems

Reduced Order Models (ROMs) form essential tools across engineering domains by virtue of their function as surrogates for computationally intensive digital twinning simulators. Although purely data-driven methods are available for ROM construction, schemes that allow to retain a portion of the physics tend to enhance the interpretability and generalization of ROMs. However, physics-based techniques can adversely scale when dealing with nonlinear systems that feature parametric dependencies. This study introduces a generative physics-based ROM that is suited for nonlinear systems with parametric dependencies and is additionally able to provide numerical error bounds associated with the respective estimates. A main contribution of this work is the conditioning of these parametric ROMs to features that can be derived from monitoring measurements, feasibly in an online fashion. This is contrary to most existing ROM schemes, which remain restricted to the prescription of the physics-based, and usually a priori unknown, system parameters. Our work utilizes conditional Variational Autoencoders to continuously map the required reduction bases to a feature vector extracted from limited output measurements, while additionally allowing for a probabilistic assessment of the ROM-estimated Quantities of Interest. An auxiliary task using a neural network-based parametrization of suitable probability distributions is introduced to re-establish the link with physical model parameters. We verify the proposed scheme on a series of simulated case studies incorporating effects of geometric and material nonlinearity under parametric dependencies related to system properties and input load characteristics.

Conditional VAEs

Lewis Acid Site Engineering in Chromite Spinels Orchestrated Surface Reconstruction and Surpasses RuO 2 in Oxygen Evolution

Atomic-scale engineering of chromite spinels featuring redox-active tetrahedral A-sites and strong Cr–O covalency offers a promising route to superior platinum-group-metal-free oxygen evolution reaction (OER) catalysts. However, comprehensive studies addressing how cation substitution influences surface chemistry and governs OER activity and durability in chromite spinels remain limited. Here, in this work, a systematic investigation of the multicationic chromite series Ni x Fe y Cr 3−x−y O 4 is presented, identifying composition-dependent Lewis acidity as a descriptor of superior OER performance. It is further demonstrated that tuning surface acidity directly controls dynamic reconstruction processes and lattice-oxygen participation during spinel-based electrocatalysis. Following activation, the optimized Ni 0.8 Fe 0.3 Cr 1.9 O 4 catalyst delivers a current density of 10 mA cm −2 at an overpotential of 235 mV, surpassing RuO 2 , with excellent long-term stability. Integrating microscopic and spectroscopic analysis with operando impedance spectroscopy, it shows that activation generates an oxyhydroxide overlayer and reveals a previously unrecognized link between surface Lewis acidity and the growth kinetics and activity of these shells. Density functional theory calculations indicate that Fe incorporation at octahedral sites raises the O 2p-band center and lowers oxygen-vacancy formation energy, promoting lattice-oxygen activation and triggering reconstruction, yielding enhanced OER. This work integrates cation-driven surface-acidity modulation, acidity-governed reconstruction, and OER activity enhancement into a unified predictive framework for designing earth-abundant spinel-based catalysts.

operando impedance spectroscopy