Search NASASearch

SEARCH · Search NASA

Results for “Generalizable models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Riverine dissolved organic matter transformations increase with watershed area, water residence time, and Damköhler numbers in nested watersheds

Abstract Quantifying the relative influence of factors and processes controlling riverine ecosystem function is essential to predicting future conditions under global change. Dissolved organic matter (DOM) is a fundamental component of riverine ecosystems that fuels microbial food webs, influences nutrient and light availability, and represents a significant carbon flux globally. The heterogeneous nature of DOM molecular composition and its propensity for interaction (i.e., functional diversity) can characterize riverine ecosystem function across spatiotemporal scales. To investigate fundamental drivers of DOM diversity, we collected seasonal water samples from 42 nested locations within five watersheds spanning multiple watershed sizes (~ 5 to 30,000 km 2 ) across the United States. Patterns in DOM molecular richness, aromaticity, relative abundance of N-containing formulas, and putative biochemical transformations derived from high-resolution mass spectrometry were assessed across gradients of explanatory variables associated with watershed characteristics (e.g., watershed area, water residence time, land cover). We found that putative biochemical transformations were more strongly related to explanatory variables across watersheds than common bulk DOM parameters and that watershed area, surface water residence time and derived Damköhler numbers representing DOM reactivity timescales were strong predictors of DOM diversity. The data also indicate that catchment-specific land cover factors can significantly influence DOM diversity in diverging directions. Overall, the results highlight the importance of considering water residence time and land cover when interpreting longitudinal patterns in DOM chemistry and the continued challenge of identifying generalizable drivers that are transferable across watershed and regional scales for application in Earth system models. This work also introduces a Findable Accessible Interoperable Reusable (FAIR) dataset (> 300 samples) to the community for future syntheses.

54 ENVIRONMENTAL SCIENCES

A statistical method for treating molecular line opacities

A method for treating atomic and molecular line opacities in cool stellar atmospheres by a statistical opacity sampling is investigated. Under the usual assumptions of plane-parallel geometry, radiative equilibrium, hydrostatic equilibrium, and LTE, each radiative quantity is computed monochromatically at each chosen frequency and depth without any averaging of the opacity. The number of frequencies needed to allow an accurate integration of the energy flux over a given spectral interval is investigated as a function of depth, including opacity for both CN and C2. This method is extended to the calculation of a model atmosphere of a star, and the effect of the number and placement of frequency points is studied. The method is applied to treating molecular lines of CO, C2, and CN in a cool carbon star. Significant advantages of the opacity sampling method are its flexibility, which permits computation of models having arbitrary variations of chemical composition and of opacity with wavelength and depth, and generalizability to include departures from LTE.

Sneden, C.

A Simulated Evaluation of Powder Flowability Through a Partially Obstructed Consumable in Blown Powder Directed Energy Deposition Systems

Abstract In the interest of continued industrialization of metal additive manufacturing in modern production environments, cost is often referenced as a primary deterrent to new adopters. Conventional economic models for additive systems, processes, and supply chains often focus on specific process applications with little generalizability, or they neglect significant costs associated with production such as machine maintenance and consumable part replacement. Compounding the latter issue are substantial knowledge gaps in consumable part wear characterization for additive and other convergent manufacturing systems. In coaxial blown powder directed energy deposition systems, gas atomized metal powder is wasted during material deposition at a rate that is partly dependent on present wear phenomena in a consumable nozzle housed in the cladding head assembly. The price and lead time required to replace the nozzle incentivizes its reuse even when visibly worn. Often this initiates a process quality decline in the form of underbuilt geometry and internal defects due to losses in powder catchment efficiency. While depositing H13 steel using a hybrid manufacturing machine tool equipped with such a deposition system, a unique partial clog with a bridge-like structure formed at the consumable nozzle exit when supporting argon gas flows failed mid-process. To further understand coaxial multi-phase powder flow in the event of support gas failure, a computational fluid dynamics simulation is tailored to relevant process parameters, H13 powder material profile, and machine operator observations collected after the incident. The resulting differences in powder flow compared to control gas flow parameters is presented and discussed. The powder flowability and performance of the clogged nozzle is then assessed by using an optical profilometer to extract the profile of the clog and recreate the clog geometry within the simulation environment. In past work this simulation has been experimentally validated for a 316L steel powder material profile and used specifically for analyzing powder stream geometry and catchment efficiency. After the initial powder flow characterization, the clog is removed, and the nozzle is reprofiled. After removing the obstructing clog, the newly unobstructed nozzle geometry, the original off the shelf nozzle geometry, and additional nozzle profiles exploring different consumable refurbishment strategies are reevaluated in the simulation. Powder catchment efficiency for all variant nozzle geometries and relevant flow variables are compared and discussed, along with potential mitigation strategies for optimizing powder flowability with worn consumables. This work expands on the known morphology of blown powder obstructions and wear defects present in consumable coaxial nozzles while discussing pragmatic simulation driven responses to unanticipated subsystem failure in hybrid manufacturing machining platforms.

DeWitte, Lisa

Automated Probabilistic Finite Element Model Calibration Tool Based on Uncertainty Quantification and Machine Learning

Qualification and certification of safety critical parts is a hurdle to the adoption of metallic additively manufactured components for aerospace vehicle applications. Challenges include variability in part properties due to inconsistent defect distribution and microstructure. Understanding of the process through finite element modeling (FEM), and process control through in-situ monitoring, may result in significant improvements; however, solutions useful to manufacturers will require large volumes of data and automated data utilization. Toward this end, a generalizable automated FEM calibration paradigm is developed. This paradigm leverages existing and novel tools from machine learning and uncertainty quantification to enable the automatic calibration of FEMs without requiring prior knowledge of the model performance across input parameter space, including meshing and solver settings, which can require time consuming manual model probing or cause noisy and inconsistent predictions. The result is a probabilistic distribution of calibrated and validated FEM input parameters targeting measured data.

Additive manufacturing model calibration finite el

Interpreting AI for fusion: An application to plasma profile analysis for tearing mode stability

Artificial intelligence models have demonstrated strong predictive capabilities for various instabilities in fusion devices such as Tokamaks, including tearing modes (TM), edge localized modes, and disruptive events, but their opaque nature raises concerns about safety and trustworthiness when applied to fusion power plants. Here, we present a physics-based interpretation framework using a TM prediction model as a demonstration that is validated through a dedicated DIII-D TM avoidance experiment. By applying Shapley analysis, we identify how profiles such as rotation, temperature, and density contribute to the model's prediction of TM stability. Our analysis shows that in our experimental scenario, core electron temperature and rotation peaking play the primary role in TM stability, while density changes have smaller effects on stability. We show that off-axis ion temperature stabilizes TMs, suggesting that off-axis neutral beam heating can further stabilize this scenario. This work presents a generalizable ML-based event prediction methodology, from training to physics-driven interpretation, bridging the gap between physics understanding and opaque ML models.

Farre-Kaga, Hiro J. [Princeton Univ., NJ (United S

Good practices for documenting AI-based studies on energy and buildings

Artificial intelligence has transformed building science research over the past decade, with applications spanning energy modeling, energy prediction, HVAC optimization and controls, fault detection, and occupancy modeling. However, many studies lack adequate documentation of datasets, algorithms, training procedures, and validation methods. Building science research faces additional challenges including inconsistent evaluation metrics, limited generalizability across building types, climates, and significant gaps between experimental studies and deployed systems. This communication provides practical guidance for good practices in documenting and publishing AI-based research following established standards from the computer science and machine learning communities. By adopting frameworks such as Datasheets for Datasets, Model Cards, and standardized reproducibility checklists, researchers can ensure their work meets the rigorous documentation standards necessary for reproducible, comparable, and impactful building science research.

Hong, Tianzhen [Lawrence Berkeley National Laborat

A non-intrusive framework using acoustic signals and deep learning for boiling diagnostics in visual-limited environments

Accurate monitoring of boiling heat transfer is critical for safeguarding high-power systems operating in environments where conventional optical diagnostics are hindered by radiation fields or restricted visual accessibility. This study presents a non-intrusive framework that integrates hydroacoustic sensing with deep learning to infer near-wall boiling characteristics and enable predictive thermal assessment without visual access. In a prototypical subcooled flow-boiling facility representative of the Isotope Production Facility (IPF) at Los Alamos, hydrophones capture boiling-induced acoustic emissions that are transformed into background-removed Short-Time Fourier Transform (STFT) spectrograms. A convolutional neural network (CNN) then regresses heat flux, wall superheat, and key bubble parameters directly from these spectrograms. The CNN achieved predictive accuracy under nominal conditions and demonstrated robustness and generalization under acoustic noise for Signal-to-Noise Ratios (SNRs) down to approximately 0 dB. When integrated into an ANSYS CFX wall-boiling model, the acoustically inferred parameters reproduced boiling curve and critical heat flux (CHF) values consistent with image-based benchmarks. Furthermore, the model retained reliable performance under moderate variations in bulk temperature, flow rate, and hydrophone placement, confirming its generalizability across practical boundary conditions. These results demonstrate the feasibility of hydroacoustic-based deep learning as a viable path toward real-time, radiation-tolerant boiling diagnostics and predictive thermal safety assessment in inaccessible systems such as the IPF.

42 ENGINEERING

Cross-functional transferability in foundation machine learning interatomic potentials

The rapid development of foundation potentials (FPs) in machine learning interatomic potentials demonstrates the possibility for generalizable learning of the universal potential energy surface. The accuracy of FPs can be further improved by bridging the model from lower-fidelity datasets to high-fidelity ones. In this work, we analyze the challenge of this transfer learning (TL) problem within the CHGNet framework. We show that significant energy scale shifts and poor correlations between GGA and r 2 SCAN hinder cross-functional transferability. By benchmarking different TL approaches on the MP-r 2 SCAN dataset, we demonstrate the importance of elemental energy referencing in the TL of FPs. By comparing the scaling law with and without the pre-training on a low-fidelity dataset, we show that significant data efficiency can still be achieved through TL, even with a target dataset of sub-million structures. We highlight the importance of proper TL and multi-fidelity learning in creating next-generation FPs on high-fidelity data.

Huang, Xu [University of California, Berkeley, CA

Real-Time, Adaptive Radiological Anomaly Detection and Isotope Identification Using Non-Negative Matrix Factorization

Spectroscopic anomaly detection and isotope identification algorithms are integral components in nuclear nonproliferation applications such as search operations. The task is especially challenging in the case of mobile detector systems because the observed gamma-ray background changes more than for a static detector system, and a pretrained background model can easily find itself out of domain. The result is that algorithms may exceed their intended false alarm rate or sacrifice detection sensitivity to maintain the desired false alarm rate. Non-negative matrix factorization (NMF) is a powerful tool for spectral anomaly detection and identification, but, like many similar algorithms that rely on data-driven background models, in its conventional implementation, it is unable to update in real time to account for environmental changes that affect the background spectroscopic signature. Here, we have developed a novel NMF-based algorithm that periodically updates its background model to accommodate changing environmental conditions. The adaptive NMF algorithm involves fewer assumptions about its environment, making it more generalizable than existing NMF-based methods while maintaining or exceeding detection performance on simulated and real-world datasets.

Anomaly detection

A methodology for the design and evaluation of user interfaces for interactive information systems

The definition of proposed research addressing the development and validation of a methodology for the design and evaluation of user interfaces for interactive information systems is given. The major objectives of this research are: the development of a comprehensive, objective, and generalizable methodology for the design and evaluation of user interfaces for information systems; the development of equations and/or analytical models to characterize user behavior and the performance of a designed interface; the design of a prototype system for the development and administration of user interfaces; and the design and use of controlled experiments to support the research and test/validate the proposed methodology. The proposed design methodology views the user interface as a virtual machine composed of three layers: an interactive layer, a dialogue manager layer, and an application interface layer. A command language model of user system interactions is presented because of its inherent simplicity and structured approach based on interaction events. All interaction events have a common structure based on common generic elements necessary for a successful dialogue. It is shown that, using this model, various types of interfaces could be designed and implemented to accommodate various categories of users. The implementation methodology is discussed in terms of how to store and organize the information.

Dominick, Wayne D.

Multi-task Parallelism for Robust Pre-training of Graph Foundation Models on Multi-source, Multi-fidelity Atomistic Modeling Data

Graph foundation models using graph neural networks promise sustainable, efficient atomistic modeling. To tackle challenges of processing multi-source, multi-fidelity data during pre-training, recent studies employ multi-task learning, in which shared message passing layers initially process input atomistic structures regardless of source, then route them to multiple decoding heads that predict data-specific outputs. This approach stabilizes pre-training and enhances a model’s transferability to unexplored chemical regions. Preliminary results on approximately four million structures are encouraging, yet questions remain about generalizability to larger, more diverse datasets and scalability on supercomputers. We propose a multi-task parallelism method that distributes each head across computing resources with GPU acceleration. Implemented in the open-source HydraGNN architecture, our method was trained on over 24 million structures from five datasets and tested on the Perlmutter, Aurora, and Frontier supercomputers, demonstrating efficient scaling on all three highly heterogeneous super-computing architectures.

Lupo Pasini, Massimiliano [ORNL] (ORCID:0000000249

Probing the evolution of fault properties during the seismic cycle with deep learning

We use seismic waves that pass through the hypocentral region of the 2016 M6.5 Norcia earthquake together with Deep Learning (DL) to distinguish between foreshocks, aftershocks and time-to-failure (TTF). Binary and N-class models defined by TTF correctly identify seismograms in test with > 90% accuracy. We use raw seismic records as input to a 7 layer CNN model to perform the classification. Here we show that DL models successfully distinguish seismic waves pre/post mainshock in accord with lab and theoretical expectations of progressive changes in crack density prior to abrupt change at failure and gradual postseismic recovery. Performance is lower for band-pass filtered seismograms (below 10 Hz) suggesting that DL models learn from the evolution of subtle changes in elastic wave attenuation. Tests to verify that our results indeed provide a proxy for fault properties included DL models trained with the wrong mainshock time and those using seismic waves far from the Norcia mainshock; both show degraded performance. Our results demonstrate that DL models have the potential to track the evolution of fault zone properties during the seismic cycle. If this result is generalizable it could improve earthquake early warning and seismic hazard analysis.

58 GEOSCIENCES

Evaluating the factors influencing accuracy, interpretability, and reproducibility in the use of machine learning classifiers in biology to enable standardization

The complexity and variability of biological data has promoted the increased use of machine learning methods to understand processes and predict outcomes. These same features complicate reliable, reproducible, interpretable, and responsible use of such methods, resulting in questionable relevance of the derived. outcomes. Here we systematically explore challenges associated with applying machine learning to predict and understand biological processes using a well- characterized in vitro experimental system. We evaluated factors that vary while applying machine learning classifers: (1) type of biochemical signature (transcripts vs. proteins), (2) data curation methods (pre- and post-processing), and (3) choice of machine learning classifier. Using accuracy, generalizability, interpretability, and reproducibility as metrics, we found that the above factors significantly mod- ulate outcomes even within a simple model system. Our results caution against the unregulated use of machine learning methods in the biological sciences, and strongly advocate the need for data standards and validation tool-kits for such studies.

59 BASIC BIOLOGICAL SCIENCES

Vehicular Re-Identification from Uncontrolled Multiple Views

Vehicle re-identification (re-ID) across disparate sensing modalities remains a fundamental challenge for transportation research. In this work, we introduce a deep multi-view vehicle re-ID framework that leverages Siamese networks to compare pairs of vehicle images and produce matching scores, enabling robust association across drastically different viewpoints such as those from UAVs, surveillance cameras, and ground sensors. The model exploits convolutional neural networks to learn features that remain discriminative under changes in angle, distance, and illumination, supporting more generalizable re-ID performance. As part of this effort, we also developed an automated pipeline to synchronize roadside and UAV video streams, producing a multi-perspective dataset that complements preexisting real collections and a synthetic dataset generated in this study. Together, these contributions advance the capability to re-identify vehicles across wide viewing baselines; establish a foundation for scalable, reproducible research in vehicle re-ID; and open pathways for future applications, such as inferring routine behaviors, movement patterns, and daily habits of the individual associated with the vehicle.

convolutional neural networks

Harnessing Machine Learning and Data Fusion for Accurate Undocumented Well Identification in Satellite Images

This study utilizes satellite data to detect undocumented oil and gas wells, which pose significant environmental concerns, including greenhouse gas emissions. Three key findings emerge from the study. Firstly, the problem of imbalanced data is addressed by recommending oversampling techniques like Rotation–GaussianBlur–Solarization data augmentation (RGS), the Synthetic Minority Over-Sampling Technique (SMOTE), or ADASYN (an extension of SMOTE) over undersampling techniques. The performance of borderline SMOTE is less effective than that of the rest of the oversampling techniques, as its performance relies heavily on the quality and distribution of data near the decision boundary. Secondly, incorporating pre-trained models trained on large-scale datasets enhances the models’ generalization ability, with models trained on one county’s dataset demonstrating high overall accuracy, recall, and F1 scores that can be extended to other areas. This transferability of models allows for wider application. Lastly, including persistent homology (PH) as an additional input improves performance for in-distribution testing but may affect the model’s generalization for out-of-distribution testing. A careful consideration of PH’s impact on overall performance and generalizability is recommended. Overall, this study provides a robust approach to identifying undocumented oil and gas wells, contributing to the acceleration of a net-zero economy and supporting environmental sustainability efforts.

SMOTE

Optimizing high energy density sulfur cathodes: A multivariate approach to electrode formulation and processing

Lithium-sulfur (Li-S) batteries involve complex solid-liquid-solid phase transformations during both discharging and charging processes, where cathode materials, formulation, and structure play a crucial role. Here, a design of experiments (DoE) methodology and an empirical model are developed to systematically explore the interactions and trade-offs among cathode factors and process variables, and to obtain generalizable effects estimates for the multivariate system. Compared to the conventional one-factor-at-a-time (OFAT) approach, this work demonstrates advantages in both efficiency and accuracy by allowing the data to guide future research and decisions. Further, an optimized cathode formulation and processing parameters are predicted and validated experimentally, achieving over 1000 mAh g -1 in discharge capacity and improved cycling under practical lean electrolyte (4 µL mg -1 S) and high S-loading cathodes (>4 mg cm -2 ) conditions. The optimized cathode was scaled up and assembled into Li-S pouch cells, achieving 316 Wh kg -1 in cell-level energy, proving that the comprehensive and rigorous framework for optimizing complex systems with DoE leads to improved performance in a practical pouch cell system.

25 ENERGY STORAGE

Predicting Airport Runway Configurations for Decision-Support Using Supervised Learning

One of the most challenging tasks for air traffic controllers is runway configuration management (RCM). It deals with the optimal selection of runways to operate on (for arrivals and departures) based on traffic, surface wind speed, wind direction, other environmental variables, noise constraints, and several other airport-specific factors. It affects the efficiency of the National Airspace System (NAS) and both surface and airspace operations can benefit from better understanding future runway configurations. In this paper, we present a comprehensive implementation of predictive models for runway configuration estimation from large volumes of historical data. Specifically, operational data from two full years (2018 and 2019) is collected, analyzed, and fused together to build the data product used in this work. The data set differs from prior work in the field in terms of its scope, resolution, and variety of factors collected and considered. Meteorological data is collected from two different sources – current weather conditions from METAR (Meteorological Terminal Aviation Routine Weather Report) and forecast weather conditions from Localized Aviation MOS Program (LAMP). Operational data from the Federal Aviation Administration (FAA) Aviation System Performance Metrics (ASPM) related to scheduled and actual number of arrivals and departures, average taxi times, etc. are collected. NASA’s Sherlock Data Warehouse is used to identify critical information such as go-arounds, and other events that might impact RCM decision-making. All data is collected and aggregated over 15-minute intervals throughout the two years. This provides a resolution like the timescales that might be necessary for runway configuration management decision-making. A variety of supervised learning algorithms are tested including Support Vector Machine, Random Forest, Gradient Boosting, etc. including tuning of the model hyperparameters. The modeling process is applied and presented on two representative U.S. airports – Charlotte Douglas International Airport (KCLT) and Denver International Airport (KDEN). The two airports present different levels of complexity in terms of the total number of configurations used and provide a balanced perspective on the generalizability of the developed approach to other airports in the NAS. Initial results are promising (F1 score of 0.91 at KCLT and 0.83 at KDEN) for data in the test set. The final paper will contain a comprehensive comparison between different models and model building strategies as well as further refined results. Most important predictors for each airport will be identified along with a discussion and recommendations on adapting the framework to other scenarios.

Tejas G Puranik

Causality-respecting adaptive refinement for PINNs: enabling precise interface evolution in phase field modeling

Physics-informed neural networks (PINNs) have emerged as a powerful tool for solving physical systems described by partial differential equations (PDEs). However, their accuracy in dynamical systems, particularly those involving sharp moving boundaries with complex initial morphologies, remains a challenge. Here, this study introduces an approach combining residual-based adaptive refinement (RBAR) with causality-informed training to enhance the performance of PINNs in solving spatio-temporal PDEs. Our method employs a three-step iterative process: initial causality-based training, RBAR-guided domain refinement, and subsequent causality training on the refined mesh. Applied to the Allen-Cahn equation, a widely-used model in phase field simulations, our approach demonstrates significant improvements in solution accuracy and computational efficiency over traditional PINNs. Notably, we observe an ‘overshoot and relocate’ phenomenon in dynamic cases with complex morphologies, showcasing the method’s adaptive error correction capabilities. This synergistic interaction between RBAR and causality training enables accurate capture of interface evolution, even in challenging scenarios where traditional PINNs fail. Our framework not only resolves the limitations of uniform refinement strategies but also provides a generalizable methodology for solving a broad range of spatio-temporal PDEs. The enhanced performance of the RBAR–causality combined framework demonstrates its strong potential for advancing PINN-based modeling of physical systems characterized by complex, evolving interfaces.

Allen-Cahn equations