Search NASASearch

SEARCH · Search NASA

Results for “training models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

OmniXAS: A universal deep-learning framework for materials x-ray absorption spectra

X-ray absorption spectroscopy (XAS) is a powerful characterization technique for probing the local chemical environment of absorbing atoms. However, analyzing XAS data presents significant challenges, often requiring extensive, computationally intensive simulations, as well as significant domain expertise. These limitations hinder the development of fast, robust XAS analysis pipelines that are essential in high-throughput studies and for autonomous experimentation. Here, we address these challenges with OmniXAS, a framework that contains a suite of transfer learning approaches for XAS prediction, each uniquely contributing to improved accuracy and efficiency, as demonstrated on the K-edge spectra database covering eight 3⁢d transition metals (Ti–Cu). The OmniXAS framework is built upon three distinct strategies. First, we use M3GNet [Nat. Comput. Sci. 2, 718 (2022)] to derive latent representations of the local chemical environment of absorption sites as input for XAS prediction, achieving significant improvements over conventional featurization techniques. Second, we employ a hierarchical transfer learning strategy, training a universal multitask model across elements before fine-tuning for element-specific predictions. Models based on this cascaded approach after elementwise fine-tuning outperform element-specific models by up to 69%. Third, we implement cross-fidelity transfer learning, adapting a universal model to predict spectra generated by simulation of a different fidelity with a much higher computational cost. This approach improves prediction accuracy by up to 11% over models trained on the target fidelity alone. Our approach significantly boosts the throughput of XAS modeling by orders of magnitude as compared to first-principles simulations and is extendable to XAS prediction for a broader range of elements. The proposed transfer learning framework is generalizable to enhance deep-learning models that target other properties in materials research.

36 MATERIALS SCIENCE

Neural network reconstruction of the DIII-D tokamak plasma boundary using a reduced set of diagnostics

This study investigates the feasibility of reconstructing the last closed flux surface in the DIII-D tokamak using neural network models trained on reduced input feature sets, addressing an ill-posed task. Two models are compared: one trained solely on coil currents and another incorporating coil currents, plasma current and loop voltage. The model trained exclusively on coil currents achieved a mean point displacement of $0.04$ m on a held-out test set, while the inclusion of plasma current and loop voltage reduced the error to $0.03$ m. This comparison highlights the trade-offs between input feature complexity and reconstruction accuracy, demonstrating the potential of machine learning algorithms to perform effectively in data-limited environments, such as those expected in fusion power plants due to diagnostic constraints imposed by the presence of blankets and shielding.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Vibrational Analysis of Engine Components Using Neural-Net Processing and Electronic Holography

The use of computational-model trained artificial neural networks to acquire damage specific information from electronic holograms is discussed. A neural network is trained to transform two time-average holograms into a pattern related to the bending-induced-strain distribution of the vibrating component. The bending distribution is very sensitive to component damage unlike the characteristic fringe pattern or the displacement amplitude distribution. The neural network processor is fast for real-time visualization of damage. The two-hologram limit makes the processor more robust to speckle pattern decorrelation. Undamaged and cracked cantilever plates serve as effective objects for testing the combination of electronic holography and neural-net processing. The requirements are discussed for using finite-element-model trained neural networks for field inspections of engine components. The paper specifically discusses neural-network fringe pattern analysis in the presence of the laser speckle effect and the performances of two limiting cases of the neural-net architecture.

Decker, Arthur J.

Vibrational Analysis of Engine Components Using Neural-Net Processing and Electronic Holography

The use of computational-model trained artificial neural networks to acquire damage specific information from electronic holograms is discussed. A neural network is trained to transform two time-average holograms into a pattern related to the bending-induced-strain distribution of the vibrating component. The bending distribution is very sensitive to component damage unlike the characteristic fringe pattern or the displacement amplitude distribution. The neural network processor is fast for real-time visualization of damage. The two-hologram limit makes the processor more robust to speckle pattern decorrelation. Undamaged and cracked cantilever plates serve as effective objects for testing the combination of electronic holography and neural-net processing. The requirements are discussed for using finite-element-model trained neural networks for field inspections of engine components. The paper specifically discusses neural-network fringe pattern analysis in the presence of the laser speckle effect and the performances of two limiting cases of the neural-net architecture.

Decker, Arthur J.

Machine Learning for Anomaly Detection in Neural Network Security and SRF Cavities

This dissertation explores the development and deployment of machine learning approaches to address critical challenges in anomaly detection across two distinct domains: neural network security in federated learning settings and cavity behavior analysis in particle accelerator operations at Jefferson Lab in Newport News, Virginia. Anomaly detection identifies deviations from expected patterns, safeguarding systems in cybersecurity, industry, and research against malicious activities and failures. This dissertation demonstrates how our machine learning approaches enhance detection accuracy and efficiency in both neural network security and industrial applications. First, we investigate vulnerabilities in deep neural networks deployed in federated learning. Although federated learning preserves user privacy by training models locally, it remains vulnerable to backdoor attacks, in which malicious participants embed hidden triggers that induce targeted misbehavior. We propose a self-supervised contrastive learning framework to detect and mitigate such backdoor attacks. In our experiments, this method achieves higher detection accuracy and lower false positive rates than existing defenses, while operating without access to local model updates or original training data and thus preserving the privacy guarantees of the federated setting. Second, we address the operational reliability of superconducting radio-frequency (SRF) cavities at the Continuous Electron Beam Accelerator Facility (CEBAF). Our research leverages an unsupervised learning approach, combined with Principal Component Analysis (PCA) and k-means clustering, to identify anomalous behaviors in SRF cavities. Our method detects subtle anomalous behavior by analyzing SRF signal data. This knowledge allows for the early detection and resolution of potential faults, significantly improving the efficiency and reliability of operations. Third, we extend these insights to time-series anomaly detection more broadly. We design a contrastive-learning based model tailored to increasingly dynamic environments and academic research. This model improves detection accuracy in settings that require real-time monitoring and predictive maintenance. Our research underscores the broader applicability and impact of advanced machine learning techniques in anomaly detection. By extracting meaningful patterns from complex data, machine learning can significantly enhance security in distributed neural networks and improve the efficiency of particle accelerator operations. This dissertation serves as a stepping stone for future investigations into the vast possibilities of anomaly detection, inspiring further exploration and development of machine learning techniques in this field.

Ferguson, Hal [Old Dominion University]

Network Anomaly Detection in Distributed Edge Computing Infrastructure

As networks continue to grow in complexity and scale, detecting anomalies has become increasingly challenging, particularly in diverse and geographically dispersed environments. Traditional approaches often struggle with managing the computational burden associated with analyzing large-scale network traffic to identify anomalies. This paper introduces a distributed edge computing framework that integrates federated learning with Apache Spark and Kubernetes to address these challenges. We hypothesize that our approach, which enables collaborative model training across distributed nodes, significantly enhances the detection accuracy of network anomalies across different network types. We show that by leveraging distributed computing and containerization technologies, our framework not only improves scalability and fault tolerance but also achieves superior detection performance compared to state-of-the-art methods. Extensive experiments on the UNSW-NB15 and ROAD datasets validate the effectiveness of our approach, demonstrating statistically significant improvements in detection accuracy and training efficiency over baseline models, as confirmed by MannWhitney U and Kolmogorov-Smirnov tests (p<0.05).

Marfo, William [University of Texas at El Paso,Dep

pvcracks: trained VAE model

The resulting model weights for the variational autoencoder for solar cell crack parametrization to be loaded into the python code for other to use

14 SOLAR ENERGY

Houston, We Have a Problem Solving Model for Training

In late 2006, the Mission Operations Directorate (MOD) at NASA began looking at ways to make training more efficient for the flight controllers who support the International Space Station. The average certification times for flight controllers spanned from 18 months to three years and the MOD, responsible for technical training, was eager to develop creative solutions that would reduce the time to 12 months. Additionally, previously trained flight controllers sometimes participated in more than 50 very costly, eight-hour integrated simulations before becoming certified. New trainees needed to gain proficiency with far fewer lessons and training simulations than their predecessors. This poster presentation reviews the approach and the process that is currently in development to accomplish this goal.

Schmidt, Lacey

Computational investigation of water glasses using machine-learning potentials

The molecular origins of water’s anomalous properties have long been a subject of scientific inquiry. The liquid–liquid phase transition hypothesis, which posits the existence of distinct low-density and high-density liquid states separated by a first-order phase transition terminating at a critical point, has gained increasing experimental and computational support and offers a thermodynamically consistent framework for many of water’s anomalies. However, experimental challenges in avoiding crystallization near the postulated liquid–liquid critical point have focused attention to water’s canonical glassy states: low-density and high-density amorphous ice. Here, we use two Deep Potential machine-learning models, trained on the Strongly Constrained and Appropriately Normed density functional and the highly accurate Many-Body Polarizable potential, to conduct an investigation of water’s glassy phenomenology based on quantum mechanical calculations. Despite not being explicitly trained on amorphous ices, both models accurately capture the structure and transformation of the water glasses, including their interconversion along different thermodynamic paths. Isobaric quenching of liquid water at various pressures generates a continuum of intermediate amorphous ices and density fluctuations increase near the liquid–liquid critical pressure. The glass transition temperatures of the amorphous ices produced at different pressures exhibit two distinct branches, corresponding to low-density and high-density amorphous ice behaviors, consistent with experiment and the liquid–liquid transition hypothesis. Extrapolating transformation pressures from isothermal compressions to experimental compression rates brings our simulations into excellent agreement with data. Our findings demonstrate that machine-learning potentials trained on equilibrium phases can effectively model nonequilibrium glassy behavior and pave the way for studying long-timescale, out-of-equilibrium processes with quantum mechanical accuracy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Quantum Hardware-Enabled Molecular Dynamics via Transfer Learning

The ability to perform ab initio molecular dynamics simulations using potential energy surfaces provided by quantum computers would open the door to virtually exact dynamics for a variety of chemical and biochemical systems, with impacts on catalysis and biophysics. Nonetheless, performing molecular dynamics on surfaces produced by quantum hardware has been hampered by the noisy energies typically produced by quantum computers and challenges associated with computing gradients and scaling to large systems interest. A recent set of advances in machine learning, known as transfer learning, provides a new path forward for molecular dynamics simulations on quantum hardware. Transfer learning offers a workaround, where one first trains models on larger, less accurate classical datasets and then refines them on smaller, more accurate quantum datasets. We explore this approach by training machine learning models to predict a molecule's potential energy based on its geometric structure using Behler-Parrinello neural networks. When successfully trained, the model enables energy gradient predictions necessary for dynamic simulations. To reduce the quantum resources needed, the model is initially trained with data derived from classical density functional theory and subsequently refined with a smaller dataset obtained from a variational quantum eigensolver optimization of the unitary coupled cluster ansatz. We show that this approach significantly reduces the size of the needed quantum training dataset while capturing the high accuracies needed within quantum chemistry simulations. The success of this two-step training method opens more opportunities to apply machine learning models on quantum data, a significant stride towards efficient quantum-classical hybrid computational models.

quantum computing

Designing a training tool for imaging mental models

The training process can be conceptualized as the student acquiring an evolutionary sequence of classification-problem solving mental models. For example a physician learns (1) classification systems for patient symptoms, diagnostic procedures, diseases, and therapeutic interventions and (2) interrelationships among these classifications (e.g., how to use diagnostic procedures to collect data about a patient's symptoms in order to identify the disease so that therapeutic measures can be taken. This project developed functional specifications for a computer-based tool, Mental Link, that allows the evaluative imaging of such mental models. The fundamental design approach underlying this representational medium is traversal of virtual cognition space. Typically intangible cognitive entities and links among them are visible as a three-dimensional web that represents a knowledge structure. The tool has a high degree of flexibility and customizability to allow extension to other types of uses, such a front-end to an intelligent tutoring system, knowledge base, hypermedia system, or semantic network.

Dede, Christopher J.

Neural Network and Regression Soft Model Extended for PAX-300 Aircraft Engine

In fiscal year 2001, the neural network and regression capabilities of NASA Glenn Research Center's COMETBOARDS design optimization testbed were extended to generate approximate models for the PAX-300 aircraft engine. The analytical model of the engine is defined through nine variables: the fan efficiency factor, the low pressure of the compressor, the high pressure of the compressor, the high pressure of the turbine, the low pressure of the turbine, the operating pressure, and three critical temperatures (T(sub 4), T(sub vane), and T(sub metal)). Numerical Propulsion System Simulation (NPSS) calculations of the specific fuel consumption (TSFC), as a function of the variables can become time consuming, and numerical instabilities can occur during these design calculations. "Soft" models can alleviate both deficiencies. These approximate models are generated from a set of high-fidelity input-output pairs obtained from the NPSS code and a design of the experiment strategy. A neural network and a regression model with 45 weight factors were trained for the input/output pairs. Then, the trained models were validated through a comparison with the original NPSS code. Comparisons of TSFC versus the operating pressure and of TSFC versus the three temperatures (T(sub 4), T(sub vane), and T(sub metal)) are depicted in the figures. The overall performance was satisfactory for both the regression and the neural network model. The regression model required fewer calculations than the neural network model, and it produced marginally superior results. Training the approximate methods is time consuming. Once trained, the approximate methods generated the solution with only a trivial computational effort, reducing the solution time from hours to less than a minute.

Patnaik, Surya N.

Field-level reconstruction from foreground-contaminated 21-cm maps

Current and upcoming 21-cm experiments will soon be able to map 21-cm spatial fluctuations in three dimensions for a wide range of redshifts. However, bright foreground contamination and the nature of radio interferometry create significant challenges, making it difficult to access rich cosmological information from the Fourier modes that lie within the “foreground wedge”. Here, in this work, we introduce two approaches aiming to reconstruct the full 21-cm density field, including the missing modes in the wedge: (a) a field-level inference under an effective field theory (EFT) framework; (b) a diffusion-based deep generative model trained on simulations. Under the EFT framework, we implement a fully differentiable forward model that maps the initial conditions of matter fluctuations to the observed, foreground-filtered 21-cm maps. This enables a gradient-based sampler to simultaneously sample the initial conditions and bias parameters, allowing a physically motivated mode reconstruction. Alternatively, we apply a variational diffusion model to perform 21-cm density reconstruction at the map level. Our model is trained on semi-numerical simulations over a wide range of astrophysical parameters. Our results from both approaches should provide improved cosmological constraints from the field level and also enable cross-correlation between experiments that have little or no overlapping modes.

cosmological perturbation theory

Integrated edge-to-exascale workflow for real-time steering in neutron scattering experiments

We introduce a computational framework that integrates artificial intelligence (AI), machine learning, and high-performance computing to enable real-time steering of neutron scattering experiments using an edge-to-exascale workflow. Focusing on time-of-flight neutron event data at the Spallation Neutron Source, our approach combines temporal processing of four-dimensional neutron event data with predictive modeling for multidimensional crystallography. At the core of this workflow is the Temporal Fusion Transformer model, which provides voxel-level precision in predicting 3D neutron scattering patterns. The system incorporates edge computing for rapid data preprocessing and exascale computing via the Frontier supercomputer for large-scale AI model training, enabling adaptive, data-driven decisions during experiments. This framework optimizes neutron beam time, improves experimental accuracy, and lays the foundation for automation in neutron scattering. Although real-time experiment steering is still in the proof-of-concept stage, the demonstrated potential of this system offers a substantial reduction in data processing time from hours to minutes via distributed training, and significant improvements in model accuracy, setting the stage for widespread adoption across neutron scattering facilities and more efficient exploration of complex material systems.

97 MATHEMATICS AND COMPUTING

Data-driven analysis of dipole strength functions using artificial neural networks

Here, we present a data-driven analysis of dipole strength functions across the nuclear chart, employing an artificial neural network to model nuclear dipole responses. We train the network on a dataset of experimentally measured dipole strength functions for 216 different nuclei. To assess its predictive capability, we test the trained model on an additional set of 10 new nuclei, where experimental data exist. We demonstrate that the artificial neural network not only accurately reproduces known data but also identifies potential inconsistencies in experimental datasets, indicating which results may warrant further review or possible rejection. For nuclei where experimental data are sparse or unavailable, the network confirms theoretical calculations, reinforcing its utility as a predictive tool in nuclear physics. Finally, utilizing the predicted electric dipole polarizability, we extract the value of the symmetry energy at saturation density and find it consistent with results from the literature.

artificial neural networks

JELC-LITE: Unconventional Instructional Design for Special Operations Training

Current special operations staff training is based on the Joint Event Life Cycle (JELC). It addresses operational level tasks in multi-week, live military exercises which are planned over a 12 to 18 month timeframe. As the military experiences changing global mission sets, shorter training events using distributed technologies will increasingly be needed to augment traditional training. JELC-Lite is a new approach for providing relevant training between large scale exercises. This new streamlined, responsive training model uses distributed and virtualized training technologies to establish simulated scenarios. It keeps proficiency levels closer to optimal levels -- thereby reducing the performance degradation inherent in periodic training. It can be delivered to military as well as under-reached interagency groups to facilitate agile, repetitive training events. JELC-Lite is described by four phases paralleling the JELC, differing mostly in scope and scale. It has been successfully used with a Theater Special Operations Command and fits well within the current environment of reduced personnel and financial resources.

Friedman, Mark

Novel Approach to PV Inverter Modeling and Simulation Leveraging Experiments, Learning Based Modeling and Co-Simulation: Preprint

Photovoltaic inverter (PV) inverter manufacturers use custom, proprietary control approaches and topologies in their inverter design. Due to this proprietary nature, it is not possible to share EMT domain models for system studies. This research work presents a novel approach in experimental design, high fidelity data collection, use of learning-based modeling, and co-simulation to enhance the PV inverter modeling. We used a 20 kW off-the-shelf grid following PV inverter and subjected the inverter to controlled tests including voltage and frequency step changes, as well as solar irradiance variations. The recorded high frequency data was used in learning-based model training. This learning-based model was imported into an Electromagnetic Transient (EMT) simulation tool using co-simulation techniques to complete the modeling effort and integrate the model into an EMT simulation tool. The three key components in this research work are the design of experimental setup, use of learning-based approach for model development and use of co-simulation to complete the approach. The proposed approach will allow users to develop a model in a really short period of time and achieve reasonable inverter models.

artificial intelligence