Search NASA⌕ Search

SEARCH · Search NASA

Results for “Representation learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Object detection with deep learning for rare event search in the GADGET II TPC

In the pursuit of identifying rare two-particle events within the GADGET II Time Projection Chamber (TPC), this paper presents a comprehensive approach for leveraging Convolutional Neural Networks (CNNs) and various data processing methods. To address the inherent complexities of 3D TPC track reconstructions, the data is expressed in 2D projections and 1D quantities. This approach capitalizes on the diverse data modalities of the TPC, allowing for the efficient representation of the distinct features of the 3D events, with no loss in topology uniqueness. Additionally, it leverages the computational efficiency of 2D CNNs and benefits from the extensive availability of pre-trained models. Given the scarcity of real training data for the rare events of interest, simulated events are used to train the models to detect real events. To account for potential distribution shifts when predominantly depending on simulations, significant perturbations are embedded within the simulations. This produces a broad parameter space that works to account for potential physics parameter and detector response variations and uncertainties. These parameter-varied simulations are used to train sensitive 2D CNN object detectors. When combined with 1D histogram peak detection algorithms, this multi-modal detection framework is highly adept at identifying rare, two-particle events in data taken during experiment 21072 at the Facility for Rare Isotope Beams (FRIB), demonstrating a 100% recall for events of interest. Here, we present the methods and outcomes of our investigation and discuss the potential future applications of these techniques.

Convolutional neural network↗

Improving thermodynamic nudging in the E3SM Atmosphere Model version 2 (EAMv2): strategy and hindcast skills on weather systems

Nudging techniques are commonly employed to constrain atmospheric simulations toward observed states, facilitating model evaluation and sensitivity studies. However, if applied improperly – particularly to thermodynamic variables such as temperature and humidity – nudging can distort physical processes and introduce spurious biases, undermining the credibility of the simulations. This study presents an improved nudging implementation that applies vertically modulated tendencies to reduce adverse impacts on model physics. The framework is tested in version 2 of the Energy Exascale Earth System Model (EAMv2) using a suite of hindcast simulations nudged toward ERA5 reanalysis. We systematically evaluate the individual and combined effects of nudging wind, temperature, and humidity fields on the model's ability to represent large-scale atmospheric states and high-impact weather systems. Results show that the revised strategy – particularly when nudging temperature and humidity at selected levels – enhances hindcast skill by improving agreement with ERA5 without degrading the hydrological cycle or precipitation processes. Additional improvements in surface temperature, outgoing longwave radiation, and precipitation biases are achieved through targeted nudging of land-surface variables. The proposed approach strengthens the representation of large-scale conditions relevant to tropical cyclones, atmospheric rivers, and extratropical cyclones in the low-resolution EAMv2. These findings demonstrate that carefully designed thermodynamic nudging, especially of temperature and humidity, improves the realism of constrained simulations and broadens the utility of nudged EAMv2 for atmospheric modeling, machine learning, and high-impact weather research.

Atmospheric river↗

Divertor Plasma Detachment Control Neural Network

DivControlNN is a state-of-the-art software tool that leverages advanced machine learning techniques to predict and control divertor plasma behavior in fusion reactors. Plasma, a highly energetic and electrically charged gas, requires meticulous management to protect reactor components and maintain optimal energy production. Conventional simulation methods, although extremely detailed, typically demand extensive computational time-making them unsuitable for real-time control scenarios. DivControlNN addresses this challenge by learning from tens of thousands of high-fidelity simulations, thereby creating a rapid surrogate model that can deliver near-instantaneous predictions. At the core of its functionality is a sophisticated technique known as latent space mapping, which condenses complex, high-dimensional plasma data into a compact, lower-dimensional representation. This streamlined representation enables the system to quickly forecast essential plasma properties and determine the precise conditions required for effective detachment. Detachment is a crucial process in which the plasma is cooled before reaching the divertor plates, thereby reducing heat loads and mitigating material erosion. In recent experiments conducted on the KSTAR tokamak in South Korea, DivControlNN successfully guided the detachment process without any fine-tuning-even when applied to a new tungsten divertor configuration. By achieving a computational speed-up of over one hundred million times compared to traditional simulation methods while maintaining low prediction errors, DivControlNN stands to significantly enhance real-time control and diagnostic capabilities in future fusion reactors. This breakthrough paves the way for safer, more reliable reactor operation and represents a major advancement toward realizing fusion energy as a practical, sustainable, and clean power source.

Xu, Xueqiao [Lawrence Livermore National Laborator↗

NASA-STD 3001 and the Human Integration Design Handbook (HIDH): Evolution of NASA-STD-3000

The Habitability & Environmental Factors and Space Medicine Divisions have developed the Space Flight Human System Standard (SFHSS) (NASA-STD-3001) to replace NASA-STD-3000 as a new NASA standard for all human spaceflight programs. The SFHSS is composed of 2 volumes. Volume 1, Crew Health, contains medical levels of care, permissible exposure limits, and fitness for duty criteria, and permissible outcome limits as a means of defining successful operating criteria for the human system. Volume 2, Habitability and Environmental Health, contains environmental, habitability and human factors standards. Development of the Human Integration Design Handbook (HIDH), a companion to the standard, is currently under construction and entails the update and revision of NASA-STD-3000 data. This new handbook will, in the fashion of NASA STD-3000, assist engineers and designers in appropriately applying habitability, environmental and human factors principles to spacecraft design. Organized in a chapter-module-element structure, the HIDH will provide the guidance for the development of requirements, design considerations, lessons learned, example solutions, background research, and assist in the identification of gaps and research needs in the disciplines. Subject matter experts have been and continue to be solicited to participate in the update of the chapters. The purpose is to build the HIDH with the best and latest data, and provide a broad representation from experts in industry, academia, the military and the space program. The handbook and the two standards volumes work together in a unique way to achieve the required level of human-system interface. All new NASA programs will be required to meet Volumes 1 and 2. Volume 2 presents human interface goals in broad, non-verifiable standards. Volume 2 also requires that each new development program prepare a set of program-specific human factors requirements. These program-specific human and environmental factors requirements must be verifiable and tailored to assure the new system meets the Volume 2 standards. Programs will use the HIDH to write their verifiable program-specific requirements.

Pickett, Lynn↗

Ongoing Breakthroughs in Convective Parameterization

While the increase of computer power mobilizes a part of the community towards models with explicit convection or based on machine learning, we review the part of the literature dedicated to convective parameterization development for large-scale forecast and climate models. Recent findings: Many developments are underway to overcome endemic limitations of traditional convective parameterizations, either in unified or multi-object frameworks: scale-aware and stochastic approaches, new prognostic equations or representations of new components such as cold pools. Understanding their impact on the emergent properties of a model remains challenging, due to subsequent tuning of parameters and the limited understanding given by traditional metrics. Summary: Further effort still needs to be dedicated to the representation of the life cycle of convective systems, in particular their mesoscale organization and associated cloud cover. The development of more process-oriented metrics based on new observations is also needed to help quantify model improvement and better understand the mechanisms of climate change.

parameterizations for large-scale models↗

High-Dimensional Similarity Search with Quantum-Assisted Variational Autoencoder

Recent progress in quantum algorithms and hardware indicates the potential importance of quantum computing in the near future. However, finding suitable application areas remains an active area of research. Quantum machine learning is touted as a potential approach to demonstrate quantum advantage within both the gate-model and the adiabatic schemes. For instance, the QVAE has been proposed as a quantum enhancement to the discrete VAE. We extend on previous work and study the real-world applicability of a QVAE by presenting a proof-of-concept for similarity search in large-scale high-dimensional datasets. While exact and fast similarity search algorithms are available for low dimensional datasets, scaling to high-dimensional data is non-trivial. We show how to construct a space-efficient search index based on the latent space representation of a QVAE. Our experiments show a correlation between the Hamming distance in the embedded space and the Euclidean distance in the original space on the MODIS dataset. Further, we find real-world speedups compared to linear search and demonstrate memory-efficient scaling to half a billion data points.

Data mining, similarity search, quantum machine le↗

AI and simulation: What can they learn from each other

Simulation and Artificial Intelligence share a fertile common ground both from a practical and from a conceptual point of view. Strengths and weaknesses of both Knowledge Based System and Modeling and Simulation are examined and three types of systems that combine the strengths of both technologies are discussed. These types of systems are a practical starting point, however, the real strengths of both technologies will be exploited only when they are combined in a common knowledge representation paradigm. From an even deeper conceptual point of view, one might even argue that the ability to reason from a set of facts (i.e., Expert System) is less representative of human reasoning than the ability to make a model of the world, change it as required, and derive conclusions about the expected behavior of world entities. This is a fundamental problem in AI, and Modeling Theory can contribute to its solution. The application of Knowledge Engineering technology to a Distributed Processing Network Simulator (DPNS) is discussed.

Colombano, Silvano P.↗

MaTableGPT: GPT‐Based Table Data Extractor from Materials Science Literature

Abstract Efficiently extracting data from tables in the scientific literature is pivotal for building large‐scale databases. However, the tables reported in materials science papers exist in highly diverse forms; thus, rule‐based extractions are an ineffective approach. To overcome this challenge, the study presents MaTableGPT, which is a GPT‐based table data extractor from the materials science literature. MaTableGPT features key strategies of table data representation and table splitting for better GPT comprehension and filtering hallucinated information through follow‐up questions. When applied to a vast volume of water splitting catalysis literature, MaTableGPT achieves an extraction accuracy (total F1 score) of up to 96.8%. Through comprehensive evaluations of the GPT usage cost, labeling cost, and extraction accuracy for the learning methods of zero‐shot, few‐shot, and fine‐tuning, the study presents a Pareto‐front mapping where the few‐shot learning method is found to be the most balanced solution owing to both its high extraction accuracy (total F1 score >95%) and low cost (GPT usage cost of 5.97 US dollars and labeling cost of 10 I/O paired examples). The statistical analyses conducted on the database generated by MaTableGPT revealed valuable insights into the distribution of the overpotential and elemental utilization across the reported catalysts in the water splitting literature.

Yi, Gyeong Hoon [Computational Science Research Ce↗

Passive mapping and intermittent exploration for mobile robots

An adaptive state space architecture is combined with diktiometric representation to provide the framework for designing a robot mapping system with flexible navigation planning tasks. This involves indexing waypoints described as expectations, geometric indexing, and perceptual indexing. Matching and updating the robot's projected position and sensory inputs with indexing waypoints involves matchers, dynamic priorities, transients, and waypoint restructuring. The robot's map learning can be opganized around the principles of passive mapping.

Engleson, Sean P.↗

Domain Knowledge Guided Bayesian Optimization For Autonomous Alignment Of Complex Scientific Instruments

Bayesian Optimization (BO) is a powerful tool for optimizing complex non-linear systems. However, its performance degrades in high-dimensional problems with tightly coupled parameters and highly asymmetric objective landscapes, where rewards are sparse. In such needle-in-a-haystack scenarios, even advanced methods like trust-region BO (TurBO) often lead to unsatisfactory results. We propose a domain knowledge guided Bayesian Optimization approach, which leverages physical insight to fundamentally simplify the search problem by transforming coordinates to decouple input features and align the active subspaces with the primary search axes. We demonstrate this approach's efficacy on a challenging 12-dimensional, 6-crystal Split-and-Delay optical system, where conventional approaches, including standard BO, TuRBO and multi-objective BO, consistently led to unsatisfactory results. When combined with an reverse annealing exploration strategy, this approach reliably converges to the global optimum. The coordinate transformation itself is the key to this success, significantly accelerating the search by aligning input co-ordinate axes with the problem's active subspaces. As increasingly complex scientific instruments, from large telescopes to new spectrometers at X-ray Free Electron Lasers are deployed, the demand for robust high-dimensional optimization grows. Our results demonstrate a generalizable paradigm: leveraging physical insight to transform high-dimensional, coupled optimization problems into simpler representations can enable rapid and robust automated tuning for consistent high performance while still retaining current optimization algorithms.

FOS: Computer and information sciences↗

Application of Machine Learning and Data Augmentation Algorithms in the Discovery of Metal Hydrides for Hydrogen Storage

The development of efficient and sustainable hydrogen storage materials is a key challenge for realizing hydrogen as a clean and flexible energy carrier. Among various options, metal hydrides offer high volumetric storage density and operational safety, yet their application is limited by thermodynamic, kinetic, and compositional constraints. In this work, we investigate the potential of machine learning (ML) to predict key thermodynamic properties—equilibrium plateau pressure, enthalpy, and entropy of hydride formation—based solely on alloy composition using Magpie-generated descriptors. We significantly expand an existing experimental dataset from ~400 to 806 entries and assess the impact of dataset size and data augmentation, using the PADRE algorithm, on model performance. Models including Support Vector Machines and Gradient Boosted Random Forests were trained and optimized via grid search and cross-validation. Results show a marked improvement in predictive accuracy with increased dataset size, while data augmentation benefits are limited to smaller datasets and do not improve accuracy in underrepresented pressure regimes. Furthermore, clustering and cross-validation analyses highlight the limited generalizability of models across different material classes, though high accuracy is achieved when training and testing within a single hydride family (e.g., AB2). The study demonstrates the viability and limitations of ML for accelerating hydride discovery, emphasizing the importance of dataset diversity and representation for robust property prediction.

augmentation↗

Improving Search Properties in Genetic Programming

With the advancing computer processing capabilities, practical computer applications are mostly limited by the amount of human programming required to accomplish a specific task. This necessary human participation creates many problems, such as dramatically increased cost. To alleviate the problem, computers must become more autonomous. In other words, computers must be capable to program/reprogram themselves to adapt to changing environments/tasks/demands/domains. Evolutionary computation offers potential means, but it must be advanced beyond its current practical limitations. Evolutionary algorithms model nature. They maintain a population of structures representing potential solutions to the problem at hand. These structures undergo a simulated evolution by means of mutation, crossover, and a Darwinian selective pressure. Genetic programming (GP) is the most promising example of an evolutionary algorithm. In GP, the structures that evolve are trees, which is a dramatic departure from previously used representations such as strings in genetic algorithms. The space of potential trees is defined by means of their elements: functions, which label internal nodes, and terminals, which label leaves. By attaching semantic interpretation to those elements, trees can be interpreted as computer programs (given an interpreter), evolved architectures, etc. JSC has begun exploring GP as a potential tool for its long-term project on evolving dextrous robotic capabilities. Last year we identified representation redundancies as the primary source of inefficiency in GP. Subsequently, we proposed a method to use problem constraints to reduce those redundancies, effectively reducing GP complexity. This method was implemented afterwards at the University of Missouri. This summer, we have evaluated the payoff from using problem constraints to reduce search complexity on two classes of problems: learning boolean functions and solving the forward kinematics problem. We have also developed and implemented methods to use additional problem heuristics to fine-tune the searchable space, and to use typing information to further reduce the search space. Additional improvements have been proposed, but they are yet to be explored and implemented.

Janikow, Cezary Z.↗

System identification in the repetition domain

Procedures for system identification using realization theory in conjunction with learning control ideas are developed. The Markov parameters of the system are identified by combining data from repeated experiments. Three approaches are discussed for identification of as many Markov parameters as sample points in the experiment. Making use of all the parameters, realization theory is then employed to determine the system order and to obtain a minimal order representation. The first two approaches are non-recursive, which in the case of noise-free data yields a one step solution. The third approach uses a recursive formulation rendered from adaptive control but modified for successive experiments. A simple example shows the numerical convergence of the identified parameters as a function of the number of experiments. The procedure presented herein is an extension of the existing Eigensystem Realization Algorithm (ERA), which has been successfully applied for system identification of large structures.

Juang, Jer-Nan↗

A View from Space: Evolution of the 1997-98 El Nino and La Nina

After the last extreme El Nino in 1982-1983, an extensive in situ observing system was deployed in the tropical Pacific Ocean in support of monitoring and predicting El Nino. Within the past ten years a series of ocean and atmosphere remote sensing satellites have been launched that serve to supplement and enhance the observations being taken at the surface, and at depth, in the equatorial Pacific Ocean. The 1997-1998 "El Nino Event of the Century" has been the best monitored El Nino on record. The 1997-1998 El Nino will be the first time a major El Nino event and subsequent La Nina will have been observed from start to finish from a combination of remotely-sensed measurements of sea surface temperature, sea surface topography, sea surface winds, ocean color, and precipitation. Among some of the lessons learned to date from the 1997-1998 event have been the need for global observations in addition to just those in the equatorial Pacific Ocean. In this presentation the evolution of the 1997-1998 El Nino will be depicted from the unique vantage point provided by these space-based observations as analyzed separately, and together as a representation of the coupled system. Comparisons and contrasts with the evolution 1982-1983 El Nino and how the in situ and space-based observations complement each other will be discussed.

Busalacchi, Antonio J.↗

A View from Space: Evolution of the 1997-1998 El Nino and La Nina

After the last extreme El Nino in 1982-1983, an extensive in situ observing system was deployed in the tropical Pacific Ocean in support of monitoring and predicting El Nino. Within the past ten years a series of ocean and atmosphere remote sensing satellites have been launched that serve to supplement and enhance the observations being taken at the surface, and at depth, in the equatorial Pacific Ocean. The 1997-1998 "El Nino Event of the Century" has been the best monitored El Nino on record. The 1997-1998 El Nino will be the first time a major El Nino event and subsequent La Nina will have been observed from start to finish from a combination of remotely-sensed measurements of sea surface temperature, sea surface topography, sea surface winds, ocean color, and precipitation. Among some of the lessons learned to date from the 1997-1998 event have been the need for global observations in addition to just those in the equatorial Pacific Ocean. In this presentation the evolution of the 1997-1998 El Nino will be depicted from the unique vantage point provided by these space-based observations as analyzed separately, and together as a representation of the coupled system. Comparisons and contrasts with the evolution 1982-1983 El Nino and how the in situ and space-based observations complement each other will be discussed.

Busalacchi, Antonio J.↗

A View From Space: Evolution of the 1997-1998 El Nino and La Nina

After the last extreme El Nino in 1982-1983, an extensive in situ observing system was deployed in the tropical Pacific Ocean in support of monitoring and predicting El Nino. Within the past ten years a series of ocean and atmosphere remote sensing satellites have been launched that serve to supplement and enhance the observations being taken at the surface, and at depth, in the equatorial Pacific Ocean. The 1997-1998 "El Nino Event of the Century" has been the best monitored El Nino on record. The 1997-1998 El Nino will be the first time a major El Nino event and subsequent La Nina will have been observed from start to finish from a combination of remotely-sensed measurements of sea surface temperature, sea surface topography, sea surface wind, ocean color, and precipitation, Among some of the lessons learned to date from the 1997-1998 event have been the need for global observation in addition to just those in the equatorial Pacific Ocean. In this presentation the evolution of the 1997-1998 El Nino will be depicted from the unique vantage point provided by these space-based observations as analyzed separately, and together as a representation of the coupled system. Comparisons and contrasts with the evolution 1982-1983 El Nino and how the in situ and space-based observations complement each other will be discussed.

Busalacchi, Antonio↗

Inference of response functions with the help of machine-learning algorithms

Response functions are a key quantity to describe the near-equilibrium dynamics of strongly interacting many-body systems. Recent techniques that attempt to overcome the challenges of calculating these ab initio have employed expansions in terms of orthogonal polynomials. We employ a neural network prediction algorithm to reconstruct a response function 𝑆⁡(𝜔) defined over a range in frequencies 𝜔. Here, we represent the calculated response function as a truncated Chebyshev series whose coefficients can be optimized to reduce the representation error. We compare the quality of response functions obtained using coefficients calculated using a neural network (NN) algorithm with those computed using the Gaussian integral transform (GIT) method. In the regime where only a small number of terms in the Chebyshev series are retained, we find that the NN scheme outperforms the GIT method.

Kurkcuoglu, Doga Murat [Fermi National Accelerator↗

Graph theory inspired anomaly detection at the LHC

Designing model-independent anomaly detection algorithms for analyzing LHC data remains a central challenge in the search for new physics, due to the high dimensionality of collider events. In this work, we develop a graph autoencoder as an unsupervised, model-agnostic tool for anomaly detection, using the LHC Olympics dataset as a benchmark. By representing jet constituents as a graph, we introduce a method to systematically control the information available to the model through sparse graph constructions that serve as physically motivated inductive biases. Specifically, (1) we construct graph autoencoders based on locally rigid Laman graphs and globally rigid unique graphs, and (2) we explore the clustering of jet constituents into subjets to interpolate between high- and low-level input representations. We obtain the best performance, measured in terms of the Significance Improvement Characteristic curve for an intermediate level of subjet clustering and certain sparse unique graph constructions. We further investigate the role of graph connectivity in jet classification tasks. Our results demonstrate the potential of leveraging graph-theoretic insights to refine and increase the interpretability of machine learning tools for collider experiments.

Automation↗