Search NASA⌕ Search

SEARCH · Search NASA

Results for “Representation learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Variational autoencoders for at-source data reduction and anomaly detection in high energy particle detectors

Detectors in next-generation high-energy physics experiments face several daunting requirements, such as high data rates, damaging radiation exposure, and stringent constraints on power, space, and latency. To address these challenges, machine learning in readout electronics can be leveraged for smart detector designs, enabling intelligent inference and data reduction at-source. Variational autoencoders (VAEs) offer a variety of benefits for front-end readout; an on-sensor encoder can perform efficient lossy data compression while simultaneously providing a latent space representation that can be used for anomaly detection. Results are presented from low-latency and resource-efficient VAEs for front-end data processing in a futuristic silicon pixel detector. Encoder-based data compression is found to preserve good performance of off-detector analysis while significantly reducing the off-detector data rate as compared to a similarly sized data filtering approach. Furthermore, the latent space information is found to be a useful discriminator in the context of real-time sensor defect monitoring. Together, these results highlight the multifaceted utility of autoencoder-based front-end readout schemes and motivate their consideration in future detector designs.

47 OTHER INSTRUMENTATION↗

Key insights from US Department of Energy Better Plants workforce development bootcamps (2022–2025)

This study examines the effectiveness of the US Department of Energy’s Better Plants Program Bootcamps, which are designed to enhance participants’ technical skills in improving energy efficiency and optimizing operations in manufacturing facilities. Through the analysis of survey data collected from 529 participants across 9 bootcamps, the research investigates the motivations, benefits, and demographic trends of attendees. The findings reveal that skill acquisition and improvement are primary drivers for participation, with key benefits including hands-on training on diagnostic equipment and software tools, networking opportunities, and access to technical resources. The analysis shows strong participation from sectors characterized by high energy consumption and employment, such as chemical and transportation equipment manufacturing. Over 50% of participants have job titles that include “EHS” or “Energy” showing their key roles in leading energy efficiency and energy management efforts in manufacturing. Furthermore, the analysis highlights the distribution of participants across managerial, engineering, and technical roles, revealing a higher representation of managers and engineers. This observation suggests a need for targeted outreach to engage technicians, equipment operators, maintenance staff, and floor workers to ensure comprehensive workforce development. The post-bootcamp survey showed that the participants highly valued the opportunities for peer learning and idea exchange, and the benefits they gained from them. This research contributes to the advancement of manufacturing education by demonstrating the efficacy of specialized training in addressing critical industry challenges and fostering a more competent and empowered workforce.

Energy efficiency↗

Overview of Advanced Stirling Convertor Dual Convertor Controller Development and Testing at NASA Glenn Research Center

For over two decades, NASA Glenn Research Center (GRC) has been supporting the development of Radioisotope Power Systems (RPS). NASA desires higher conversion efficiency RPS options that are reliable and robust with long life design. Dynamic conversion, such as Stirling offer the potential for higher conversion efficiencies than current RPS but have yet to be demonstrated in a flight application. GRC developed a non-nuclear representation of a RPS consisting of a pair of Advanced Stirling Convertors (ASC), a Dual Convertor Controller (DCC), and associated support equipment. The DCC was designed by the Johns Hopkins University/Applied Physics Laboratory (JHU/APL) to actively control a pair of ASCs. Based on lessons learned, three generations of DCCs were developed over the past decade. After each generation of the DCC was completed, tests were performed to verify functionality. These tests include acceptance testing, fault testing, characterization testing, and testing with the Radioisotope Power System Systems Integration Laboratory (RSIL). Acceptance testing verifies that the DCC can control a pair of ASCs to produce full convertor power output. The fault testing verifies that the DCC can maintain control of the ASCs during either internal or external faults. Characterization testing verifies the functionality of the DCC over a range of bus voltage values. RSIL provides insight into the electrical interactions between a representative radioisotope power generator, its associated control schemes, and realistic electric system loads. The DCC design, development, lessons learned, test results, and future work are presented in this paper.

Stirling↗

Map learning with indistinguishable locations

Nearly all spatial reasoning problems involve uncertainty of one sort or another. Uncertainty arises due to the inaccuracies of sensors used in measuring distances and angels. This is inferred as directional uncertainty. Uncertainty also arises in combining spatial information when one location is mistakenly identified with another. This is referred to as recognition uncertainty. Most problems in constructing spatial representations (maps) for the purpose of navigation involve both directional and recognition uncertainty. It is shown that a particular class of spatial reasoning problems involving the construction of representations of large-scale space can be solved efficiently even in the presence of directional and recognition uncertainty. Particular attention is paid to the problems that arise due to recognition uncertainty. The results described are applicable to the construction of global maps from satellite data as well as the construction of local navigation maps from measurements made by a rover in exploring a planetary surface.

Basye, Kenneth↗

Semi-supervised permutation invariant particle-level anomaly detection

The development of analysis methods to distinguish potential beyond the Standard Model phenomena in a model-agnostic way can significantly enhance the discovery reach in collider experiments. However, the typical machine learning (ML) algorithms employed for this task require fixed length and ordered inputs that break the natural permutation invariance in collision events. To address this, a semi-supervised anomaly detection tool is presented that takes a variable number of particle-level inputs and leverages a signal model to encode this information into a permutation invariant, event-level representation via supervised training with a Particle Flow Network (PFN). Data events are then encoded into this representation and given as input to an autoencoder for unsupervised ANomaly deTEction on particLe flOw latent sPacE (ANTELOPE), classifying anomalous events based on a low-level and permutation invariant input modeling. Performance of the ANTELOPE architecture is evaluated on simulated samples of hadronic processes in a high energy collider experiment, showing good capability to distinguish disparate models of new physics.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

An explainable variational autoencoder model for three-dimensional acoustic emission source localization in hollow cylindrical structures

We introduce an explainable variational autoencoder for three-dimensional (3D) localization of acoustic emission sources in hollow cylindrical structures, with an unsupervised approach. This research capitalizes on multi-arrival waveforms generated by helical path propagation in cylindrical geometries to enable efficient two-receiver localization. By integrating the modal characteristics of Lamb modes under multi-path conditions, we demonstrate that two sets of time-of-arrival differences and peak amplitudes extracted from one receiver can serve as effective localization features. This initial approach identifies four potential source locations, highlighting the feasibility of two-receiver source localization using traditional feature extraction methods. However, direct extraction can be challenging when mode overlaps occur, complicating the localization process. To address this, our work proposes a novel waveform-based method. This method leverages the consistent dispersion characteristics within isotropic materials, where each unique combination of mode arrival times and peak amplitudes constructs a distinct waveform. This distinctiveness overcomes the ambiguities associated with mode overlaps, significantly enhancing the method’s precision and robustness. Our approach adopts a data-driven strategy for waveform-based localization using variational autoencoder (VAE). VAE discerns waveform patterns for localization, while also addressing data uncertainties. The VAE’s encoder and decoder networks capture the localization process and the source’s influence on waveform generation, respectively, guiding latent variables to segregate waveforms by source in the latent space. The design of the learning process focuses on specific localization characteristics to enhance result explainability. Localization predictions are generated by projecting test waveforms, not included in the training set, onto a trained latent space. The prediction is determined using a nearest-neighbor approach based on the closest latent representation of a source. Validation with pencil-lead-break tests on a metallic pipe confirmed our method’s effectiveness, achieving an averaged 3D localization accuracy of 0.84.

Lee, Guan-Wei↗

Digitizing Today’s Buildings in the Real World: Lessons from Field Demonstrations

Digital twins, created by generating a virtual replica of a building, enable safe evaluation of operational scenarios and applications like fault detection and diagnosis and advanced controls. However, a prerequisite is the creation of a machine-readable digital representation of a building, currently hindered by fragmented information scattered across mechanical drawings, point lists, and natural language sequences. As a result, digital twin development remains labor-intensive, error-prone, and difficult to validate. To address these challenges, two efforts from ASHRAE aim to support the digitalization of buildings. ASHRAE s223 establishes a semantic model of buildings, representing system components, configuration, and data sources. ASHRAE s231 defines a vendor-neutral programming language for expressing their control logic. As the industry evaluates implementing them in their products, understanding the challenges that vendors and implementers may face is crucial. In this paper, we present findings and lessons learned from field demonstrations in five buildings that implemented control applications using ASHRAE s223 and s231. The demonstrations highlight how semantic modeling and formalized control descriptions can significantly reduce software development time, manual point mapping, and hard-coding. Beyond time efficiency, they enable reliable automation by minimizing human interpretation and providing a means for consistency across projects. We describe the processes and best practices for model creation and model usage, from translating heterogeneous building documentation into semantic representations to implementing control logic in real-world systems. Finally, we discuss the challenges that persist, including integration with legacy software environments, gaps in interoperability, and the level of expertise still required to effectively leverage semantic models.

Prakash, Anand Krishnan↗

MOSAIC-CONUS: A Multimodal, Multi-Temporally Paired Dataset for Earth Sciences

Earth embeddings—vector representations of geographic locations indexed in space and time—are emerging as a unifying interface for geospatial AI. However, their quality depends not only on model design, but on how multimodal Earth observation (EO) data are spatially indexed, temporally aligned, and cross-modally associated during pretraining. We introduce MOSAIC-CONUS (Multimodal Observations with Spatially Aligned Imagery, Urban Points of Interest, In-Situ Measurements and Text Captions), a large-scale EO dataset over the contiguous United States, organized around 250,000 stratified point indices that serve as stable spatial keys across seven modalities: active radar, passive optical imagery, lidar-derived elevation, land cover, functional context, hydrometeorological measurements, and textual summaries. Unlike existing EO datasets, MOSAIC-CONUS introduces four contributions not jointly addressed in prior work: 1. an open-source, large-scale multimodal EO corpus structured around point-indexed data designed to support Earth embedding learning; 2. explicit radar-optical pairing tables spanning twelve temporal alignment regimes, formalizing cross-sensor alignment as a controllable variable for analyzing how temporal mismatch across modalities influences learned embeddings quality; 3. a benchmark suite spanning cross-modal retrieval, annual nightlights regression, and basin-held-out streamflow prediction, positioning MOSAIC-CONUS as a benchmark-ready resource for multimodal AI systems; and 4. a language-based embedding layer through co-registered textual summaries, enabling Earth embeddings to function as a queryable interface for agentic AI systems. The dataset and pairing protocols are publicly released.

54 ENVIRONMENTAL SCIENCES↗

ICAT: The Interactive Corpus Analysis Tool

The Interactive Corpus Analysis Tool (ICAT) is a Python library for creating dashboards to explore textual datasets and build simple binary classification models to help filter through them and focus on entries of interest. This tool uses a form of interactive machine learning (IML), a paradigm of “machine teaching” (Simard et al., 2017) that sits at the intersection of the fields of human computer interaction (HCI), visual analytics, and machine learning. The intent of ICAT is to allow subject matter experts (SME) with limited to no experience in machine learning to benefit from an iterative human-in-the-loop (HITL) approach to building their own model without needing to understand the details of the underlying algorithm. This interactivity is achieved by allowing the user to create features, label data points, and visually manipulate a representation of the features to manually cluster and investigate data, while a model is trained on the fly based on these actions. ICAT is built on top of the Panel (Holoviz, 2018) library, using a combination of Vega, a custom IPyWidget using D3, and ipyvuetify, and is intended to be used inside of a Jupyter environment.

Martindale, Nathan [Oak Ridge National Laboratory ↗

Exploiting Multi-Domain Features for Detection of Unclassified Electromagnetic Signals

Deep Learning based classification techniques have shown excellent performance in static environments, where the training and testing samples are drawn from the same distribution. However, real world scenarios often present samples that do not belong to the known set of classes chosen during training. This is quite common for electromagnetic signals, where it is impractical to assume that all possible waveforms are known a-priori, specially in scenarios like warfare. To address this problem, we propose a deep learning based adversarial model where the generator learns to generate waveform features that can deceive the discriminator model as true samples. We introduce domain knowledge of wireless signals by decomposing the signal into a lower dimensional unique feature set, which is used for classifying known versus unknown signals. We further introduce multiple domain representations of the signal to extract features and combine them together to accurately classify new waveforms as an unknown class. Our results show that combined features from multiple domains outperform any single domain representation, especially at low SNR regimes with fewer number of samples to classify.

99 - GENERAL AND MISCELLANEOUS↗

Exploiting Multi-Domain Features for Detection of Unclassified Electromagnetic Signals (Presentation)

Deep Learning based classification techniques have shown excellent performance in static environments, where the training and testing samples are drawn from the same distribution. However, real world scenarios often present samples that do not belong to the known set of classes chosen during training. This is quite common for electromagnetic signals, where it is impractical to assume that all possible waveforms are known a-priori, specially in scenarios like warfare. To address this problem, we propose a deep learning based adversarial model where the generator learns to generate waveform features that can deceive the discriminator model as true samples. We introduce domain knowledge of wireless signals by decomposing the signal into a lower dimensional unique feature set, which is used for classifying known versus unknown signals. We further introduce multiple domain representations of the signal to extract features and combine them together to accurately classify new waveforms as an unknown class. Our results show that combined features from multiple domains outperform any single domain representation, especially at low SNR regimes with fewer number of samples to classify.

99 - GENERAL AND MISCELLANEOUS↗

Near real-time air quality forecasts using the NASA GEOS model

NASA's Global Modeling and Assimilation Office (GMAO) produces high-resolution global forecasts for weather, aerosols, and air quality. The NASA Global Earth Observing System (GEOS) model has been expanded to provide global near-real-time5-day forecasts of atmospheric composition at unprecedented horizontal resolution of 0.25 degrees (~25 km). This composition forecast system (GEOS-CF) combines the operational GEOS weather forecasting model with the state-of-the-science GEOS-Chem chemistry module (version 12) to provide detailed analysis of a wide range of air pollutants such as ozone, carbon monoxide, nitrogen oxides, and fine particulate matter (PM2.5). Satellite observations are assimilated into the system for improved representation of weather and smoke. The assimilation system is being expanded to include chemically reactive trace gases. We discuss current capabilities of the GEOS Constituent Data Assimilation System (CoDAS) to improve atmospheric composition modeling and possible future directions, notably incorporating new observations (TROPOMI, geostationary satellites) and machine learning techniques. We show how machine learning techniques can be used to correct for sub-grid-scale variability, which further improves model estimates at a given observation site.

Air Quality Forecast↗

The Mars 2020 Ground Data System Architecture

The Mars 2020 Mission’s primary objective is to collect 20 geographically unique samples during its prime mission of one and a quarter Martian years, or just over 2 Earth years. Mission planners determined the project needed to develop a system that would enable the operations team to analyze engineering and science data, make science decisions, select viable rover targets at a millimeter resolution and validate an uplink bundle for a car sized rover with more complex science instruments than any previous Mars surface mission. All this had to be done within a five hour time frame. Doing this with a small team would be a challenge, but this had to be accomplished by a large team of engineers and scientists located across North America and Europe. Achieving this level of operational efficiency was unheard of in the prime mission. In addition, the mission had another set of requirements that had nothing to do with surface operations; the Mars 2020 Ground Data System (GDS) was also expected to comply with a new set of security requirements to keep up with the ever changing cybersecurity landscape. The Mars 2020 Ground Data System (GDS) is a re-architected version of the Mars Science Laboratory GDS. The primary goal was to integrate the lessons learned from previous Mars surface missions, accommodate a set of new requirements and capabilities required to ensure mission success, and comply with a new set of cybersecurity controls. The new architecture includes several unique qualities including a data lake, language-agnostic system-wide event-based operations, containerization, automated deployment, network segmentation, infrastructure-as-code, API-driven interfaces, and the first Mars surface GDS to operate primarily in the cloud. The new architecture enabled greater access to the system’s data, tighter integration with the operations team, and a higher level of traceability. The availability of the data also enabled a new set of capabilities previously not possible on surface missions. These new capabilities include an autonomous data to information, pipeline for downlink analysis, horizontal scaling of science data processing capabilities, autonomous round trip data tracking of science and engineering data, integration of flight system state into the tactical planning cycle, high fidelity targeting utilizing kinematic data, and hierarchical image and 3d meshes data representations. This paper will introduce the requirements for the Mars 2020 Mission, the heritage architecture, and the rationale for the changes to achieve the new architecture. The paper will continue to describe the fundamental changes made to the GDS architecture, how these changes enabled a more tightly integrated GDS, and the new capabilities that were enabled by the new architecture. The paper will conclude with the lessons learned from the process of rearchitecting a heritage GDS system and from the first 200 days of operations supporting over 800 users from around the world.

Lopez-Roig, Reynaldo↗

Clifford Neural Operators on Atmospheric Data Influenced Partial Differential Equations

Mathematical representations of the atmosphere are key to forecasting and research tasks across Earth science. Numerically solving the underlying partial differential equations(PDEs) of the atmosphere, however, can be difficult and computationally expensive with numerous trade-offs between computing efficiency and accuracy. Utilizing neural net-works to learn approximations of the PDE solutions from the data can help us model complex phenomena more efficiently than traditional numerical schemes. Here, we have applied Clifford algebra-based neural operators for predicting atmospheric variables. Clifford Fourier neural operators are used with two different backbone architectures, ResNet and UNet, on custom data of U10, V10 and surface pressure as well as U500, V500 and Z500. Clifford Fourier neural operators, coupled with ResNet and UNet architectures, are applied to a key reanalysis dataset. Model performance is initially strong, but we observe increasing errors, resulting in the model becoming highly unstable.

Sujit Roy↗

Universal Workflow Language and Software Enable Geometric Learning and FAIR Scientific Protocol Reporting

Written language and conventional data structures for representing scientific procedures suffer from low process detail, often fail to accurately represent protocols, and lack universality. New strategies for the handling of experimental data are needed to provide viable process information for both humans and machines. In this work, we present the universal workflow language (UWL) and interface (UWLi). UWL is a findable, accessible, interoperable, and reusable (FAIR)-compatible, graph-based data architecture that can capture arbitrary scientific procedures through workflow representation, and UWLi is an accompanying software package for building, manipulating, and interpreting UWL entries. The UWL format was found to be highly effective in identifying deficiencies in the reported process details of high-impact, peer-reviewed scientific journals, and in simulated scenarios, the graph format was shown to be more effective than conventional methods in predictively modeling the outcome of diverse scientific protocols. Implementation of UWL could enable more accurate scientific communication and more impactful process datasets.

14 SOLAR ENERGY↗

Acoustic-based monitoring and machine learning of component status for microreactor applications

This report provides a description and assessment of recent efforts to couple acoustic-based experimental measurements and characterization with machine learning models in order to enhance structural health monitoring capabilities for nuclear microreactors. With resilient embedded sensors in development by others supported by programs funded by the US Department of Energy’s Office of Nuclear Energy, the work described herein builds upon ongoing efforts to improve non-destructive testing technology that relates measured acoustic signatures to component stresses and/or structural defects, using a combination of new experimental measurements and machine learning architectures. The experimental procedure remained similar to that developed for the previous year’s demonstration of damage detection by the authors, with the same damaged sample tested under similar applied stress conditions. Notably, a new mounting fixture was designed and implemented to improve measurement consistency and a more sophisticated laser Doppler vibrometer was employed to make high-fidelity vibration measurements. Two nominally identical sets of training data were collected for each experimental setup to better understand the repeatability of the experiment and to better test the generality of trained neural network models. Additionally, we obtained new high-quality 3D mode shapes of the damaged test article at various stress and excitation levels, providing greater insights into the physical response of the sample during testing. Previously, we demonstrated that a machine learning model based on a convolutional neural network can predict structural details of an artificially introduced interface (intact, rough cut, smooth cut), and the applied torque level. In this study, we have transitioned to graph-based neural network architectures to better develop and test a flexible framework that is more suitable to being transferred away from controlled benchtop experiments and into more applied settings where less-structured data inputs may be expected. In general, performance testing of a graph neural network on frequency-domain representations of the data indicates strong and consistent identification of test conditions for datasets recorded on damaged components. With goals of predicting damage location and other changing experimental conditions using limited datasets, predictive models using a graph neural network architecture correctly predicted the applied torque level with an accuracy of 85% using only a single measurement point and predicted within one torque level in 95% of test windows. Predictions of damage location had limited success due to the symmetry and minimal number of the damage scenarios presented during model training. Results were ambiguous as to whether the model could detect the location of the artificial damage, or if it was instead learning the location of a given measurement point on the part and subsequently detecting which points were closest to the location of the damage. This finding will be factored into upcoming planned work on damaged graphite components, where new experimental tests with a larger number and variety of damage scenarios are expected to provide improved validation of recent developments in monitoring methodology.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Abstraction and Problem Reformulation

In work done jointly with Toby Walsh, the author has provided a sound theoretical foundation to the process of reasoning with abstraction (GW90c, GWS9, GW9Ob, GW90a). The notion of abstraction formalized in this work can be informally described as: (property 1), the process of mapping a representation of a problem, called (following historical convention (Sac74)) the 'ground' representation, onto a new representation, called the 'abstract' representation, which, (property 2) helps deal with the problem in the original search space by preserving certain desirable properties and (property 3) is simpler to handle as it is constructed from the ground representation by "throwing away details". One desirable property preserved by an abstraction is provability; often there is a relationship between provability in the ground representation and provability in the abstract representation. Another can be deduction or, possibly inconsistency. By 'throwing away details' we usually mean that the problem is described in a language with a smaller search space (for instance a propositional language or a language without variables) in which formulae of the abstract representation are obtained from the formulae of the ground representation by the use of some terminating rewriting technique. Often we require that the use of abstraction results in more efficient .reasoning. However, it might simply increase the number of facts asserted (eg. by allowing, in practice, the exploration of deeper search spaces or by implementing some form of learning). Among all abstractions, three very important classes have been identified. They relate the set of facts provable in the ground space to those provable in the abstract space. We call: TI abstractions all those abstractions where the abstractions of all the provable facts of the ground space are provable in the abstract space; TD abstractions all those abstractions wllere the 'unabstractions' of all the provable facts of the abstract space are provable in the ground space; and TC abstractions all those abstractions where a fact is provable in the ground space if and only if its abstraction is provable in the abstract space.

Giunchiglia, Fausto↗

Improving Grid Awareness by Empowering Utilities with Machine Learning and Artificial Intelligence

Gap filling time series data typically depends on linear interpolation. More recently gap filling advancements include machine learning techniques. However, none leverage advanced learning approach that uses cohort training or a neighborhood informed approach, which is described in this report. The report also describes a physics informed approach using Reduced Order Models (ROM). There are several methods to capture the nature of the detailed system in aggregated models, however there is a trade-off for these methods developed for multiple applications. These methods have specific requirements and applications that includes consideration of dynamics or covering a larger range of operating conditions, etc. The various methods of aggregation are: 1) Thevenin equivalents for downstream networks 2) Equivalent feeder representation to capture downstream network losses accurately 3) Structured reduced order models for dynamics 4) System identification-based ROM (abstract dynamical model) Methods described in items 1 and 2 above are ideal for steady-state models and useful for this application. Of these two methods, based on the data availability, the targeted application, the reduced order model that is proposed to be developed is the equivalent feeder model representation. This includes a structure of the reduced order model whose parameters can be determined by the system load and losses with the meter measurements.

14 SOLAR ENERGY↗