Search NASASearch

SEARCH · Search NASA

Results for “Statistical Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Accurate and efficient parameterization of an atomic cluster expansion (ACE) potential for ammonia under extreme conditions

We present a machine learning interatomic potential for ammonia designed to capture its complex multiphase behavior, including both molecular and superionic phases. The potential is based on the atomic cluster expansion (ACE) formulation and has been parameterized to facilitate high-fidelity molecular dynamics simulations of ammonia under extreme conditions, for pressures up to 100 GPa and for temperatures above 500 K and up to 6000 K. A diverse range of configurations was generated through high-quality ab initio molecular dynamics simulations, covering insulating and superionic ice phases, liquid ammonia, molecular nitrogen (N 2 ) and hydrogen (H 2 ), and metastable compounds that form upon dissociation, including $NH^{+}_{4}$, $H^{+}_{3}$, N 2 H 4 , and N 3 H. We demonstrate that the ammonia ACE potential accurately reproduces experimental and density functional theory predicted isotherms and Hugoniots. Crucially, the potential is able to capture the intricate phase behavior of ammonia, including the transition from insulating molecular fluid to the superionic phase. This work provides a robust interatomic potential that can be used for large-scale, accurate simulations of ammonia under extreme thermodynamic conditions, offering a powerful tool for investigating its behavior in various phases and applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

An interactive machine learning platform for analyzing multi-particle coincidence data from cold target recoil ion momentum spectroscopy

We present SCULPT (Supervised Clustering and Uncovering Latent Patterns with Training), a comprehensive software platform for analyzing tabulated high-dimensional multi-particle coincidence data from Cold Target Recoil Ion Momentum Spectroscopy (COLTRIMS) experiments. The software addresses critical challenges in modern momentum spectroscopy by integrating advanced machine learning techniques with physics-informed analysis in an interactive web-based environment. SCULPT implements uniform manifold approximation and projection for non-linear dimensionality reduction to reveal correlations in high-dimensional data. We also discuss potential extensions to deep autoencoders for feature learning and genetic programming for automated discovery of physically meaningful observables. A novel adaptive confidence scoring system provides quantitative reliability assessments by evaluating user-selected clustering quality metrics with predefined weights that reflect each metric’s robustness. The platform features configurable molecular profiles for different experimental systems, interactive visualization with selection tools, and comprehensive data filtering capabilities. Utilizing a subset of SCULPT’s capabilities, we analyze photo-double-ionization data measured using the COLTRIMS method for three-body dissociation of the D 2 O molecule, revealing distinct fragmentation channels and their correlations with physics parameters. The software’s modular architecture and web-based implementation make it accessible to the broader atomic and molecular physics community, significantly reducing the time required for complex multi-dimensional analyses. This opens the door to finding and isolating rare events exhibiting non-linear correlations on the fly during experimental measurements, which can help steer exploration and improve the efficiency of experiments.

Artificial neural networks

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will produce unprecedented volumes of heterogeneous astronomical data (images, catalogs, and alerts) that challenge traditional analysis pipelines. The LSST Dark Energy Science Collaboration (DESC) aims to derive robust constraints on dark energy and dark matter from these data, requiring methods that are statistically powerful, scalable, and operationally reliable. Artificial intelligence and machine learning (AI/ML) are already embedded across DESC science workflows, from photometric redshifts and transient classification to weak lensing inference and cosmological simulations. Yet their utility for precision cosmology hinges on trustworthy uncertainty quantification, robustness to covariate shift and model misspecification, and reproducible integration within scientific pipelines. This white paper surveys the current landscape of AI/ML across DESC's primary cosmological probes and cross-cutting analyses, revealing that the same core methodologies and fundamental challenges recur across disparate science cases. Since progress on these cross-cutting challenges would benefit multiple probes simultaneously, we identify key methodological research priorities, including Bayesian inference at scale, physics-informed methods, validation frameworks, and active learning for discovery. With an eye on emerging techniques, we also explore the potential of the latest foundation model methodologies and LLM-driven agentic AI systems to reshape DESC workflows, provided their deployment is coupled with rigorous evaluation and governance. Finally, we discuss critical software, computing, data infrastructure, and human capital requirements for the successful deployment of these new methodologies, and consider associated risks and opportunities for broader coordination with external actors.

Aubourg, Eric [APC, Paris] (ORCID:000000025592023X

Bayesian Approach to the Joint Inversion of Gravity and Magnetic Data, with Application to the Ismenius Area of Mars

This viewgraph presentation reviews a Bayesian approach to the inversion of gravity and magnetic data with specific application to the Ismenius Area of Mars. Many inverse problems encountered in geophysics and planetary science are well known to be non-unique (i.e. inversion of gravity the density structure of a body). In hopes of reducing the non-uniqueness of solutions, there has been interest in the joint analysis of data. An example is the joint inversion of gravity and magnetic data, with the assumption that the same physical anomalies generate both the observed magnetic and gravitational anomalies. In this talk, we formulate the joint analysis of different types of data in a Bayesian framework and apply the formalism to the inference of the density and remanent magnetization structure for a local region in the Ismenius area of Mars. The Bayesian approach allows prior information or constraints in the solutions to be incorporated in the inversion, with the "best" solutions those whose forward predictions most closely match the data while remaining consistent with assumed constraints. The application of this framework to the inversion of gravity and magnetic data on Mars reveals two typical challenges - the forward predictions of the data have a linear dependence on some of the quantities of interest, and non-linear dependence on others (termed the "linear" and "non-linear" variables, respectively). For observations with Gaussian noise, a Bayesian approach to inversion for "linear" variables reduces to a linear filtering problem, with an explicitly computable "error" matrix. However, for models whose forward predictions have non-linear dependencies, inference is no longer given by such a simple linear problem, and moreover, the uncertainty in the solution is no longer completely specified by a computable "error matrix". It is therefore important to develop methods for sampling from the full Bayesian posterior to provide a complete and statistically consistent picture of model uncertainty, and what has been learned from observations. We will discuss advanced numerical techniques, including Monte Carlo Markov

data analysis

Evaluation of Classifier Complexity for Delay Tolerant Network Routing

The growing popularity of small cost effective satellites (SmallSats, CubeSats, etc.) creates the potential for a variety of new science applications involving multiple nodes functioning together or independently to achieve a task, such as swarms and constellations. As this technology develops and is deployed for missions in Low Earth Orbit and beyond, the use of delay tolerant networking (DTN) techniques may improve communication capabilities within the network. In this paper, a network hierarchy is developed from heterogeneous networks of SmallSats, surface vehicles, relay satellites and ground stations which form an integrated network. There is a tradeoff between complexity, flexibility, and scalability of user defined schedules versus autonomous routing as the number of nodes in the network increases. To address these issues, this work proposes a machine learning classifier based on DTN routing metrics. A framework is developed which will allow for the use of several categories of machine learning algorithms (decision tree, random forest and deep learning) to be applied to a dataset of historical network statistics, which allows for the evaluation of algorithm complexity versus performance to be explored. We develop the emulation of a hierarchical network, consisting of tens of nodes which form a cognitive network architecture. CORE (Common Open Research Emulator) is used to emulate the network using bundle protocol and DTN IP neighbor discovery.

Dudukovich, Rachel

NASA GeneLab Multi-study Visualization Portal

NASA GeneLab has helped advance the field of Space Biology by providing a public repository where researchers can store, share, analyze and visualize the results of space flight related omics experiments. The GeneLab data visualization portal allows any user, regardless of bioinformatics knowledge or access to computational resources, to interact with the experimental data, draw their own conclusions, and gain insights about the effects of space on living systems. These tools help democratize scientific research and foster the NASA Open Science initiative. The new multi-study feature of the GeneLab visualization platform allows users to mine study metadata from RNA sequencing (RNA-seq) experiments to identify samples of interest by filtering datasets based on organism, tissue, assay technology type, and/or factor. Once samples are selected from multiple datasets, users can combine and normalize the sample data, then utilize the visualization displays, including Principal Component Analysis (PCA) plots, to assess sample distributions. Finally, users can perform differential gene expression analysis on the combined data and visualize the results through PCA plots, Volcano plots, Pair plots, Heatmap, Ideogram and Gene Set Enrichment Analysis. All user-generated results and visualizations will be available for download. Here, we present a biological study using samples from multiple GeneLab RNA-seq datasets and analyzed using the multi-study visualization platform to demonstrate inter- and intra-study variability, as well as commonly differentially expressed genes between spaceflight and ground control conditions across datasets. This new feature opens a wide range of possibilities and opportunities for further development including combining other assay technology types and integration with batch effect correction techniques and machine learning applications. Overall, this tool allows users to increase the statistical power of individual experiments, validate hypothesis, identify patterns, and opens the door to new and exciting research.

space biology

Contrasting Time-Frequency Representations for Unknown Waveform Detection

Identifying unseen electromagnetic waveforms is critical for many applications, like interference management, electronic warfare and spectrum management. Traditionally this is done using statistical methods for anomaly detection, which has evolved to deep learning models for identifying the unseen data, formally termed as open set recognition. Some prior methods use a generative model to emulate open set data, which face challenges in generating synthetic samples for open set while simultaneously selecting an optimal discriminator for accurate classification. To alleviate this issue, we propose a discriminative model that effectively combines time and frequency domain features of communication signals for accurate predictions. We further introduce a cosine similarity loss that makes the domain specific features unique to enhance the prediction rate. Additionally, our model avoids generic feature vectors by extracting class-specific features during training, resulting in improved class representation. The experiment results show that this combined feature approach with cosine loss outperforms single-domain models and improves accuracy by 10% over models without cosine loss.

99 - GENERAL AND MISCELLANEOUS

Improving statistical precision in Monte Carlo samples with negative weights via reweighting and uncertainty quantification

High statistical precision is critical for Monte Carlo (MC) samples in high energy physics and is degraded by negatively weighted events. This paper investigates a procedure to learn the relationship between the negative and positive weight distributions of any sample, allowing the reduction of statistical uncertainty by reweighting kinematically equivalent events with the same sign. A robust uncertainty quantification method is required for the practical application of such method. Two methods for the estimation of the reweighting uncertainty are developed: one at the event and another one at the final observable level. The latter method is strongly favored. The gains in statistical precision are then quantified. The method is demonstrated on Sherpa vector boson plus jets samples when using all generated events and when restricted to the signal region of a mock analysis. It is demonstrated to significantly reduce stochastic behavior in sparse MC samples while decreasing the overall uncertainty with a sufficiently well-known reweighting function.

Monte Carlo methods

Variance Preserving Spectral Subsampling

Generating statistically faithful short-duration gamma-ray spectra from a single long measurement is essential in nuclear safeguards, supporting tasks such as algorithm development and machine-learning applications, especially when list-mode data are unavailable. Existing subsampling methods often distort the statistical characteristics of genuine short-duration measurements, leading to biased or unreliable analytical outcomes and thereby undermining downstream tasks. In this work, we compare five subsampling approaches using a benchmark set of 156 genuine replicate spectra collected with a high-purity germanium detector. We evaluate each method with respect to run-to-run variance, channel-to-channel variance, and preservation of total counts (losslessness). Across a wide range of subsampling ratios, only binomial subsampling without replacement consistently reproduces the statistical properties of genuine short-duration spectra, maintaining proper dispersion even in sparse spectral regions and perfectly preserving total counts. These results provide a mathematically principled and practically validated framework for generating synthetically shortened spectra when true short-duration measurements are unavailable.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P

EFIT‐AI: Machine Learning and Artificial Intelligence Assisted Equilibrium Reconstruction for Tokamak Experiments and Burning Plasmas (Final Report)

The EFIT-AI project is creating a modern advanced equilibrium reconstruction code suitable for tokamak experiments of burning plasmas. EFIT [1,2] was the first and is the most extensively used equilibrium reconstruction code in the world. This project builds on the production-level experience and adds key elements as follows. 1. A Model Order Reduction (MOR) version of the two-dimensional (2D) Grad-Shafranov equation solver (EFIT-MORNN) using physics-informed neural networks. 2. Improved optimization and data analysis capabilities using a Bayesian framework enhanced with machine learning. 3. A MOR version of the three-dimensional (3D) perturbed equilibrium reconstruction tool.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Anomalous electroweak physics unraveled via evidential deep learning

The ever-growing ecosystem of beyond standard model (BSM) calculations and parametrizations has motivated the development of systematic methods for making quantitative cross-comparisons over the wide range of possible models, especially with controllable uncertainties. In this setting, the language of uncertainty quantification (UQ) furnishes useful metrics for assessing statistical overlaps and discrepancies among BSM and related models. In this study, we leverage recent machine learning (ML) developments in evidential deep learning (EDL) for UQ to separate data (aleatoric) and knowledge (epistemic) uncertainties in a model-discrimination setting. We construct several potentially BSM-motivated scenarios for the anomalous electroweak interaction (AEWI) of neutrinos with nucleons in deep inelastic scattering ( v DIS). These scenarios are then quantitatively mapped, as a demonstration, alongside Monte Carlo replicas of the CT18 PDFs used to calculate the $\varDelta \chi ^{2}$ statistic for a typical multi-GeV v DIS experiment, CDHSW. Our framework effectively highlights areas of model agreement and provides a classification of out-of-distribution (OOD) samples. By offering the opportunity to quantitatively understand model overlaps, the approach presented in this work can help facilitate efficient BSM model exploration and exclusion for future New Physics searches.

AI

Snow Distribution Patterns Revisited: A Physics-Based and Machine Learning Hybrid Approach to Snow Distribution Mapping in the Sub-Arctic

Snowpack distribution in Arctic and alpine landscapes often occurs in repeating, year-to-year patterns due to local topographic, weather, and vegetation characteristics. Previous studies have suggested that with years of observational data, these snow distribution patterns can be statistically integrated into a snow process modeling workflow. Recent advances in snow hydrology and machine learning (ML) have increased our ability to predict snowpack distribution using in-situ observations, remote sensing data sets, and simple landscape characteristics that can be easily obtained for most environments. Here, we propose a hybrid approach to couple a ML snow distribution pattern (MLSDP) map with a physics-based, snow process model. We trained a random forest ML algorithm on tens of thousands of snow survey observations from a subarctic study area on the Seward Peninsula, Alaska, collected during peak snow water equivalent (SWE). We validated hybrid model outputs using in-situ snow depth and SWE observations, as well as a light detection and ranging data set and a distributed temperature profiling sensor data set. When the hybrid results were compared with the physics-based method, the hybrid method more accurately depicted the spatial patterns of the snowpack, areas of drifting snow, and years when no in-situ observations were used in the random forest ML training data set. The hybrid method also showed improvements in root mean squared error at 61% of locations where time-series estimations of snow depth were observed. These results can be applied to any physics-based model to improve the snow distribution patterning to reflect observed conditions in high latitude and high elevation cold region environments.

54 ENVIRONMENTAL SCIENCES

Full event particle-level unfolding with variable-length latent variational diffusion

The measurements performed by particle physics experiments must account for the imperfect response of the detectors used to observe the interactions. One approach, unfolding, statistically adjusts the experimental data for detector effects. Recently, generative machine learning models have shown promise for performing unbinned unfolding in a high number of dimensions. However, all current generative approaches are limited to unfolding a fixed set of observables, making them unable to perform full-event unfolding in the variable dimensional environment of collider data. A novel modification to the variational latent diffusion model (VLD) approach to generative unfolding is presented, which allows for unfolding of high- and variable-dimensional feature spaces. The performance of this method is evaluated in the context of semi-leptonic t\bar{t} t t ‾ production at the Large Hadron Collider.

Shmakov, Alexander

NASA Tech Briefs, June 2014

Topics include: Real-Time Minimization of Tracking Error for Aircraft Systems; Detecting an Extreme Minority Class in Hyperspectral Data Using Machine Learning; KSC Spaceport Weather Data Archive; Visualizing Acquisition, Processing, and Network Statistics Through Database Queries; Simulating Data Flow via Multiple Secure Connections; Systems and Services for Near-Real-Time Web Access to NPP Data; CCSDS Telemetry Decoder VHDL Core; Thermal Response of a High-Power Switch to Short Pulses; Solar Panel and System Design to Reduce Heating and Optimize Corridors for Lower-Risk Planetary Aerobraking; Low-Cost, Very Large Diamond-Turned Metal Mirror; Very-High-Load-Capacity Air Bearing Spindle for Large Diamond Turning Machines; Elevated-Temperature, Highly Emissive Coating for Energy Dissipation of Large Surfaces; Catalyst for Treatment and Control of Post-Combustion Emissions; Thermally Activated Crack Healing Mechanism for Metallic Materials; Subsurface Imaging of Nanocomposites; Self-Healing Glass Sealants for Solid Oxide Fuel Cells and Electrolyzer Cells; Micromachined Thermopile Arrays with Novel Thermo - electric Materials; Low-Cost, High-Performance MMOD Shielding; Head-Mounted Display Latency Measurement Rig; Workspace-Safe Operation of a Force- or Impedance-Controlled Robot; Cryogenic Mixing Pump with No Moving Parts; Seal Design Feature for Redundancy Verification; Dexterous Humanoid Robot; Tethered Vehicle Control and Tracking System; Lunar Organic Waste Reformer; Digital Laser Frequency Stabilization via Cavity Locking Employing Low-Frequency Direct Modulation; Deep UV Discharge Lamps in Capillary Quartz Tubes with Light Output Coupled to an Optical Fiber; Speech Acquisition and Automatic Speech Recognition for Integrated Spacesuit Audio Systems, Version II; Advanced Sensor Technology for Algal Biotechnology; High-Speed Spectral Mapper; "Ascent - Commemorating Shuttle" - A NASA Film and Multimedia Project DVD; High-Pressure, Reduced-Kinetics Mechanism for N-Hexadecane Oxidation; Method of Error Floor Mitigation in Low-Density Parity-Check Codes; X-Ray Flaw Size Parameter for POD Studies; Large Eddy Simulation Composition Equations for Two-Phase Fully Multicomponent Turbulent Flows; Scheduling Targeted and Mapping Observations with State, Resource, and Timing Constraints;

Source record

At Risk Population Estimates for Belarus, Poland and Slovakia with Machine Learning

High-resolution gridded population modeling is crucial for various applications, including disaster response planning, infectious disease spread modeling, climate change impact estimation, policy development, and more. Multiple gridded population datasets have been developed, each tailored to meet specific objectives. Among them, LandScan Global dataset is designed to represent ambient and unwarned population distributions. However, this dataset relies on a statistical approach that requires manual adjustments, making it time consuming and labour intensive. Existing machine learning (ML) methods often train and test at different spatial resolutions, potentially leading to inflated results, and they rely on Census population totals for disaggregation. To address these limitations, in this study we developed population estimates using ML models trained and tested at a consistent 30 arc-second resolution (≈1 square kilometer), specifically using Random Forest (RF) and XGBoost. These models were trained on 2020 datum to predict for 2021 for three countries: Belarus, Poland, and Slovakia. Our findings show that both RF (MAE varies from 5.75 to 13.25) and XGBoost (MAE varies from 8.15 to 23.44) model performance is close to LandScan Global estimates. Furthermore, neither of the models performed the best across all grid cells: the RF model was more effective in areas with lower populations, while XGBoost excelled in more densely populated regions. The proposed approach can be used for countries where the Census data is not available.

Lebakula, Viswadeep [ORNL] (ORCID:0000000152935914

Ellicott City Disasters II: Enhancing a Statistical Flood Risk Model to Continue Improving Early Warning Systems and Public Safety in Ellicott City, Maryland

As flooding events in the United States grow in frequency and intensity, the use of technological advancements and applied science are increasingly necessary for effective flood monitoring and warning systems. The NASA DEVELOP Ellicott City Disasters II project investigated the use of machine learning for applications in flood risk detection to support the improvement of early warning systems. To strengthen the efforts of the Howard County Office of Emergency Management (OEM) in building a more robust flood monitoring system, the project improved the original statistical flood risk model, FLuME (Flood Learning Model Environment), programmed by the first DEVELOP term. The enhancements incorporated an additional six years of precipitation and soil moisture data from the North American Land Data Assimilation System (NLDAS), modeled using Aqua Advanced Microwave Scanning Radiometer for EOS and Tropical Rainfall Measuring Mission (TRMM) Microwave Imager. These Earth observations were supplemented by stream gauge data from the OEM and the US Geological Survey. The resultant flood risk model FLASH (Flood Learning Environment and Severity Assessment Hub) was trained to evaluate input variables and predict stage height in Ellicott City in real time. The addition of an advanced deep learning framework known as long short-term memory improved the model’s ability to capture relationships between variables. To assess the effectiveness of the new model, FLASH produced a model efficiency metric of 0.99, a significant improvement over the 0.85 value produced by the previous model. The project assisted the OEM in pursuing the integration of open data and NASA Earth observations into a threat matrix capable of informing near real-time decision making.

Disasters

Ellicott City Disasters II: Enhancing a Statistical Flood Risk Model to Continue Improving Early Warning Systems and Public Safety in Ellicott City, Maryland

As flooding events in the United States grow in frequency and intensity, the use of technological advancements and applied science are increasingly necessary for effective flood monitoring and warning systems. The NASA DEVELOP Ellicott City Disasters II project investigated the use of machine learning for applications in flood risk detection to support the improvement of early warning systems. To strengthen the efforts of the Howard County Office of Emergency Management (OEM) in building a more robust flood monitoring system, the project improved the original statistical flood risk model, FLuME (Flood Learning Model Environment), programmed by the first DEVELOP term The enhancements incorporated an additional six years of precipitation and soil moisture data from the North American Land Data Assimilation System (NLDAS), modeled using Aqua Advanced Microwave Scanning Radiometer for EOS and Tropical Rainfall Measuring Mission TRMM Microwave Imager. These Earth observations were supplemented by stream gauge data from the OEM and the US Geological Survey. The resultant flood risk model FLASH (Flood Learning Environment and Severity Assessment Hub) was trained to evaluate input variables and predict stage height in Ellicott City in real time. The addition of an advanced deep learning framework known as long short-term memory improved the model’s ability to capture relationships between variables. To assess the effectiveness of the new model, FLASH produced a model efficiency metric of 0.99, a significant improvement over the 0.85 value produced by the previous model. The project assisted the OEM in pursuing the integration of open data and NASA Earth observations into a threat matrix capable of informing near real-time decision making.

Disasters