Search NASASearch

SEARCH · Search NASA

Results for “sparse data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

A fast and accurate domain decomposition nonlinear manifold reduced order model

Here, this paper integrates nonlinear-manifold reduced order models (NM-ROMs) with domain decomposition (DD). NM ROMs approximate the full order model (FOM) state in a nonlinear-manifold by training a shallow, sparse autoencoder using FOM snapshot data. These NM-ROMs can be advantageous over linear-subspace ROMs (LS-ROMs) for problems with slowly decaying Kolmogorov n-width. However, the number of NM-ROM parameters that need to be trained scales with the size of the FOM. Moreover, for “extreme-scale” problems, the storage of high-dimensional FOM snapshots alone can make ROM training expensive. To alleviate the training cost, this paper applies DD to the FOM, computes NM-ROMs on each subdomain, and couples them to obtain a global NM-ROM. This approach has several advantages: Subdomain NM-ROMs can be trained in parallel, involve fewer parameters to be trained than global NM-ROMs, require smaller subdomain FOM dimensional training data, and can be tailored to subdomain specific features of the FOM. The shallow, sparse architecture of the autoencoder used in each subdomain NM-ROM allows application of hyper-reduction (HR), reducing the complexity caused by nonlinearity and yielding computational speedup of the NM-ROM. This paper provides the first application of NM-ROM (with HR) to a DD problem. In particular, this paper details an algebraic DD reformulation of the FOM, training a NM-ROM with HR for each sub domain, and a sequential quadratic programming (SQP) solver to evaluate the coupled global NM-ROM. Theoretical convergence results for the SQP method and a priori and a posteriori error estimates for the DD NM-ROM with HR are provided. The proposed DD NM-ROM with HR approach is numerically compared to a DD LS-ROM with HR on the 2D steady-state Burgers’ equation, showing an order of magnitude improvement in accuracy of the proposed DD NM-ROM over the DD LS-ROM.

97 MATHEMATICS AND COMPUTING

Bayesian learning with Gaussian processes for low-dimensional representations of time-dependent nonlinear systems

This work presents a data-driven method for learning low-dimensional time-dependent physics-based surrogate models whose predictions are endowed with uncertainty estimates. We use the operator inference approach to model reduction that poses the problem of learning low-dimensional model terms as a regression of state space data and corresponding time derivatives by minimizing the residual of reduced system equations. Standard operator inference models perform well with accurate training data that are dense in time, but producing stable and accurate models when the state data are noisy and/or sparse in time remains a challenge. Another challenge is the lack of uncertainty estimation for the predictions from the operator inference models. Our approach addresses these challenges by incorporating Gaussian process surrogates into the operator inference framework to (1) probabilistically describe uncertainties in the state predictions and (2) procure analytical time derivative estimates with quantified uncertainties. The formulation leads to a generalized least-squares regression and, ultimately, reduced-order models that are described probabilistically with a closed-form expression for the posterior distribution of the operators. The resulting probabilistic surrogate model propagates uncertainties from the observed state data to reduced-order predictions. Furthermore, we demonstrate the method is effective for constructing low-dimensional models of two nonlinear partial differential equations representing a compressible flow and a nonlinear diffusion–reaction process, as well as for estimating the parameters of a low-dimensional system of nonlinear ordinary differential equations representing compartmental models in epidemiology.

Data-driven model reduction

Predicting High‐Resolution Spatial and Spectral Features in Mass Spectrometry Imaging with Machine Learning and Multimodal Data Fusion

Recent advancements in molecular Mass Spectrometry Imaging have sparked interest in integrating high spatial resolution methods with molecular mass-spectrometry-based chemical imaging. Fusion-based algorithms have proven effective in generating high spatial-resolution molecular mass spectra. However, a significant challenge stems from the differing physical mechanisms underlying image generation and data upsampling techniques, potentially leading to discrepancies in integrated information channels. Integrating physical constraints into data processing workflows is essential to tackle this issue. In this study, we propose an innovative approach that merges data from Fourier transform ion cyclotron resonance (FTICR), time-of-flight matrix-assisted laser desorption/ionization, and time-of-flight secondary ion mass spectrometry imaging techniques. By leveraging FT-ICR's unparalleled spectral resolution and ToF-SIMS's exceptional spatial resolution, we achieve submicron spatial resolution, enabling the observation of intact molecular species with remarkable spectral precision. Canonical correlation analysis is employed to incorporate physical constraints. Through sophisticated image processing and machine learning techniques, the results of this fusion hold significant promise for advancing our comprehension of complex systems and unveiling concealed molecular intricacies.

canonical correlation analysis

Increasing phosphorus loss despite widespread concentration decline in US rivers

The loss of phosphorous (P) from the land to aquatic systems has polluted waters and threatened food production worldwide. Systematic trend analysis of P, a nonrenewable resource, has been challenging, primarily due to sparse and inconsistent historical data. Here, we leveraged intensive hydrometeorological data and the recent renaissance of deep learning approaches to fill data gaps and reconstruct temporal trends. We trained a multitask long short-term memory model for total P (TP) using data from 430 rivers across the contiguous United States (CONUS). Trend analysis of reconstructed daily records (1980–2019) shows widespread decline in concentrations, with declining, increasing, and insignificantly changing trends in 60%, 28%, and 12% of the rivers, respectively. Concentrations in urban rivers have declined the most despite rising urban population in the past decades; concentrations in agricultural rivers however have mostly increased, suggesting not-as-effective controls of nonpoint sources in agriculture lands compared to point sources in cities. TP loss, calculated as fluxes by multiplying concentration and discharge, however exhibited an overall increasing rate of 6.5% per decade at the CONUS scale over the past 40 y, largely due to increasing river discharge. Results highlight the challenge of reducing TP loss that is complicated by changing river discharge in a warming climate.

Science & Technology - Other Topics

Influence of initial conditions on data-driven model identification and information entropy for ideal mhd problems

Data-driven methods of model identification are able to discern governing dynamics of a system from data. Such methods are well suited to help us learn about systems with unpredictable evolution or systems with ambiguous governing dynamics given our current understanding. Many plasma problems of interest fall into these categories as there are a wide range of models that exist, however each model is only useful in a certain regime and often limited by computational complexity. To ensure data-driven methods align with theory, they must be consistent and predictable when acting on data whose governing dynamics are known. Weak Sparse Identification of Nonlinear Dynamics (WSINDy) is a recently developed data-driven method that has shown promise in learning governing dynamics from data with high noise levels [1]. This work examines how WSINDy acts on ideal MHD test problems as the initial conditions are varied and specifies limiting requirements for successful equation identification. Furthermore, it is hard to recover the governing dynamics from data that emphasize a single dominant behavior. In these low information cases, Shannon information entropy is able to pick up on the redundancies in the data that affect recoverability.

97 MATHEMATICS AND COMPUTING

Chemical signature characterization with hyperspectral imagery: novel deep learning model architectures and physically-motivated data augmentation techniques

The high spectral resolution afforded by Hyperspectral Imaging (HSI) sensors is poised to bring unprecedented advancements to signature characterization applications. Thus far, much of the research in the machine learning field devoted to HSI applications has focused on a few specific tasks like land-use land-cover classification. In land classification tasks, spatial information is very important, and model architectures are often designed to leverage spatial contexts. However, it is unclear how well these spatially-tuned models will translate to tasks where spectral information is critical, like the detection and characterization of chemicals. In this work, we compare spectral models (inputs are 1D spectra) and spatial-spectral models (inputs are 3D cubes) in the context of predicting chemical concentration maps. We find that spatial-spectral models perform the best, though we find a wide range in performance across the different architectures tested. Additionally, we find that model performance is impacted by the availability of training data, particularly in scenarios where the training data doesn't fully capture the true variance of real-world conditions. We find that data augmentation can help mitigate sparse coverage of observed parameter space (e.g., seasonal or geographic variability in ground cover), and present augmentation strategies that are tailored to hyperspectral data.

• Artificial intelligence (AI) / machine learning

$\overline{TKE}$ Parameterization and $\bar{v}$ Uncertainty Analysis for CGMF

Previous work was performed on tuning CGMF parameters for 235 U, 238 U, and Plutonium isotopes. Now work is being done to tune minor uranium isotopes. However, uranium isotopes like 232 U and 236 U have almost no experimental data. We are applying cross-isotope models to extrapolate and tune CGMF on isotopes that lack experimental data. There exist several internal CGMF physics quantities that affect the output of CGMF—multi-chance fission probability, excitation energy sharing, spin-cutoff factor, spin scaling, and fragment total kinetic energy to name a few. The mean fragment total kinetic energy, $\overline{TKE}$, is particularly interesting because of its strong anti-correlation with $\bar{v}$. We are most interested in the mean fragment total kinetic energy before neutron emissions. $\overline{TKE}$ is assumed to be pre-neutron emission unless otherwise stated. Currently in CGMF, the $\overline{TKE}$ model for 233,234,235,238 U are tuned independently to reproduce ν for the associated isotopes. In this report, we will tune a cross-isotope $\overline{TKE}$ model to experimental $\overline{TKE}$ data for 232,233,234,235,236,238 U. Because of the unreliable and sparse nature of $\overline{TKE}$ experimental data, future work will use more reliable experimental $\bar{v}$ data to infer the $\overline{TKE}$ model (and likely other internal CGMF parameters) for uranium isotopes. Such work has been performed previously using a sensitivity analysis and Kalman filter methods.

07 ISOTOPE AND RADIATION SOURCES

A First Search for Argon-Bound Neutron-Antineutron Oscillation using the MicroBooNE LArTPC

The use of Liquid Argon Time Projection Chambers (LArTPCs) as a detector technology in neutrino experiments has grown considerably over the past two decades. The excellent spatial and calorimetric resolution offered by LArTPCs enable precise neutrino oscillation measurements as well as beyond-Standard Model searches. One such search, which is the focus of this note, is the search for nucleus-bound neutron-antineutron (n ₋ n̄) oscillation. The n ₋ n̄ oscillation process is a baryon number violating process that produces a unique, star-like topology as a result of multiple final state pions. This unique signature is a key feature that may be used to search for this signal process. This note describes a machine learning-based analysis of MicroBooNE data, making use of a sparse convolutional neural network to search for n ₋ n̄ oscillation-like signals in MicroBooNE. While the future DUNE LArTPC can search for this signature with high sensitivity, existing MicroBooNE data can be used to demonstrate and validate methodologies that can be used as part of the DUNE search. This document presents the first-ever search for n ₋ n̄ oscillation in a LArTPC, using MicroBooNE off-beam data (data collected when the neutrino beam was not running).

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

A first search for argon-bound neutron-antineutron oscillation using the MicroBooNE LArTPC

The use of Liquid Argon Time Projection Chambers (LArTPCs) as a detector technology in neutrino experiments has grown considerably over the past two decades. The excellent spatial and calorimetric resolution offered by LArTPCs enable precise neutrino oscillation measurements as well as beyond-Standard Model searches. One such search, which is the focus of this note, is the search for nucleus-bound neutron-antineutron (n – n̄) oscillation. The n – n̄ oscillation process is a baryon number violating process that produces a unique, star-like topology as a result of multiple final state pions. This unique signature is a key feature that may be used to search for this signal process. This note describes a machine learning-based analysis of MicroBooNE data, making use of a sparse convolutional neural network to search for n – n̄ oscillation-like signals in MicroBooNE. While the future DUNE LArTPC can search for this signature with high sensitivity, existing MicroBooNE data can be used to demonstrate and validate methodologies that can be used as part of the DUNE search. This document presents the first-ever search for n – n̄ oscillation in a LArTPC, using MicroBooNE off-beam data (data collected when the neutrino beam was not running).

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Atomistic Simulations for Thermophysical Properties of Uranium-Containing Halide Molten Salts

Characterizing the thermophysical properties in both fuel and coolant salts are critical in modeling, developing, process optimizing and utilizing molten salt reactors (MSRs), as these properties directly relate to operation metrics and can inform on the selection of candidate salts. The demand for consistent, accurate and publicly available thermophysical property data has become more apparent in recent years as interests have increased from molten salt reactor developers. There are a number of challenges in experimentally measuring properties such as thermal conductivity, viscosity, density and heat capacity , which have led to sparse and often times conflicting data points or molten salts in general. Additionally, there are a number of hazards to consider when synthesizing, storing, using, treating and disposing of molten salts. With the advances in computational capabilities over the last 10 years, the use of atomistic simulations can be implemented to support these efforts. The primary objective of this work is characterize the thermophysical transport properties in a number of molten chloride salts, and in particular NaCl-UCl 3 using ab-initio molecular dynamic (AIMD) simulations. In this binary salt the UCl 3 acts as the primary fissile material and NaCl acts as a carrier salt due with its’ high solubility for actinides A number of studies on the thermophysical properties of NaCl-UCl 3 have been published but there is not a vast amount of viscosity data for this system. In 1975, Desyatnik, et al published a study reporting dynamic viscosities that were calculated from kinematic viscosity measurements, and using the coefficients provided the viscosity in a 70:30 NaCl:UCl 3 mixture is 2.29 cP and 2.88 for a 60:40 mixture. Termini et al. recently reported viscosities in the range of 2.75 – 3 cP for the 63:37 NaCl-UCl 3 mixture in the same temperature range using rolling ball viscosity measurements. Computational viscosity of a similar mixture (64:36) can be obtained from the work Andersson et al. using the reported diffusion coefficients, and the hydrodynamic radius from the pair-radial distribution functions (RDFs). Using Eq (1) (vida infra), the viscosity would be 2.50 cP at 1100K. This is not to say that these values are incorrect due to the varying reported values, but aims to highlight the necessity of this work. The data reported in this ongoing work are computations on a 64:36 mixture of NaCl-UCl 3 at 987K. This work is likely to be expanded into varying concentrations of this mixture along with the inclusion of other salt candidate mixtures.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA

Physics-Informed Active Learning With Simultaneous Weak-Form Latent Space Dynamics Identification

The parametric greedy latent space dynamics identification (gLaSDI) framework has demonstrated promising potential for accurate and efficient modeling of high-dimensional nonlinear physical systems. However, it remains challenging to handle noisy data. Here, to enhance robustness against noise, we incorporate the weak-form estimation of nonlinear dynamics (WENDy) into gLaSDI. In the proposed weak-form gLaSDI (WgLaSDI) framework, an autoencoder and WENDy are trained simultaneously to discover intrinsic nonlinear latent-space dynamics of high-dimensional data. Compared with the standard sparse identification of nonlinear dynamics (SINDy) employed in gLaSDI, WENDy enables variance reduction and robust latent space discovery, therefore leading to more accurate and efficient reduced-order modeling. Furthermore, the greedy physics-informed active learning in WgLaSDI enables adaptive sampling of optimal training data on the fly for enhanced modeling accuracy. The effectiveness of the proposed framework is demonstrated by modeling various nonlinear dynamical problems, including viscous and inviscid Burgers' equations, time-dependent radial advection, and the Vlasov equation for plasma physics. With data that contains 5%–10% Gaussian white noise, WgLaSDI outperforms gLaSDI by orders of magnitude, achieving 1%–7% relative errors. Compared with the high-fidelity models, WgLaSDI achieves 121 to 1779x speed-up.

97 MATHEMATICS AND COMPUTING

Multimodal super-resolution: discovering hidden physics and its application to fusion plasmas

Understanding complex physical systems often requires integrating data from multiple diagnostics, each with limited resolution or coverage. We present a machine learning framework that reconstructs synthetic high-temporal-resolution data for a target diagnostic using information from other diagnostics, without direct target measurements during the inference. This multimodal super-resolution technique improves diagnostic robustness and enables monitoring even in case of measurement failures or degradation. Applied to fusion plasmas, our method targets edge-localized modes (ELMs), which can damage plasma-facing materials. By reconstructing super-resolution Thomson Scattering data from complementary diagnostics, we uncover fine-scale plasma dynamics and validate the role of resonant magnetic perturbations (RMPs) in ELM suppression through magnetic island formation. The approach provides new observation supporting the plasma profile flattening due to these islands. Our results demonstrate the framework’s ability to generate high-fidelity synthetic diagnostics, offering a powerful tool for ELM control development in future reactors like ITER. The approach is broadly transferable to other domains facing sparse, incomplete, or degraded diagnostic data, opening new avenues for discovery.

Jalalvand, Azarakhsh [Princeton Univ., NJ (United

The Observed and Projected Changes of Global Monsoons: Current Status and Future Perspectives

The global monsoon system, encompassing the Asian-Australian, African, and American monsoons, sustains two-thirds of the world’s population by regulating water resources and agriculture. Monsoon anomalies pose severe risks, including floods and droughts. Recent research associated with the implementation of the Global Monsoons Model Intercomparison Project under the umbrella of CMIP6 has advanced our understanding of its historical variability and driving mechanisms. Observational data reveal a 20th-century shift: increased rainfall pre-1950s, followed by aridification and partial recovery post-1980s, driven by both internal variability (e.g., Atlantic Multidecadal Oscillation) and external forcings (greenhouse gases, aerosols), while ENSO drives interannual variability through ocean-atmosphere interactions. Future projections under greenhouse forcing suggest long-term monsoon intensification, though regional disparities and model uncertainties persist. Models indicate robust trends but struggle to quantify extremes, where thermodynamic effects (warming-induced moisture rise) uniformly boost heavy rainfall, while dynamical shifts (circulation changes) create spatial heterogeneity. Volcanic eruptions and proposed solar radiation modification (SRM) further complicate predictions: tropical eruptions suppress monsoons, whereas high-latitude events alter cross-equatorial flows, highlighting unresolved feedbacks. The emergent constraint approach is booming in terms of correcting future projections and reducing uncertainty with respect to the global monsoons. Critical challenges remain. Model biases and sparse 20th-century observational data hinder accurate attribution. The interplay between natural variability and anthropogenic forcings, along with nonlinear extreme precipitation risks under warming, demands deeper mechanistic insights. Additionally, SRM’s regional impacts and hemispheric monsoon interactions require systematic evaluation. Addressing these gaps necessitates enhanced observational networks, refined climate models, and interdisciplinary efforts to disentangle multiscale drivers, ultimately improving resilience strategies for monsoon-dependent regions.

climate extreme events

Data from a multi-year targeted proteomics study of a longitudinal birth cohort of type 1 diabetes

The deployment of liquid chromatography-mass spectrometry-based plasma proteomics experiments in a large cohort is sparse, leading to a lack of data available for benchmarking, method development or validation. Comprised of 6,426 plasma analyses, The Environmental Determinants of Diabetes in the Young (TEDDY) proteomics validation study constitutes one of the largest targeted proteomics experiments in the literature to date. The proteomics data from this study were generated over the course of 2.5 years from over 900 study subjects, each providing up to 29 longitudinal samples. The data also includes 916 quality control samples. The targeted mass spectrometry assay was comprised of 694 peptides mapping to 167 proteins and the panel was measured in each subject and QC sample. The targeted proteomic dataset presented here can be used as a resource for new computational method development, such as for batch correction, as well as for benchmarking and comparing the performance of different methods/tools.

60 APPLIED LIFE SCIENCES

Sedimentation and Nonlinear Trapping in Texas Reservoirs Identified Using Remote Sensing and Bathymetric Survey Records

Decreasing reservoir storage capacity due to sedimentation poses great challenges to aging U.S. reservoirs, as it reduces the efficacy and reliability of their socio‐economic services. However, systematic assessments of reservoir sedimentation rates and associated issues remain limited because of sparse and infrequent bathymetry survey data. In this study, we use remote sensing‐driven estimates of sediment concentrations to identify regions experiencing rapid reservoir capacity loss, as observed in repeated bathymetry surveys. Our analysis focuses on Texas, where one of the most reliable state‐level reservoir capacity loss data sets is available through a unique long‐term monitoring program by the Texas Water Development Board. We find that reservoirs with large storage capacities and high sedimentation rate are concentrated in Northeast Texas. We also show that reduced forest, increased barren land, and erosive soil properties are co‐varying with high reservoir sedimentation rates. Temporal changes in the longitudinal gradient of the remotely sensed sediment flux highlight the nonlinear nature of sediment trapping processes, and can be used to estimate the reservoir storage capacity loss over time. In combination with standard bathymetry surveys, our approach shows potential for more cost‐effective and frequent assessment of reservoir storage loss.

Lee, Jiyong [ORNL] (ORCID:0000000198957406)

Bayesian inference of anisotropic 2D small-angle scattering from sparse measurement

Here, we present a Bayesian inference framework for reconstructing anisotropic two-dimensional small-angle scattering (2D SAS) patterns from sparse, noisy, or partially missing data. The method combines a symmetry-aware angular basis with radial Gaussian process priors to enable accurate, training-free interpolation and denoising. Computational benchmarks demonstrate reliable recovery of both isotropic and high-order anisotropic features under severe data reduction. Experimental validations on stretched polymers, sheared wormlike micelles, and carbon fibers show improved fidelity and resolution compared to raw measurements, achieving comparable accuracy with up to 50-fold fewer detected neutrons. This approach enables quantitative structural analysis under low-flux, time-limited, or single-shot conditions, extending the applicability of 2D SAS techniques to compact neutron sources and mechanically driven soft matter systems undergoing transient structural changes.

Tung, Chi-Huan [Oak Ridge National Laboratory (ORN

Designing an Optimal Sensor Network via Minimizing Information Loss

Optimal experimental design is a classic topic in statistics, with many well-studied problems, applications, and solutions. The design problem we study is the placement of sensors to monitor spatiotemporal processes, explicitly accounting for the temporal dimension in our modeling and optimization. We observe that recent advancements in computational sciences often yield large datasets based on physics-based simulations, which are rarely leveraged in experimental design. We introduce a novel model-based sensor placement criterion, along with a highly-efficient optimization algorithm, which integrates physics-based simulations and Bayesian experimental design principles to identify sensor networks that “minimize information loss” from simulated data. Our technique relies on sparse variational inference and (separable) Gauss-Markov priors, and thus may adapt many techniques from Bayesian experimental design. We validate our method through a case study monitoring air temperature in Phoenix, Arizona, using state-of-the-art physics-based simulations. Our results show our framework to be superior to random or quasi-random sampling, particularly with a limited number of sensors. We conclude by discussing practical considerations and implications of our framework, including more complex modeling tools and real-world deployments.

54 ENVIRONMENTAL SCIENCES

Improving Reliability of Large Language Models for Nuclear Power Plant Diagnostics [Poster]

Large Language Models (LLMs) struggle out of the box when answering factually about detailed questions, especially in domains that are sparsely represented in their training data. This causes hallucinations and reduces reliability making it difficult for them to be used in practice. This work shows that using RAG techniques can improve factual accuracy and reliability, allowing for the application of LLMs in specialized areas, even when those areas that aren’t extensively covered in their initial training.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND