Search NASA⌕ Search

SEARCH · Search NASA

Results for “High dimensional data,”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Accurate data-driven surrogates of dynamical systems for forward propagation of uncertainty

Stochastic collocation (SC) is a well-known non-intrusive method of constructing surrogate models for uncertainty quantification. In dynamical systems, SC is especially suited for full-field uncertainty propagation that characterizes the distributions of the high-dimensional solution fields of a model with stochastic input parameters. However, due to the highly nonlinear nature of the parameter-to-solution map in even the simplest dynamical systems, the constructed SC surrogates are often inaccurate. Here, this work presents an alternative approach, where we apply the SC approximation over the dynamics of the model, rather than the solution. By combining the data-driven sparse identification of nonlinear dynamics framework with SC, we construct dynamics surrogates and integrate them through time to construct the surrogate solutions. We demonstrate that the SC-over-dynamics framework leads to smaller errors, both in terms of the approximated system trajectories as well as the model state distributions, when compared against full-field SC applied to the solutions directly. We present numerical evidence of this improvement using three test problems: a chaotic ordinary differential equation, and two partial differential equations from solid mechanics.

42 ENGINEERING↗

Dimensional Reduction for Sampled Priors and Application to Photometric Redshift Distributions

A typical Bayesian inference on the values of some parameters of interest q from some data D involves running a Markov Chain (MC) to sample from the posterior $p$($q$,$n$|$D$) $\propto$ $\mathcal{L}$($D$|$q$,$n$)$p$(q)$p$($n$), where n are some nuisance parameters with a separable prior. In some cases, the nuisance parameters are high-dimensional, and their prior p(n) is itself defined only by a set of samples that have been drawn from some other MC. The MC for the posterior will typically require evaluation of p(n) at arbitrary values of n, i.e., one needs to provide a density estimator over the full n space from the provided samples. But the high dimensionality of n hinders both the density estimation and the efficiency of the MC for the posterior. We describe a solution to this problem: a linear compression of the n space into a much lower-dimensional space u, which projects away directions in n space that cannot appreciably alter $\mathcal{L}$. The algorithm for doing so is a slight modification to principal components analysis, and is less restrictive on p(n) than other proposed solutions to this issue. We demonstrate this “mode projection” technique using the analysis of 2-point correlation functions of weak lensing fields and galaxy density in the Dark Energy Survey, where n is a binned representation of the redshift distribution n(z) of the galaxies.

79 ASTRONOMY AND ASTROPHYSICS↗

Deep learning-based predictive models for laser direct drive at the Omega Laser Facility

The rich and complex physics of inertial confinement fusion provides a unique and challenging space for high-fidelity first-principles modeling. Consequently, simulation codes that are used to design experiments are computationally expensive and lack the predictive capability required for extensive parameter exploration in search of a high-performing design for laser direct drive. In this article, we present two deep-learning-based predictive models intended to address these difficulties. The first model (TL DNN) acts as a fast emulator of simulations as well as experiments at the Omega Laser Facility. This model is trained on a simulation database and subsequently calibrated on experimental data using transfer learning. To facilitate the development of this model, an autoencoder is developed to reduce the dimensionality of the input space by compressing the laser pulse input. The model predicts key experimental scalar observables of Omega experiments with high accuracy and minimal computational cost. This deep neural net enables rapid exploration of a high-dimensional input parameter space for an optimal implosion design. The second model (DNN SM+) aims to extend the statistical modeling work of Lees et al. [Phys. Rev. Lett. 127, 105001 (2021)], by increasing the complexity of the model space and allowing for coupling between degradation terms. Since the model capacity of DNN SM+ is higher than the model of Lees et al., DNN SM+ can potentially provide an improvement in predictive capability, and we use this model to provide insight into complicated degradation dependencies.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Active Learning‐Driven Inkless Additive Nanomanufacturing for Printed Electronics

Inkless additive nanomanufacturing for printed electronics promises broad material and substrate versatility, yet the high-dimensional print parameter space makes tuning print parameters time-intensive. We present a Bayesian optimization study that constructs a digital twin from printed-silver data to benchmark surrogate models, acquisition functions, and batch sizes head-to-head to achieve user-specified target resistance. Tested surrogate models included Gaussian process, random forest, and Bayesian neural network surrogates with expected improvement and confidence bound acquisition functions. In total, we evaluate 48 unique model configurations alongside a random sampling baseline for comparison. For printed silver, the Bayesian neural network with a batch size of one achieved the lowest average cumulative regret, approximately four times more efficient on average than random sampling. To balance performance and substrate space, a random forest model with expected improvement and a batch size of four was chosen as the model for validation testing. Applying this chosen configuration to copper with an additional print parameter, the model achieved a resistance within 0.15 Ω of a 1 Ω target in fewer than 30 printed lines across five validation sets. Altogether, the workflow yields a tuned and validated model that efficiently guides experiments toward the target while simultaneously learning the parameter space.

Bevel, Colton [Auburn University, AL (United State↗

Modulated Thermomechanical Analysis of Compression-Molded High-Density Polyethylene

Thermomechanical analysis (TMA) experiments conducted on high-density polyethylene (HDPE) show both reversible and irreversible dimensional changes. To further explore these reversible and irreversible processes, modulated thermomechanical analysis (MTMA) was used. Before reliable data on compression-molded HDPE was collected, a parameter optimization was performed to obtain a suitable MTMA method. Once a suitable method was obtained, several MTMA experiments were conducted on compression-molded HDPE. This work highlights the steps taken during the MTMA parameter optimization and the results obtained from MTMA experiments conducted on pristine compression-molded HDPE samples.

36 MATERIALS SCIENCE↗

A Gigaparsec-scale Hydrodynamic Volume Reconstructed with Deep Learning

The next generation of spectroscopic surveys will map the large-scale structure of the Universe at high redshifts (2 ≤ z ≤ 5) using millions of quasar spectra, enabling major advances in constraining both the standard cosmological model and its extensions. Robust cosmological analyses of these data sets require numerical simulations that both cover gigaparsec volumes and resolve features on ∼10 kpc scales and smaller. However, running such large-volume, high-resolution hydrodynamic simulations is computationally prohibitive. We present a generative deep learning model that enhances a low-resolution, gigaparsec-scale (960 h −1 Mpc) hydrodynamic simulation using a smaller (80 h −1 Mpc) high-resolution input hydrodynamic simulation as training data. The resulting enhanced simulation reproduces the line-of-sight power spectrum to within ∼10% and the three-dimensional power spectrum at the ∼20% level at intermediate to small scales (k ≲ 2 h Mpc −1 ). Our method shows strong promise for producing realistic simulations for cosmological analyses with current surveys such as the Dark Energy Spectroscopic Instrument and upcoming next-generation experiments, but further improvements are needed to accurately recover the large-scale modes. We publicly release the enhanced hydrodynamic simulation, along with a halo catalog from a companion N-body dark matter simulation to support the calibration of data analysis pipelines for these large-scale surveys.

Convolutional neural networks↗

Reduced-Order Modeling of Multigroup Neutron Cross Sections for High-Temperature Gas-cooled Reactors

Abstract – Deterministic neutronics calculations rely on multigroup neutron cross section libraries, which consist of databases of tabulated values, used to calculate the neutron cross sections through multivariate linear interpolation. However, interpolation of the multidimensional cross section data becomes memory inefficient and time consuming as the number of tabulations increases, significantly slowing down the neutronics calculation, especially in the case of microscopic cross section libraries where every isotope (on the order of hundreds) has its own set of specific reactions and cross sections. In order to address this challenge, this work constructs efficient and robust reduced-order models (ROMs) of the multi-group cross sections to support the Griffin simulation of high-temperature gas-cooled reactors (HTGRs). The first part of the study investigates the linearity of the multigroup cross section data across isotopes, reaction types, and energy groups on pre-generated datasets for the purpose of dimensionality reduction. Secondly, a down-selection of ROM techniques is presented on representative classical machine learning (ML) techniques, including variants of linear regression, kernel-based methods, tree-based algorithms, and artificial neural networks. The selection criteria jointly consider the memory efficiency, predictive accuracy, prediction speed, and scalability in comparison to the multidimensional interpolation. Among all the ML techniques, deep neural networks (DNNs) have proven to be the best selection with sufficient accuracy, high robustness, good memory efficiency, great scalability, and superior flexibility. DNNs have been trained for all isotopes in this work and systematic Griffin testing is ongoing to ensure the feasibility of this ROM technique for predicting cross section and reducing memory requirements without a significant sacrifice in computational performance.

42 - ENGINEERING↗

Advanced Cross Section Library Generation using Reduced Order Models

Deterministic neutronics calculations rely on multigroup neutron cross section libraries, which consist of databases of tabulated values, used to calculate the neutron cross sections through multivariate linear interpolation. However, interpolation of the multidimensional cross section data becomes memory inefficient and time consuming as the number of tabulations increases, significantly slowing down the neutronics calculation, especially in the case of microscopic cross section libraries where every isotope (on the order of hundreds) has its own set of specific reactions and cross sections. In order to address this challenge, this work constructs efficient and robust reduced-order models (ROMs) of the multi-group cross sections to support the Griffin simulation of high-temperature gas-cooled reactors (HTGRs). The first part of the study investigates the linearity of the multigroup cross section data across isotopes, reaction types, and energy groups on pre-generated datasets for the purpose of dimensionality reduction. Secondly, a down-selection of ROM techniques is presented on representative classical machine learning (ML) techniques, including variants of linear regression, kernel-based methods, tree-based algorithms, and artificial neural networks. The selection criteria jointly consider the memory efficiency, predictive accuracy, prediction speed, and scalability in comparison to the multidimensional interpolation. Among all the ML techniques, deep neural networks (DNNs) have proven to be the best selection with sufficient accuracy, high robustness, good memory efficiency, great scalability, and superior flexibility. DNNs have been trained for all isotopes in this work and systematic Griffin testing is ongoing to ensure the feasibility of this ROM technique for predicting cross section and reducing memory requirements without a significant sacrifice in computational performance.

42 - ENGINEERING↗

Sieving Hydrogen Isotopes via Machine Learning Assisted Chemical Vapor Deposition (CVD) of High‐Quality Monolayer Hexagonal Boron Nitride (h‐BN) on Iron Foils

Atomically thin two-dimensional (2D) ceramics, such as monolayer hexagonal boron nitride (h-BN), present potential for disruptive advances in separations. However, sub-atomic scale separation of hydrogen isotopes (H + /D + ) require near pristine 2D material membranes, and scalable synthesis of such high-quality h-BN comparable to mechanically exfoliated crystals remains a significant challenge. Here, we report a scalable Fe-catalyzed chemical vapor deposition (CVD) process for bottom-up synthesis of large-area, high-quality monolayer h-BN films, overcoming key limitations of conventional ammonia-based routes. By leveraging mechanistic insights and higher CVD temperatures, we suppress multilayer formation and achieve uniform monolayer h-BN coverage on commercially available Fe foils. Machine learning enables systematic exploration of the complex, multi-dimensional CVD parameter space (growth time, temperature, precursor temperature, multilayer faction, coverage), providing data-driven approaches to visualize and identify process regimes facilitating predominantly monolayer h-BN growth with minimal secondary nuclei/ad-layers. The optimized Fe-catalyzed CVD h-BN membranes show high-quality as observed by proton/deuteron (H + /D + ) selectivity ≈8.45, approaching the highest quality benchmark of mechanically exfoliated h-BN (H + /D + selectivity ≈10) as well as significantly outperforming Cu-catalyzed CVD h-BN membranes (H + /D + selectivity ≈3.62, control selectivity ≈1.7). Our work provides a scalable cost-effective route for high-quality monolayer h-BN synthesis for sub-atomic scale separations (H + /D + ) and demonstrates the broader potential of machine learning-guided optimization of CVD for advancing synthesis of 2D materials.

36 MATERIALS SCIENCE↗

Automated segmentation of soft X-ray tomography: Native cellular structure with submicron resolution at high-throughput for whole-cell quantitative imaging in yeast

Soft X-ray tomography (SXT) is an invaluable tool for quantitatively analyzing cellular structures at suboptical isotropic resolution. However, it has traditionally depended on manual segmentation, limiting its scalability for large datasets. Here, we leverage a deep learning-based autosegmentation pipeline to segment and label cellular structures in hundreds of cells across three Saccharomyces cerevisiae strains. This task-based pipeline uses manual iterative refinement to improve segmentation accuracy for key structures, including the cell body, nucleus, vacuole, and lipid droplets, enabling high-throughput and precise phenotypic analysis. Using this approach, we quantitatively compared the three-dimensional (3D) whole-cell morphometric characteristics of wild-type, VPH1-GFP, and vac14 strains, uncovering detailed strain-specific cell and organelle size and shape variations. We show the utility of SXT data for precise 3D curvature analysis of entire organelles and cells and detection of fine morphological features using surface meshes. Our approach facilitates comparative analyses with high spatial precision and statistical throughput, uncovering subtle morphological features at the single-cell and population level. This workflow significantly enhances our ability to characterize cell anatomy and supports scalable studies on the mesoscale, with applications in investigating cellular architecture, organelle biology, and genetic research across diverse biological contexts.

Chen, Jianhua [Lawrence Berkeley National Laborato↗

DP-TwoLevel: two-stage gradient subspace learning for differentially private federated learning

Federated learning (FL) enables collaborative model training across distributed data sources without sharing raw data, but faces fundamental challenges in communication efficiency and privacy. Differentially private (DP) training mitigates information leakage but introduces noise that degrades model performance, especially in high-dimensional settings. We propose DP-TwoLevel, a hierarchical gradient projection method that improves utility under fixed DP constraints by exploiting low-dimensional structure in model updates. Our approach learns a two-level PCA-based representation of gradients and applies DP noise in a reduced-dimensional subspace, thereby lowering the effective noise magnitude while preserving dominant signal components. We evaluate the method across three datasets (MNIST, Fashion-MNIST, CIFAR-10) and three privacy regimes (ϵ∈0.5, 1.0, 2.0). Across nine experimental settings, DP-TwoLevel consistently outperforms DP-FedAvg, achieving an average accuracy improvement of 9.44%, with larger gains observed in lower ϵ(higher-noise) regimes (up to +22.31%). We further analyze scalability across models ranging from 100K to 1.49M parameters and identify a variance-based success criterion: performance remains strong when the projection preserves more than 75% of gradient variance, degrades in a marginal regime (65–75%), and fails below this threshold. Our results demonstrate that structure-aware dimensionality reduction can significantly improve the privacy–utility tradeoff in FL without modifying formal privacy guarantees. We also provide empirical evidence of scaling limitations for global projections and motivate per-layer extensions for larger models.

Kotevska, Olivera [ORNL] (ORCID:0000000316772243)↗

Reduce-Order Modeling of Multigroup Neutron Cross Sections for High-Temperature Gas-cooled Reactors

Deterministic neutronics calculations rely on multigroup neutron cross section libraries, which usually consists of a database of tabulated values, used to calculate the cross sections through multivariate linear interpolation. However, interpolation of the multidimensional cross section data becomes memory inefficient and time consuming as the number of tabulations increases, significantly slowing down the neutronics calculation, especially in the case of micro cross section libraries where every isotope (on the order of hundreds) has its own set of specific reactions and cross sections. To address this challenge, this work constructs efficient and robust reduced-order models (ROMs) of the multi-group cross sections to support the Griffin simulation of high-temperature gas-cooled reactors (HTGRs). The first part of the study investigates the linearity of the multi-group cross section data across isotopes, reaction types and energy groups on pre-generated datasets for the purpose of dimensionality reduction. Secondly, a down-selection of ROM techniques is presented on representative classical machine learning (ML) techniques, including variants of linear regression, kernel-based methods, tree-based algorithms, and artificial neural networks. The selection criteria jointly consider the memory efficiency, predictive accuracy, prediction speed, and scalability in comparison to the multidimensional interpolation. Among all the ML techniques, deep neural networks (DNNs) have proven to be the best selection with sufficient accuracy, high robustness, good memory efficiency, great scalability, and superior flexibility. DNNs for have been trained for all isotopes in this work and systematic Griffin testing is ongoing at this moment to ensure the feasibility of this ROM technique for cross section predictions.

42 - ENGINEERING↗

A continuous calibration of the ATLAS flavour-tagging classifiers via optimal transportation maps

A calibration of the ATLAS flavour-tagging algorithms using a new calibration procedure based on optimal transportation maps is presented. Simultaneous, continuous corrections to the b-jet, c-jet, and light-flavour jet classification probabilities from jet-tagging algorithms in simulation are derived for b-jets using $t\bar{t} \rightarrow e\mu \nu \nu bb$ data. After application of the derived calibration maps, closure between simulation and observation is achieved for jet flavour observables used in ATLAS analyses of Large Hadron Collider (LHC) Run 2 proton-proton collision data. This continuous calibration opens up new possibilities for the future use of jet flavour information in LHC analyses and also serves as a guide for deriving high-dimensional corrections to simulation via transportation maps, an important development for a broad range of inference tasks.

Aad, G. [Aix-Marseille Université] (ORCID:00000002↗

A hybrid CNN-LSTM surrogate model for hyper-resolution spatiotemporal flood forecasting in Norfolk, Virginia

Study region: Norfolk, Virginia, United States Study focus: Accurate and timely flood forecasting is essential for enhancing resilience in coastal urban areas in the context of increasing frequency and intensity of rainfall, sea level rise and urbanization. This study presents a hybrid deep learning-based surrogate model that integrates Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM) networks to enable real-time spatiotemporal flood forecasting. The model leverages CNN to capture spatial features from inputs such as elevation and Topographic Wetness Index (TWI), while LSTM processes time-series inputs of rainfall and tide data to capture temporal features. New hydrologic insights for the region: The hybrid CNN-LSTM model was trained using the physics-based hydrodynamic model simulations obtained from the Two-dimensional Unsteady FLOW (TUFLOW) model for Norfolk, Virginia, and achieved high predictive accuracy across diverse flood-prone areas. The reduced computational time from four to six hours using TUFLOW to 3.2 min per event using CNN-LSTM enables rapid flood inundation mapping and early warning applications. The model effectively captured both spatial flood extents and their temporal evolution across different flooding scenarios, providing forecasts at a 2.5-m spatial resolution and 15-min temporal resolution and a one-hour-ahead prediction horizon. While challenges remain in terms of transferability to new regions and real-time data assimilation, this approach demonstrates strong potential for supporting operational flood risk management in coastal urban environments.

Coastal urban flooding↗

Mapping wall-to-wall fractional cover of Arctic tundra plant functional types in Alaska using 20-m spatial resolution satellite imagery and harmonized plot observations

Estimates of fractional cover (fCover) across given land surfaces are used to assess, and often model, vegetation composition and diversity, which are crucial for understanding the health and functioning of terrestrial ecosystems. Remote sensing provides a useful means for scaling local, plot-measured fCover estimates to regional scales. Leveraging a recently synthesized and harmonized plot database, this study generated wall-to-wall maps of fCover for six Alaskan-Arctic plant functional types (PFT), including non-vascular plants, forbs, graminoids, and deciduous and evergreen shrubs, using 20-m satellite data (Sentinel-1, Sentinel-2, ArcticDEM) using a machine learning regression approach, specifically the random forest (RF) algorithm, which is well-suited for handling nonlinear relationships and high-dimensional satellite datasets. This study additionally addressed the spatio-temporal inconsistencies e.g., sampling scale, plot size, and collection year in plot measured fCover by adopting a multivariate outlier detection approach—Cook’s distance—to identify high-quality plots for model training and validation. Our approach achieves high accuracy (R 2 = 0.59–0.93, root mean squared errors = 0.02–0.10 for all PFTs) between plot-observed and satellite-derived fCover when using high-quality plot samples. The mapped fCover characterizes the spatial patterns of different PFTs across the tundra biome at a 20-m resolution, providing key information needed for improved representation of Arctic tundra vegetation in terrestrial biosphere models to better understand climate-vegetation feedback across the Arctic tundra.

Arctic tundra↗

Machine learning inversion of interatomic force constants from single-crystal inelastic neutron scattering

Atomic vibrations govern many macroscopic properties of materials, but experiments to comprehensively probe them remain challenging. Inelastic neutron scattering (INS) is a powerful technique to map phonon dispersions in crystals, especially when leveraging modern time-of-flight (ToF) spectrometers with large detectors. However, efficiently and robustly extracting interatomic force constants (FCs) parameterizing phonon dynamics from experimental spectra remains a bottleneck due to the complexity and high dimensionality of ToF INS datasets. Here, we present a machine learning approach for the direct inversion of FCs from single-crystal INS measurements. The framework leverages synthetic training data generated using universal machine-learned force fields and an efficient physics-based forward model. We benchmark two neural architectures–one emphasizing structured latent representation learning and the other direct, supervised spectral regression–across simulated datasets for two materials under idealized and noisy conditions. The latent-representation model is subsequently applied to experimental single-crystal INS data on germanium. The model is shown to reproduce FCs derived from both first-principles simulations and from iterative optimization, and furthermore achieves reliable inference even from sparse, single-orientation measurements representing short data acquisitions. Analysis of the learned latent space reveals semantically continuous and physically interpretable encodings that support strong cross-domain generalization. By bridging theoretical and experimental domains, we establish a path toward rapid inversion of experimental spectra and data-driven interpretation of temperature-dependent lattice dynamics.

42 ENGINEERING↗

Thermodynamics of the dipole-octupole pyrochlore magnet Ce 2 Hf 2 O 7 in applied magnetic fields

The recently discovered dipole-octupole pyrochlore magnet Ce 2 Hf 2 O 7 is a promising three-dimensional quantum spin liquid candidate which shows no signs of ordering at low temperature. Here we investigate the thermodynamic response to magnetic fields applied along the global [110] direction using specific heat measurements and fits using numerical methods, and solve the corresponding magnetic structure using neutron diffraction. Specific heat data in moderate fields are reproduced well, however, at high fields the agreement is not satisfactory. We especially observe a two-step release of entropy, a finding that demands a review of both theory and experiment. We address it within the framework of three possible scenarios, including an analysis of the crystal field Hamiltonian not restricted to the two-dimensional single-ion doublet subspace. We conclusively rule out two of these scenarios and find qualitative agreement with a simple model of field misalignment with respect to the crystalline direction. As a result, we discuss the implications of our findings for [111] applied fields and for future experiments on Ce 2 Hf 2 O 7 and its sister compounds.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Comparing Gravity Waves in a Kilometer‐Scale Run of the IFS to AIRS Satellite Observations and ERA5

Abstract Atmospheric gravity waves (GWs) impact the circulation and variability of the atmosphere. Sub‐grid scale GWs, which are too small to be resolved, are parameterized in weather and climate models. However, some models are now available at resolutions at which these waves become resolved and it is important to test whether these models do this correctly. In this study, a GW resolving run of the European Center for Medium‐Range Weather Forecasts (ECMWF) Integrated Forecasting System (IFS), run with a 1.4 km average grid spacing (TCo7999 resolution), is compared to observations from the Atmospheric Infrared Sounder (AIRS) instrument, on NASA's Aqua satellite, to test how well the model resolves GWs that AIRS can observe. In this analysis, nighttime data are used from the first 10 days of November 2018 over part of Asia and surrounding regions. The IFS run is resampled with AIRS's observational filter using two different methods for comparison. The ECMWF ERA5 reanalysis is also resampled as AIRS, to allow for comparison of how the high resolution IFS run resolves GWs compared to a lower resolution model that uses GW drag parametrizations. Wave properties are found in AIRS and the resampled models using a multi‐dimensional S‐Transform method. Orographic GWs can be seen in similar locations at similar times in all three data sets. However, wave amplitudes and momentum fluxes in the resampled IFS run are found to be significantly lower than in the observations. This could be a result of horizontal and vertical wavelengths in the IFS run being underestimated.

Meteorology & Atmospheric Sciences↗