Search NASA⌕ Search

SEARCH · Search NASA

Results for “sparse data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Spike-and-Slab Shrinkage Priors for Structurally Sparse Bayesian Neural Networks

Network complexity and computational efficiency have become increasingly significant aspects of deep learning. Sparse deep learning addresses these challenges by recovering a sparse representation of the underlying target function by reducing heavily overparameterized deep neural networks. Specifically, deep neural architectures compressed via structured sparsity (e.g., node sparsity) provide low-latency inference, higher data throughput, and reduced energy consumption. In this article, we explore two well-established shrinkage techniques, Lasso and Horseshoe, for model compression in Bayesian neural networks (BNNs). To this end, we propose structurally sparse BNNs, which systematically prune excessive nodes with the following: 1) spike-and-slab group Lasso (SS-GL) and 2) SS group Horseshoe (SS-GHS) priors, and develop computationally tractable variational inference, including continuous relaxation of Bernoulli variables. We establish the contraction rates of the variational posterior of our proposed models as a function of the network topology, layerwise node cardinalities, and bounds on the network weights. Furthermore, we empirically demonstrate the competitive performance of our models compared with the baseline models in prediction accuracy, model compression, and inference latency.

97 MATHEMATICS AND COMPUTING↗

Non-intrusive reduced-order modeling for dynamical systems with spatially localized features

This work presents a non-intrusive reduced-order modeling framework for dynamical systems with spatially localized features characterized by slow singular value decay. The proposed approach builds upon two existing methodologies for reduced and full-order non-intrusive modeling, namely Operator Inference (OpInf) and sparse Full-Order Model (sFOM) inference. We decompose the domain into two complementary subdomains that exhibit fast and slow singular value decay. The dynamics of the subdomain exhibiting slow singular value decay are learned with sFOM while the dynamics with intrinsically low dimensionality on the complementary subdomain are learned with OpInf. The resulting, coupled OpInf-sFOM formulation leverages the computational efficiency of OpInf and the high resolution of sFOM, and thus enables fast non-intrusive predictions for conditions beyond those sampled in the training data set. A novel regularization technique with a closed-form solution based on the Gershgorin disk theorem is introduced to promote stable sFOM and OpInf models. We also provide a data-driven indicator for subdomain selection and ensure solution smoothness over the interface via a post-processing interpolation step. We evaluate the efficiency of the approach in terms of offline and online speedup through a quantitative, parametric computational cost analysis. We demonstrate the coupled OpInf-sFOM formulation for two test cases: a one-dimensional Burgers’ model for which accurate predictions beyond the span of the training snapshots are presented, and a two-dimensional parametric model for the Pine Island Glacier ice thickness dynamics, for which the OpInf-sFOM model achieves an average prediction error on the order of 1% with an online speedup factor of approximately 8$\times$ compared to the numerical simulation.

42 ENGINEERING↗

Anomaly Detection for Online Monitoring of Thermocouple Sensors in the Advanced Test Reactor

This study explores data-driven anomaly detection methods to analyze sensor fail- ures in the Advanced Gas Reactor (AGR) nuclear fuel irradiation experiments. Specifically, we examine failures of thermocouples (TCs), which are critical for mon- itoring and controlling in-reactor temperatures during operation. Failures were pri- marily observed during abrupt power transitions and manifested as sensor drop-outs, drifts, or unexplained behavior. We applied three time-series analysis techniques— rolling mean smoothing, matrix profile, and vector auto-regression (VAR)—to de- tect anomalies in TC data prior to failure events. The rolling mean method effec- tively highlighted deviations aligned with reported failures, while the matrix profile provided partial early warning but sometimes flagged normal fluctuations during power-down periods. VAR shows potential in capturing multivariate dependencies but requires further calibration. A rare case of TC drift was also documented, which did not result in failure, underscoring the challenge of building predictive models with sparse positive examples. Our findings demonstrate that traditional statistical tools can aid anomaly detection but have limited predictive power without richer training data. We propose future directions including synthetic data generation, real- time surrogate modeling, and multi-modal feature integration. This work provides a foundation for applying robust anomaly detection frameworks to mission-critical sensor systems in experimental settings.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Accelerated Constrained Sparse Tensor Factorization on Massively Parallel Architectures

This study presents the first constrained sparse tensor factorization (cSTF) framework that optimizes and fully offloads computation to massively parallel GPU architectures, and the first performance characterization of cSTF on GPU architectures. In contrast to prior work on tensor factorization, where the matricized tensor times Khatri-Rao product (MTTKRP) is the primary performance bottleneck, our systematic analysis of the cSTF algorithm on GPUs reveals that adding constraints creates an additional bottleneck in the update operation for many real-world sparse tensors. While executing the update operation on the GPU brings significant speedup over its CPU counterpart, it remains a significant bottleneck. To further accelerate the update operation, we propose cuADMM, a new update algorithm that leverages algorithmic and code optimization strategies to minimize both computation and data movement on GPUs. As a result, our framework delivers significantly improved performance compared to prior state-of-the-art. On 10 real-world sparse tensors, our framework achieves geometric mean speedup of 5.1 × (max 41.59 ×) and 7.01 × (max 58.05 ×) on the NIVIDA A100 and H100 GPUs, respectively, over the state-of-the-art SPLATT library running on a 26-core Intel Ice Lake Xeon CPU.

Soh, Yongseok↗

Collaborative Research: Enabling multi-scale studies of magnetic reconnection with interpretable data-driven models

The development of accurate reduced descriptions and improved closures for magnetic reconnection is an important and a long‐standing challenge in plasma physics. The four‐fluid approach, and associated closures, that were investigated have the potential to improve the accuracy of plasma fluid models, capturing physical effects which would otherwise require a kinetic description. If successful, this approach could have an important impact for the modeling of laboratory and space plasmas. The major goals of this project were to develop new machine learning (ML) tools based on sparse and symbolic regression techniques, and to extract interpretable and generalizable reduced models (e.g., in the form of partial differential equations - PDEs) from data generated by first principles plasma simulations. Preserving interpretability of such data‐driven models is key to addressing the long‐standing theoretical and numerical challenges. Prior proof‐of‐principle studies have demonstrated the enormous potential of this approach, by recovering the well‐established hierarchy of plasma equations (from Vlasov to MHD) from data produced by particle‐in‐cell (PIC) simulations. Our goal in this project was to extend and apply these new tools to construct better kinetic closures for magnetic reconnection; to derive better models of particle injection and acceleration by this fundamental plasma process; and to use this understanding to accelerate the development of multi‐scale plasma algorithms. While our immediate focus was on the problem of magnetic reconnection, the tools that were will developed are general and applicable to other areas of plasma physics, and more broadly to many‐body phenomena. We anticipate that the development of these multi‐scale models will have a significant impact across different areas of plasma science, from fusion to space and astrophysical plasmas.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Effect of Biomass Water Dynamics in Cosmic-Ray Neutron Sensor Observations: A Long-Term Analysis of Maize–Soybean Rotation in Nebraska

Precise soil water content (SWC) measurement is crucial for effective water resource management. This study utilizes the Cosmic-Ray Neutron Sensor (CRNS) for area-averaged SWC measurements, emphasizing the need to consider all hydrogen sources, including time-variable plant biomass and water content. Near Mead, Nebraska, three field sites (CSP1, CSP2, and CSP3) growing a maize–soybean rotation were monitored for 5 (CSP1 and CSP2) and 13 (CSP3) years. Data collection included destructive biomass water equivalent (BWE) biweekly sampling, epithermal neutron counts, atmospheric meteorological variables, and point-scale SWC from a sparse time domain reflectometry (TDR) network (four locations and five depths). In 2023, dense gravimetric SWC surveys were collected eight (CSP1 and CSP2) and nine (CSP3) times over the growing season (April to October). The N0 parameter exhibited a linear relationship with BWE, suggesting that a straightforward vegetation correction factor may be suitable (fb). Results from the 2023 gravimetric surveys and long-term TDR data indicated a neutron count rate reduction of about 1% for every 1 kg m−2 (or mm of water) increase in BWE. This reduction factor aligns with existing shorter-term row crop studies but nearly doubles the value previously reported for forests. This long-term study contributes insights into the vegetation correction factor for CRNS, helping resolve a long-standing issue within the CRNS community.

Chemistry↗

Particle hit clustering and identification using point set transformers in liquid argon time projection chambers

Liquid argon time projection chambers are often used in neutrino physics and dark-matter searches because of their high spatial resolution. The images generated by these detectors are extremely sparse, as the energy values detected by most of the detector are equal to 0, meaning that despite their high resolution, most of the detector is unused in a particular interaction. Instead of representing all of the empty detections, the interaction is usually stored as a sparse matrix, a list of detection locations paired with their energy values. Traditional machine learning methods that have been applied to particle reconstruction such as convolutional neural networks (CNNs), however, cannot operate over data stored in this way and therefore must have the matrix fully instantiated as a dense matrix. Operating on dense matrices requires a lot of memory and computation time, in contrast to directly operating on the sparse matrix. We propose a machine learning model using a point set neural network that operates over a sparse matrix, greatly improving both processing speed and accuracy over methods that instantiate the dense matrix, as well as over other methods that operate over sparse matrices. Compared to competing state-of-the-art methods, our method improves classification performance by 14%, segmentation performance by more than 22%, while taking 80% less time and using 66% less memory. Compared to state-of-the-art CNN methods, our method improves classification performance by more than 86%, segmentation performance by more than 71%, while reducing runtime by 91% and reducing memory usage by 61%.

calibration and fitting methods↗

Smooth trends in fermium charge radii and the impact of shell effects

The quantum-mechanical nuclear-shell structure determines the stability and limits of the existence of the heaviest nuclides with large proton numbers Z ≳ 100. Shell effects also affect the sizes and shapes of atomic nuclei, as shown by laser spectroscopy studies in lighter nuclides. However, experimental information on the charge radii and the nuclear moments of the heavy actinide elements, which link the heaviest naturally abundant nuclides with artificially produced superheavy elements, is sparse. Here we present laser spectroscopy measurements along the fermium (Z = 100) isotopic chain and an extension of data in the nobelium isotopic chain (Z = 102) across a key region. Multiple production schemes and different advanced techniques were applied to determine the isotope shifts in atomic transitions, from which changes in the nuclear mean-square charge radii were extracted. A range of nuclear models based on energy density functionals reproduce well the observed smooth evolution of the nuclear size. Both the remarkable consistency of model prediction and the similarity of predictions for different isotopes suggest a transition to a regime in which shell effects have a diminished effect on the size compared with lighter nuclei.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Experimental validation of a collision-radiation dataset for molecular hydrogen in plasmas

Quantitative spectroscopy of molecular hydrogen has generated substantial demand, leading to the accumulation of diverse elementary process data encompassing radiative transitions, electron-impact transitions, predissociations, and quenching. However, their rates currently available are still sparse, and there are inconsistencies among those proposed by different authors. In this study, we demonstrate an experimental validation of such a molecular dataset by composing a collisional-radiative model (CRM) for molecular hydrogen and comparing experimentally obtained vibronic populations across multiple levels. From the population kinetics of molecular hydrogen, the importance of each elementary process in various parameter space is studied. In low-density plasmas (electron density ne≲1017 m−3) the excitation rates from the ground states and radiative decay rates, both of which have been reported previously, determine the excited state population. The inconsistency in the excitation rates affects the population distribution the most significantly in this parameter space. However, in higher density plasmas (ne≳1018 m−3), the excitation rates from excited states become important, which have never been reported in the literature, and may need to be approximated in some way. In order to validate these molecular datasets and approximated rates, we carried out experimental observations for two different hydrogen plasmas; a low-density radio frequency heated plasma (ne≈1016 m−3) and the Large Helical Device (LHD) divertor plasma (ne≳1018 m−3). The visible emission lines from EF1Σg+, HH¯1Σg+, D1Πu±, GK1Σg+, I1Πg±, J1Δg±, h3Σg+, e3Σu+, d3Πu±,g3Σg+, i3Πg±, and j3Δg± states were observed simultaneously and their population distributions were obtained from their intensities. We compared the observed population distributions with the CRM prediction, in particular the CRM with the rates compiled by Janev et al., Miles et al., and those calculated with the molecular convergent close-coupling (MCCC) method. The MCCC prediction gives the best agreement with the experiment, particularly for the emission from the low-density plasma. However, the population distribution in the LHD divertor shows a worse agreement with the CRM than those from low-density plasma, indicating the necessity of the precise excitation rates from excited states. We also found that the rates for the electron attachment is inconsistent with experimental results. This requires further investigation.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A convergence metric for counting statistics in time-resolved small angle neutron scattering

Here, this work introduces a model-independent, dimensionless metric for predicting optimal measurement duration in time-resolved small-angle neutron scattering using early-time data. Built on a Gaussian process regression framework, the method reconstructs scattering profiles with quantified uncertainty, even from sparse or noisy measurements. Demonstrated on the EQ-SANS instrument at the Spallation Neutron Source, the approach generalizes to general SANS instruments with a two-dimensional detector. A key result is the discovery of a dimensionless convergence metric revealing a universal power-law scaling in profile evolution across soft matter systems. When time is normalized by a system-specific characteristic time t*, the variation in inferred profiles collapses onto a single curve with an exponent between −2 and −1. This trend emerges within the first ten time steps, enabling early prediction of measurement sufficiency. The method supports real-time experimental optimization and is especially valuable for maximizing efficiency in low-flux environments such as compact accelerator-based neutron sources.

Tung, Chi-Huan [Oak Ridge National Laboratory (ORN↗

Data-Efficient Methods for Determining Flory–Huggins χ Parameters in Multicomponent Polymer Formulations

Polymer formulations are essential in diverse applications including personal care products, coatings, paints, adhesives, and plastic materials. Designing these formulations requires navigating large, complex design spaces, where phase and self-assembly behavior critically impact performance. The Flory–Huggins χ parameter, which quantifies segmental miscibility, is widely used to parametrize the excess free energy of mixing in formulation models. In this work, we introduce two data-efficient, top-down methods for estimating χ parameters using the Random Phase Approximation (RPA): (i) Boundary Nonlinear Regression (Boundary-NLR), which fits theoretical spinodal boundaries to experimental phase boundaries, and (ii) Surrogate Model Inverse Parameter Estimation (SMIPE), which uses a Gaussian Process Classifier to fit sparse phase maps via a surrogate model. Both methods allow rapid parametrization of polymer field-theoretic models without the need for additional experiments. We evaluate these approaches on data sets involving polymer–solvent–nonsolvent ternary mixtures and block copolymer–solvent systems, demonstrating their robustness to experimental noise and their relevance for real-world formulation design.

copolymers↗

Multi-Level Structural Damage Characterization Using Sparse Acoustic Sensor Networks and Knowledge Transferred Deep Learning

Standard structural health monitoring techniques face well-known difficulties for comprehensive defect diagnosis in real-world structures that have structural, material, or geometric complexity. This motivates the exploration of machine-learning-based structural health monitoring methods in complex structures. However, creating sufficient training data sets with various defects is an ongoing challenge for data-driven machine (deep) learning algorithms. The ability to transfer the knowledge of a trained neural network from one component to another or to other sections of the same component would drastically reduce the required training data set. Also, it would facilitate computationally inexpensive machine learning based inspection systems. In this work, a machine-learning-based multi-level damage characterization is demonstrated with the ability to transfer trained knowledge within the sparse sensor network. A novel network spatial assistance and an adaptive convolution technique are proposed for efficient knowledge transfer within the deep learning algorithm. Proposed structural health monitoring method is experimentally evaluated on an aluminum plate with artificially induced defects. It was observed that the method improves the performance of knowledge transferred damage characterization by 50% during localization and 24% during severity assessment. Further, experiments using time windows with and without multiple edge reflections are studied. Results reveal that multiply scattered waves contain rich and deterministic defect signatures that can be mined using deep learning neural networks, improving the accuracy of both identification and quantification. In the case of a fixed sensor network, using multiply scattered waves shows 100% prediction accuracy at all levels of damage characterization.

36 MATERIALS SCIENCE↗

Uncovering multiscale structure-property correlations via active learning in scanning tunneling microscopy

Atomic arrangements and local sub-structures fundamentally influence emergent material functionalities. These structures are conventionally probed using spatially resolved studies and the property correlations are deciphered by a researcher based on sequential explorations, thereby limiting the efficiency and scope. Here we demonstrate a multi-scale Bayesian deep-learning based framework that automatically correlates material structure with its electronic properties using scanning tunneling microscopy (STM) measurements in real-time. Its predictions are used to autonomously direct exploration toward regions of the sample that optimize a given material property. This method is deployed on a low-temperature ultra-high vacuum STM to understand the structure-property relationship in a europium-based semimetal, EuZn 2 As 2 , a promising candidate relevant to magnetism-driven topological phenomena. The framework employs a sparse-sampling approach to efficiently construct the scalar-property space using minimal measurements, about 1–10% of the data required in standard hyperspectral methods. Moreover, we formulate the problem hierarchically across length scales, implementing autonomous workflow to locate mesoscopic and atomic structures that correspond to a target material property. This framework offers the choice to design scalar-property from the spectroscopic data to steer sample exploration. Our findings reveal correlations of the electronic properties unique to surface terminations, local defect density, and point defects.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

CORE-BFS: Communication-Optimized REctangular-partitioned BFS Achieving 160.845 TeraTEPS on Frontier Supercomputer

Distributed Breadth-First Search (BFS) is fundamental to many large-scale graph applications, but its performance on parallel systems is often limited by high communication overhead. This paper presents CORE-BFS, an extremely scalable GPU-based BFS implementation that introduces a unique rectangular 2D partitioning-based design for Frontier supercomputer. To further improve performance, we propose four key optimizations: (1) Rectangular 2D-partition specific data formats that use two compressed row and one compressed column status array bitmaps combined with a Double Compressed Sparse Row (DCSR) format per partition, reducing memory footprint and inter-rank traffic; (2) Adaptive frontier & communication strategy that unifies top-down and bottom-up traversal on the rectangular layout, uses lazy synchronization in top-down levels, and switches variants based on frontier size to minimize communication overhead; (3) Frontier-split degree-aware update that maps frontier vertices to thread-centric, wavefront-centric, and block-centric kernels based on their degree to improve GPU utilization and memory coalescing; (4) Row-reduction pipeline that overlaps bottom-up adjacency list processing with row-wise bitmap reduction to hide inter-rank latency. Together, these techniques increase parallelism while reducing memory and communication overhead. On the Graph500 benchmark, CORE - BFS scales up to 9,248 Frontier nodes with scale-42 graphs and reaches 160.845 TTEPS, delivering a 5.42 × speedup over our previous Frontier implementation.

Yang, Haoshen [Rutgers University]↗

AmeriFlux CA-RSB Resolute Bay Polar Desert

This is the AmeriFlux version of the carbon flux data for the site CA-RSB Resolute Bay Polar Desert. Site Description - Polar desert on a gentle slope. Sparse grases and forbs in lower lying areas between microtopography in ice wedge polygon formations.

Arndt, Kyle [Woodwell Climate Research Center]↗

Monthly Mean In Situ Surface Flux Observations Paired with Satellite-Derived and Reanalysis-Based Flux Data for the Great Lakes Region, 2001–2020

Surface radiative and turbulent heat fluxes over the Great Lakes strongly influence regional hydrological and meteorological processes, and their accurate representation is critical for numerical weather prediction and coupled atmosphere–lake modeling. However, direct flux observations are spatially sparse across the region, so gridded reanalysis and satellite-derived products are often used for climatological analyses and model evaluation despite differences in their flux representations. This dataset provides processed, quality-controlled, monthly mean surface flux observations from the Great Lakes Evaporation Network (GLEN), AmeriFlux, and the National Data Buoy Center, paired with spatiotemporally matched flux estimates from two reanalysis products, the fifth generation European Centre for Medium-Range Weather Forecasts (ECMWF) reanalysis dataset (ERA5) and the Modern Era Reanalysis for Research and Applications, version 2 (MERRA-2), and two satellite-derived products, the Clouds and Earth's Radiant Energy Systems Energy Balanced and Filled (CERES-EBAF) and the Cloud, Albedo and Surface Radiation dataset from AVHRR data - Edition 3 (CLARA-A3). The dataset includes sixteen observational stations with variable temporal coverage within 2001–2020. For each station, a CSV file contains monthly time series of available flux variables, including surface downwelling shortwave radiation (SW), surface downwelling longwave radiation (LW), sensible heat (SH) flux, and latent heat flux (LH), alongside matched gridded product values where available. Columns in the CSV file correspond to different variables sourced from each dataset, with column titles structured as "{dataset}_{variable}". Columns with relevant metadata are also provided in each CSV file, including station latitude and longitude, monthly timestamps, and the name of the sourced observational data. These files are structured for direct use in common analysis tools, including Microsoft Excel, Python pandas, and Python matplotlib. This dataset supports climatological analysis of the Great Lakes regional surface energy budget, evaluation of satellite-derived and reanalysis-based flux products, and development or validation of flux representations in numerical weather prediction and coupled atmosphere–lake models.

Great Lakes↗

Evaluation of Station Performance of the Idaho National Laboratory Seismic Monitoring Network Using Network Detection Thresholds

The Idaho National Laboratory (INL) Seismic Monitoring Network is located in eastern Idaho and monitors a portion of the intermountain seismic belt. It has been in place for 50 yr and has undergone several major changes, the most recent of which has been the transition to the Antelope real‐time acquisition system and the implementation of automatic phase picking algorithms to aid in analysis. This study discusses the efforts to evaluate the performance of the INL seismic monitoring network (and other surrounding stations) using the new real‐time acquisition system. The method outlined by Wilson et al. (2021) is used to develop an empirical relationship between the observability of local earthquakes as a function of magnitude and distance. This relationship is used to produce detection thresholds for Pwaves for all stations of interest. The INL seismic network has two main goals: monitor tectonic‐and volcanic‐related events and measure ground motions for input into seismic hazard analysis. Because of these two overall objectives, several seismic stations have been installed near critical facilities and, therefore, are not as quiet as stations that are used primarily for earthquake detection. This is reflected in their detection thresholds, which are much smaller for stations away from facilities. This study shows that the INL Seismic Monitoring Network is able to detect earthquakes near INL facilities with M L > 1.2, with redundancies built in to ensure this sensitivity even if data became unavailable from some stations. This study also shows “holes” in the monitoring network where the detection of smaller earthquakes is highly dependent on sparsely placed seismic stations. In conclusion, the results of this study will be used to govern plans for expansion of earthquake monitoring in Idaho and the surrounding region and to fine‐tune the detection thresholds for individual stations.

58 - GEOSCIENCES↗

The 3D Lyman- α forest power spectrum from eBOSS DR16

We measure the three-dimensional power spectrum (P3D) of the transmitted flux in the Lyman-α (Ly α) forest using the complete extended Baryon Oscillation Spectroscopic Survey data release 16 (eBOSS DR16). This sample consists of ~205 000 quasar spectra in the redshift range 2 ≤ z ≤ 4 at an effective redshift z = 2.334. We propose a pair-count spectral estimator in configuration space, weighting each pair by exp( i k ∙ r), for wave vector k and pixel pair separation r, effectively measuring the anisotropic power spectrum without the need for fast Fourier transforms. This accounts for the window matrix in a tractable way, avoiding artefacts found in Fourier-transform based power spectrum estimators due to the sparse sampling transverse to the line of sight of Ly α skewers. We extensively test our pipeline on two sets of mocks: (i) idealized Gaussian random fields with a sparse sampling of Ly α skewers, and (ii) log-normal LyaCoLoRe mocks including realistic noise levels, the eBOSS survey geometry and contaminants. On eBOSS DR16 data, the Kaiser formula with a non-linear correction term obtained from hydrodynamic simulations yields a good fit to the power spectrum data in the range $(0.02 ≤ k ≤ 0.35)$ h Mpc -1 at the 1–2σ level with a covariance matrix derived from LyaCoLoRe mocks. We demonstrate a promising new approach for full-shape cosmological analyses of Ly α forest data from cosmological surveys such as eBOSS, the currently observing Dark Energy Spectroscopic Instrument and future surveys such as the Prime Focus Spectrograph, WEAVE-QSO, and 4MOST.

79 ASTRONOMY AND ASTROPHYSICS↗