Search NASA⌕ Search

SEARCH · Search NASA

Results for “data subsampling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Scalable statistical inference of photometric redshift via data subsampling

Handling big data has largely been a major bottleneck in traditional statistical models. Consequently, when accurate point prediction is the primary target, machine learning models are often preferred over their statistical counterparts for bigger problems. But full probabilistic statistical models often outperform other models in quantifying uncertainties associated with model predictions. We develop a data-driven statistical modeling framework that combines the uncertainties from an ensemble of statistical models learned on smaller subsets of data carefully chosen to account for imbalances in the input space. We demonstrate this method on a photometric redshift estimation problem in cosmology, which seeks to infer a distribution of the redshift—the stretching effect in observing the light of far-away galaxies—given multivariate color information observed for an object in the sky. Our proposed method performs balanced partitioning, graph-based data subsampling across the partitions, and training of an ensemble of Gaussian process models.

data subsampling↗

Bayesian Fit to NOvA Data Subsamples for Three Flavor Oscillation Analysis

NOvA (NuMI Off-Axis $\nu_e$ Appearance) is a long baseline neutrino experiment designed to measure the oscillation of muon neutrinos to electron neutrinos over a distance of 810 km. NOvA uses a near and far detector to observe $\nu_\mu$ disappearance and $\nu_e$ appearance of neutrinos produced by the NuMI beam at Fermilab. NOvA uses a Bayesian analysis framework in addition to its Frequentist method to measure neutrino oscillation parameters such as the mixing angles, mass ordering, and CP-violating phase. We report preliminary results of Bayesian fits to representative NOvA datasets. Comparison of fits to $\nu_\mu$ disappearance and $\nu_e$ appearance enables a cross-check of NOvA results with reactor $\bar{\nu_e}$ disappearance measurements. NOvA also searches for violation of Lorentz invariance by analyzing fits of forward horn current (FHC) versus reverse horn current (RHC) samples. The results validate and advance NOvA's contributions to precision measurements of neutrino properties.

Zhao, Larry [Fermilab]↗

Maximizing efficiency of dataset compression for machine learning potentials with information theory

Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.

36 MATERIALS SCIENCE↗

Dashboard for Visualizing Molecular Property Prediction Machine Learning Results

This is a dashboard for exploring the results of machine learning models for predicting molecular properties from molecular structure. It includes tools for: 1. Modifying molecules to observe the change in predicted properties 2. Exploring the relationship between molecular structure and predicted properties 3. Recommending structurally similar molecules with improved properties 4. Exploring the impact of data subsampling on model performance metrics

Xu, Audrey↗

Framework of compressive sensing and data compression for 4D-STEM

Four-dimensional Scanning Transmission Electron Microscopy (4D-STEM) is a powerful technique for high-resolution and high-precision materials characterization at multiple length scales, including the characterization of beam-sensitive materials. However, the field of view of 4D-STEM is relatively small, which in absence of live processing is limited by the data size required for storage. Furthermore, the rectilinear scan approach currently employed in 4D-STEM places a resolution- and signal-dependent dose limit for the study of beam sensitive materials. Improving 4D-STEM data and dose efficiency, by keeping the data size manageable while limiting the amount of electron dose, is thus critical for broader applications. Here we introduce a general method for reconstructing 4D-STEM data with subsampling in both real and reciprocal spaces at high fidelity. The approach is first tested on the subsampled datasets created from a full 4D-STEM dataset, and then demonstrated experimentally using random scan in real-space. The same reconstruction algorithm can also be used for compression of 4D-STEM datasets, leading to a large reduction (100 times or more) in data size, while retaining the fine features of 4D-STEM imaging, for crystalline samples.

4D-STEM↗

A genome-scale phylogeny of the kingdom Fungi

Phylogenomic studies using genome-scale amounts of data have greatly improved understanding of the tree of life. Despite the diversity, ecological significance, and biomedical and industrial importance of fungi, evolutionary relationships among several major lineages remain poorly resolved, especially those near the base of the fungal phylogeny. To examine poorly resolved relationships and assess progress toward a genome-scale phylogeny of the fungal kingdom, we compiled a phylogenomic data matrix of 290 genes from the genomes of 1,644 species that includes representatives from most major fungal lineages. We also compiled 11 data matrices by subsampling genes or taxa from the full data matrix based on filtering criteria previously shown to improve phylogenomic inference. Analyses of these 12 data matrices using concatenation- and coalescent-based approaches yielded a robust phylogeny of the fungal kingdom, in which ~85% of internal branches were congruent across data matrices and approaches used. We found support for several historically poorly resolved relationships as well as evidence for polytomies likely stemming from episodes of ancient diversification. By examining the relative evolutionary divergence of taxonomic groups of equivalent rank, we found that fungal taxonomy is broadly aligned with both genome sequence divergence and divergence time but also identified lineages where current taxonomic circumscription does not reflect their levels of evolutionary divergence. Finally, our results provide a robust phylogenomic framework to explore the tempo and mode of fungal evolution and offer directions for future fungal phylogenetic and taxonomic studies.

59 BASIC BIOLOGICAL SCIENCES↗

DESI Data Release 2 ELGs: Property-dependent subsamples, imaging systematics, and clustering

Using emission-line galaxies (ELGs) from the Dark Energy Spectroscopic Instrument (DESI) Data Release 2, we evaluate a property-dependent correction to imaging systematics. We derive systematic weights following the same linear regression method used for other DESI tracers, but do so separately on ELG subsamples to provide a physically-informed alternative to the fiducial, neural-network-based approach. In doing so, we show that the deeper imaging in the Dark Energy Survey (DES) footprint leads to a higher overall number density but a lack of targets with extreme $g-r$ and $r-z$ colors. ELGs in the DES region also show a distinct redshift distribution when subsampled by position in the $g-r$ vs. $r-z$ plane. To address these effects, we implement a separate treatment of the DES footprint within the DESI catalog production pipeline, which is generally well-motivated and, in some cases, imperative for accurate clustering measurements. With DES treated separately, we find that property-dependent systematic weights further mitigate spurious clustering signal in $\sim$10% of subsamples, while the fiducial scheme remains optimal for the full sample.

Hagen, T. [Utah U.]↗

The Binary Fraction of Stars in the Dwarf Galaxy Ursa Minor via Dark Energy Spectroscopic Instrument

We utilize multi-epoch line-of-sight velocity measurements from the Milky Way Survey of the Dark Energy Spectroscopic Instrument to estimate the binary fraction for member stars in the dwarf spheroidal galaxy Ursa Minor. Our dataset comprises 670 distinct member stars, with a total of more than 2,000 observations collected over approximately one year. We constrain the binary fraction for UMi to be $0.61^{+0.16}_{-0.20}$ and $0.69^{+0.19}_{-0.17}$, with the binary orbital parameter distributions based on solar neighborhood observation from Duquennoy & Mayor (1991) and Moe & Di Stefano (2017), respectively. Furthermore, by dividing our data into two subsamples at the median metallicity, we identify that the binary fraction for the metal-rich ([Fe/H]>-2.14) population is slightly higher than that of the metal-poor ([Fe/H]<-2.14) population. Based on the Moe & Di Stefano model, the best-constrained binary fractions for metal-rich and metal-poor populations in UMi are $0.86^{+0.14}_{-0.24}$ and $0.48^{+0.26}_{-0.19}$, respectively. After a thorough examination, we find that this offset cannot be attributed to sample selection effects. We also divide our data into two subsamples according to their projected radius to the center of UMi, and find that the more centrally concentrated population in a denser environment has a lower binary fraction of $0.33^{+0.30}_{-0.20}$, compared with $1.00^{+0.00}_{-0.32}$ for the subsample in more outskirts.

Qiu, Tian [Shanghai Jiao Tong Univ. (China); DESI ↗

A global gridded dataset for cloud vertical structure from combined CloudSat and CALIPSO observations

Abstract. The vertical structure of clouds has a profound effect on the global energy budget, the global circulation, and the atmospheric hydrological cycle. The CloudSat and Cloud-Aerosol Lidar and Infrared Pathfinder Satellite Observations (CALIPSO) missions have taken complementary, colocated observations of cloud vertical structure for over a decade. However, no globally gridded dataset is available to the public for the full length of this unique combined data record. Here we present the 3S-GEOPROF-COMB product (Bertrand et al. 2023, https://doi.org/10.5281/zenodo.8057791), a globally gridded (level 3S) community data product summarizing geometrical profiles (GEOPROF) of hydrometeor occurrence from combined (COMB) CloudSat and CALIPSO data. Our product is calculated from the latest release (R05) of per-orbit (level-2) combined cloud mask profiles. We process a set of cloud cover, vertical cloud fraction, and sampling variables at 2.5, 5, and 10° spatial resolutions and monthly and seasonal temporal resolutions. We address the 2011 reduction in CloudSat data collection with Daylight-Only Operations (DO-Op) mode by subsampling pre-2011 data to mimic DO-Op collection patterns, thereby allowing users to evaluate the impact of the reduced sampling on their analyses. We evaluate our data product against CloudSat-only and CALIPSO-only global-gridded data products as well as four comparable surface-based sites, underscoring the added value of the combined product. Interest in the product is anticipated for the study of cloud processes, cloud–climate interactions, and as a candidate baseline climate data record for comparison to follow-up satellite missions, among other uses.

Bertrand, Leah (ORCID:0009000001607558)↗

Performance and calibration of quark/gluon-jet taggers using 140 fb -1 of pp collisions at $\sqrt{s}$ = 13 TeV with the ATLAS detector

The identification of jets originating from quarks and gluons, often referred to as quark/gluon tagging, plays an important role in various analyses performed at the Large Hadron Collider, as Standard Model measurements and searches for new particles decaying to quarks often rely on suppressing a large gluon-induced background. This paper describes the measurement of the efficiencies of quark/gluon taggers developed within the ATLAS Collaboration, using $\sqrt{s}$ = 13 TeV proton–proton collision data with an integrated luminosity of 140 fb -1 collected by the ATLAS experiment. Two taggers with high performances in rejecting jets from gluon over jets from quarks are studied: one tagger is based on requirements on the number of inner-detector tracks associated with the jet, and the other combines several jet substructure observables using a boosted decision tree. A method is established to determine the quark/gluon fraction in data, by using quark/gluon-enriched subsamples defined by the jet pseudorapidity. Differences in tagging efficiency between data and simulation are provided for jets with transverse momentum between 500 GeV and 2 TeV and for multiple tagger working points.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Characterization and differentiation of aluminum powders used in improvised explosive devices – Part 1: Proof of concept of the utility of particle micromorphometry

Abstract Aluminum (Al) powders are commonly used in improvised explosive devices as metallic fuels, a component of explosive mixtures. These powders can be obtained readily from industrial‐scale and consumer products, and produced using unsophisticated “kitchen chemistry” techniques. This research demonstrates the potential of automated particle micromorphometry for comparisons between known source and questioned Al powders recovered from IEDs, as well as for insight into the method of Al powder manufacture. Al powder samples were obtained from legitimate manufacturers, and 56 samples were produced “in‐house” from Al‐containing spray paints and ball‐milled Al foils. Transmitted light microscope images of Al powder particles were acquired using an automated stage with automated z‐focus; 17 size and shape parameters were measured for all particles. Approximately 37,000–2,500,000 particles/sample were analyzed using an open‐source statistical package with customized code. Dimensionality reduction was required for processing the large datasets: eight of the 17 measured variables were selected based on inspection of the correlation matrix. Data from four subsamples from each of the 56 samples produced using “in‐house” methods were analyzed using ANOVA to assess the within‐ and between‐sample variation. High within‐sample variation was noted; however, ANOVA and post‐hoc Tukey's honestly significant difference (HSD) tests demonstrated that the between‐sample variation was substantially larger than the within‐sample variation. Each sample could be differentiated from all other samples in the test set. Future experiments will focus on ways to reduce the within‐sample variation, and additional statistical and microanalytical methods to classify sources and confidently constrain the method of Al powder manufacture.

Baldaino, JenaMarie↗

Intelligent Sampling of Extreme-Scale Turbulence Datasets for Accurate and Efficient Spatiotemporal Model Training

With the end of Moore’s law and Dennard scaling, efficient training increasingly requires rethinking data volume. Can we train better models with significantly less data via intelligent subsampling? To explore this, we develop SICKLE, a sparse intelligent curation framework for efficient learning, featuring a novel maximum entropy (MaxEnt) sampling approach, scalable training, and energy benchmarking. We compare MaxEnt with random and phase-space sampling on large direct numerical simulation (DNS) datasets of turbulence. Evaluating SICKLE at scale on Frontier, we show that subsampling as a preprocessing step can, in many cases, improve model accuracy and substantially lower energy consumption, with observed reductions of up to 38×.

Brewer, Wes [ORNL] (ORCID:0000000236393956)↗

Exploration with Scalable Gaussian Process Reinforcement Learning

Exploration is a challenging problem in reinforcement learning (RL), especially in environments with sparse rewards. Quantifying and utilizing the parametric uncertainty has been shown to be paramount for successful exploration [Osband et al., 2018]. Bayesian, or approximately Bayesian, methods present a principled means of estimating the parametric uncertainty in RL problems. Gaussian processes, nonparametric Bayesian models, are often impractical due to poor scalability and computational bottlenecks. We introduce a scalable Gaussian process RL (GPRL) method which directly induces sparsity in the covariance matrix to facilitate faster computation. This is a departure from previous GPRL methods which instead rely on data reduction and subsampling. We compare various covariance-based exploration techniques (Thompson sampling, upper confidence bound, and probabilistic maximum variance) which leverage our scalable GP framework in sparse reward environments. Finally, we show favorable comparison against the bootstrapped deep Q-Network.

97 MATHEMATICS AND COMPUTING↗

Temporal Subsampling Diminishes Small Spatial Scales in Recurrent Neural Network Emulators of Geophysical Turbulence

The immense computational cost of traditional numerical weather and climate models has sparked the development of machine learning (ML) based emulators. Because ML methods benefit from long records of training data, it is common to use data sets that are temporally subsampled relative to the time steps required for the numerical integration of differential equations. Here, we investigate how this often overlooked processing step affects the quality of an emulator's predictions. We implement two ML architectures from a class of methods called reservoir computing: (a) a form of Nonlinear Vector Autoregression (NVAR), and (b) an Echo State Network (ESN). Despite their simplicity, it is well documented that these architectures excel at predicting low dimensional chaotic dynamics. We are therefore motivated to test these architectures in an idealized setting of predicting high dimensional geophysical turbulence as represented by Surface Quasi-Geostrophic dynamics. In all cases, subsampling the training data consistently leads to an increased bias at small spatial scales that resembles numerical diffusion. Interestingly, the NVAR architecture becomes unstable when the temporal resolution is increased, indicating that the polynomial based interactions are insufficient at capturing the detailed nonlinearities of the turbulent flow. The ESN architecture is found to be more robust, suggesting a benefit to the more expensive but more general structure. Spectral errors are reduced by including a penalty on the kinetic energy density spectrum during training, although the subsampling related errors persist. Future work is warranted to understand how the temporal resolution of training data affects other ML architectures.

58 GEOSCIENCES↗

Variance Preserving Spectral Subsampling

Generating statistically faithful short-duration gamma-ray spectra from a single long measurement is essential in nuclear safeguards, supporting tasks such as algorithm development and machine-learning applications, especially when list-mode data are unavailable. Existing subsampling methods often distort the statistical characteristics of genuine short-duration measurements, leading to biased or unreliable analytical outcomes and thereby undermining downstream tasks. In this work, we compare five subsampling approaches using a benchmark set of 156 genuine replicate spectra collected with a high-purity germanium detector. We evaluate each method with respect to run-to-run variance, channel-to-channel variance, and preservation of total counts (losslessness). Across a wide range of subsampling ratios, only binomial subsampling without replacement consistently reproduces the statistical properties of genuine short-duration spectra, maintaining proper dispersion even in sparse spectral regions and perfectly preserving total counts. These results provide a mathematically principled and practically validated framework for generating synthetically shortened spectra when true short-duration measurements are unavailable.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Cosmic ray spectrum of protons plus helium nuclei between 6 and 158 TeV from HAWC data

Here, a measurement with high statistics of the differential energy spectrum of light elements in cosmic rays, in particular, of primary H plus He nuclei, is reported. The spectrum is presented in the energy range from 6 to 158 TeV per nucleus. Data was collected with the High Altitude Water Cherenkov (HAWC) Observatory between June 2015 and June 2019. The analysis was based on a Bayesian unfolding procedure, which was applied on a subsample of vertical HAWC data that was enriched to 82% of events induced by light nuclei. To achieve the mass separation, a cut on the lateral age of air shower data was set guided by predictions of CORSIKA/QGSJET-II-04 simulations. The measured spectrum is consistent with a broken power-law spectrum and shows a kneelike feature at around E = 24.0$^{+3.6}_{-3.1}$ TeV , with a spectral index γ = -2.51 ± 0.02 before the break and with γ = -2.83 ± 0.02 above it. The feature has a statistical significance of 4.1σ. Within systematic uncertainties, the significance of the spectral break is 0.8σ.

79 ASTRONOMY AND ASTROPHYSICS↗

Predicting Rare Earth Element Potential in Produced and Geothermal Waters of the United States via Emergent Self-Organizing Maps

This work applies emergent self-organizing map (ESOM) techniques, a form of machine learning, in the multidimensional interpretation and prediction of rare earth element (REE) abundance in produced and geothermal waters in the United States. Visualization of the variables in the ESOM trained using the input data shows that each REE, with the exception of Eu, follows the same distribution patterns and that no single parameter appears to control their distribution. Cross-validation, using a random subsample of the starting data and only using major ions, shows that predictions are generally accurate to within an order of magnitude. Using the same approach, an abridged version of the U.S. Geological Survey Produced Waters Database, Version 2.3 (which includes both data from produced and geothermal waters) was mapped to the ESOM and predicted values were generated for samples that contained enough variables to be effectively mapped. Results show that in general, produced and geothermal waters are predicted to be enriched in REEs by an order of magnitude or more relative to seawater, with maximum predicted enrichments in excess of 1000-fold. Cartographic mapping of the resulting predictions indicates that maximum REE concentrations exceed values in seawater across the majority of geologic basins investigated and that REEs are typically spatially co-associated. The factors causing this co-association were not determined from ESOM analysis, but based on the information currently available, REE content in produced and geothermal waters is not directly controlled by lithology, reservoir temperature, or salinity.

Engle, Mark A. (ORCID:0000000152587374)↗