Search NASA⌕ Search

SEARCH · Search NASA

Results for “data subsampling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Scalable statistical inference of photometric redshift via data subsampling

Handling big data has largely been a major bottleneck in traditional statistical models. Consequently, when accurate point prediction is the primary target, machine learning models are often preferred over their statistical counterparts for bigger problems. But full probabilistic statistical models often outperform other models in quantifying uncertainties associated with model predictions. We develop a data-driven statistical modeling framework that combines the uncertainties from an ensemble of statistical models learned on smaller subsets of data carefully chosen to account for imbalances in the input space. We demonstrate this method on a photometric redshift estimation problem in cosmology, which seeks to infer a distribution of the redshift—the stretching effect in observing the light of far-away galaxies—given multivariate color information observed for an object in the sky. Our proposed method performs balanced partitioning, graph-based data subsampling across the partitions, and training of an ensemble of Gaussian process models.

data subsampling↗

Bayesian Fit to NOvA Data Subsamples for Three Flavor Oscillation Analysis

NOvA (NuMI Off-Axis $\nu_e$ Appearance) is a long baseline neutrino experiment designed to measure the oscillation of muon neutrinos to electron neutrinos over a distance of 810 km. NOvA uses a near and far detector to observe $\nu_\mu$ disappearance and $\nu_e$ appearance of neutrinos produced by the NuMI beam at Fermilab. NOvA uses a Bayesian analysis framework in addition to its Frequentist method to measure neutrino oscillation parameters such as the mixing angles, mass ordering, and CP-violating phase. We report preliminary results of Bayesian fits to representative NOvA datasets. Comparison of fits to $\nu_\mu$ disappearance and $\nu_e$ appearance enables a cross-check of NOvA results with reactor $\bar{\nu_e}$ disappearance measurements. NOvA also searches for violation of Lorentz invariance by analyzing fits of forward horn current (FHC) versus reverse horn current (RHC) samples. The results validate and advance NOvA's contributions to precision measurements of neutrino properties.

Zhao, Larry [Fermilab]↗

Maximizing efficiency of dataset compression for machine learning potentials with information theory

Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.

36 MATERIALS SCIENCE↗

Observations on the long-period variability of the Gulf Stream downstream of Cape Hatteras

To examine the long-period variability of the Gulf Stream, sea level residuals relative to a 2-year mean sea level in the Gulf Stream downstream of Cape Hatteras (between 75 deg W and 60 deg W longitude) are used. Residuals, as derived from Geosat altimetry between November 1986 and December 1988, were gridded in space and time at a temporal resolution of 10 days and spatial resolution of 1/4 deg. Complex empirical orthogonal function (CEOF) analysis was applied to the data set to extract the spatially correlated signal with the original data subsampled to 1/2 deg. In addition to determining the space-time scales and propagation characterisitics of the different modes, wavenumber-frequency spectral techniques were used to separate the variability into propagating and stationary components. The CEOF technique applied to the data set indicated that the first four CEOF modes accounted for 60% of the variability and were found to be above the noise leve 99% of the time. CEOF 1 was associated with westward propagation at 5 km/d at a wavelength of 2000 km and eastward propagation at 1-2 km/d centered at a 500-km wavelength. This first CEOF is in good agreement with thin-jet equivalent barotropic models which predict westward propagation for wavelengths greater than 1130 km. A deflection of the wavelike pattern at 65 deg W also indicates a possible topographic effect. A simple scaling of the effect of topography indicates that for length scales longer than the internal Rossby radius of deformation, the topographic term is at least of the same order of magnitude as the beta effect. The second CEOF was more broadbanded in wavenumber space, with eastward propagation occurring in a wavenumber-frequency band between 300 and 1400 km and 0.5 and 2.0 cycles/yr. The third CEOF is similar in structure to the first, but with less energy. CEOF 4 was clearly identifiable with higher frequencies than the first three with westward propagation at 4 km/d. The spatial location of this mode along with the westward propagation indicates possible influences from eddy-stream interactions. Thus topography, Rossby wave dynamics and eddy-stream interactions all appear to have a significant role in determining the space-time scales and propagation properties of the long-period response of sea level in the Gulf Stream.

Vazquez, Jorge↗

Dashboard for Visualizing Molecular Property Prediction Machine Learning Results

This is a dashboard for exploring the results of machine learning models for predicting molecular properties from molecular structure. It includes tools for: 1. Modifying molecules to observe the change in predicted properties 2. Exploring the relationship between molecular structure and predicted properties 3. Recommending structurally similar molecules with improved properties 4. Exploring the impact of data subsampling on model performance metrics

Xu, Audrey↗

Velocity dispersions for X-ray-emitting clusters of galaxies

Combining spectroscopy for five clusters of galaxies reported to contain X-ray sources with previously published data for 21 X-ray clusters, suggested correlations between the cluster's velocity dispersions and their X-ray properties have been tested. Unlike previous investigations, it is found that, for all reasonable data subsets, the cluster radial velocity dispersions sigma(r) are correlated with the cluster X-ray luminosities L/x/ at a confidence level exceeding 99%. The best-fit slope of the (log sigma/r/, log L/x/)-relation is somewhat larger than theoretically predicted, but accurate determination of that relation requires further X-ray and optical observations. For the 'most reliable' data subsample a correlation between sigma(r) and the X-ray source temperatures is also found but at a much lower confidence level (85%) than derived by previous investigators from smaller samples.

Hintzen, P.↗

Errors of five-day mean surface wind and temperature conditions due to inadequate sampling

Surface meteorological reports of wind components, wind speed, air temperature, and sea-surface temperature from buoys located in equatorial and midlatitude regions are used in a simulation of random sampling to determine errors of the calculated means due to inadequate sampling. Subsampling the data with several different sample sizes leads to estimates of the accuracy of the subsampled means. The number N of random observations needed to compute mean winds with chosen accuracies of 0.5 (N sub 0.5) and 1.0 (N sub 1,0) m/s and mean air and sea surface temperatures with chosen accuracies of 0.1 (N sub 0.1) and 0.2 (N sub 0.2) C were calculated for each 5-day and 30-day period in the buoy datasets. Mean values of N for the various accuracies and datasets are given. A second-order polynomial relation is established between N and the variability of the data record. This relationship demonstrates that for the same accuracy, N increases as the variability of the data record increases. The relationship is also independent of the data source. Volunteer-observing ship data do not satisfy the recommended minimum number of observations for obtaining 0.5 m/s and 0.2 C accuracy for most locations. The effect of having remotely sensed data is discussed.

Legler, David M.↗

Onboard Science and Applications Algorithm for Hyperspectral Data Reduction

An onboard processing mission concept is under development for a possible Direct Broadcast capability for the HyspIRI mission, a Hyperspectral remote sensing mission under consideration for launch in the next decade. The concept would intelligently spectrally and spatially subsample the data as well as generate science products onboard to enable return of key rapid response science and applications information despite limited downlink bandwidth. This rapid data delivery concept focuses on wildfires and volcanoes as primary applications, but also has applications to vegetation, coastal flooding, dust, and snow/ice applications. Operationally, the HyspIRI team would define a set of spatial regions of interest where specific algorithms would be executed. For example, known coastal areas would have certain products or bands downlinked, ocean areas might have other bands downlinked, and during fire seasons other areas would be processed for active fire detections. Ground operations would automatically generate the mission plans specifying the highest priority tasks executable within onboard computation, setup, and data downlink constraints. The spectral bands of the TIR (thermal infrared) instrument can accurately detect the thermal signature of fires and send down alerts, as well as the thermal and VSWIR (visible to short-wave infrared) data corresponding to the active fires. Active volcanism also produces a distinctive thermal signature that can be detected onboard to enable spatial subsampling. Onboard algorithms and ground-based algorithms suitable for onboard deployment are mature. On HyspIRI, the algorithm would perform a table-driven temperature inversion from several spectral TIR bands, and then trigger downlink of the entire spectrum for each of the hot pixels identified. Ocean and coastal applications include sea surface temperature (using a small spectral subset of TIR data, but requiring considerable ancillary data), and ocean color applications to track biological activity such as harmful algal blooms. Measuring surface water extent to track flooding is another rapid response product leveraging VSWIR spectral information.

Chien, Steve A.↗

Portable Airborne Laser System Measures Forest-Canopy Height

(PALS) is a combination of laser ranging, video imaging, positioning, and data-processing subsystems designed for measuring the heights of forest canopies along linear transects from tens to thousands of kilometers long. Unlike prior laser ranging systems designed to serve the same purpose, the PALS is not restricted to use aboard a single aircraft of a specific type: the PALS fits into two large suitcases that can be carried to any convenient location, and the PALS can be installed in almost any local aircraft for hire, thereby making it possible to sample remote forests at relatively low cost. The initial cost and the cost of repairing the PALS are also lower because the PALS hardware consists mostly of commercial off-the-shelf (COTS) units that can easily be replaced in the field. The COTS units include a laser ranging transceiver, a charge-coupled-device camera that images the laser-illuminated targets, a differential Global Positioning System (dGPS) receiver capable of operation within the Wide Area Augmentation System, a video titler, a video cassette recorder (VCR), and a laptop computer equipped with two serial ports. The VCR and computer are powered by batteries; the other units are powered at 12 VDC from the 28-VDC aircraft power system via a low-pass filter and a voltage converter. The dGPS receiver feeds location and time data, at an update rate of 0.5 Hz, to the video titler and the computer. The laser ranging transceiver, operating at a sampling rate of 2 kHz, feeds its serial range and amplitude data stream to the computer. The analog video signal from the CCD camera is fed into the video titler wherein the signal is annotated with position and time information. The titler then forwards the annotated signal to the VCR for recording on 8-mm tapes. The dGPS and laser range and amplitude serial data streams are processed by software that displays the laser trace and the dGPS information as they are fed into the computer, subsamples the laser range and amplitude data, interleaves the subsampled data with the dGPS information, and records the resulting interleaved data stream.

Nelson, Ross↗

Determination of Local Slope on the Greenland Ice Sheet Using a Multibeam Photon-Counting Lidar in Preparation for the ICESat-2 Mission

The greatest changes in elevation in Greenland and Antarctica are happening along the margins of the ice sheets where the surface frequently has significant slopes. For this reason, the upcoming Ice, Cloud, and land Elevation Satellite-2 (ICESat-2) mission utilizes pairs of laser altimeter beams that are perpendicular to the flight direction in order to extract slope information in addition to elevation. The Multiple Altimeter Beam Experimental Lidar (MABEL) is a high-altitude airborne laser altimeter designed as a simulator for ICESat-2. The MABEL design uses multiple beams at fixed angles and allows for local slope determination. Here, we present local slopes as determined by MABEL and compare them to those determined by the Airborne Topographic Mapper (ATM) over the same flight lines in Greenland. We make these comparisons with consideration for the planned ICESat-2 beam geometry. Results indicate that the mean slope residuals between MABEL and ATM remain small (< 0.05 degrees) through a wide range of localized slopes using ICESat-2 beam geometry. Furthermore, when MABEL data are subsampled by a factor of 4 to mimic the planned ICESat-2 transmit-energy configuration, the results are indistinguishable from the full-data-rate analysis. Results from MABEL suggest that ICESat-2 beam geometry and transmit-energy configuration are appropriate for the determination of slope on approx. 90-m spatial scales, a measurement that will be fundamental to deconvolving the effects of surface slope from the ice-sheet surface change derived from ICESat-2.

Photon-Counting↗

Determination of Local Slope on the Greenland Ice Sheet Using a Multibeam Photon-Counting Lidar in Preparation for the ICESat-2 Mission

The greatest changes in elevation in Greenland and Antarctica are happening along the margins of the ice sheets where the surface frequently has significant slopes. For this reason, the upcoming Ice, Cloud, and land Elevation Satellite-2 (ICESat-2) mission utilizes pairs of laser altimeter beams that are perpendicular to the flight direction in order to extract slope information in addition to elevation. The Multiple Altimeter Beam Experimental Lidar (MABEL) is a high-altitude airborne laser altimeter designed as a simulator for ICESat-2. The MABEL design uses multiple beams at fixed angles and allows for local slope determination. Here, we present local slopes as determined by MABEL and compare them to those determined by the Airborne Topographic Mapper (ATM) over the same flight lines in Greenland. We make these comparisons with consideration for the planned ICESat-2 beam geometry. Results indicate that the mean slope residuals between MABEL and ATM remain small (< 0.05◦) through a wide range of localized slopes using ICESat-2 beam geometry. Furthermore, when MABEL data are subsampled by a factor of 4 to mimic the planned ICESat-2 transmit-energy configuration, the results are indistinguishable from the full-data-rate analysis. Results from MABEL suggest that ICESat-2 beam geometry and transmit-energy configuration are appropriate for the determination of slope on ∼90-m spatial scales, a measurement that will be fundamental to deconvolving the effects of surface slope from the ice-sheet surface change derived from ICESat-2.

Greenland↗

Framework of compressive sensing and data compression for 4D-STEM

Four-dimensional Scanning Transmission Electron Microscopy (4D-STEM) is a powerful technique for high-resolution and high-precision materials characterization at multiple length scales, including the characterization of beam-sensitive materials. However, the field of view of 4D-STEM is relatively small, which in absence of live processing is limited by the data size required for storage. Furthermore, the rectilinear scan approach currently employed in 4D-STEM places a resolution- and signal-dependent dose limit for the study of beam sensitive materials. Improving 4D-STEM data and dose efficiency, by keeping the data size manageable while limiting the amount of electron dose, is thus critical for broader applications. Here we introduce a general method for reconstructing 4D-STEM data with subsampling in both real and reciprocal spaces at high fidelity. The approach is first tested on the subsampled datasets created from a full 4D-STEM dataset, and then demonstrated experimentally using random scan in real-space. The same reconstruction algorithm can also be used for compression of 4D-STEM datasets, leading to a large reduction (100 times or more) in data size, while retaining the fine features of 4D-STEM imaging, for crystalline samples.

4D-STEM↗

A Comprehensive Machine Learning Study to Classify Precipitation Type over Land from Global Precipitation Measurement Microwave Imager (GPM-GMI) Measurements

Precipitation type is a key parameter used for better retrieval of precipitation characteristics as well as to understand the cloud–convection–precipitation coupling processes. Ice crystals and water droplets inherently exhibit different characteristics in different precipitation regimes (e.g., convection, stratiform), which reflect on satellite remote sensing measurements that help us distinguish them. The Global Precipitation Measurement (GPM) Core Observatory’s microwave imager (GMI) and dual-frequency precipitation radar (DPR) together provide ample information on global precipitation characteristics. As an active sensor, the DPR provides an accurate precipitation type assignment, while passive sensors such as the GMI are traditionally only used for empirical understanding of precipitation regimes. Using collocated precipitation type flags from the DPR as the “truth”, this paper employs machine learning (ML) models to train and test the predictability and accuracy of using passive GMI-only observations together with ancillary information from a reanalysis and GMI surface emissivity retrieval products. Out of six ML models, four simple ones (support vector machine, neural network, random forest, and gradient boosting) and the 1-D convolutional neural network (CNN) model are identified to produce 90–94% prediction accuracy globally for five types of precipitation (convective, stratiform, mixture, no precipitation, and other precipitation), which is much more robust than previous similar effort. One novelty of this work is to introduce data augmentation (subsampling and bootstrapping) to handle extremely unbalanced samples in each category. A careful evaluation of the impact matrices demonstrates that the polarization difference (PD), brightness temperature (Tc) and surface emissivity at high-frequency channels dominate the decision process, which is consistent with the physical understanding of polarized microwave radiative transfer over different surface types, as well as in snow and liquid clouds with different microphysical properties. Furthermore, the view-angle dependency artifact that the DPR’s precipitation flag bears with does not propagate into the conical-viewing GMI retrievals. This work provides a new and promising way for future physics-based ML retrieval algorithm development.

machine learning/artificial intelligence↗

A genome-scale phylogeny of the kingdom Fungi

Phylogenomic studies using genome-scale amounts of data have greatly improved understanding of the tree of life. Despite the diversity, ecological significance, and biomedical and industrial importance of fungi, evolutionary relationships among several major lineages remain poorly resolved, especially those near the base of the fungal phylogeny. To examine poorly resolved relationships and assess progress toward a genome-scale phylogeny of the fungal kingdom, we compiled a phylogenomic data matrix of 290 genes from the genomes of 1,644 species that includes representatives from most major fungal lineages. We also compiled 11 data matrices by subsampling genes or taxa from the full data matrix based on filtering criteria previously shown to improve phylogenomic inference. Analyses of these 12 data matrices using concatenation- and coalescent-based approaches yielded a robust phylogeny of the fungal kingdom, in which ~85% of internal branches were congruent across data matrices and approaches used. We found support for several historically poorly resolved relationships as well as evidence for polytomies likely stemming from episodes of ancient diversification. By examining the relative evolutionary divergence of taxonomic groups of equivalent rank, we found that fungal taxonomy is broadly aligned with both genome sequence divergence and divergence time but also identified lineages where current taxonomic circumscription does not reflect their levels of evolutionary divergence. Finally, our results provide a robust phylogenomic framework to explore the tempo and mode of fungal evolution and offer directions for future fungal phylogenetic and taxonomic studies.

59 BASIC BIOLOGICAL SCIENCES↗

DESI Data Release 2 ELGs: Property-dependent subsamples, imaging systematics, and clustering

Using emission-line galaxies (ELGs) from the Dark Energy Spectroscopic Instrument (DESI) Data Release 2, we evaluate a property-dependent correction to imaging systematics. We derive systematic weights following the same linear regression method used for other DESI tracers, but do so separately on ELG subsamples to provide a physically-informed alternative to the fiducial, neural-network-based approach. In doing so, we show that the deeper imaging in the Dark Energy Survey (DES) footprint leads to a higher overall number density but a lack of targets with extreme $g-r$ and $r-z$ colors. ELGs in the DES region also show a distinct redshift distribution when subsampled by position in the $g-r$ vs. $r-z$ plane. To address these effects, we implement a separate treatment of the DES footprint within the DESI catalog production pipeline, which is generally well-motivated and, in some cases, imperative for accurate clustering measurements. With DES treated separately, we find that property-dependent systematic weights further mitigate spurious clustering signal in $\sim$10% of subsamples, while the fiducial scheme remains optimal for the full sample.

Hagen, T. [Utah U.]↗

The Binary Fraction of Stars in the Dwarf Galaxy Ursa Minor via Dark Energy Spectroscopic Instrument

We utilize multi-epoch line-of-sight velocity measurements from the Milky Way Survey of the Dark Energy Spectroscopic Instrument to estimate the binary fraction for member stars in the dwarf spheroidal galaxy Ursa Minor. Our dataset comprises 670 distinct member stars, with a total of more than 2,000 observations collected over approximately one year. We constrain the binary fraction for UMi to be $0.61^{+0.16}_{-0.20}$ and $0.69^{+0.19}_{-0.17}$, with the binary orbital parameter distributions based on solar neighborhood observation from Duquennoy & Mayor (1991) and Moe & Di Stefano (2017), respectively. Furthermore, by dividing our data into two subsamples at the median metallicity, we identify that the binary fraction for the metal-rich ([Fe/H]>-2.14) population is slightly higher than that of the metal-poor ([Fe/H]<-2.14) population. Based on the Moe & Di Stefano model, the best-constrained binary fractions for metal-rich and metal-poor populations in UMi are $0.86^{+0.14}_{-0.24}$ and $0.48^{+0.26}_{-0.19}$, respectively. After a thorough examination, we find that this offset cannot be attributed to sample selection effects. We also divide our data into two subsamples according to their projected radius to the center of UMi, and find that the more centrally concentrated population in a denser environment has a lower binary fraction of $0.33^{+0.30}_{-0.20}$, compared with $1.00^{+0.00}_{-0.32}$ for the subsample in more outskirts.

Qiu, Tian [Shanghai Jiao Tong Univ. (China); DESI ↗