Search NASA⌕ Search

SEARCH · Search NASA

Results for “data set”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

pixelvar79/ESGAN-Flowering-Detection-paper

Machine learning (ML) can accelerate biological research. However, the adoption of such tools to facilitate phenotyping based on sensor data has been limited by (i) the need for a large amount of human-annotated training data for each context in which the tool is used and (ii) phenotypes varying across contexts defined in terms of genetics and environment. This is a major bottleneck because acquiring training data is generally costly and time-consuming. This study demonstrates how a ML approach can address these challenges by minimizing the amount of human supervision needed for tool building. A case study was performed to compare ML approaches that examine images collected by an uncrewed aerial vehicle to determine the presence/absence of panicles (i.e. “heading”) across thousands of field plots containing genetically diverse breeding populations of 2 Miscanthus species. Automated analysis of aerial imagery enabled the identification of heading approximately 9 times faster than in-field visual inspection by humans. Leveraging an Efficiently Supervised Generative Adversarial Network (ESGAN) learning strategy reduced the requirement for human-annotated data by 1 to 2 orders of magnitude compared to traditional, fully supervised learning approaches. The ESGAN model learned the salient features of the data set by using thousands of unlabeled images to inform the discriminative ability of a classifier so that it required minimal human-labeled training data. This method can accelerate the phenotyping of heading date as a measure of flowering time in Miscanthus across diverse contexts (e.g. in multistate trials) and opens avenues to promote the broad adoption of ML tools.

Varela, Sebastian↗

Performance Year 1 Technical Report - OPEN COG Grid: Extendable Coherent Models-Datasets for Cognitive Power Grids

The OPEN COG Grid project is a collaborative effort between LLNL, NREL, and Texas A&M University (TAMU) to develop synthetic power system datasets that (i) contain all technical information that would be available in a real system, allowing to conduct studies ranging from dynamic simulation to long term planning studies; ii) are accessible to researchers from the broader data sciences community, as oppossed to power system experts only; and (iii) This report summarizes the work conducted during the first 15 months of execution of the project. These activities encompassed: 1. Conduct a survey of existing open data sets and open source power systems simulators, their supported use cases, and accessibility (Chapter 1). 2. Define a new extensible specification for power system data, covering all parameters necessary for most computational use cases (Chapter 2). 3. Collecting real technical system data to complete missing parameters in existing open source datasets (Chapter 3). 4. Develop models that capture the behavior of emergent actors in power grids, neglected by existing datasets; aggregated residential demand response (Chapter 4) and demand response of cryptocurrency miners (Chapter 5). 5. Collect detailed spatial information on distributed energy resources, particular, solar photovoltaic facilities (Chapter 6). The following chapters provide detailed descriptions of these tasks, the assumptions taken, and their findings. In conducting these tasks, the project team produced: two (accepted) conference papers; one journal paper under submission; one draft journal paper pending submission; released one repository with the developed power system data specification, with documentation and examples; and one extended dataset for the Texas power grid under review for release. The team hopes these contributions will enhance access to power system data and remove barriers to the development of new computational techniques for power systems, particularly, those inspired by cognitive sciences.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Multitaper Magnitude‐Squared Coherence for Time Series With Missing Data: Understanding Oscillatory Processes Traced by Multiple Observables

To explore the hypothesis of a common source of variability in two time series, observers may estimate the magnitude-squared coherence (MSC), which is a frequency-domain view of the cross correlation. For time series that do not have uniform observing cadence, MSC can be estimated using Welch's overlapping segment averaging. However, multitaper has superior statistical properties to Welch's method in terms of the tradeoff between bias, variance, and bandwidth. The classical multitaper technique has recently been extended to accommodate time series with underlying uniform observing cadence from which some observations are missing. This situation is common for solar and geomagnetic data sets, which may have gaps due to breaks in satellite coverage, instrument downtime, or poor observing conditions. We demonstrate the scientific use of missing-data multitaper magnitude-squared coherence by detecting known solar mid-term oscillations in simultaneous, missing-data time series of solar Lyman α flux and geomagnetic Disturbance Storm Time index. Due to their superior statistical properties, we recommend that multitaper methods be used for all heliospheric time series with underlying uniform observing cadence.

Astro-statistics techniques (1886)↗

Bioenergy Feedstock Library Annual Summary Report 2024

The Bioenergy Feedstock Library (BFL), part of the Biomass Feedstock National User Facility (BFNUF) located at Idaho National Laboratory (INL), is a physical sample repository and a web-accessible electronic database. The BFL stores physical and chemical characteristics of biomass and waste carbon sources for energy use, as well as samples generated from U.S. Department of Energy (DOE) Bioenergy Technologies Office (BETO) and U.S. Department of Agriculture-funded projects. The objective of this Bioenergy Feedstock Library Annual Summary Report for 2024, similar to the 2023 Annual Summary Report , is to focus on the updates to: (1) publicly available analytical data and equipment tracked through the BFNUF, (2) significant increases in the physical samples available for request, (3) sample and data archival progress from recent BETO-funded projects, and (4) publicly available data sets created upon request from BETO, INL projects, or outside entities compared to the previous annual summary reports. This report highlights key statistics and available data and information important for INL, BFL users, academics, and industry.

09 BIOMASS FUELS↗

Constituent Data Replacement Tool

The purpose of this tool is to estimate key parameters that may be missing in public wastewater composition datasets. The tool can be applied to develop complete treatment and critical mineral extraction profiles for leachate, produced water and other aqueous waste streams. The tool applies machine learning algorithms to replace missing data in a user’s water data set that are adjusted based on user preferences for options including algorithm type, number of features, and classification variables. The tool can use the user’s data alone or combine user data with the NEWTS USGS Produced Water Database for more robust training. This research was funded by the U.S. Department of Energy’s Office Fossil Energy and Carbon Management (FECM) through National Energy Technology Laboratory’s ongoing research under the Water Management for Power System Field Work Proposal, DE-FECM 1022428 and Critical Minerals Field Work Proposal, DE-FECM 1022420.

Aqueous Chemistry↗

Frameworks, Algorithms, and Scalable Technologies for Mathematics (FASTMath) SciDAC Institute

As computational models scale to larger computers, the rate at which they produce data has far outstripped the same computers ability to write that data and further the file systems ability to store that data. Almost all of the SciDAC applications, but especially those related to fusion solve very large scale PDEs whose scientific output his impacted by this problem. To gain access to dynamics in an exascale simulation that are not identifiable a priori and to make that dynamical data available to machine learning requires fundamental research in the area of in situ data data analytics. Here data analytics includes compression, visualization, uncertainty quantification, and machine learning. This in situ data analytics will enable on-the-fly spatial and temporal compression of solution dynamics, expose that space-time compressed field to machine learning algorithms that have been specialized to work with dynamically evolving data (existing machine learning algorithms treat data sets as static), greatly improving the opportunity for machine learning to provide feedback to the compression, all within an ongoing simulation, without the need to write data to files. The same concepts are also being applied to uncertainty quantification and multi-fidelity modeling which have similar needs for spatial and temporal compression of the ongoing exascale simulation to perform either without the typical, unacceptable writing of data to files.

97 MATHEMATICS AND COMPUTING↗

Probing Anisotropic Cosmic Birefringence with Foreground-Marginalised SPT B-mode Likelihoods

In this work, we construct foreground-marginalised versions of the SPT-3G D1 and SPTpol cosmic microwave background (CMB) B-mode polarisation likelihoods. The compression is performed using the CMB-lite framework and we use the resulting data sets to constrain anisotropic cosmic birefringence, parametrised by the amplitude of a scale-invariant anisotropic birefringence spectrum, A CB . Using the new SPT-3G data we report a upper limit on of 95% upper limit on A CB of 1.2 x 10 -4 , which tightens to 0.53 x 10 -4 when imposing a prior on the amplitude of gravitational lensing based on CMB lensing reconstruction analyses. These are the tightest constraints on anisotropic birefringence from BB power spectrum measurements to-date, demonstrating the constraining power of the South Pole Telescope. The likelihoods used in this work are made publicly available at https://github.com/lbalkenhol/candl_data

Balkenhol, Lennart [Sorbonne Université, Paris (Fr↗

Bridging the time scale in exascale computing of chemical systems (Final Technical Report)

This report summarizes the work carried out with support of the United States Department of Energy under Award DE-SC0019441. The theme of this project was to develop and apply methods that allowed for the acceleration of atomistic calculations, particularly in challenging areas such as multiphase systems, electrified interfaces, uncertainty estimation, and applications requiring chemical accuracy, which tend to be applications where simulation time is severely bottlenecked by the computational time requirements. Much of the focus was on the application of emerging machine-learning methodologies, although a wide range of methodologies were employed. This report has two major sections. The first focuses on the methodological advances themselves. Within this part, we report a number of major advances, a few examples of which are described here. We report the first machine-learning scheme for the acceleration of electronically grand-canonical calculations (that is, those applicable to electrochemistry). We report new methods of performing transfer learning, in which physics-based priors can be used to provide predictions, often with uncertainty estimates, of images well outside of training sets; we also offer ways to fine-tune these transfer-learning models. We provide a new systematic means to generate and apply minimal training data sets to very large (10,000’s of atoms) systems, with only small training sets appropriate for electronic structure. We developed new methodologies to integrate surface vibrations into surface adsorption calculations. We made advances to the applicability of diffusion Monte Carlo methods to allow (learned) force prediction, finite-size error correction, and force-free means of searching for transition states. We integrated machine-learned atomistic predictions into mechanism generation codes. Additionally, we released new software including AmpTorch, a modernized version of our original atomistic machine-learning code Amp. The second part of this report focuses on the scientific applications that accompanied, and were often enabled by, the methodological advances described earlier. A few examples follow, but full details are in the individual chapters of the report. For example, we developed a general theory of phonon-induced friction on molecular adsorbates. We showed fundamentally how solvent influences the adsorption and desorption process and how it differs from the processes typically involved at the solid–gas interface, making aqueous-phase and electrocatalysis different from traditional thermocatalysis. We examined how metal–insulator and magnetic transitions can be probed, and accelerated exciton dynamics via Frenkel Hamiltonian parameters. We showed that the nearsighted force-training approach, developed within this project, can predict both the stability and reactivity of large nanoparticles, and can also lead to insights on catalyst coverage on binding energies and entropies. These applied studies, which generally integrated with our method development, allowed us to push forward the theoretical understanding of several reaction classes.

08 HYDROGEN↗

CO2 and CH4 leaf-level fluxes and soil porewater concentrations from common vegetation patches in Louisiana’s coastal wetlands

This dataset contains leaf-level flux and soil porewater concentration measurements of carbon dioxide (CO2) and methane (CH4 ) in plots in the footprint of Ameriflux sites US-LA2 and US-LA3. Leaf fluxes in US-LA2 were measured on patches dominated by Sagittaria lancifolia and co-dominated by Sagittaria lancifolia and Typha latifolia. In US-LA3, fluxes were measured from distinct Juncus roemerianus and Spartina alterniflora patches. The porewater concentrations were collected across a vertical profile (~50 cm depth) at centric locations within 25 m2 plots where we measured the leaf fluxes. US-LA3 included an additional set of measurements in open water spots. We aimed to evaluate differences in leaf fluxes and porewater pools of CO2 and CH4 of representative ecohydrological patches across a salinity gradient. We also used this dataset to help develop ELM-Wet, a more realistic representation of wetland carbon biogeochemical processes within the U.S. Department of Energy’s Energy Exascale Earth System Model (E3SM) Land Model version 1 (ELM v.1). The files can be opened with regular text editors or spreadsheet programs. Version 2.0 (8/26/2025): This is the latest version of this dataset. The update includes additional samples of soil porewater CH4/CO2 concentrations from June-2021 to November-2022, as well as minor adjustments made to V1 samples via changing Henry's solubility to account for porewater salinity. Additionally leaf-level measurments of spectral indices, PSRI, NDVI, and PRI have been added to complement Leaf-level flux measurements. All V1 data sets have been integrated into V2 sheets, ensuring data from the previous version is contained with the additional samples and consistent with V2 metadata.

54 ENVIRONMENTAL SCIENCES↗

Search for vector-like leptons with long-lived particle decays in the CMS muon system in proton-proton collisions at $$\sqrt{\text{s}}$$ = 13 TeV

Abstract A first search is presented for vector-like leptons (VLLs) exclusively decaying into a light long-lived pseudoscalar boson and a standard model τ lepton. The pseudoscalar boson is assumed to have a mass below the τ + τ − threshold, so that it decays exclusively into two photons. It is identified using the CMS muon system. The analysis is carried out using a data set of proton-proton collisions at a center-of-mass energy of 13 TeV collected by the CMS experiment in 2016–2018, corresponding to an integrated luminosity of 138 fb −1. Selected events contain at least one pseudoscalar boson decaying electromagnetically in the muon system and at least one hadronically decaying τ lepton. No significant excess of data events is observed compared to the background expectation. Upper limits are set at 95% confidence level on the vector-like lepton production cross section as a function of the VLL mass and the pseudoscalar boson mean proper decay length. The observed and expected exclusion ranges of the VLL mass extend up to 700 and 670 GeV, respectively, depending on the pseudoscalar boson lifetime.

Chekhovsky, V. [Yerevan Physics Institute]↗

Search for vector-like leptons with long-lived particle decays in the CMS muon system in proton-proton collisions at $\sqrt{s}$ = 13 TeV

A first search is presented for vector-like leptons (VLLs) exclusively decaying into a light long-lived pseudoscalar boson and a standard model τ lepton. The pseudoscalar boson is assumed to have a mass below the τ + τ − threshold, so that it decays exclusively into two photons. It is identified using the CMS muon system. The analysis is carried out using a data set of proton-proton collisions at a center-of-mass energy of 13 TeV collected by the CMS experiment in 2016–2018, corresponding to an integrated luminosity of 138 fb −1 . Selected events contain at least one pseudoscalar boson decaying electromagnetically in the muon system and at least one hadronically decaying τ lepton. No significant excess of data events is observed compared to the background expectation. Upper limits are set at 95% confidence level on the vector-like lepton production cross section as a function of the VLL mass and the pseudoscalar boson mean proper decay length. The observed and expected exclusion ranges of the VLL mass extend up to 700 and 670 GeV, respectively, depending on the pseudoscalar boson lifetime.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Simultaneous inference of equation of state parameters and unknown data errors with uncertainty quantification via hierarchical Bayesian posterior maximization

Equations of state (EOSs) are a key component in running hydrodynamic simulations as they relate the thermodynamic states for the material. The Davis reactants EOS is commonly used for modeling high explosives (HEs), and the EOS model parameters are calibrated using material specific data. The calibrations are often performed with uncertainty quantification via Bayesian inference to account for uncertainty in the data and generate ensembles of likely parameters. However, there are relatively few HE data sets to use for calibration and many are historical and lack error information. In this work, we simultaneously calibrate the Davis reactants EOS model parameters and unknown data error terms for the high explosive PBX 9501. To quantify the uncertainty in the models and the data, we use a Bayesian framework for the calibration and compute the hierarchical Bayesian posterior distribution with both a posteriori maximization approach and Markov Chain Monte Carlo. In general, we find that, given our assumptions, the two approaches result in similar calibrated parameters, posterior covariance matrices, and insights about the parameters but that the posterior maximization requires far less computational resources.

97 MATHEMATICS AND COMPUTING↗

Test of lepton flavor universality in B ± → K ± μ + μ – and B ± → K ± e + e – decays in proton-proton collisions at $\sqrt{s}$ = 13 TeV

A test of lepton flavor universality in B ± → K ± μ + μ – and B ± → K ± e + e – decays, as well as a measurement of differential and integrated branching fractions of a nonresonant B ± → K ± μ + μ – decay are presented. The analysis is made possible by a dedicated data set of proton-proton collisions at $\sqrt{s}$ = 13 TeV recorded in 2018, by the CMS experiment at the LHC, using a special high-rate data stream designed for collecting about 10 billion unbiased b hadron decays. The ratio of the branching fractions B(B ± → K ± μ + μ – ) to B(B ± → K ± e + e – ) is determined from the measured double ratio R(K) of these decays to the respective branching fractions of the B ± → J/ψK ± with J/ψ → μ + μ – and e + e – decays, which allow for significant cancellation of systematic uncertainties. The ratio R(K) is measured in the range 1.1 < q 2 < 6.0 GeV 2 , where q is the invariant mass of the lepton pair, and is found to be R(K) = 0.78$^{+0.47}_{-0.23}$, in agreement with the standard model expectation R(K) ≈ 1. This measurement is limited by the statistical precision of the electron channel. The integrated branching fraction in the same q 2 range, B(B ± → K ± μ + μ – ) = (12.42 ± 0.68) x 10 -8 , is consistent with the present world-average value and has a comparable precision.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Code for the manuscript "Mori-Zwanzig Modal Decomposition"

We would like to create an open source repository in LANL's github on code written in Julia, in which we implement and extend the data-driven Mori-Zwanzig method for extracting large-scale spatio-temporal structures from data, which we call MZMD. This method is an extension of Dynamic Mode Decomposition (DMD) in which Mori-Zwanzig memory kernels are included into the associated companion matrix. In the code we would like to release, we apply MZMD to a flow over a cylinder with Reynolds number 100 rather than the much larger data set used in the associated manuscript. DMD is used extensively in the fluid dynamics community mainly for extracting large scale spatio-temporal structures (patters) from flow data. This is useful for understanding the key mechanisms that generate certain complex dynamical process relevant in engineering design. In MZMD, we improve upon DMD by adding the Mori-Zwanzig memory kernels, and show this improvement is especially important in strongly nonlinear regions of the flow.

Woodward, Michael↗

Data reduction for low energy nuclear physics experiments using data frames

Low energy nuclear physics experiments are transitioning towards fully digital data acquisition systems. Realizing the gains in flexibility afforded by these systems relies on equally flexible data reduction techniques. In this paper, methods utilizing data frames and in-memory techniques to work with data, including data from self-triggering, digital data acquisition systems, are discussed within the context of a Python package, sauce. It is shown that data frame operations can encompass common analysis needs and allow interactive data analysis. Two event building techniques, dubbed referenced and referenceless event building, are shown to provide a means to transform raw list mode data into correlated multi-detector events. These techniques are demonstrated in the analysis of two example data sets.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Antarctic ice sheet model comparison with uncurated geological constraints shows that higher spatial resolution improves deglacial reconstructions

Accurately reconstructing past changes to the shape and volume of the Antarctic ice sheet relies on the use of physically based and thus internally consistent ice sheet modeling, benchmarked against spatially limited geologic data. The challenge in model benchmarking against geologic data is diagnosing whether model-data misfits are the result of an inadequate model, inherently noisy or biased geologic data, and/or incorrect association between modeled quantities and geologic observations. In this work we address this challenge by (i) the development and use of a new model-data evaluation framework applied to an uncurated data set of geologic constraints, and (ii) nested high-spatial-resolution modeling designed to test the hypothesis that model resolution is an important limitation in matching geologic data. While previous approaches to model benchmarking employed highly curated datasets, our approach applies an automated screening and quality control algorithm to an uncurated public dataset of geochronological observations (specifically, cosmogenic-nuclide exposure-age measurements from glacial deposits in ice-free areas). This optimizes data utilization by including more geological constraints, reduces potential interpretive bias, and allows unsupervised assimilation of new data as they are collected. We also incorporate a nested model framework in which high-resolution domains are downscaled from a continent-wide ice sheet model. We highlight the application of this framework by applying these methods to a small ensemble of deglacial ice-sheet model simulations, and demonstrate that the nested approach improves the ability of model simulations to match exposure age data collected from areas of complex topography and ice flow. We develop a range of diagnostic model-data comparison metrics to provide more insight into model performance than possible from a single-valued misfit statistic, showing that different metrics capture different aspects of ice sheet deflation.

Geosciences↗

High-precision Measurement of the 16 O($n, n'γ$) Cross Section using $γ$-ray Detection in Liquid Scintillators with H 2 O and BeO Targets

The 16 O($n, n'γ$) reaction was measured at the Los Alamos Neutron Science Center white neutron source using γ-ray detection in liquid scintillators present in the upper hemisphere of the Correlated Gamma-Neutron Array for sCattering (CoGNAC). Separate measurements of this reaction were performed using H 2 O and BeO targets in successive years. The unique high energies of γ rays emitted from the 16 O($n, n'γ$) reaction facilitated a clean selection of this reaction from threshold to 9.8 MeV incident neutron energy without the need for precise measurements of the γ-ray energy or the scattered neutrons. The precise time resolution of the liquid scintillator detectors was then exploited to obtain high-resolution incident neutron energy measurements, and good agreement was obtained between the H 2 O and BeO results reported here. The dominant literature data sets for this reaction have systematic differences between them, but the present results improve upon the neutron energy resolution of earlier measurements and show important discrepancies in recent data. Finally, tentative data are also shown up to 20 MeV incident neutron energy but are potentially subject to improved understanding of the relative γ-ray and α decay branches from 16 O excited states.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

AuriDESI: mock catalogues for the DESI Milky Way Survey

The Dark Energy Spectroscopic Instrument Milky Way Survey (DESI MWS) will explore the assembly history of the Milky Way by characterizing remnants of ancient dwarf galaxy accretion events and improving constraints on the distribution of dark matter in the outer halo. We present mock catalogues that reproduce the selection criteria of MWS and the format of the final MWS data set. These catalogues can be used to test methods for quantifying the properties of stellar halo substructure and reconstructing the Milky Way’s accretion history with the MWS data, including the effects of halo-to-halo variance. The mock catalogues are based on a phase-space kernel expansion technique applied to star particles in the Auriga suite of six high-resolution lambda-cold dark matter magnetohydrodynamic zoom-in simulations. They include photometric properties (and associated errors) used in DESI target selection and the outputs of the MWS spectral analysis pipeline (radial velocity, metallicity, surface gravity, and temperature). They also include information from the underlying simulation, such as the total gravitational potential and information on the progenitors of accreted halo stars. We discuss how the subset of halo stars observable by MWS in these simulations corresponds to their true content and properties. These mock Milky Ways have rich accretion histories, resulting in a large number of substructures that span the whole stellar halo out to large distances and have substantial overlap in the space of orbital energy and angular momentum.

dynamics↗