Search NASA⌕ Search

SEARCH · Search NASA

Results for “data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Deep Learning Reconstruction of Daily Soil CO 2 Efflux Reveals Biogeochemical Insights and Reduces Annual Estimate Uncertainty Despite Limited Daily Predictability

Soil CO 2 efflux is commonly measured monthly or seasonally, leaving daily dynamics poorly resolved and contributing to global estimation uncertainty. We trained a single Long Short-Term Memory (LSTM) model to predict daily soil CO 2 efflux across 82 globally distributed sites in COSORE, with 0.2%–46.9% daily data coverage from 2003 to 2020. Despite using far fewer sites than are typically used to train a single deep learning model, with observations biased toward temperate mesic sites, the LSTM model performed well at approximately one-third of sites, reconstructed nearly 2 decades of daily efflux, and outperformed commonly used approaches for estimating daily efflux when applied to the same data set. Performance was weakest at pronounced peaks and troughs and at non-temperate sites with <1.5 years of observations and irregular data patterns. Nevertheless, annual efflux from reconstructed daily data had <40% error even at underperforming sites, substantially improving estimates derived from monthly and seasonal sampling (maximum errors of 95% and 136%, respectively). Temperature sensitivity (Q 10 ) estimated from reconstructed daily predictions closely matched estimates from daily observations, whereas Q 10 values derived from monthly or seasonal observations deviated substantially, suggesting that coarse temporal sampling may contribute to uncertainty in reported Q 10 values. Consistent daily reconstructions further enabled trend analyses for well-performing, predominantly temperate sites and showed increasing soil CO 2 efflux at most sites from 2003 to 2020, with more variable summer trends. Despite limitations, these results demonstrate the potential of LSTM models to reconstruct daily soil CO 2 efflux and reduce estimation uncertainties from sparse observations.

Smykalov, Valerie [Pennsylvania State University, ↗

DUNE – Simulation Validation of Fermilab Detector Reconstruction

DUNE (Deep Underground Neutrino Experiment) is Fermilab’s flagship international experiment designed to study neutrinos by sending an intense beam from Illinois to detectors located 1,300 kilometers away at the Sanford Underground Research Facility (SURF) in South Dakota. To prepare for such a large-scale experiment, physicists develop detailed simulations to produce mock data sets which are analyzed by the CAFAna framework. During my internship, I developed software using the CAFAna framework to analyze simulated detector data and generated plots to make data trends easier to interpret and identify patterns. My analysis has uncovered inconsistencies in reconstructed neutrino tracks, duplicated reconstructed tracks causing sporadic spikes in the data, and unnatural differences in energy levels between interaction types. These analyses help verify that the improvements to detector simulations do not introduce unintended resolution errors and ensure proper reconstruction performance, supporting DUNE’s goal of making precise neutrino measurements and advancing the Department of Energy’s mission of fundamental scientific discovery.

Vershaw, Andre [Unlisted, US, IL; Fermilab] (ORCID↗

A Tutorial Set to Prepare for Science with the Vera C. Rubin Observatory

In this poster the Rubin Observatory's Community Science team (CST) presents its current suite of tutorials, which are designed to help people make use of simulated data sets in preparation for the upcoming Legacy Survey of Space and Time (LSST). We will show examples of the tutorial contents, provide custom learning modules for different astronomical fields, and describe the online environment for data analysis (the Rubin Science Platform; RSP). We will also supply a checklist for how to obtain an RSP account and access the tutorials. All are welcome to drop by the poster or the Rubin booth in the exhibit hall with questions.

79 ASTRONOMY AND ASTROPHYSICS↗

A Tutorial Set to Prepare for Science with the Vera C. Rubin Observatory

In this poster the Rubin Observatory's Community Science team (CST) presents its current suite of tutorials, which are designed to help people make use of simulated data sets in preparation for the upcoming Legacy Survey of Space and Time (LSST). We will show examples of the tutorial contents, provide custom learning modules for different astronomical fields, and describe the online environment for data analysis (the Rubin Science Platform; RSP). We will also supply a checklist for how to obtain an RSP account and access the tutorials. All are welcome to drop by the poster or the Rubin booth in the exhibit hall with questions.

79 ASTRONOMY AND ASTROPHYSICS↗

Modeling of the metal–insulator transition temperature in alio-valently doped VO 2 through symbolic regression

The correlated semiconductor vanadium dioxide (VO 2 ) exhibits an insulator–metal transition (IMT) near room temperature, which is of interest in various device applications. Precise IMT temperature control is crucial to determine the use cases across technologies such as thermochromic windows, actuators for robots or neuronal oscillators. Doping the cation or anion sites can modulate the IMT by several tens of degrees and control hysteresis. However, modeling the effects of control parameters (e.g., doping concentration, type of dopants) is challenging due to complex experimental procedures and limited data, hindering the use of traditional data-driven machine learning approaches. Symbolic regression (SR) can bridge this gap by identifying nonlinear expressions connecting key input parameters to target properties, even with small data sets. In this work, we develop SR models to capture the IMT trends in VO 2 influenced by different dopant parameters. Using experimental data from the literature, our study reveals a dual nature of the IMT temperature with varying tungsten (W) doping concentrations. The symbolic model captures data trends and accounts for experimental variability, providing a complementary approach to first-principles calculations. Our feature-driven analysis across a broader class of dopants informs selectivity and provides qualitative insights into tuning phase transition properties valuable for neuromorphic computing and thermochromic windows.

36 MATERIALS SCIENCE↗

Light Meson Spectroscopy with GlueX and Beyond

The GlueX experiment at Jefferson Lab was specifically designed for precision studies of the light-meson spectrum. For this purpose, a photon beam with energies up to 12 GeV is directed onto a liquid hydrogen target contained within a hermetic detector with near-complete neutral and charged particle coverage. Linear polarization of the photon beam with a maximum around 9 GeV provides additional information about the production process. In 2018, the experiment completed its first phase, recording data with a total integrated luminosity above 400 pb?1. We highlight a selection of results from this world-leading data set with emphasis on the search for light hybrid mesons. In the mean time, the detector underwent significant upgrades and is currently recording data with an even higher luminosity. The future plans of the GlueX experiment to explore the meson spectrum with unprecedented precision are summarized.

Austregesilo, Alexander↗

Using FIPD and OPTD to Benchmark Metallic Fuel Performance

This report serves as an introduction, tutorial, and benchmark specification for out-of-pile tests on metallic fuel. It introduces a new user to the EBR-II legacy fuel performance test program and the fast reactor fuel performance databases built to preserve the records. It then details the information stored in each database and how to find it. A benchmark specification is included for a small set of out-of-pile tests on U-10Zr fuel to function as a tutorial demonstrating how the legacy fuel performance data sets stored in the FIPD and OPTD databases can be used together to benchmark fuel performance models for steady-state and transient performance.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Neural Networks for Prediction of Complex Chemistry in Water Treatment Process Optimization

Water chemistry plays a critical role in the design and operation of water treatment processes. Detailed chemistry modeling tools use a combination of advanced thermodynamic models and extensive databases to predict phase equilibria and reaction phenomena. The complexity and formulation of these models preclude their direct integration in equation-oriented modeling platforms, making it difficult to use their capabilities for rigorous water treatment process optimization. Neural networks (NN) can provide a pathway for integrating the predictive capability of chemistry software into equation-oriented models and enable optimization of complex water treatment processes across a broad range of conditions and process designs. Herein, we assess how NN architecture and training data impact their accuracy and use in equation-oriented water treatment models. We generate training data using PhreeqC software and determine how data generation and sample size impact the accuracy of trained NNs. The effect of NN architecture on optimization is evaluated by optimizing hypothetical black-box desalination processes using a range of feed compositions from USGS brackish water data set, tracking the number of successful optimizations, and testing the impact of initial guess on the final solution. Our results clearly demonstrate that data generation and architecture impact NN accuracy and viability for use in equation-oriented optimization problems.

Dudchenko, Alexander V↗

Multiwavelength study of OT 081: broadband modelling of a transitional blazar

ABSTRACT OT 081 is a well-known, luminous blazar that is remarkably variable in many energy bands. We present the first broadband study of the source, which includes very high energy (VHE, $E\gt $ 100 GeV) $\gamma$-ray data taken by the MAGIC (Major Atmospheric Gamma-ray Imaging Cherenkov telescopes) and H.E.S.S. (High Energy Stereoscopic System) imaging Cherenkov telescopes. The discovery of VHE $\gamma$-ray emission happened during a high state of $\gamma$-ray activity in July 2016, observed by many instruments from radio to VHE $\gamma$-rays. We identify four states of activity of the source, one of which includes VHE $\gamma$-ray emission. Variability in the VHE domain is found on daily time-scales. The intrinsic VHE spectrum can be described by a power law with index $3.27\pm 0.44_{\rm stat}\pm 0.15_{\rm sys}$ (MAGIC) and $3.39\pm 0.58_{\rm stat}\pm 0.64_{\rm sys}$ (H.E.S.S.) in the energy range of 55–300 and 120–500 GeV, respectively. The broadband emission cannot be successfully reproduced by a simple one-zone synchrotron self-Compton model. Instead, an additional external Compton component is required. We test a lepto-hadronic model that reproduces the data set well and a proton-synchrotron-dominated model that requires an extreme proton luminosity. Emission models that are able to successfully represent the data place the emitting region well outside of the broad-line region to a location at which the radiative environment is dominated by the infrared thermal radiation field of the dusty torus. In the scenario described by this flaring activity, the source appears to be a flat spectrum radio quasar (FSRQ), in contrast with past categorizations. This suggests that the source can be considered to be a transitional blazar, intermediate between BL Lac and FSRQ objects.

Abe, H.↗

Hawai‘i Supernova Flows: a peculiar velocity survey using over a Thousand Supernovae in the near-infrared

ABSTRACT We introduce the Hawai‘i Supernova Flows project and present summary statistics of the first 1217 astronomical transients observed, 668 of which are spectroscopically classified Type Ia Supernovae (SNe Ia). Our project is designed to obtain systematics-limited distances to SNe Ia while consuming minimal dedicated observational resources. To date, we have performed almost 5000 near-infrared (NIR) observations of astronomical transients and have obtained spectra for over 200 host galaxies lacking published spectroscopic redshifts. In this survey paper, we describe the methodology used to select targets, collect/reduce data, calculate distances, and perform quality cuts. We compare our methods to those used in similar studies, finding general agreement or mild improvement. Our summary statistics include various parametrizations of dispersion in the Hubble diagrams produced using fits to several commonly used SN Ia models. We find the lowest dispersions using the SNooPy package’s EBV_model2, with a root mean square deviation of 0.165 mag and a normalized median absolute deviation of 0.123 mag. The full utility of the Hawai‘i Supernova Flows data set far exceeds the analyses presented in this paper. Our photometry will provide a valuable test bed for models of SN Ia incorporating NIR data. Differential cosmological studies comparing optical samples and combined optical and NIR samples will have increased leverage for constraining chromatic effects like dust extinction. We invite the community to explore our data by making the light curves, fits, and host galaxy redshifts publicly accessible.

Do, Aaron (ORCID:0000000334297845)↗

Breaking the barrier of human-annotated training data for machine learning-aided plant research using aerial imagery

Machine learning (ML) can accelerate biological research. However, the adoption of such tools to facilitate phenotyping based on sensor data has been limited by (i) the need for a large amount of human-annotated training data for each context in which the tool is used and (ii) phenotypes varying across contexts defined in terms of genetics and environment. This is a major bottleneck because acquiring training data is generally costly and time-consuming. This study demonstrates how a ML approach can address these challenges by minimizing the amount of human supervision needed for tool building. A case study was performed to compare ML approaches that examine images collected by an uncrewed aerial vehicle to determine the presence/absence of panicles (i.e. “heading”) across thousands of field plots containing genetically diverse breeding populations of 2 Miscanthus species. Automated analysis of aerial imagery enabled the identification of heading approximately 9 times faster than in-field visual inspection by humans. Leveraging an Efficiently Supervised Generative Adversarial Network (ESGAN) learning strategy reduced the requirement for human-annotated data by 1 to 2 orders of magnitude compared to traditional, fully supervised learning approaches. The ESGAN model learned the salient features of the data set by using thousands of unlabeled images to inform the discriminative ability of a classifier so that it required minimal human-labeled training data. This method can accelerate the phenotyping of heading date as a measure of flowering time in Miscanthus across diverse contexts (e.g. in multistate trials) and opens avenues to promote the broad adoption of ML tools.

59 BASIC BIOLOGICAL SCIENCES↗

pixelvar79/ESGAN-Flowering-Detection-paper

Machine learning (ML) can accelerate biological research. However, the adoption of such tools to facilitate phenotyping based on sensor data has been limited by (i) the need for a large amount of human-annotated training data for each context in which the tool is used and (ii) phenotypes varying across contexts defined in terms of genetics and environment. This is a major bottleneck because acquiring training data is generally costly and time-consuming. This study demonstrates how a ML approach can address these challenges by minimizing the amount of human supervision needed for tool building. A case study was performed to compare ML approaches that examine images collected by an uncrewed aerial vehicle to determine the presence/absence of panicles (i.e. “heading”) across thousands of field plots containing genetically diverse breeding populations of 2 Miscanthus species. Automated analysis of aerial imagery enabled the identification of heading approximately 9 times faster than in-field visual inspection by humans. Leveraging an Efficiently Supervised Generative Adversarial Network (ESGAN) learning strategy reduced the requirement for human-annotated data by 1 to 2 orders of magnitude compared to traditional, fully supervised learning approaches. The ESGAN model learned the salient features of the data set by using thousands of unlabeled images to inform the discriminative ability of a classifier so that it required minimal human-labeled training data. This method can accelerate the phenotyping of heading date as a measure of flowering time in Miscanthus across diverse contexts (e.g. in multistate trials) and opens avenues to promote the broad adoption of ML tools.

Varela, Sebastian↗

Performance Year 1 Technical Report - OPEN COG Grid: Extendable Coherent Models-Datasets for Cognitive Power Grids

The OPEN COG Grid project is a collaborative effort between LLNL, NREL, and Texas A&M University (TAMU) to develop synthetic power system datasets that (i) contain all technical information that would be available in a real system, allowing to conduct studies ranging from dynamic simulation to long term planning studies; ii) are accessible to researchers from the broader data sciences community, as oppossed to power system experts only; and (iii) This report summarizes the work conducted during the first 15 months of execution of the project. These activities encompassed: 1. Conduct a survey of existing open data sets and open source power systems simulators, their supported use cases, and accessibility (Chapter 1). 2. Define a new extensible specification for power system data, covering all parameters necessary for most computational use cases (Chapter 2). 3. Collecting real technical system data to complete missing parameters in existing open source datasets (Chapter 3). 4. Develop models that capture the behavior of emergent actors in power grids, neglected by existing datasets; aggregated residential demand response (Chapter 4) and demand response of cryptocurrency miners (Chapter 5). 5. Collect detailed spatial information on distributed energy resources, particular, solar photovoltaic facilities (Chapter 6). The following chapters provide detailed descriptions of these tasks, the assumptions taken, and their findings. In conducting these tasks, the project team produced: two (ac&#x2;cepted) conference papers; one journal paper under submission; one draft journal paper pending submission; released one repository with the developed power system data specifi&#x2;cation, with documentation and examples; and one extended dataset for the Texas power grid under review for release. The team hopes these contributions will enhance access to power system data and remove barriers to the development of new computational techniques for power systems, particularly, those inspired by cognitive sciences.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Multitaper Magnitude‐Squared Coherence for Time Series With Missing Data: Understanding Oscillatory Processes Traced by Multiple Observables

To explore the hypothesis of a common source of variability in two time series, observers may estimate the magnitude-squared coherence (MSC), which is a frequency-domain view of the cross correlation. For time series that do not have uniform observing cadence, MSC can be estimated using Welch's overlapping segment averaging. However, multitaper has superior statistical properties to Welch's method in terms of the tradeoff between bias, variance, and bandwidth. The classical multitaper technique has recently been extended to accommodate time series with underlying uniform observing cadence from which some observations are missing. This situation is common for solar and geomagnetic data sets, which may have gaps due to breaks in satellite coverage, instrument downtime, or poor observing conditions. We demonstrate the scientific use of missing-data multitaper magnitude-squared coherence by detecting known solar mid-term oscillations in simultaneous, missing-data time series of solar Lyman α flux and geomagnetic Disturbance Storm Time index. Due to their superior statistical properties, we recommend that multitaper methods be used for all heliospheric time series with underlying uniform observing cadence.

Astro-statistics techniques (1886)↗

Bioenergy Feedstock Library Annual Summary Report 2024

The Bioenergy Feedstock Library (BFL), part of the Biomass Feedstock National User Facility (BFNUF) located at Idaho National Laboratory (INL), is a physical sample repository and a web-accessible electronic database. The BFL stores physical and chemical characteristics of biomass and waste carbon sources for energy use, as well as samples generated from U.S. Department of Energy (DOE) Bioenergy Technologies Office (BETO) and U.S. Department of Agriculture-funded projects. The objective of this Bioenergy Feedstock Library Annual Summary Report for 2024, similar to the 2023 Annual Summary Report , is to focus on the updates to: (1) publicly available analytical data and equipment tracked through the BFNUF, (2) significant increases in the physical samples available for request, (3) sample and data archival progress from recent BETO-funded projects, and (4) publicly available data sets created upon request from BETO, INL projects, or outside entities compared to the previous annual summary reports. This report highlights key statistics and available data and information important for INL, BFL users, academics, and industry.

09 BIOMASS FUELS↗

Constituent Data Replacement Tool

The purpose of this tool is to estimate key parameters that may be missing in public wastewater composition datasets. The tool can be applied to develop complete treatment and critical mineral extraction profiles for leachate, produced water and other aqueous waste streams. The tool applies machine learning algorithms to replace missing data in a user’s water data set that are adjusted based on user preferences for options including algorithm type, number of features, and classification variables. The tool can use the user’s data alone or combine user data with the NEWTS USGS Produced Water Database for more robust training. This research was funded by the U.S. Department of Energy’s Office Fossil Energy and Carbon Management (FECM) through National Energy Technology Laboratory’s ongoing research under the Water Management for Power System Field Work Proposal, DE-FECM 1022428 and Critical Minerals Field Work Proposal, DE-FECM 1022420.

Aqueous Chemistry↗

Frameworks, Algorithms, and Scalable Technologies for Mathematics (FASTMath) SciDAC Institute

As computational models scale to larger computers, the rate at which they produce data has far outstripped the same computers ability to write that data and further the file systems ability to store that data. Almost all of the SciDAC applications, but especially those related to fusion solve very large scale PDEs whose scientific output his impacted by this problem. To gain access to dynamics in an exascale simulation that are not identifiable a priori and to make that dynamical data available to machine learning requires fundamental research in the area of in situ data data analytics. Here data analytics includes compression, visualization, uncertainty quantification, and machine learning. This in situ data analytics will enable on-the-fly spatial and temporal compression of solution dynamics, expose that space-time compressed field to machine learning algorithms that have been specialized to work with dynamically evolving data (existing machine learning algorithms treat data sets as static), greatly improving the opportunity for machine learning to provide feedback to the compression, all within an ongoing simulation, without the need to write data to files. The same concepts are also being applied to uncertainty quantification and multi-fidelity modeling which have similar needs for spatial and temporal compression of the ongoing exascale simulation to perform either without the typical, unacceptable writing of data to files.

97 MATHEMATICS AND COMPUTING↗