Search NASA⌕ Search

SEARCH · Search NASA

Results for “data set”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Non-linear relationships between daily temperature extremes and US agricultural yields uncovered by global gridded meteorological datasets

Global agricultural commodity markets are highly integrated among major producers. Prices are driven by aggregate supply rather than what happens in individual countries in isolation. Furthermore, estimating the effects of weather-induced shocks on production, trade patterns and prices hence requires a globally representative weather data set. Recently, two data sets that provide daily or hourly records, GMFD and ERA5-Land, became available. Starting with the US, a data rich region, we formally test whether these global data sets are as good as more fine-scaled country-specific data in explaining yields and whether they estimate similar response functions. While GMFD and ERA5-Land have lower predictive skill for US corn and soybeans yields than the fine-scaled PRISM data, they still correctly uncover the underlying non-linear temperature relationship. All specifications using daily temperature extremes under any of the weather data sets outperform models that use a quadratic in average temperature. Correctly capturing the effect of daily extremes has a larger effect than the choice of weather data. In a second step, focusing on Sub Saharan Africa, a data sparse region, we confirm that GMFD and ERA5-Land have superior predictive power to CRU, a global weather data set previously employed for modeling climate effects in the region.

54 ENVIRONMENTAL SCIENCES↗

Ab Initio-Based Bond Order Potential for Arsenene Polymorphs Developed via Hierarchical Reinforcement Learning

Arsenene, a less-explored two-dimensional material, holds the potential for applications in wearable electronics, memory devices, and quantum systems. This study introduces a bond-order potential model with Tersoff formalism, the ML-Tersoff, which leverages multireward hierarchical reinforcement learning (RL), trained on an ab initio data set. This data set covers a spectrum of properties for arsenene polymorphs, enhancing our understanding of its mechanical and thermal behaviors without the complexities of traditional models requiring multiple parameter sets. Our RL strategy utilizes decision trees coupled with a hierarchical reward strategy to accelerate convergence in high-dimensional continuous search spaces. Unlike the Stillinger-Weber approach, which demands separate formalisms for buckled and puckered forms, the ML-Tersoff model concurrently captures multiple properties of the two polymorphs by effectively representing the local environment, thereby avoiding the need for different atomic types. Here, we apply the ML model to understand the mechanical and thermal properties of the arsenene polymorphs and nanostructures. We observe an inverse relationship between the critical strain and temperature in arsenene. Thermal conductivity calculations in nanosheets show good agreement with ab initio data, reflecting a decrease in thermal conductivity attributable to increased anharmonic effects at higher temperatures. We also apply the model to predict the thermal behavior of arsenene nanotubes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Ligand-Based Compound Activity Prediction via Few-Shot Learning

Predicting the activities of new compounds against biophysical or phenotypic assays based on the known activities of one or a few existing compounds is a common goal in early stage drug discovery. This problem can be cast as a “few-shot learning” challenge, and prior studies have developed few-shot learning methods to classify compounds as active versus inactive. However, the ability to go beyond classification and rank compounds by expected affinity is more valuable. We describe Few-Shot Compound Activity Prediction (FS-CAP), a novel neural architecture trained on a large bioactivity data set to predict compound activities against an assay outside the training set, based on only the activities of a few known compounds against the same assay. Our model aggregates encodings generated from the known compounds and their activities to capture assay information and uses a separate encoder for the new compound whose activity is to be predicted. The new method provides encouraging results relative to traditional chemical-similarity-based techniques as well as other state-of-the-art few-shot learning methods in tests on a variety of ligand-based drug discovery settings and data sets.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Cambium 2024 Scenario Descriptions and Documentation

The National Renewable Energy Laboratory's (NREL's) Cambium data sets are annually released sets of simulated hourly data for a range of modeled futures of the U.S. electric sector with metrics designed to be useful for long-term decision- making. The 2024 Cambium data set is the fifth annual release. The data sets are a companion product to NREL's Standard Scenarios, which are likewise released annually and are a set of projections of how the U.S. electric sector could evolve across a suite of different potential futures, but covering more scenarios with less temporal granularity. Information about Cambium and related publications can be found at https://www.nrel.gov/analysis/cambium.html, and the Cambium data sets can be viewed and downloaded at https://scenarioviewer.nrel.gov/. In this documentation, we describe Cambium 2024's scenarios, define the metrics, and document the Cambium-specific methods for calculating those metrics.

24 POWER TRANSMISSION AND DISTRIBUTION↗

High-Fidelity, Large-Scale, Realistic Dataset Development

The final report summarizes the work performed for supporting the ARPA-E Grid Optimization Competition (Challenge 2 and Challenge 3) within the stated period. Challenge 2 For the challenge period, the main responsibility of the team is to investigate, gen- erate, and deliver parts of the data sets for the competition, based on the competition model for Challenge 2, existing data sets from Challenge 1, and data source supplied by other data set teams. Challenge 3 For the challenge period, the main responsibility of the team is to propose, create, deliver, and maintain the data format during the competition period. The data format will specify how the benchmark data will be represented and communicated to competitors. It will also specify how competitors should report back the solutions. The data format will be closely aligned with the problem formulation (maintained by the formulation team) and the solution validation process (maintained by the validation team). Our team is also responsible in investigating, generating, and delivering parts of the data sets for the competition. The data sets will be created based on the competition model for Challenge 3, existing data sets from Challenge 1 and Challenge 2, and data source supplied by other data set teams.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Oak Ridge National Laboratory EAGLE-I TM : Modeling Electric Utility County Customers for Situational Awareness

During natural hazard events (hurricanes, wildfires, earthquakes, etc.) and recent man-made events (e.g., cyber attacks), the exchange of near real-time, spatially refined data within the response community is critical. The EAGLE-I$^{TM}$ platform is one tool that facilitates this data for decision makers within the energy sector. While much information can be collected and integrated into the system directly, other pertinent data must be augmented by other derived data products to enhance the information and allow for a consistent evaluation of on-the-ground conditions. One such data set that requires the addition of other derived data is the electric utility customer outage data that is aggregated to the county level within the EAGLE-I application. Without a county customer data set, outages can only be compared on total counts, which gives greater importance to higher population outages. Including an electric utility customer data set at the county level allows for these outage counts to be converted to percent outages and brings a consistent classification of outages and equal importance to all outages. To achieve this, several available data sets were combined and spatial disaggregation techniques were employed to model customer estimates at the county scale. This paper presents the approach to produce this data for the United States and lessons learned from working with these disparate data sets. Data validation is provided, where possible, and limitations of the model and possible improvements are discussed.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Experimental data for damage mechanics simulation challenge

While there are many computational approaches for simulating damage in rock and other materials, few have been ground truth tested with either known experimental data or with blind data sets. Here, in this work, we present a bench-mark laboratory data set for a damage mechanics challenge to compare computational approaches on damage evolution in brittle-ductile materials. The samples were fabricated through additive manufacturing to produce repeatable specimens designed to fail in controlled ways. The failure was induced in the samples using a 3-point bending test to produce different Modes such as Mode I and mixed Modes including I-II, I-III and I-II-III Modes to generate a calibration data set and a blind challenge data set. Data collected included spatial and temporal measurements from traditional digital load–displacement sensors, 2D digital image correlation measurement to map surface deformations, 3D X-ray microscopy to ground-truth the crack-failure geometry, and laser profilometry to capture surface roughness. The data sets are available, on a data repository, to the community to advance computational models to improve our ability to predict damage in brittle-ductile materials.

3-point bending↗

Quality Ranking of Unary Chloride Salt Property Data Included in MSTDB-TP

Molten salt reactor developers rely on thermal property data to design, license and operate the reactors. The Molten Salt Thermal Database-Thermophysical Properties (MSTDB-TP) was established under the DOE Nuclear Energy Advanced Modeling and Simulation (NEAMS) program and is managed by Oak Ridge National Laboratory to serve as a single source of thermophysical property values measured for a wide variety of molten salt systems for use by researchers, molten salt reactor developers, and regulators. These properties include density, viscosity and thermal diffusivity and conductivity. Published measurements of molten salt properties are lacking for many salts of interest and the data that are available are often inconsistent. This creates a challenge for MSR developers when determining which property values to use when designing their reactors. It is the purpose of this work to apply a consistent ranking system to all data entries that indicates the quality of property values listed in the database. These rankings will be the technical basis for down-selections by the database developers and alert users about the quality of the available property values. MSTDB-TP collects all available property data and indicates preferred data sets or correlations. However, all available data sets are included in the database. Quality assessments and rankings are being applied to data in MSTDB-TP to provide an indication of the quality of each data set independent of consistency with other data. Previous reports detailed the ranking system that was followed and assessments of unary fluoride data sets. Documentation of the quality of data in MSTDB-TP was continued by reviewing and assessing all available sources of density, viscosity and thermal diffusivity or conductivity values for unary chloride salts in MSTDB-TP V3.0 using the same criteria.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Evaluating the 238 U PFNS Including Chi-Nu Experimental Data

This report documents an evaluation of 238 U prompt fission neutron spectra (PFNS) which is a deliverable for a FY2024 Q4 NCSP (Nuclear Criticality Safety Program) milestone. This evaluation is new; its prior input is based on extended Los Alamos and exciton models implemented in the code CoH. Experimental covariances were estimated for five experimental data sets. One of these data sets that was measured by the Chi-Nu team of LANL and LLNL. It covers the 238 U PFNS for continuous incident-neutron energies of 1–20 MeV and outgoing-neutron energies from 10 keV– 10 MeV with high precision. Contrary to Chi-Nu data, previous data sets were measured in a limited energy range. The resulting evaluated data correspond well to the experimental PFNS taken into account for the evaluation. The evaluated PFNS also produce average mean energies in agreement with associated Chi-Nu data. If one uses the new evaluated data to predict the neutron multiplication factor, k eff , of the Flattop, Flattop-Pu and BigTen ICSBEP critical assemblies (which all have thick reflectors with high percentages of 238 U), the differences of simulated values compared to those using ENDF/B-VIII.1β3 is modest (less than 25 pcm). In addition to that, the new PFNS predict on average 238 U LLNL pulsed-sphere neutron-leakage spectra slightly better than ENDF/BVIIII.0 and ENDF/B-VIII.1β3 PFNS. The differences are, however, well within the experimental uncertainties.

238U↗

Hourly PM 2.5 Estimates across California from 2018 to 2023

This study presents a new data set of hourly PM 2.5 concentrations across California from 2018 to 2023 at a three-kilometer resolution. This data set was developed by assimilating observations from PurpleAir and the U.S. EPA Air Quality System monitors into wildfire smoke forecasts from the High-Resolution Rapid Refresh Smoke (HRRR-Smoke) model using the Gridpoint Statistical Interpolation (GSI) three-dimensional variational data assimilation framework. Archived forecasts of modeled wildfire smoke PM 2.5 from HRRR-Smoke create the background field for assimilation, which is then corrected using surface observations of total PM 2.5 . The resulting reanalysis from GSI provides an estimate of total PM 2.5 that minimizes error from both the observational and the model data. Validation results indicate strong performance, with monthly R 2 values ranging from 0.73 to 0.91 across the six-year data set, comparable to other PM 2.5 data sets. Case studies are presented for three major fire events, the 2018 Camp Fire, 2019 Kincade Fire, and 2020 Lightning Complex Fires to demonstrate the data set’s fidelity in resolving plume dynamics and local exposure patterns. Root-mean-squared error averaged over each month scales with average PM 2.5 concentrations, resulting in a low error under typical conditions but higher absolute errors during extreme smoke events. This is the first long-term, hourly PM 2.5 data set of its kind for California and enables the generation of subdaily exposure metrics, such as peak hourly concentrations, exceedance durations, and time-of-day exposure peaks. The novelty and strong validation of this data set make it a compelling resource for future studies on the impact and significance of subdaily PM 2.5 exposure.

PM2.5↗

BENEFIT with Northeastern University: HVAC Hardware-in-the-Loop Experimental Testing of a Heat Pump and Air Conditioner

This dataset includes HVAC Hardware-in-the-Loop (HIL) experimental results for a single stage, SEER 16, HSPF 9.5, 3-ton single-speed air source heat pump with 15 kW of backup auxiliary heating tested in both cooling and heating mode, and a two stage, SEER 21, 2-ton central air conditioner tested in cooling mode for a set of outdoor temperatures and indoor setpoint temperatures. In addition to these tests, experimental tests focused on the operation of auxiliary heating for the heat pump for winter condition were also conducted. The laboratory experiments for transient testing of the heat pump and air conditioner were conducted using the two HIL systems in the Systems Performance Laboratory (SPL) at NREL’s Energy Systems Integration Facility (ESIF). Further information on laboratory design and capabilities of the SPL along with the architecture of HVAC HIL system can be found in: Sparn, B. F. 2018. Laboratory Resources and Techniques to Evaluate Smart Home Technology (No. NREL/CP-5500-71696). National Renewable Energy Laboratory (NREL), Golden, CO (United States). https://www.nrel.gov/docs/fy18osti/71696.pdf and the experimental setup and validation of HVAC HIL platform can be found in: Ramaraj, S. and Sparn, B. 2022. Validation of HVAC Hardware-In-the-Loop Simulation for Advanced Control Strategies in Smart Homes (No. NREL/CP-5500-82562). National Renewable Energy Lab (NREL), Golden, CO (United States). https://www.nrel.gov/docs/fy22osti/82562.pdf. These experimental results can be used to validate how we currently model the cycling behavior of heat pumps and air conditioners. Additionally, many demand response programs implement heat pump and air conditioner control by changing the thermostat set point – these data may also be used to verify our models for heat pump and air conditioner demand response control are implemented correctly. The Test_Matrix file describes all the indoor and outdoor test conditions for heat pump and air conditioner and the file names of data sets include information about the test conditions. A wide range of outdoor air temperatures were chosen to accommodate summer and winter conditions. In addition to operating the HVAC equipment with different outdoor temperatures, we also operate the system with different indoor temperature set points to represent different grid signals or different operating conditions. For cooling conditions, the baseline set point is 72°F. To represent Load Up signals, the setpoint is changed to 68°F. The Load Shed set point is 76°F. For heating conditions, the baseline set point was assumed to be 68°F. The Load add set point is 72°F and the Load shed set point is 64°F. The starting indoor temperature for cooling conditions was set ~2°F above the indoor setpoint temperature so that the equipment turned on quickly. Similarly, the initial indoor temperature was set ~2°F lower than setpoint for heating mode tests to ensure that heating began quickly. The return air temperature was assumed to be equal to the indoor setpoint temperature in all cases. The experimental data are sampled at 1-second intervals. The data from ecobee thermostat at 5-minute interval are resampled and added to the corresponding file. The content of each data set is as follows: • T_Return (C): Measured return air temperature [C] • T_Return_SP (C): Return air temperature setpoint from E+ model, sent to HIL [C] • T_Supply (C): Measured supply air temperature at evaporator outlet [C] • T_Outdoor (C): Measured outdoor air temperature [C] • T_Outdoor_SP (C): Outdoor air temperature setpoint from weather file, sent to HIL [C] • T_Indoor (C): Measured indoor air temperature [C] • T_Indoor_SP (C): Indoor air temperature setpoint from E+ model, sent to HIL [C] • Outdoor Unit Power (W): Measured power of the outdoor unit [W] • Indoor Unit Power (W): Measured power of the indoor unit [W] • Evaporator Airflow Rate (CFM): Measured evaporator or indoor unit airflow rate sent to E+ model [CFM] • Cooling/Heating Capacity (kW): Calculated cooling/heating capacity sent to E+ model [kW] • T_SP_Thermostat (C): Thermostat cooling/heating setpoint temperature [C] • T_Indoor_Thermostat (C): Thermostat indoor air temperature [C]

24 POWER TRANSMISSION AND DISTRIBUTION↗

AI-Batt (Autonomous Identification of Battery Life Models) [SWR 21-36]

Autonomous Identification of Battery Life Models (AI-Batt) AI-Batt is a MATLAB code base for developing lifetime models for batteries from accelerated aging data. The code base provides many functions for processing, visualizing, and modeling battery aging data, making the data processing, exploration, and modeling workflow substantially faster. These tools are tailored for working with battery aging data sets, which usually consist of many separate time-series for each cell, with many test conditions and possible replicates at each condition, which makes it difficult to simply process or visualize the data set. Complex modeling tasks, such as cross-validation, sensitivity analysis, and uncertainty quantification have been implemented to enable thorough statistical investigation of model predictions. Additionally, several machine-learning algorithms are implemented to autonomously identify suitable models via symbolic regression. Data processing functions automatically cast data from the struct data type, which is commonly used to store experimental data, but is not an acceptable input for most algorithms, to the table data type, which can be easily used as input to any optimization algorithm. Also, the data can be separated into time-invariant and time-variant data tables, which is helpful for exploring the data set as well as developing separate models for time-variant and time-invariant aging mechanisms. For example, in aging tests with constant temperature, temperature is a time-invariant experimental condition. Visualization tools enable plotting of data, model fits, and model simulations possible with single-line function calls, empowering data exploration of complex data sets with both time-varying and time-invariant trends. Plots can be automatically generated for the whole data set, or separated by data group (groups of test replicates) or individual data series. Data points or data series can be automatically colored by the value of a variable with a variety of color maps, and model predictions can also be colored by the value of a fit statistic. Comparisons between data sets and the predictions/simulations of different models on the same data set can be easily plotted as well. Distributions of parameter values from bootstrap resampling can be plotted to visualize the reliability of parameter estimation, or determine any correlations between parameters. Modeling tools handle the complex task of creating and parsing symbolic equations for modeling battery lifetime. Equations are parsed to grab relevant data variables, parameter values, or specified sub-models for input into optimization, evaluation, or simulation functions. Models can be optimized locally (one set of parameters for each data series), bi-level (some parameters shared across the data set), or globally (single set of parameters for all data). Functions implementing symbolic regression algorithms help users to discover effective model equations, even in poorly sampled, high-dimensional data.

Smith, Kandler [National Renewable Energy Lab. (NR↗

Measuring Neutron Polarisation in Deuteron Photo-disintegration with the CLAS Start Counter [Thesis]

Deuteron photo-disintegration (γd → γp) is a reaction that represents the simplest case in which nuclear and hadron physics models can be tested. Despite this, associated polarization analyses are limited in terms of angular coverage and energy ranges, especially in observables related to the recoil neutron. This is largely due to a lack in dedicated polarimetry equipment, and represents a roadblock in global progress to understand high-energy phenomena such as hexaquarks, and quark-gluon degrees of freedom. To address this problem, this PhD thesis pioneers a new methodology for the parasitic measurement of nucleon polarization using kinematic reconstruction of (spin-dependent) nucleon-nucleus scattering of reaction products, prior to their detection in large acceptance particle detector apparatus. Following this novel approach, which requires no dedicated polarimeter, a determination of the double polarization observable, $C^n_{x'}$, from deuteron photo-disintegration is presented, using Jefferson Lab’s CLAS detector. The analysis utilizes the (n,p) charge exchange reaction in CLAS’s "start counter" (plastic scintillator) to determine the final state neutron polarizations. The results present the first ever data for this observable above 0.7 GeV (photon beam energy) and significantly extend the angular range of the world data set. This new data is largely statistically consistent with the previous measurement of $C^n_{x'}$ by Bashkanov et al . in the overlapping energy range of 0.4-0.7 GeV. It is planned for the statistical accuracy of the presented result to be increased by the inclusion of additional data. The analysis herein serves as a key proof of concept for future applications, including a recommended similar analysis to be implemented with data from the more modern CLAS12 detector. This paves the way for a plethora of additional analyses using existing data sets that would provide crucial new constraints for hadron and nuclear physics.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Satellite Imagery of PV Site Storm Damage

"This repository contains multiple data sets focused on visible damage to photovoltaic (PV) installations following extreme weather events such as hailstorms and hurricanes. Data sets are split into two categories: the first category, the ‘manually labeled’ data, was compiled by researchers manually, and contains manually identified PV sites exposed to storms. The second data set, the ‘aggregated’ data, is a compilation of the manually labeled PV sites and deep learning-identified PV sites. The hail damage data set focuses on post-storm PV damage following a September 24, 2023 hailstorm in Austin, TX, which caused over $600 million in damages in the Austin metro area. The hurricane damage data set focuses on post-storm PV damage following Hurricanes Irma and Maria in Puerto Rico and the US Virgin Islands. Hurricanes Irma and Maria were back-to-back category 5 hurricanes, which pummeled the Caribbean and southeastern United States in September 2017, causing an estimated $115.2 billion in damages."

14 SOLAR ENERGY↗