Search NASASearch

SEARCH · Search NASA

Results for “data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

A cross-dimensional analysis of data-driven short-term load forecasting methods with large-scale smart meter data

Electricity load forecasting is essential to utility operation and power grid stability. A wide spectrum of data-driven methods, ranging from linear regression models to more recent deep learning models have been adopted to forecast electric load over the years. However, there still lacks a holistic evaluation of the applicability of conventional statistical and machine learning based algorithms with respect to different temporal and spatial scopes, computational requirements, and sensitivity of model-tuning. Enabled by a large-scale electricity load profile dataset of over 40,000 residential customers in a utility region, we conducted a cross-dimensional analysis of data-driven load forecasting methods. Three regression-based and seven deep learning algorithms with different model configurations were evaluated in terms of their overall and peak load prediction accuracy, and training burdens, across spatial aggregation levels ranging from the transformer, feeder, substation, to neighborhood. We found, first, the load forecasting accuracy is constrained by a predictability boundary, influenced by the forecasting horizon and spatial aggregation level. Specifically, RandomForest, XGBoost, TFT, TSMixer, and TiDE models achieved less than 10 % prediction error for up to 96-h ahead forecasting for district, substation, and feeder levels, while other models struggle at long-horizon predictions; Second, for winter and summer peak load dates, most models were able to predict the peak demand timing within ± 1 h, but the prediction percentage error varied by models, with TFT and TiDE models being the top performers; Third, models with similar prediction accuracy can differ in training burden by an order of magnitude. Therefore, choosing model configurations that balance prediction performance and computational resource is an important practical consideration for large-scale deployment of the machine learning based load forecasting. The outcome of this study can guide researchers and practitioners to choose the proper load forecasting algorithms based on their problem scope, required accuracy, and available resources. The predictability boundary can serve as a benchmark for electricity load forecasting problems with new algorithms and datasets.

Li, Han

SIMS Data Correction Procedure for Quasi‐Simultaneous Arrival ( QSA ) Under‐counting and Ramifications of Misapplication

Isotope geochemistry requires isotope ratios measured using secondary ion mass spectrometry (SIMS) to be made with optimal precision and accuracy. Under some analytical conditions when using electron multiplier detectors, secondary ions may be under‐counted because of quasi‐simultaneous arrival (QSA) at the first dynode. The relative magnitude of the associated QSA correction to raw measured isotopic ratios can be up to seventy permil or more. Therefore, not applying the correction, or misapplication of it could lead to significant inaccuracies in published isotope ratio data. Examples and ramifications of the latter are described in addition to a straightforward procedure for QSA under‐counting correction.

Geochemistry & Geophysics

Data for The Stem Cell-Type Transcriptome of Bioenergy Sorghum Reveals the Spatial Regulation of Secondary Cell Wall Networks

Bioenergy sorghum is a low-input, drought-resilient, deep-rooting annual crop that has high biomass yield potential enabling the sustainable production of biofuels, biopower, and bioproducts. Bioenergy sorghum’s 4-5 m stems account for ~80% of the harvested biomass. Stems accumulate high levels of sucrose that could be used to synthesize bioethanol and useful biopolymers if information about stem cell-type gene expression and regulation was available to enable engineering. To obtain this information, Laser Capture Microdissection (LCM) was used to isolate and collect transcriptome profiles from five major cell types that are present in stems of the sweet sorghum Wray. Transcriptome analysis identified genes with cell-type specific and cell-preferred expression patterns that reflect the distinct metabolic, transport, and regulatory functions of each cell type. Analysis of cell-type specific gene regulatory networks (GRNs) revealed that unique TF families contribute to distinct regulatory landscapes, where regulation is organized through various modes and identifiable network motifs. Cell-specific transcriptome data was combined with a stem developmental transcriptome dataset to identify the GRN that differentially activates the secondary cell wall (SCW) formation in stem xylem sclerenchyma and epidermal cells. The cell-type transcriptomic dataset provides a valuable source of information about the function of sorghum stem cell types and GRNs that will enable the engineering of bioenergy sorghum stems.

Software

Data for Propagation Method and Planting Density Influence Canopy Developmental Transition and Biomass Productivity in Miscanthus × giganteus

Understanding how establishment practices influence the mechanisms underlying Miscanthus × giganteus (miscanthus) productivity and canopy development is critical for optimizing management. Data was collected during the juvenile (2011–2013) and mature (2024) phases of a long-term field experiment established in Urbana, Illinois, to evaluate the effects of propagation method (plug propagation [PP] and rhizome propagation [RP]), planting density (1.0, 0.75, and 0.25 plants m⁻²), and nitrogen application (0 and 67 kg N ha⁻¹) on end-of-season biomass yield, tiller mass, tiller density, and tiller height. Linear regression models identified the dominant predictors of yield across stand ages and management regimes. Planting density, nitrogen (N) application, and propagation method significantly influenced early yield and canopy development. During the juvenile phase, biomass yield was driven by tiller density due to canopy expansion; in the mature phase, yield became driven by tiller mass. The PP plots produced higher tiller density than the RP plots, resulting in faster canopy closure and higher juvenile-phase yields. Rhizome-propagated (RP) plots produced lower tiller density, but individual tillers were 3.3–6.4 g tiller−1 heavier than PP tillers. After the canopy reached equilibrium, the PP and RP yields were similar because greater RP tiller mass compensated for its lower tiller density. Higher planting density resulted in greater yield and tiller density during the second year (2012), but this effect was absent from the third year (2013) onward. In the juvenile phase, N fertilization enhanced yield by 1.6–3.4 Mg ha−1. Initiating fertilization in 2013 on unfertilized plots produced biomass similar to that in fertilized plots, suggesting yield recovery in the mature phase. These findings revealed that establishment strategies, including propagation method and planting density, influence juvenile miscanthus canopy development and productivity, transitioning from tiller-density- to mass-dominated yields, but not mature phase productivity.

Miscanthus

Data for Creating Yellow Seed Camelina sativa with Enhanced Oil Accumulation by CRISPR-Mediated Disruption of Transparent Testa 8

Camelina ( Camelina sativa L.), a hexaploid member of the Brassicaceae family, is an emerging oilseed crop being developed to meet the increasing demand for plant oils as biofuel feedstocks. In other Brassicas, high oil content can be associated with a yellow seed phenotype, which is unknown for camelina. We sought to create yellow seed camelina using CRISPR/Cas9 technology to disrupt its Transparent Testa 8 (TT8) transcription factor genes and to evaluate the resulting seed phenotype. We identified three TT8 genes, one in each of the three camelina subgenomes, and obtained independent CsTT8 lines containing frameshift edits. Disruption of TT8 caused seed coat colour to change from brown to yellow reflecting their reduced flavonoid accumulation of up to 44%, and the loss of a well-organized seed coat mucilage layer. Transcriptomic analysis of CsTT8-edited seeds revealed significantly increased expression of the lipid-related transcription factors LEC1, LEC2, FUS3, and WRI1 and their downstream fatty acid synthesis-related targets. These changes caused metabolic remodelling with increased fatty acid synthesis rates and corresponding increases in total fatty acid (TFA) accumulation from 32.4% to as high as 38.0% of seed weight, and TAG yield by more than 21% without significant changes in starch or protein levels compared to parental line. These data highlight the effectiveness of CRISPR in creating novel enhanced-oil germplasm in camelina. The resulting lines may directly contribute to future net-zero carbon energy production or be combined with other traits to produce desired lipid-derived bioproducts at high yields.

Biofuels

Data for NB6 HBRR Science Design ORNL/TM-2025/3807

Data for the report (ORNL/TM-2025/3807) that describes the calculations and the Monte Carlo Ray Tracing simulations performed using the McStas package to determine the coatings and geometry for the NB-6 guide. It provides the information to inform the mechanical design, validation tests and verification that it meets the science requirements.

47 OTHER INSTRUMENTATION

Dual-Fuel Ammonia Equivalence Ratio Sweep Data

Ammonia (NH3) has garnered significant interest as an alternative fuel for meeting international emissions reduction mandates in sectors with high weight and distance requirements, such as shipping. Technical barriers and unanswered questions remain on the combustion strategies that can maximize ammonia utilization and minimize emissions. Prior research studies at the US Department of Energy’s Oak Ridge National Laboratory have shown strong performance with NH3 under dual-fuel mode using conventional diesel combustion (CDC) manifold air pressure settings. Diesel airflow was initially used to simplify retrofitting (no turbocharger modification), which resulted in air-fuel equivalence ratios (λ) greater than 1.5. To characterize potential improvements in dual-fuel NH3 combustion performance at richer in-cylinder conditions, a global λ sweep compared the use of early (E-pilot) and late (L-pilot) single diesel injections. The experiments were conducted at 1200 RPM and 12.8 ± 0.2 bar (75 % load), and λ was varied by decreasing the commanded air flow to the engine at greater than 90 % ammonia energy substitution level. A diesel injection timing sweep was conducted for both the injection strategies at fixed λ, and the timing with the lowest engine-out N2O emissions was identified. The results indicated an optimal balance between CO2,eq and thermal efficiency benefits both E-pilot and L-pilot injection strategy cases compared with CDC at a λ of 1.4. The indicated nitrogen-based emissions exhibited a strong correlation to the ratio of CA5–50 and ignition delay for L-pilot, but no apparent trend emerged for the E-pilot injection strategy at the tested boundary conditions. This dataset includes the raw experimental data and documentation of the experiment conditions and methods.

30 DIRECT ENERGY CONVERSION

EAGLE-I Power Outage Data 2025

The provided EAGLE-I historic dataset includes power outage information at the county level for 2025 at 15-minute intervals collected by the EAGLE-I program at ORNL. The data has been collected from utility's public outage maps using an ETL process. The dataset details FIPS code, county name, state name, total number of customers without power, and a date/timestamp. For detailed metadata, refer to the linked metadata DOI.

EAGLE-I

Open-Source Data for MAC-POSTS: Mobility Data Analytics Center - Prediction, Optimization, and Simulation Toolkit for Transportation Systems

MAC-POSTS (Mobility Data Analytics Center - Prediction, Optimization, and Simulation toolkit for Transportation Systems) is a toolkit for dynamic transportation network modeling. Developed by the Mobility Data Analytics Center (MAC) at Carnegie Mellon University, this package implements many classic dynamic transportation network models, as well as new models proposed by MAC members. It has served as one building block for many other models and research projects. As such, this package used to be treated as an internal research project of the MAC lab, and admittedly, the code base is messy, and the interface is hard to use. However, we are working hard to make it a generally usable and useful toolkit for dynamic transportation network modeling. We would really appreciate any feedback, comments, suggestions, or criticisms.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Technical Track on Biomass Carbon Removal and Storage (BiCRS): Mapping bioresources, phase 1 - Consistency check comparing Mission Innovation’s Data Visualization Tool for Bioresources and the Clean Energy Ministerial Biofuture Initiative Global Biomass data accessible via the US Department of Energy’s Bioenergy Knowledge Discovery Framework (KDF)

The Mission Innovation (MI) Carbon Dioxide Removal (CDR) Mission, Technical Track on Biomass Carbon Dioxide Removal and Storage (BiCRS), has produced a biomass resource database for its members. In parallel, Oak Ridge National Laboratory (ORNL) developed the International Feedstock Reporting data portal—herein referred to as the CEM Biofuture-KDF data—on behalf of the Clean Energy Ministerial Biofuture Initiative (CEM Biofuture), as a specific task under Biofuture’s 2024–25 Action Plan. This work was conducted at the request of CEM Biofuture and funded by the U.S. Department of Energy in support of that initiative, and it is hosted within DOE’s Knowledge Discovery Framework (KDF).

09 BIOMASS FUELS

U.S. Hydropower Market Report Data and Metadata (2025 update)

This database complements the U.S. Hydropower Market Report (2025 update). This update focuses on data and trends in 2024 and contextualizes this information compared to evolving high-level trends over the past 10–20 years. It contains data on U.S. hydropower (and pumped storage hydropower) development pipeline, relicenses, license surrenders, performance metrics, and supply chain.

Johnson, Megan [ORNL] (ORCID:0000000290141741)

Creating Accurate Methane Emission Inventories through Data-Driven Airborne Survey Strategies

Because natural gas emits less carbon than other fossil fuels, it holds promise as a green energy transition fuel. However, the overall carbon footprint of natural gas is significantly elevated by methane emissions that occur during its production and transmission (Cusworth et al. 2022). Methane “super-emitters,” while comprising only about 1% of sites, are responsible for the majority of oil- and gas-sourced methane emissions, making their detection and mitigation critical in reducing the climate impact of natural gas and in meeting national and global sustainability goals (Sherwin et al. 2024). Yet, despite advancements in detection, significant uncertainties remain regarding the size, frequency, and duration distributions of methane emissions (e.g., Frankenberg et al. 2016, Cusworth et al. 2022, Chen, Sherwin et al. 2022, Conrad et al. 2023, Johnson et al. 2023, Sherwin et al. 2024) underscoring the need for comprehensive emissions inventories segmented by basin across the US. Airborne surveys are well-suited for collecting data to build these comprehensive, basin-level inventories because they allow for extensive spatial coverage, and have the spatial resolution, and the sensitivity to pinpoint individual methane sources. As remote sensing technologies enable rapid basin-scale surveys, it is imperative to establish scientifically and statistically robust standards to generate reliable and actionable emissions inventories. Recent work has shown that differences in airborne sampling strategies, detection technologies, and analysis can lead to large differences between survey conclusions if not correctly accounted for (Chen et al. 2024). This elevates the importance of incorporating proper sampling and analysis techniques when designing a methane emissions monitoring campaign to produce accurate results and facilitate cross-study comparisons. In this paper, we describe a survey strategy designed using the latest conclusions from the literature to align results from different aerial surveys. We identify several sampling and analysis principles, including large sample sizes, balanced sampling across oil and gas production, careful survey area definition, and a unified protocol for analysis, to be vital to producing an unbiased estimate of basin-scale emissions. We present results from a Department of Energy-funded project that deployed this survey strategy in two understudied oil and gas- producing regions in the United States: the Haynesville Basin in Texas and Louisiana, and the Woodford Shale in the Anadarko Basin in Oklahoma.

03 NATURAL GAS