Search NASA⌕ Search

SEARCH · Search NASA

Results for “Research data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Data for Yield from Iowa’s first commercial miscanthus fields: implications of spatial variability for productivity and sustainability beyond research plots

This dataset contains biomass yield measurements and associated vegetation index data collected from commercial Miscanthus × giganteus fields in eastern Iowa during the 2022–2023 growing seasons. The data support the analyses presented in the article: “Yield From Iowa's First Commercial Miscanthus Fields: Implications of Spatial Variability for Productivity and Sustainability Beyond Research Plots.” We collected 105 ground-truth biomass samples from four mature commercial fields (>4 years old) covering 92.81 ha. Samples were taken from 3 m² quadrats that were hand-harvested in alignment with commercial harvest timing. Stem biomass (excluding leaves) was weighed, moisture-corrected, and converted to dry-matter yield expressed in Mg DM ha⁻¹. Sampling locations were selected to capture spatial variability visible in aerial imagery and were recorded using RTK GPS. Each biomass observation was paired with vegetation indices derived from high-resolution PlanetScope satellite imagery (3 m resolution). Images were acquired throughout the growing season, and indices were calculated to evaluate their ability to predict end-of-season biomass yield. Statistical and machine learning approaches were used to identify key predictors, and a linear regression model based on end-of-July Green Normalized Difference Vegetation Index (GNDVI) was developed and evaluated. This repository includes the data used in that modeling workflow. Management practices, economic data, full imagery time series, and additional methodological details are described in the associated publication and are not included here. The dataset consists of three comma-separated value (CSV) files: 1. Combine_Groundtruth_Yield_VI_22_23.csv This file contains ground-truth biomass yield measurements and associated key vegetation index values collected during the 2022 and 2023 growing seasons. Rows: 105 observations Columns: Year — Year of observation (2022 or 2023) Field — Field location identifier Sample_number — Unique sample identifier GNDVI_End_Jul — Green Normalized Difference Vegetation Index calculated at end of July GNDVI_End_Aug — Green Normalized Difference Vegetation Index calculated at end of August NDRE_End_Aug — Normalized Difference Red Edge index calculated at end of August Biomass_Stem_Yield_MgDM/ha — Measured stem biomass yield (megagrams dry matter per hectare) 2. trainData_GNDVI.csv This file contains the subset of observations used to train the predictive relationship between July GNDVI and biomass yield. Rows: 76 observations Columns: Unnamed: 0 — Row index retained from the original data processing workflow GNDVI_End_Jul — GNDVI at end of July Stem_Yield_MgDM/ha — Observed stem biomass yield (Mg DM ha⁻¹) 3. testData_GNDVI.csv This file contains the test dataset used to evaluate model performance. Rows: 29 observations Columns: Unnamed: 0 — Row index retained from the original data processing workflow GNDVI_End_Jul — GNDVI at end of July Predicted_Yield_MgDM/ha — Model-predicted stem biomass yield (Mg DM ha⁻¹) Observed_Yield_MgDM/ha — Measured stem biomass yield (Mg DM ha⁻¹)

Potential yield, yield gap, in-field management, y↗

Towards the next generation of Geospatial Artificial Intelligence

Geospatial Artificial Intelligence (GeoAI), as the integration of geospatial studies and AI, has become one of the fastest-developing research directions in spatial data science and geography. This rapid change in the field calls for a deeper understanding of the recent developments and envision where the field is going in the near future. In this work, we provide a quantitative analysis of the GeoAI literature from the spatial, temporal, and semantic aspects. We briefly discuss the history of AI and GeoAI by highlighting some pioneering work. Then we discuss the current landscape of GeoAI by selecting five representative subdomains including remote sensing, urban computing, Earth system science, cartography, and geospatial semantics. Finally, we highlight several unique future research directions of GeoAI which are classified into two groups: GeoAI method development challenges and GeoAI Ethics challenges. Topics include heterogeneity-aware GeoAI, knowledge-guided GeoAI, spatial representation learning, geo-foundation models, fairness-aware GeoAI, privacy-aware GeoAI, as well as interpretable and explainable GeoAI. We hope our review of GeoAI’s past, present, and future is comprehensive and can enlighten the next generation of GeoAI research.

58 GEOSCIENCES↗

Deployment and Evaluation of SciStream on OLCF's Advanced Computing Ecosystem (ACE)

The growing demand for real-time analysis, experimental steering, and decision-making in scientific workflows has created a need for tightly coupled integrations between experimental facilities and high-performance computing (HPC) systems. The Department of Energy’s Integrated Research Infrastructure (IRI) initiative highlights data streaming as a key capability for enabling memory-to-memory data transfers, bypassing the limitations of traditional store-and-forward models. SciStream is a toolkit developed by researchers at Argonne National Laboratory (ANL) to support such streaming by addressing cross-domain security, delegated authentication, and application transparency. We deployed and evaluated SciStream on the Oak Ridge Leadership Computing Facility’s (OLCF) Advanced Computing Ecosystem (ACE) infrastructure, leveraging the Olivine OpenShift cluster and its high-bandwidth Data Streaming Nodes (DSNs) as gateway nodes. Our evaluation included synthetic streaming workloads derived from IRI science workflows, a streaming simulator, and integration with RabbitMQ to handle low-level messaging. This report documents the deployment process, performance evaluation, and challenges encountered, along with opportunities for future improvements.

97 MATHEMATICS AND COMPUTING↗

AmeriFlux FLUXNET-1F US-Aud Audubon Research Ranch

This is the AmeriFlux Management Project (AMP) created FLUXNET-1F version of the carbon flux data for the site US-Aud Audubon Research Ranch. This is the FLUXNET version of the carbon flux data for the site US-Aud Audubon Research Ranch produced by applying the standard ONEFlux (1F) software. Site Description - None supplied.

Krishnan, Praveena [NOAA/ARL]↗

Data as a Key Resource in Catalysis: A Community Account

The deployment of artificial intelligence (AI) is transforming the scientific fields central to interdisciplinary catalysis research. By enabling more effective use of data, AI (including simpler machine learning and data science tools) holds great promise for accelerating discoveries. However, progress has so far been modest, largely due to the lack of standardized, machine-readable, and openly shared catalysis data. This perspective, accounting for community insights emerging at conferences, analyses the underlying reasons for these challenges and proposes solutions to a future whereFAIR data management becomes an integral part of research in catalysis. In the short-term, we deem that mandatory FAIR data depositing prior to scientific publications along with consensualized top-down guidelines on data sharing powered by ease-to-use tools can make the necessary step change happen to catalyse data as key resource in our community.

36 - MATERIALS SCIENCE↗

Leafweb: Leaf Gas Exchange and Pulse-Amplitude Modulated Fluorometry for C4 Species, June 2026 Release

This dataset contains leaf gas exchange and Pulse-Amplitude Modulated (PAM) fluorometry for 98 C4 species. The C4 photosynthetic pathway employs specialized CO2 concentration mechanisms and Kranz anatomy to enrich CO2 concentration around Rubisco, the enzyme that catalyzes carbon fixation in the Calvin-Benson cycle to suppress photorespiration and increase the use efficiencies of light, nitrogen, and water as compared to the C3 photosynthetic pathways. Large-scale C4 photosynthetic datasets are relatively scarce, which has affected C4 photosynthesis research. To improve C4 photosynthetic data availability, Leafweb organized an effort to systematically collect, compile, standardize, and organize measurements of leaf gas exchange and/or Pulse-Amplitude Modulated (PAM) fluorometry of C4 species. This derived a C4 photosynthetic dataset containing measurements made by independent researchers in multiple countries in various environments (field, garden, or greenhouse). It covers three biochemical subtypes – the nicotinamide adenine dinucleotide phosphate-malic enzyme (NADP-ME), nicotinamide adenine dinucleotide-malic enzyme (NAD-ME), and phosphoenolpyruvate carboxykinase (PEP-CK) subtypes. This dataset is useful for using Artificial Intelligence / Machine Learning and mechanistic models to study C4 photosynthesis and compare across different biochemical subtypes. This dataset contains 3 compressed (*.zip) folders containing 1,892 data files in comma-separate values (*.csv) format. Additional metadata are provided: one data dictionary and a file-level metadata file in comma-separate values (*.csv) format and a user guide in PDF (*.pdf) format.

Zhou, Haoran [Tianjin University, China]↗

Leveraging generative artificial intelligence to bridge domain gaps in wind turbine research

A central challenge in wind turbine health monitoring is the scarcity of real-world data due to limited instrumentation, leading researchers to rely on simulation models that often suffer from reduced fidelity. However, even within simulation environments, discrepancies arise because of modeling assumptions, and configuration fidelities, creating domain gaps that limit the transferability of learned representations. Here, to investigate domain translation under controlled conditions, this project explores the use of generative artificial intelligence, specifically cycle-consistent generative adversarial networks (CGANs), to bridge the gap between OpenFAST simulation models representing 1.5 MW and 5 MW wind turbines. A physics-informed CGAN architecture is introduced, where a simplified turbine tower dynamics model is incorporated into the training loss to ensure physically consistent outputs. Quantitative results showed moderate to high agreement in frequency-domain features. Incorporating the physics-informed loss function improved the R 2 values by 30%, reduced the RMSE from 1.39 to 1.1 m/s 2 , and reduced training time by 82%. Furthermore, under increased turbulence intensity (IEC Category A), the RMSE remained stable at approximately 1.1 m/s 2 . While the present study is entirely simulation-based, it establishes a pipeline for evaluating physics-informed generative domain translation, which may serve as a foundation for future simulation-to-reality validation studies.

17 WIND ENERGY↗

Solar Resource Measurements in Eugene, OR: Cooperative Research and Development Final Report, CRADA Number CRD-07-00252

Site-specific, long-term, continuous, and high-resolution measurements of solar irradiance are important for developing renewable resource data. These data are used for several research and development activities consistent with the NLR mission: establish a national 3-year climatological database of measured solar irradiances; provide high quality ground-truth data for satellite remote sensing validation; support development of radiative transfer models for estimating solar irradiance from available meteorological observations; provide solar resource information needed for technology deployment and operations. Data acquired under this agreement will be available to the public through NLR's Measurement & Instrumentation Data Center – MIDC (http://www.nlr.gov/midc) Or the Renewable Resource Data Center - RReDC (http://rredc.nlr.gov). The MIDC offers a variety of standard data display, access, and analysis tools designed to address the needs of a wide user audience (e.g., industry, academia, and government interests).

14 SOLAR ENERGY↗

AmeriFlux CA-AF1 Acadia Research Forest

This is the AmeriFlux version of the carbon flux data for the site CA-AF1 Acadia Research Forest. Site Description - The site is located in the Acadia Research Forest (managed by the Canadian Forest Service). The research forest is representative of the Acadian Forest Region in New Brunswick and Nova Scotia and comprises of mixed forest containing softwood, hardwood, and mixed-wood stands. The forest is managed and features different forest composition depending on management history.

Helbig, Manuel [Dalhousie University]↗

Understanding Isomeric Effects on Properties of Aviation Fuels via a Group Contribution Method: Preprint

The molecular composition of aviation fuels, including conventional and sustainable aviation fuels (SAFs), significantly influences their performance, safety, and environmental impact. This study examines the effect of isomeric variation for compounds with the same carbon number and chemical family on key fuel properties, focusing on compounds commonly found in conventional jet fuels and SAFs. A group contribution method (GCM) is employed to predict thermophysical and combustion properties, providing an efficient analytical approach to evaluate the contributions of individual compounds to overall fuel mixture behavior. As part of this work, we introduce FuelLib, an open-source Python tool built around the GCM, to calculate individual compound and fuel mixture properties using various mixing rules. Our work evaluates whether the GCM can capture isomeric effects, that are often overlooked in traditional fuel property estimation. This is particularly important for SAFs, which often are composed of a more limited set of compound classes than conventional fuels, making isomeric differences more critical. Two-dimensional gas chromatography (GCxGC) data, which can be obtained from small fuel samples, provides weight percentages of compounds grouped by chemical family and carbon number rather than detailed information about individual compounds. As a result, assumptions must be made when decomposing GCxGC data into functional groups for GCM applications. Using GCxGC data, we show that the FuelLib tool can be used to predict the fuel properties of conventional jet fuels, with validation against experimental data. The provided tool enables researchers to predict fuel properties of candidate fuels and supports the design of new SAFs at both the component and mixture levels. This capability provides a foundation for studying fuel and combustion properties during SAF development, reducing reliance on costly experimental methods and advancing progress toward certification of new SAFs.The molecular composition of aviation fuels, including conventional and sustainable aviation fuels (SAFs), significantly influences their performance, safety, and environmental impact. This study examines the effect of isomeric variation for compounds with the same carbon number and chemical family on key fuel properties, focusing on compounds commonly found in conventional jet fuels and SAFs. A group contribution method (GCM) is employed to predict thermophysical and combustion properties, providing an efficient analytical approach to evaluate the contributions of individual compounds to overall fuel mixture behavior. As part of this work, we introduce FuelLib, an open-source Python tool built around the GCM, to calculate individual compound and fuel mixture properties using various mixing rules. Our work evaluates whether the GCM can capture isomeric effects, that are often overlooked in traditional fuel property estimation. This is particularly important for SAFs, which often are composed of a more limited set of compound classes than conventional fuels, making isomeric differences more critical. Two-dimensional gas chromatography (GCxGC) data, which can be obtained from small fuel samples, provides weight percentages of compounds grouped by chemical family and carbon number rather than detailed information about individual compounds. As a result, assumptions must be made when decomposing GCxGC data into functional groups for GCM applications. Using GCxGC data, we show that the FuelLib tool can be used to predict the fuel properties of conventional jet fuels, with validation against experimental data. The provided tool enables researchers to predict fuel properties of candidate fuels and supports the design of new SAFs at both the component and mixture levels. This capability provides a foundation for studying fuel and combustion properties during SAF development, reducing reliance on costly experimental methods and advancing progress toward certification of new SAFs.

33 ADVANCED PROPULSION SYSTEMS↗

dCache: The Storage System of Choice for Data-Intensive Applications

The ever-increasing volumes of data produced by modern scientific facilities like EuXFEL and LHC put significant stress on data management infrastructure operated by laboratories and research centers. The challenges to be addressed span the entire data life cycle, from ingest and efficient data analysis to long-term preservation, typically involving large tape libraries. dCache, a storage system developed in collaboration between the Deutsches Elektronen-Synchrotron (DESY), Fermi National Accelerator Laboratory, and Nordic e-Infrastructure Collaboration (NeIC), is designed to manage a large number of disk servers and to facilitate transparent data migration to and from archival storage. Its multifaceted approach offers a unified method to support a variety of scientific use cases with the same storage infrastructure, including high-throughput data ingest, data sharing over wide area networks, efficient access from HPC clusters, and long-term data preservation on tertiary storage. Initially developed for high energy physics (HEP) experiments, dCache is now used by various scientific communities, including astrophysics, biomedical research, and life sciences, each having specific requirements. This paper presents architecture, deployment strategies, performance and scalability enhancements, and recent advancements in dCache addressing the needs of scientific communities. Finally, we touch on the development and release process, ensuring the software’s high quality.

DCache↗

Development of Real-Time High-Density Pulsar Data Transmission and Processing for Grid Synchronization

Taking advantage of the extreme stability of the pulsar period, it can serve as the timing source for grid synchronization to compensate for the timing drift instigated by the loss of GPS signal. Nevertheless, the real-time transmission and processing of the pulsar data suffer from its high-frequency data rate, varying from megahertz to gigahertz, resulting in reduced computing speed and increased time delay. To mitigate this issue, the hardware and software frameworks are implemented for the high-density pulsar data transmission and processing for grid synchronization in this research. Initially, the high-density pulsar data is transferred using open-source software. The complementary duty cycle timing module is designed to coordinate the operation of the dual-channel high-speed interface and software. Subsequently, the multiple-threading is applied to the receiving, parsing, and splicing pulsar data. Next, the pulsar signal extraction method is implemented based on the polyphase filterbank and time of arrival estimation. Ultimately, real-time performance verification experiments are carried out for different components under two hardware platforms. Finally, the results demonstrate that only 0.482 s is required for processing 4 Gigabyte data through multiple-threading, which is 3.8 times faster than the single thread. The pulsar signal extraction can also be executed within 707 ms for 4.8 seconds of data, thereby indicating that real-time requirements can be met.

24 POWER TRANSMISSION AND DISTRIBUTION↗

High-Resolution WRF-Based Downscaling of Earth System Model Projections for Energy Applications across CONUS

Evaluating energy resources under future scenarios requires meteorological information that adequately resolves regional-scale variability and is suitable for regional energy system studies. Although Earth system model (ESM) outputs provide essential large-scale context, their coarse resolution and inherent biases limit direct use in energy system applications. This work presents a high-resolution dynamical downscaling framework using the Weather Research and Forecasting (WRF) model to generate energy-relevant regional fields for future scenarios over the contiguous United States (CONUS). The framework first identifies an optimal WRF configuration through sensitivity experiments, then evaluates raw and bias-corrected ESM initial and boundary conditions, with soil moisture and soil temperature bias correction implemented as an integral component of the bias-corrected ESM atmospheric forcing prior to WRF dynamical downscaling to improve land-atmosphere interactions. Simulations performed at 4-km resolution show that uncorrected ESM forcing leads to systematically dry and cold soil states, which propagate into elevated near-surface air temperature and solar irradiance biases, particularly during summer for the period 2000-2014. Incorporating bias-corrected atmospheric forcing together with soil state bias correction substantially reduces these errors and improves the representation of surface energy processes in WRF simulations. The results highlight the importance of bias-aware initialization strategies in high-resolution dynamical downscaling for future energy system analysis and planning.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Common Column Identification for Table Similarity Detection in Electrified Transportation Data Lakes

Electrified transportation often requires researchers and operators to interact with datasets from a wide range of sources and disciplines, such as transportation, power systems, public health, policies, and regulations. These datasets vary in quality and format, making it difficult to understand, preprocess, and identify key columns representing real-world entities or values for indexing and joining, which can negatively impact downstream analysis and operation. Existing solutions are limited, requiring extensive manual customization or data expertise to utilize. In this article, we propose a multi-layered approach to automatically identify key columns to expedite preprocessing and aid in analysis of electrified transportation data. Our method leverages a dynamic ontology to identify common fields and an information theory-based strategy for edge cases that are difficult to generalize. Evaluations on a number of datasets from data.gov and kaggle.com show improved performance of our methods over several baseline techniques, and our ablation analyses illustrate the efficacy of individual components of our method. Our case studies also demonstrate that our methods have the potential to improve analysis of electrified transportation data and aid in automatic integration of such datasets.

33 ADVANCED PROPULSION SYSTEMS↗

Probabilistic Error Bounds for Low-Rank Tensor Decompositions Used in Large-Scale Data Analysis Applications (LDRD Final Report)

This report documents a research project on analyzing low-rank tensor models for data analysis that took place at Sandia National Laboratories from October 2023–September 2025. The focus of this work was to extend theoretical frameworks from statistics and probability theory for use with models for scalar, vector, and matrix data to models with tensor, or general multi-dimensional array, data. Through this work, we have provided a new set of tools for bounding errors on low-rank tensor models of both complete and sampled data. The remainder of this report is organized as follows. In Section 1, we describe the proposed work at the start of the project. Section 2 describes the research advances made as part of the project. Other research contributions in the form of conference presentations and software development is provided in Section 3. Workforce development at Sandia and Florida Atlantic University (via a subcontract on this project) is provided in Section 4.

97 MATHEMATICS AND COMPUTING↗

Advancing an integrated understanding of land–ocean connections in shaping the marine ecosystems of coastal temperate rainforest ecoregions

Land and ocean ecosystems are strongly connected and mutually interactive. As climate changes and other anthropogenic stressors intensify, the complex pathways that link these systems will strengthen or weaken in ways that are currently beyond reliable prediction. In this review we offer a framework of land–ocean couplings and their role in shaping marine ecosystems in coastal temperate rainforest (CTR) ecoregions, where high freshwater and materials flux result in particularly strong land–ocean connections. Using the largest contiguous expanse of CTR on Earth—the Northeast Pacific CTR (NPCTR)—as a case study, we integrate current understanding of the spatial and temporal scales of interacting processes across the land–ocean continuum, and examine how these processes structure and are defining features of marine ecosystems from nearshore to offshore domains. We look ahead to the potential effects of climate and other anthropogenic changes on the coupled land–ocean meta-ecosystem. Finally, we review key data gaps and provide research recommendations for an integrated, transdisciplinary approach with the intent to guide future evaluations of and management recommendations for ongoing impacts to marine ecosystems of the NPCTR and other CTRs globally. In the light of extreme events including heatwaves, fire, and flooding, which are occurring almost annually, this integrative agenda is not only necessary but urgent.

54 ENVIRONMENTAL SCIENCES↗

Optimizing high energy density sulfur cathodes: A multivariate approach to electrode formulation and processing

Lithium-sulfur (Li-S) batteries involve complex solid-liquid-solid phase transformations during both discharging and charging processes, where cathode materials, formulation, and structure play a crucial role. Here, a design of experiments (DoE) methodology and an empirical model are developed to systematically explore the interactions and trade-offs among cathode factors and process variables, and to obtain generalizable effects estimates for the multivariate system. Compared to the conventional one-factor-at-a-time (OFAT) approach, this work demonstrates advantages in both efficiency and accuracy by allowing the data to guide future research and decisions. Further, an optimized cathode formulation and processing parameters are predicted and validated experimentally, achieving over 1000 mAh g -1 in discharge capacity and improved cycling under practical lean electrolyte (4 µL mg -1 S) and high S-loading cathodes (>4 mg cm -2 ) conditions. The optimized cathode was scaled up and assembled into Li-S pouch cells, achieving 316 Wh kg -1 in cell-level energy, proving that the comprehensive and rigorous framework for optimizing complex systems with DoE leads to improved performance in a practical pouch cell system.

25 ENERGY STORAGE↗

Investigating the Determinants of Household Capabilities Burden During Power Outages: The Case of Winter Storm Uri

Existing research primarily uses census data to identify the vulnerability of communities to hazards. These vulnerability indices provide aggregated data and are not hazard-specific nor well-validated with post-event data. In contrast, our study uses household survey data (n=1065) to understand which Texan households suffered the greatest loss of their capabilities due to power outages and other utility service disruptions during Winter Storm Uri. Inspired by the Capabilities Approach, our measures of burden include the number of household capability types disrupted during the outages (e.g., cooking, heating, refrigeration), the severity of impact for each disrupted capability, and the additional time and financial costs of coping with these disruptions. We perform a clustering analysis, and find two distinct groups in our data, consisting of ‘lesser burden' and ‘heavier burden' households. Results indicate that the households experiencing the heaviest capabilities burden were most likely to experience longer power outages and the loss of water services. They were also more likely to have a Hispanic-Latino household member, lack access to a generator, live in a rented home, have larger households with more young children, fewer adults over 65, lower household incomes, been impacted by the COVID-19 pandemic, and more family characteristics that made life harder. We also fit a logistic regression model to assess the role of outage, household, and community characteristics in predicting differences in capabilities burden. Our results offer insights into enumerating the consequences of utility service disruptions on households, which can inform more targeted and equitable resilience strategies.

24 POWER TRANSMISSION AND DISTRIBUTION↗