Search NASA⌕ Search

SEARCH · Search NASA

Results for “open datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

ORBIT-2 Dataset for Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling

This dataset release corresponds to the work conducted in ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling, where large-scale AI methods were applied to improve climate and weather resolution. The collection integrates four widely used, publicly available datasets: ERA5, PRISM, DAYMET, and IMERG. To prepare the data for ORBIT-2 model training and evaluation, we applied a preprocessing pipeline that generates paired low-resolution and high-resolution samples, enabling supervised downscaling experiments. The transformation from coarse to fine scales was performed using bilinear regridding, consistent with the procedures described in WeatherBench2, a community benchmark for weather and climate AI models. This dataset supports the development and evaluation of foundation models designed for weather and climate downscaling at exascale. Additional details on methodology and applications can be found in Wang et al., ORBIT-2 (arXiv:2505.04802, 2025).

54 ENVIRONMENTAL SCIENCES↗

Roadmap and Benchmarking: Privacy in Federated Load Forecasting

Data-driven techniques for energy demand forecasting continue to emerge with promising impacts on distribution grid planning. However, the development of robust and generalizable machine learning models requires that representative high quality training data are available. Distributed energy resources have begun to embed intelligence, gathering large amounts of data on customer demand, behavior, and household devices that are connected to the grid. Though utilities aggregate meter-level demand data for load shaping, demand response, outage management, reliability planning, and billing applications, there lies an inherent privacy concern in sharing consumption data that may identify individual consumer behavioral patterns. Hence, while sharing the data is crucial, the private sensitive customer data must be safeguarded from being exposed or manipulated. In this study, we propose a roadmap for implementing a based privacy preserving framework to support the advancement of data-driven analytics in data-sensitive distributed energy resources environments. The roadmap incorporates federated learning–a distributed training framework, differential privacy–a statistical framework that provides guarantees to safeguard the leakage of sensitive data, secure multiparty computation and homomorphic encryption– techniques for encrypting model gradients and applying secure aggregation on the server. Moreover, we perform baseline experiments on the federated short-term load forecasting (STLF) task using open-source residential load profile datasets, offering insights into the challenges of integrating differential privacy into federated learning.

Abebe, Waqwoya [Oak Ridge National Laboratory (ORN↗

High-n Rydberg transition spectroscopy for heavy impurity transport studies in W7-X (invited)

Here, we present a novel spectroscopy approach to investigate impurity transport by analyzing line-radiation following high-n Rydberg transitions. While high-n Rydberg states of impurity ions are unlikely to be populated via impact excitation, they can be accessed by charge exchange (CX) reactions along the neutral beams in high-temperature plasmas. Hence, localized radiation of highly ionized impurities, free of passive contributions, can be observed at multiple wavelengths in the visible range. For the analysis and modeling of the observed Rydberg transitions, a technique for calculating effective emission coefficients is presented that can well reproduce the energy dependence seen in datasets available on the OPEN-ADAS database. By using the rate coefficients and comparing modeling results with the new high-n Rydberg CX measurements, impurity transport coefficients are determined with well-documented 2σ confidence intervals for the first time. This demonstrates that high-n Rydberg spectroscopy provides important constraints on the determination of impurity transport coefficients. By additionally considering Bolometer measurements, which provide constraints on the overall impurity emissivity and, therefore, impurity densities, error bars can be reduced even further.

Instruments & Instrumentation↗

Synthesizing land use and demographic change in Southeast Asia’s smaller urbanized areas from 2000–2015

The majority of the human population now reside in urban areas today. The United Nations estimates that nearly half of all urban dwellers currently live in cities smaller than 500 000 persons and the majority of future urban growth will take place in Asia and Africa, likely in these smaller urban areas, not mega cities. Thus, understanding the factors that influence urban demographic trajectories in small urban areas is critical to address sustainable and equitable policy initiatives related to food security, changing climate hazard exposure, and economic opportunities. Here we focus on Southeast Asia—a region historically characterized by lower urban population proportions, yet with a rapidly shifting dynamic demographic—to examine correlates of demographic change among smaller cities. We combine two open-source satellite-informed datasets: GHS urban center database (2015) and age-sex gridded data from WorldPop to calculate socio-demographic characteristics to model drivers of change in annualized urban population growth from 2000–2015 for 505 urbanized places. We find a general pattern of decreasing dependency ratios as city-size increases for most urban areas in Southeast Asia. Higher rates of growth and more variation is observed for smaller cities—those with fewer than 300 000 persons, the lowest population limit for UN data on urbanization. When examining covariates of urban population growth, we find significant statistical associations of population change in smaller urbanized areas with climatic, economic, and land cover/land use variables, but with country-specific variations. Characterizing a continuum of urban population development in the context of changing environmental, economic and climate conditions has been an important sustainable development and equity issue for decades, but newer analysis of city-level drivers allows for systematic inquiry thus moving beyond total population counts for policy-relevant insight.

Southeast Asia synthesis↗

SO(3)-invariant PCA with application to molecular data

Principal component analysis (PCA) is a fundamental technique for dimensionality reduction and denoising; however, its application to three-dimensional data with arbitrary orientations -- common in structural biology -- presents significant challenges. A naive approach requires augmenting the dataset with many rotated copies of each sample, incurring prohibitive computational costs. In this paper, we extend PCA to 3D volumetric datasets with unknown orientations by developing an efficient and principled framework for SO(3)-invariant PCA that implicitly accounts for all rotations without explicit data augmentation. By exploiting underlying algebraic structure, we demonstrate that the computation involves only the square root of the total number of covariance entries, resulting in a substantial reduction in complexity. We validate the method on real-world molecular datasets, demonstrating its effectiveness and opening up new possibilities for large-scale, high-dimensional reconstruction problems.

Fraiman, Michael [Tel Aviv Univ., Tel Aviv (Israel↗

New model for the ion collection by cylindrical probes over a wide range of collisionality

Langmuir probes remain one of the most important diagnostic tools for plasma processing applications. Modern probe analysis usually relies on the electron current part of the Langmuir probe characteristic using the Druyvesteyn method. However, for electronegative plasmas or for discharges containing dust the analysis of the ion current attracted by the probe can be desirable to determine the ion density. But, even at low pressures of a few Pa, the ion current is affected by collisions due to the large cross section for charge exchange. Available theories for collisional or collision-enhanced ion currents onto probes are complex and not well validated. Thus, in this contribution, we compare available collisional probe theories for the ion current to results of particle-in-cell (PIC) simulations. To this end, the probe surrounded by a semi-infinite plasma is simulated using a modified version of the open-source code EDIPIC. A dataset of currents for different neutral gas pressures is obtained and compared to the different theories from the literature. Based on these results, we propose a simpler and more intuitive model for the ion current collected by the probe, based on the model of Gatti and Kortshagen (Phys. Rev. E 78, 046402, 2008), developed for the charging of dust particles.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Hydrologic Model Data for the East Fork Poplar Creek Watershed Simulated with the Advanced Terrestrial Simulator (ATS): Streamflow and Network Expansion–Contraction Dynamics

This dataset supports hydrologic modeling and stream network expansion–contraction analysis for the East Fork Poplar Creek (EFPC) Watershed in Tennessee. It includes a Jupyter notebook for model setup, model configuration files, simulation outputs, and derived products used to evaluate model performance and investigate stream dynamics under varying hydrologic conditions. The dataset was generated using the Watershed Workflow Python package and the Advanced Terrestrial Simulator (ATS), enabling integrated surface–subsurface hydrologic simulations using a stream-aligned mesh. Outputs include high-resolution time series of streamflow, active network length, water table depth, and related hydrologic variables. Also included are spatially explicit stream persistency indices and classifications of reaches as perennial or non-perennial. These data facilitate reproducibility and support further research on stream intermittency and variability in network extent.The model data archive is organized in following directories:1) model_setup_inputsContains the Watershed Workflow Jupyter notebooks (accessed through any open source code editor), selected input datasets, and resulting ATS input files, including XML files (access through any open source code editor), computational mesh (.exo files can be viewed using Paraview), and meteorological forcing files (.h5 files can be accessed through h5py python package and HDFView open source software). 2) model_outputsIncludes ATS simulation outputs relevant to this study. Time series of spatially integrated or averaged variables (e.g., streamflow, water table depth) are provided as CSV files. Select spatial fields (e.g., ponded depth and water table depth) are saved as pickled Python objects to reduce file size, and can be accessed through pickle package in Python. Key geometry objects from Watershed Workflow—such as the surface mesh and river tree—are also included to support analysis of streamflow persistency and expansion–contraction dynamics. These files can also be accessed through Watershed Workflow Python package.3) model_evaluationProvides observed streamflow time series and field survey-based flow regime classifications used to evaluate model performance. Jupyter notebooks for processing ATS outputs and comparing model predictions with observations to build confidence in the model prior to scientific analysis are also included.4) Q_L_relationshipsContains workflows for generating time series of discharge, active network length, and related hydrologic variables used in the stream network expansion–contraction analysis. Includes routines for delineating baseflow-dominated periods. For each catchment, notebooks and processed data (as pickled DataFrames accessed through Pandas Python package) are provided. 5) figure_scriptsProvides the Jupyter notebooks used to generate the figures presented in the paper.

54 ENVIRONMENTAL SCIENCES↗

Videos and front speeds of frontal ring-opening metathesis polymerization (FROMP) of DCPD/ENB with norbornene-functionalized PDMS comonomers

This dataset contains videos and front speed measurements for 16 frontal ring-opening metastasis polymerization experiments of dicyclopentadiene (DCPD)/5-ethylidene-2-norbornene (ENB) resins and norbornene-functionalized polydimethylsiloxane (nor-PDMS) comonomers. Each run was carried out in a 10 mm diameter glass test tube and recorded to quantify front propagation behavior. Reported front speeds were extracted by video tracking and reported maximum front temperatures were measured with a thermocouple.

Clarke, Brandon R.↗

Wind Turbine Sound Setbacks and Supply Curves: Ordinances and Extrapolated Trends, 110 Hub Height, 130 Rotor Diameter

This dataset provides a comprehensive set of wind turbine sound setbacks from every residential structure in the contiguous United States (CONUS). A sound setback is defined as the minimum required distance between a residential structure and a hypothetical turbine installation site to ensure that modeled sound levels received at the residence do not exceed local sound ordinances, which are commonly expressed in A-weighted decibels (dBA). Therefore, sound setbacks are a local spatial assessment combining multiple factors, including the sound pressure curve as a function of the observer location (distance and direction) relative to the turbine, local sound regulations, and the geographical distribution of residential structures. The dataset is organized into multiple scenario-based products, detailed as follows: 1. Existing and extrapolated sound setbacks. An existing scenario characterizes sound setbacks only in states or counties that have implemented sound regulations as of 2022. The extrapolated scenarios extend a constant sound threshold to counties that lack explicit sound regulations, with thresholds ranging from 35 to 60 dBA, in 5-dBA increments reflecting the variation observed in current sound ordinances. 2. Sound setbacks in directional and worst scenarios. The directional scenario accounts for the distance and orientation of residential structures relative to a hypothetical turbine location, utilizing the turbine's sound emissions in that specific direction. In contrast, the worst scenario takes loudest sound level at each distance step from the turbine, irrespective of directional considerations, which aligns with current industry practice. 3. Supply curves for Open and Reference Access scenarios. This dataset includes supply curves generated by the reV model, which integrates each of the above sound setbacks into both Open and Reference siting scenarios. In addition, two Open and Reference baselines scenarios were included which do not consider sound setbacks for comparative analysis. All sound setback data are stored in TIF files, with partial maps of the data provided in PNG format. The values in the sound setback raster range from 0 to 1, representing the fraction of developable land within a 90 meter by 90 meter pixel due to sound ordinances. A value of 0 indicates areas where wind energy development is prohibited, while a value of 1 signifies areas fully permissible. The wind turbine parameters used in the sound modeling are based on the land-based turbine from International Energy Agency (IEA), featuring a rated electrical power of 3.4 MW, a rotor diameter of 130 meters, and a hub height of 110 meters. The atmospheric conditions, including wind speed/direction, turbulence, air temperature, relative humidity, and air pressure, that drive the sound generation are obtained from the WIND Toolkit dataset.

Array↗

An Exploratory Data Mining Investigation for Constructing a Publicly Sourced Dataset of Foreign Hypersonic Tests

This document details a data mining exercise that resulted in an exploratory dataset of publicly reported foreign (non-US) hypersonic vehicle test events. Using a combination of targeted English language searches and country-specific queries, the study aggregates information from digital news media, official press releases, and social media posts. The resulting list of events captures the publicly available accounts of foreign hypersonic tests, although it does not represent an exhaustive record. Limitations such as inconsistent reporting, translation challenges, and the inherently provisional nature of open-source data are acknowledged. This dataset serves as an initial reference point for further inquiries into high-speed atmospheric phenomena and may facilitate future efforts to correlate these events with geophysical measurements.

33 ADVANCED PROPULSION SYSTEMS↗

Livewire: A Model Platform for Data Quality Assessment and AI Readiness Across DOE Missions

High-quality, well-governed data is essential for accelerating discovery and achieving operational excellence across DOE and national laboratory missions. The Livewire Data Platform is a DOE-supported platform that offers automated assessments of data quality, standardization, provenance, and Artificial Intelligence (AI) readiness. It allows researchers and data practitioners to systematically and easily evaluate datasets against established governance criteria and prepare them for advanced analytics. Livewire addresses critical challenges in DOE's data ecosystem with integrated capabilities for metadata validation, provenance tracking, and schema alignment. This platform's automated workflows assist users in identifying data quality gaps, enhancing interoperability between datasets collected from various stakeholders, and ensuring compliance with DOE data standards, all while reducing manual curation efforts. Additionally, we will discuss its AI readiness framework, which is being developed to prepare datasets for training models, developing advanced analytic tools, and machine learning applications. Using some of the more than one hundred tabular datasets on Livewire, processed with this open-source methodology, we will demonstrate how Livewire can serve as a model for scalable, standards-driven data management. This approach provides a pathway to leverage existing and future datasets within the DOE, boosting innovation and efficiency across national laboratories.

33 - ADVANCED PROPULSION SYSTEMS↗

Advances in building data management for building performance standards using the SEED platform

Reducing energy consumption and greenhouse gas emissions in the built environment is a critical step in achieving emission goals to mitigate climate change impacts. Local, federal, and international jurisdictions are deploying several methods to reduce energy and emissions such as voluntary and mandatory benchmarking and building performance standards, requiring building owners to reach energy and emission targets. Jurisdictions leveraging benchmarking and building performance standards require knowledge of the buildings covered; which is a large task due to staffing constraints, limited information on building characteristics and tax parcel data, and the need for advanced data management techniques to align datasets. This paper describes an open-source platform's recent advances to create consistent taxonomies, identify erroneous data, enable auditability, and track building performance. The paper concludes with two use cases on how the platform has been used by jurisdictions.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Atomistic Simulation of Glasses and Amorphous Materials: Challenges and Opportunities for the Next Decade

Atomistic simulations have become indispensable tools for understanding glass structure, dynamics, and properties, yet persistent challenges limit their predictive power. This perspective examines three interconnected issues, namely glass formation procedures, interatomic potential development, and machine learning applications, which emerged from the 5th International Workshop on Challenges of Atomistic Simulations of Glasses and Amorphous Materials. We identify convergent community priorities for (i) standardized validation protocols, (ii) curated benchmark datasets with complete metadata, and (iii) open repositories for glasses. A systematic was forward is provided by a hierarchical validation framework for assessing the structural fidelity, property prediction, and behavioral realism of simulation techniques. Looking ahead, transformative advances are promised by the fusion of classical techniques with machine learning based approaches, for instance, by integrating swap Monte Carlo with machine-learning (ML) potentials, leveraging foundation models through transfer learning, and finetuning ML potentials with experimental data. Progress depends on the community committing to validated models, reproducible protocols, and sustained data sharing.

Krishnan, N. M. Anoop↗

A Million Person Study Innovation: Evaluating Cognitive Impairment and other Morbidity Outcomes from Chronic Radiation Exposure Through Linkages with the Centers for Medicaid and Medicare Services Assessment and Claims Data

Here, the study of One Million U.S. Radiation Workers and Veterans, the Million Person Study (MPS), examines the health consequences, both cancer and non-cancer, of exposure to ionizing radiation received gradually over time. Recently the MPS has focused on mortality patterns from neurological and behavioral conditions, e.g., Parkinson's disease, Alzheimer's disease, dementia, and motor neuron disease such as amyotrophic lateral sclerosis. A fuller picture of radiation-related late effects comes from studying both mortality and the occurrence (incidence) of conditions not leading to death. Accordingly, the MPS is identifying neurocognitive diagnoses from fee-for-service insurance claims from the Centers for Medicare and Medicaid Services (CMS), among Medicare beneficiaries beginning in 1999 (the earliest date claims data are available). Linkages to date have identified ∼540,000 workers with available health information. Such linkages provide individual information on important co-factor and confounding variables such as smoking, alcohol consumption, blood pressure, obesity, diabetes and many other health and demographic characteristics. The total person-level set of time-dependent variables, outcomes, organ-specific dose measures, co-factors, and demographics will be massive and much too large to be evaluated with standard software. Thus, development of specialized open-source software designed for large datasets (Colossus) is nearly complete. The wealth of information available from CMS claims data, coupled with individual dose reconstructions, will thus greatly enhance the quality and precision of health evaluations for this new field of low-dose radiation and neurocognitive effects.

Dauer, Lawrence T.↗

A reproducible study design for the MIMIC-IV in-hospital mortality task

Open, tabular electronic health record (EHR) datasets such as MIMIC-III and MIMIC-IV have become critical resources for developing machine learning (ML) models addressing clinical prediction tasks, including hospital readmission, length of stay, and in-hospital mortality (IHM). While MIMIC-III has benefited from well-established preprocessing pipelines and standardized feature sets, MIMIC-IV remains comparatively challenging to work with because there are no standardized benchmarks to support reproducibility and comparability across studies. To address this limitation, we present a rigorously curated MIMIC-IV custom feature set optimized for IHM prediction, constructed through a reproducible preprocessing pipeline and feature selection strategy.

97 MATHEMATICS AND COMPUTING↗

Reliable Integration of AI Data Centers at Scale – Analysis, Modeling and Synthetic Data Generation

This report analyzes the power consumption of large dynamic digital loads using the open-source MIT supercloud and SURF datasets. With an emphasis on the MIT data, we calculate important power consumption characteristics to help system operators improve generation planning and resource allocation. We also introduce a rudimentary model for generating synthetic load profiles.

97 MATHEMATICS AND COMPUTING↗

Machine Learning meets Algebraic Combinatorics: A Suite of Benchmark Datasets to Accelerate AI for Mathematics Research

The use of benchmark datasets has become an important engine of progress in machine learning (ML) over the past 15 years. Recently there has been growing interest in utilizing machine learning to drive advances in research-level mathematics. However, off-the-shelf solutions often fail to deliver the types of insights required by mathematicians. This suggests the need for new ML methods specifically designed with mathematics in mind. The question then is: what benchmarks should the community use to evaluate these? On the one hand, toy problems such as learning the multiplicative structure of small finite groups have become popular in the mechanistic interpretability community whose perspective on explainability aligns well with the needs of mathematicians. While toy datasets are a useful benchmark for initial work, they lack the scale, complexity, and sophistication of many of the principal objects of study in modern mathematics. To address this, we introduce a new collection of benchmark datasets, Algebraic Combinatorics Benchmarks (ACBench), representing either classic or open problems in algebraic combinatorics, a subfield of mathematics that studies discrete structures arising from abstract algebra. After describing the datasets, we discuss the challenges involved in constructing “good” mathematics benchmarks, describe baseline model performance, and discuss some of the insights these datasets can provide that may be of interest even to those who are not interested in mathematics research itself.

97 MATHEMATICS AND COMPUTING↗

Dataset of Generative AI Workload Power Profiles

This dataset provides a collection of high-resolution (5/10 Hz or every 0.2/0.1 seconds) power consumption profiles for generative artificial intelligence (GenAI) workloads executed on NLR's High Performance Computing (HPC) platform Kestrel. The dataset also includes examples of representative whole-facility power profiles generated using a bottom-up, event-driven, data center energy model . This dataset is designed to support research in energy modeling, infrastructure planning, energy system integration, and sustainability analysis for AI-driven computing systems. The dataset captures time-resolved electrical power measurements across a diverse set of configurations, including variations in job type (inference vs. training), workload (LLM vs. image generation), datasets, and number of compute nodes. Power traces are provided in a standardized format and include both raw/instantaneous and aggregated files. Each profile is accompanied by metadata describing workload parameters, enabling reproducibility and cross-study comparison. The dataset is intended for use in applications such as data center infrastructure planning, energy modeling, demand response and grid impact studies, and development and validation of system-level simulation tools. By making these workload-specific power profiles publicly available, this dataset aims to address the current lack of open, empirical energy data for generative AI systems and to facilitate transparent, reproducible research on the energy and environmental impacts of large-scale AI deployment. If you use this dataset, please cite the associated publication: Vercellino et al., “Measurement of Generative AI Workload Power Profiles for Whole-Facility Data Center Infrastructure Planning,” arXiv:2604.07345 (2026).

97 MATHEMATICS AND COMPUTING↗