Search NASA⌕ Search

SEARCH · Search NASA

Results for “open datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

ORBITaL-Net: A labeled training library for large-scale building feature extraction

Over the course of several years, nearly 1.5 million building outlines have been created from approximately 128,000 training tiles covering roughly 7,000 km 2 of very high-resolution multispectral overhead imagery, primarily dated between 2010 and 2020. This dataset, dubbed the Oak Ridge Building Image and TrAining Label Net (ORBITaL-Net), is designed for machine learning applications and is global in scope, with samples drawn from 72 countries across North America, South America, Africa, Europe, and Asia. ORBITaL-Net captures a great diversity in geographic setting, structural characteristics, land use (urban and rural), terrain, and imagery conditions. While the labeled building outlines are themselves valuable, the dataset’s true strength lies in the pairing of these labels with corresponding reference imagery, which is being released for open source use. Similar to SpaceNet and Replicable AI For Microplanning (ramp), this building outline dataset will allow the larger computer vision community from academia, government, and industry the opportunity to develop robust, scalable, and generalizable geospatial machine learning techniques. Unlike SpaceNet and ramp, which offer high resolution labels and imagery primarily for large urban cities, ORBITaL-Net is not focused on training samples from heavily populated areas but instead aims to capture the innate variability of conditions present in both the physical environment and imagery collections.

Geography↗

Data for Zheng et al. (2025), "AquaMEND: Reconciling multiple impacts of salinization on soil carbon biogeochemistry"

Soil salinization, exacerbated by climate change, poses a global threat to coastal ecosystems and soil function. Salinity affects soil carbon cycling by directly impacting microbial activity and indirectly altering soil physicochemical properties, but current models inadequately represent these complexities. This dataset contains the observational and modeling data from Zheng et al. (2025), which described a process-based modeling framework that couples soil solution chemistry with microbial carbon cycling reactions to study the impacts of soil salinization. This conceptual model is implemented numerically into the open-source geochemical program PHREEQC 3.0 (Parkhurst and Appelo, 2013). This dataset consists of: - Figure2_AquaMEND_salinity_buffer: Contains model simulation outputs to assess the impact of three different cation exchange and surface complexation processes on salinity buffering (Fig. 2 from Zheng et al. 2025). - Figure3_Salinity_function: Contains salinity function fitting for literature data (Fig. 3 from Zheng et al. 2025). - Figure4_AquaMEND_microbial_mechanisms: Contains model simulation outputs for testing various microbial process-based hypotheses related to soil salinization, including microbial mortality, carbon use efficiency (CUE), extracellular enzyme activity, and other microbial mechanisms (Fig. 4 from Zheng et al. 2025). - Figure5_AquaMEND_Redox: Contains on model simulation outputs to evaluate shifts among key redox processes, such as aerobic respiration, sulfate reduction, and methanogenesis (Fig.5 from Zheng et al. 2025). - Figure6_AquaMEND_sorption: Contains on model simulation outputs for investigating the effects of salinity on dissolved organic matter (DOM) sorption and desorption processes (Fig. 6 from Zheng et al. 2025). - Figure7_AquaMEND_process_couple: Contains on model simulation outputs for exploring coupled biotic-abiotic processes and their interactions (Fig. 7 from Zheng et al. 2025). - data: Includes datasets used to develop salinity response functions and evaluate salinity buffering capacity. Datasets for MEND model calibration. - database: Contains the `.dat` file required by PHREEQC for model execution. - README.md: A Markdown plain text file describing the computational tools and directories. Files are a mixture of plain text CSV (comma-separated value) and plain text *.dat files written by the model; no special software is required to read them.

EARTH SCIENCE > AGRICULTURE > SOILS > SOIL SALINIT↗

Fast and Invertible Simplicial Approximation of Magnetic‐Following Interpolation for Visualizing Fusion Plasma Simulation Data

We introduce a fast and invertible approximation for fusion plasma simulation data represented as 2D planar meshes with connectivities approximating magnetic field lines along the toroidal dimension in deformed 3D toroidal spaces. Scientific variables (e.g., density and temperature) in these fusion data are interpolated following a complex magnetic-field-line-following scheme in the toroidal space represented by a cylindrical coordinate system. This deformation in the 3D space poses challenges for root-finding and interpolation. To this end, we propose a novel paradigm for visualizing and analyzing such data based on a newly developed algorithm for constructing a 3D simplicial mesh within the deformed 3D space. Our algorithm generates a tetrahedral mesh that connects the 2D meshes using tetrahedra while adhering to the constraints on node connectivities imposed by the magnetic field-line scheme. Specifically, we first divide the space into smaller partitions to reduce complexity based on the input geometries and constraints on connectivities. Then, we independently search for a feasible tetrahedralization of each partition, considering nonconvexity. We demonstrate our method with two X-Point Gyrokinetic Code (XGC) simulation datasets on the International Thermonuclear Experimental Reactor (ITER) and Wendelstein 7-X (W7-X), and use an ocean simulation dataset to substantiate broader applicability of our method. An open source implementation of our algorithm is available at https://github.com/rcrcarissa/DeformedSpaceTet.

Ren, Congrong [The Ohio State Univ., Columbus, OH ↗

Smart Meter Data: A Gateway for Reducing Solar Soft Costs with Model-Free Hosting Capacity Maps

Public-facing solar hosting capacity (HC) maps, which show the maximum amount of solar energy that can be installed at a location without adverse effects, have proven to be a key driver of solar soft cost reductions through a variety of pathways (e.g., streamlining interconnection, siting, and customer acquisition processes). However, current methods for generating HC maps require detailed grid models and time-consuming simulations that limit both their accuracy and scalability—today, only a handful out of almost 2,000 utilities provide these maps. This project developed and validated data-driven algorithms for calculating solar HC using data from AMI without the need of detailed grid models or simulations. The algorithms were validated on utility datasets and incorporated as an application into NRECA’s Open Modeling Framework (OMF.coop) for the over 260 coops and vendors throughout the US to use. The OMF is free and open-source for everyone.

14 SOLAR ENERGY↗

Livewire User Guide

The Livewire Data Platform houses a catalog of transportation- and mobility-related project data, as well as a publications database, making it easy to search and share data. It allows transportation researchers, industry, and academic partners to increase the visibility of their projects within the research community, securely share and preserve data, and leverage datasets from other projects. Public data on Livewire are open to anyone with a Livewire account. This guide will help Livewire users understand how to store project data as a data steward, as well as access data as a data consumer.

33 ADVANCED PROPULSION SYSTEMS↗

Subhourly Clipping Correction Model Comparison

This work will compare the Allen method and Walker method of accounting for subhourly inverter clipping power losses in hourly PV performance models. The Allen method uses a matrix lookup based on DNI clearness and clipping potential to assign a clipping correction loss at each simulation timestep. The Walker method models the PV DC power input to the inverter as a distribution over the hourly timestep and uses integration over the timestep to determine the amount of clipping that occurs within the timestep. Both these models have been recently implemented in the System Advisor Model's (SAM) open-source code, and will be applied to hourly SURFRAD datasets to analyze the subhourly clipping loss predicted by each model for different system designs and inverter loading conditions. Both models will be compared to "true" 1-minute SURFRAD data simulations to see their accuracy against more accurate 1-minute clipping correction loss predictions. This model comparisons will be investigated in more detail at the PVSC conference in Seattle, Washington June 2024.

ENERGY PLANNING, POLICY, AND ECONOMY,MATHEMATICS A↗

From 2D to 4D: a containerized workflow and browser to explore dynamic chromatin architecture

Background Characterizing the physical organization of the genome is essential for understanding long-range gene regulation, chromatin compartmentalization, and epigenetic accessibility. Hi-C experiments generate two-dimensional (2D) genome-wide contact maps of chromatin interactions by capturing the spatial proximity between genomic loci, which reveal interaction frequencies but lack the spatial resolution needed to interpret the three-dimensional (3D) genome structure(s). Emerging evidence suggests that epigenetic regulation is closely linked to 3D genome architecture, and that structural changes over time (4D) drive key biological processes in development, disease, and environmental response. Thus, integrating 3D structure with functional data is critical for a more complete understanding of genome regulation. Previous work, most notably the 4DHiC chromosome modeling framework, has shown that physical multi-dimensional modeling approaches rooted in polymer physics and molecular dynamics can resolve these structures at biologically meaningful resolutions by integrating temporal Hi-C data with physical constraints to uncover dynamic chromosome reorganization. Thus, molecular dynamics simulations, constrained by Hi-C contact matrices, can resolve fine-scale structural changes and reveal functionally significant transitions in chromatin conformation. Results Herein, we present the 4D Genome Browser Workflow (4DGBWorkflow) and the 4D Genome Browser (4DGB). The algorithm is based on the 4DHiC method, and the containerized tool is an end-to-end workflow that can transform, filter, and view 4D epigenomics and chromatin datasets, allowing non-specialists to apply three-dimensional modeling principles to diverse datasets and experimental conditions. The software executes on a laptop running macOS, Linux or Windows. From input Hi-C files (.hic), the 4DGBWorkflow produces 3D reconstructions of chromosomes, integrates the reconstruction with track data (e.g., epigenetic marks, transcriptome profiles), and provides comparative visualization of the results in a single workflow. Conclusions The 4DGBWorkflow and 4D Genome Browser are open-source tools for comparative analysis and visualization of 4D chromosome datasets, including chromatin architecture and epigenomic signals. Automatic integration of Hi-C data with molecular dynamics democratizes the construction of time resolved 3D genome structures, simplifying complex simulations and data integration schemes.

3D Genome Browser↗

2019 Meander C and Meander Z floodplain groundwater chemistry from the East River Watershed, CO, USA

This dataset includes groundwater geochemistry data from floodplain piezometers collected as a part of the Watershed Function Scientific Focus Area (SFA) located in the Upper Colorado River Basin. The data were collected in order to investigate the role of hyporheic exchange and other river corridor processes on riverine export of solutes. Data includes samples from two intra-meander zones: Meander C, in the Pumphouse vicinity, and Meander Z, just upstream of the confluence with Brush Creek. Floodplain piezometers installed along two transects across Meander C (MCP and MCB wells) and Meander Z (MZA and MZB wells) were sampled on daily to weekly time scales during summer-fall 2019. Some river water grab samples are also included. Data includes in-field measurements (pH, electrical conductivity [EC], oxidation reduction potential [ORP], dissolved oxygen [DO], and groundwater level) along with laboratory measurements (dissolved inorganic carbon [DIC], dissolved organic carbon [DOC], metals and major cations, anions [chloride, sulfate, nitrate], and dissolved ammonium). Files are included in this dataset include: sample locations and depths in both a kmz file which can be opened in Google Earth and a csv file, aqueous geochemistry data in a csv files for Meander C and Meander Z, and analytical detection limits in a csv file. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.

54 ENVIRONMENTAL SCIENCES↗

Multi-task Parallelism for Robust Pre-training of Graph Foundation Models on Multi-source, Multi-fidelity Atomistic Modeling Data

Graph foundation models using graph neural networks promise sustainable, efficient atomistic modeling. To tackle challenges of processing multi-source, multi-fidelity data during pre-training, recent studies employ multi-task learning, in which shared message passing layers initially process input atomistic structures regardless of source, then route them to multiple decoding heads that predict data-specific outputs. This approach stabilizes pre-training and enhances a model’s transferability to unexplored chemical regions. Preliminary results on approximately four million structures are encouraging, yet questions remain about generalizability to larger, more diverse datasets and scalability on supercomputers. We propose a multi-task parallelism method that distributes each head across computing resources with GPU acceleration. Implemented in the open-source HydraGNN architecture, our method was trained on over 24 million structures from five datasets and tested on the Perlmutter, Aurora, and Frontier supercomputers, demonstrating efficient scaling on all three highly heterogeneous super-computing architectures.

Lupo Pasini, Massimiliano [ORNL] (ORCID:0000000249↗

OpenCRUMS USA: An Open Machine Learning Framework for Characterizing Variability in Aerosol Reanalysis Data

Advances in artificial intelligence (AI) have called for exploring how these techniques can be used for exploring patterns in large climate datasets. To that regard, the U.S. Department of Energy AI for Earth System Predictability (AI4ESP) supported a pilot initiative called the Open Classification of Regimes in the Southeast USA (OpenCRUMS USA) project to explore how AI can be used to characterize modes of spatial variability in large climate datasets. For this study, we focus on comparing two methods for characterizing the modes of spatial variability of surface aerosol concentration over the Houston region: empirical orthogonal functions (EOFs) and layerwise relevance propagation (LRP) applied to a convolutional neural network (CNN) classifier. We show that EOF analysis typically attributes spatial variability modes that span all of southeast Texas, prohibiting the attribution of spatial variability to localized regions. However, using LRP on the CNN classifier resolves the explanatory parameters at a finer spatial resolution than EOFs. This allows for the attribution of the spatial variability of surface aerosols to local regions of organic carbon which was not possible using EOFs. In addition, the LRP analysis also suggests that synoptic-scale transport of dust is most prevalent during anticyclonic and pretrough synoptic conditions as categorized by self-organizing maps.

54 ENVIRONMENTAL SCIENCES↗

NanoPSD: A software for automatic detection of Nano-Particle Shape Distribution in electron microscopy images

Accurate quantification of the size and morphology of nanoparticles from electron microscopy (EM) images is essential to understand growth mechanisms, surface reactivity, and functional behavior in nanoscale materials. Manual analysis remains slow, subjective, and difficult to reproduce in large datasets. We introduce NanoPSD (Nano-Particle Shape Distribution), an open-source and fully automated framework for quantitative particle detection and morphology analysis from EM images. NanoPSD integrates adaptive contrast enhancement, polarity-agnostic scale-bar detection, Optical Character Recognition (OCR)-based calibration, and classical segmentation via Otsu thresholding with morphological refinement. Particle contours are used to extract geometric descriptors, including equivalent circular diameter, aspect ratio, circularity, and solidity, enabling automated classification into spherical, rod-like, and aggregate morphologies. The framework supports both single-image and batch processing, generating publication-quality visualizations, LaTeX-ready tables, and structured comma-separated values (CSV) datasets. As a demonstration, we applied NanoPSD to plasma-synthesized nanoparticle samples diagnosed via transmission electron microscopy (TEM). The code produced statistically robust size and morphology distributions spanning a few to tens of nanometers with minimal user supervision. The pipeline demonstrates high reproducibility and scalability, processing large image collections with consistent calibration and output formatting. Its modular design enables seamless integration of future deep-learning-based segmentation models, providing a pathway toward intelligent, data-driven electron microscopy analysis.

36 MATERIALS SCIENCE↗

Multi-Artifact Analysis of Self-Admitted Technical Debt in Scientific Software

Context: Self-admitted technical debt (SATD) occurs when developers acknowledge shortcuts in code. In scientific software (SSW), such debt poses unique risks to the validity and reproducibility of results. Objective: This study aims to identify, categorize, and evaluate scientific debt, a specialized form of SATD in SSW, and assess the extent to which traditional SATD categories capture these domain-specific issues. Method: We conduct a multi-artifact analysis across code comments, commit messages, pull requests, and issue trackers from 23 open-source SSW projects. We construct and validate a curated dataset of scientific debt, develop a multi-source SATD classifier to guide SATD management, and conduct a practitioner validation to assess the practical relevance of scientific debt. Results: Our classifier performs strongly across 900,358 artifacts from 23 SSW projects. SATD is most prevalent in pull requests and issue trackers, underscoring the value of multi-artifact analysis. Models trained on traditional SATD often miss scientific debt, emphasizing the need for its explicit detection in SSW. Practitioner validation confirmed that scientific debt is both recognizable and useful in practice. Conclusions: Scientific debt represents a unique form of SATD in SSW that that is not adequately captured by traditional categories and requires specialized identification and management. Our dataset, classification analysis, and practitioner validation results provide the first formal multi-artifact perspective on scientific debt, highlighting the need for tailored SATD detection approaches in SSW.

Melin, Eric [Boise State University]↗

EV-ELM (Electric Vehicle Policies with the Energy Language Model) [SWR-25-156]

Electric Vehicle Policies with the Energy Language Model (EV-ELM) leverages previous work using Large Language Models (LLMs) to find, download, and parse policy information related to energy infrastructure. In this application, we use LLMs to find policy documents related to the permitting and installation of electric vehicle charging infrastructure. This software contains the code to find, download, and parse these documents, while a related data record in the Open Energy Data Initiative (OEDI) will include the resulting output dataset that can be used for downstream analysis. The EV-ELM repository contains code for the EV-ELM project, which focuses on retrieving and processing EV permitting processes using large language models. The project is composed of two pipelines: (1) a web scraping pipeline for discovering and downloading EV permitting documents, and (2) a document parsing and extraction pipeline that processes the downloaded files to produce structured data. The web scraping pipeline is designed to extract relevant information from various websites, while the document parsing pipeline processes and analyzes the extracted documents to derive meaningful insights. Both pipelines depend on the NLR elm repository, which provides essential tools and functionalities for handling and processing the data. The web scraping pipeline is a modified version of the ordinance_gpt example within the elm repository. It has been adapted to fit the specific requirements of the EV-ELM project, ensuring that it effectively captures and processes the necessary information related to EV permitting.

Olson, Reid [National Laboratory of the Rockies (N↗

2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction Study

# 2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction Study The 2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction (NCIR) Study provided insight into the travel patterns and associated energy consumption of participants. Study results helped researchers identify opportunities for the development of policies that could incentivize the use of alternative modes of travel such as transit and micromobility. Such travel modes reduce congestion by reducing the miles traveled by privately owned vehicles in urban and rural areas. The [National Institute for Congestion Reduction](https://nicr.usf.edu/) provides multimodal congestion reduction strategies through real-world deployments that leverage advances in technology, big data science, and innovative transportation options to optimize the efficiency and reliability of the transportation system for all users. ## Data Collection Agency The University of Puerto Rico at Mayagüez conducted the study. ## Survey Methodology The study was conducted in Spanish. Data collection was enabled via the open-source [NREL OpenPATH platform](https://www.nrel.gov/transportation/openpath). The resulting dataset consists of partially automated travel diaries—combining sensed and surveyed data reflecting patterns of multimodal, end-to-end, individual human mobility—as well as demographic and socioeconomic information from the 17 participants. ## Survey Records, Data, and Documentation Study records include 17 participants. The total number of trips was 458 and total non-air-miles traveled was approximately 1,469.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

2024 University of Puerto Rico at Mayagüez Civic Innovation Challenge Study

# 2024 University of Puerto Rico at Mayagüez Civic Innovation Challenge Study The 2024 University of Puerto Rico at Mayagüez Civic Innovation Challenge (CIVIC) Study provided insight into the travel patterns and associated energy consumption of participants. Study results helped researchers identify opportunities for the development of shared mobility strategies—such as collaborative ride-sharing programs—that could address the mobility needs of rural communities in Puerto Rico. The Civic Innovation Challenge is a multiagency, federal government research and action competition that funds ready-to-implement, research-based pilot projects that have the potential for scalable, sustainable, and transferable impact on community-identified priorities. ## Data Collection Agency The University of Puerto Rico at Mayagüez conducted the study. ## Survey Methodology The study was conducted in Spanish. Data collection was enabled via the open-source [NREL OpenPATH platform](https://www.nrel.gov/transportation/openpath). The resulting dataset consists of partially automated travel diaries—combining sensed and surveyed data reflecting patterns of multimodal, end-to-end, individual human mobility—as well as demographic and socioeconomic information from the 31 participants. ## Survey Records, Data, and Documentation Survey records include 31 participants. The total number of trips was 1,373 and the total non-air-miles traveled was approximately 8,260.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction Study

# 2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction Study The 2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction (NCIR) Study provided insight into the travel patterns and associated energy consumption of participants. Study results helped researchers identify opportunities for the development of policies that could incentivize the use of alternative modes of travel such as transit and micromobility. Such travel modes reduce congestion by reducing the miles traveled by privately owned vehicles in urban and rural areas. The [National Institute for Congestion Reduction](https://nicr.usf.edu/) provides multimodal congestion reduction strategies through real-world deployments that leverage advances in technology, big data science, and innovative transportation options to optimize the efficiency and reliability of the transportation system for all users. ## Data Collection Agency The University of Puerto Rico at Mayagüez conducted the study. ## Survey Methodology The study was conducted in Spanish. Data collection was enabled via the open-source [NREL OpenPATH platform](https://www.nrel.gov/transportation/openpath). The resulting dataset consists of partially automated travel diaries—combining sensed and surveyed data reflecting patterns of multimodal, end-to-end, individual human mobility—as well as demographic and socioeconomic information from the 17 participants. ## Survey Records, Data, and Documentation Study records include 17 participants. The total number of trips was 458 and total non-air-miles traveled was approximately 1,469.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

2024 University of Puerto Rico at Mayagüez Civic Innovation Challenge Study

# 2024 University of Puerto Rico at Mayagüez Civic Innovation Challenge Study The 2024 University of Puerto Rico at Mayagüez Civic Innovation Challenge (CIVIC) Study provided insight into the travel patterns and associated energy consumption of participants. Study results helped researchers identify opportunities for the development of shared mobility strategies—such as collaborative ride-sharing programs—that could address the mobility needs of rural communities in Puerto Rico. The Civic Innovation Challenge is a multiagency, federal government research and action competition that funds ready-to-implement, research-based pilot projects that have the potential for scalable, sustainable, and transferable impact on community-identified priorities. ## Data Collection Agency The University of Puerto Rico at Mayagüez conducted the study. ## Survey Methodology The study was conducted in Spanish. Data collection was enabled via the open-source [NREL OpenPATH platform](https://www.nrel.gov/transportation/openpath). The resulting dataset consists of partially automated travel diaries—combining sensed and surveyed data reflecting patterns of multimodal, end-to-end, individual human mobility—as well as demographic and socioeconomic information from the 31 participants. ## Survey Records, Data, and Documentation Survey records include 31 participants. The total number of trips was 1,373 and the total non-air-miles traveled was approximately 8,260.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction Study

# 2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction Study The 2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction (NCIR) Study provided insight into the travel patterns and associated energy consumption of participants. Study results helped researchers identify opportunities for the development of policies that could incentivize the use of alternative modes of travel such as transit and micromobility. Such travel modes reduce congestion by reducing the miles traveled by privately owned vehicles in urban and rural areas. The [National Institute for Congestion Reduction](https://nicr.usf.edu/) provides multimodal congestion reduction strategies through real-world deployments that leverage advances in technology, big data science, and innovative transportation options to optimize the efficiency and reliability of the transportation system for all users. ## Data Collection Agency The University of Puerto Rico at Mayagüez conducted the study. ## Survey Methodology The study was conducted in Spanish. Data collection was enabled via the open-source [NREL OpenPATH platform](https://www.nrel.gov/transportation/openpath). The resulting dataset consists of partially automated travel diaries—combining sensed and surveyed data reflecting patterns of multimodal, end-to-end, individual human mobility—as well as demographic and socioeconomic information from the 17 participants. ## Survey Records, Data, and Documentation Study records include 17 participants. The total number of trips was 458 and total non-air-miles traveled was approximately 1,469.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗