Search NASA⌕ Search

SEARCH · Search NASA

Results for “open data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

PDEHats

This is code used to train and evaluate neural partial differential equation solvers on an open source fluid flow data. We evaluate two standard deep learning algorithms for their ability to generalize, a desirable capability for trusthworthy and performant models.

Amarel, James↗

xCDAT: A Python Package for Simple and Robust Analysis of Climate Data

xCDAT (Xarray Climate Data Analysis Tools) is an open-source Python package that extends Xarray (Hoyer & Hamman, 2017) for climate data analysis on structured grids. xCDAT streamlines analysis of climate data by exposing common climate analysis operations through a set of straightforward APIs. Some of xCDAT’s key features include spatial averaging, temporal averaging, and regridding. These features are inspired by the Community Data Analysis Tools (CDAT) library (Dean N. Williams et al., 2009) (D. N. Williams, 2014) (Doutriaux et al., 2019) and leverage powerful packages in the Xarray ecosystem including xESMF (Zhuang et al., 2023), xgcm (Abernathey et al., 2022), and CF xarray (Cherian et al., 2023). To ensure general compatibility across various climate models, xCDAT operates on datasets that are compliant with the Climate and Forecast (CF) metadata conventions (Hassell et al., 2017).

54 ENVIRONMENTAL SCIENCES↗

WELLBASE - An Interactive Platform for Wellbore Material Assessment

This project seeks to build an open-source wellbore material data repository with adequate material performance and contextual data to support Geological Carbon Storage (GCS). By appropriately evaluating the data types as mentioned earlier made available by the WELLBASE tool, stakeholders can make more informed decisions regarding well selections, risk assessment, and economic analysis for geologic carbon storage projects. Advanced Natural Language Processing models and other custom python scripts will be deployed in an automated process to extract unstructured data from documents, reports, and web applications and subsequently parse to more usable formats. The processed data will then be integrated into a robust and comprehensive database architecture, optimizing data accessibility, and usability for analytical purposes. The final data products will be accessible through a user-friendly visualization platform that will allow users to query and visualize the data, as well as download data in usable formats.

Tetteh, Daniel A.↗

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity↗

Outdoor Deployment Data for a Four-Terminal GaAs//Si Tandem Solar Mini-Module

This dataset contains the complete outdoor measurement and analysis data for a mechanically stacked, four-terminal (4T) gallium arsenide (GaAs)//silicon (Si) tandem solar mini-module deployed from October 2019 to January 2021 at the Solar Radiation Research Laboratory (SRRL) in Golden, Colorado, USA. The data support a performance modeling and degradation analysis framework for tandem photovoltaic devices, as described in the accompanying publication. The dataset includes: (1) current–voltage (J–V) characteristics of each sub-cell measured approximately every five minutes, with extracted performance parameters; (2) spectral irradiance from an EKO MS-710 WISER spectroradiometer, along with derived spectral mismatch ratios (SMR) and average photon energy (APE); (3) one-minute resolution meteorological data from the co-located SRRL weather station and GPS-derived precipitable water vapor (PWV); (4) pre-deployment laboratory characterization (external quantum efficiency, J–V curves, standard test conditions parameters); (5) outdoor-extracted temperature and PWV correction coefficients; and (6) PVcircuit equivalent-circuit simulation outputs used for model validation. Degradation rates of −4.1 ± 0.2 %/year (GaAs) and −2.5 ± 0.9 %/year (Si) were determined using a filtering and normalization methodology adapted for fixed-tilt tandem modules. All data are provided in open, portable formats (Apache Parquet, CSV, JSON) to enable full reproducibility of the published analysis.

14 SOLAR ENERGY↗

Infrastructure-Based Cooperative Perception at a Traffic Intersection: Overview and Challenges: Preprint

Recent advancement in autonomous driving vehicles and V2X communication has attracted increasing attention towards Intelligent Transportation Systems to build a safe and reliable traffic intersection. However, most of the systems are still at the initial stages and require significant progress to become a reality. This paper presents an overview of NREL Infrastructure Perception and Control (IPC) framework which is an open-source track-data fusion engine which takes input from infrastructure-based perception sensors and cooperatively shared messages from Connected Autonomous Vehicles (CAVs) and Connected Vehicle (CVs) and the challenges associated with deploying such cooperative perception framework at a four-way traffic intersection in the city of Colorado Springs, CO, USA. The sensor data is collected by deploying two radars and two LiDAR sensors on the IPC mobile lab and two radars on diagonally opposite traffic poles at the proposed intersection. The sensor output results imply the need for rapid sensor calibration to bring the collective perception to a common coordinate frame, the importance of time synchronization between the sensors in order to capture accurate spatial and temporal alignment of the objects, and the need for a health monitoring system with fail safe closed-loop detection model for real-time deployment.

camera↗

REDI – Readiness Engine for Data Integration

The Readiness Engine for Data Integration (REDI) is an open-source framework for automating, standardizing, and assessing the process of preparing scientific data for AI training. REDI implements a five-stage pipeline (ingest, preprocess, transform, structure, output) with per-stage provenance instrumentation via Flowcept, domain-aware transformation logic (PII anonymization, regridding, graph encoding, and more), and built-in readiness assessment and validation modes. REDI has been evaluated across climate, proteomics, materials science, and nuclear fusion datasets, demonstrating near-ideal parallel scaling to 100 nodes on OLCF's Frontier system. REDI is deployable as an agent-callable skill in coding environments such as Claude Code and OpenAI Codex, and is complemented by SetGo for FAIR compliance and catalog publication.

Brewer, Wesley [Oak Ridge National Laboratory (ORN↗

sdt (Solar Data Tools) [SWR-25-130]

Solar Data Tools (sdt) is an open-source Python library for analyzing PV power (and irradiance) time-series data. It was developed to enable analysis of unlabeled PV data, i.e. with no model, no meteorological data, and no performance index required, by taking a statistical signal processing approach in the algorithms used in the package’s main data processing pipeline. Solar Data Tools empowers PV system fleet owners or operators to analyze system performance a hundred times faster even when they only have access to the most basic data stream—power output of the system.

Meyers-Im, Bennet [National Laboratory of the Rock↗

User Guide for Sample Reduction at GP-SANS

This manual is intended as a quick guide for data reduction of GP-SANS data. It includes all necessary steps to do the data reduction based on absolute calibration using the open beam method and how to transfer the reduced data to the personal computer system. If any errors are coming up so that the reduction script is not functioning as intended, please contact the instrument scientist.

97 MATHEMATICS AND COMPUTING↗

Datasets of Faults in Variable Air Volume Terminal Units in a Multi-Zone Commercial Building

Faults in HVAC systems can decrease system efficiency and equipment lifespan, leading to 5%–30% of energy consumption being wasted in commercial buildings. We identified two common faults in HVAC variable air volume systems: a stuck damper fault in the variable air volume terminal unit and a discharge airflow sensor fault. We conducted three sets of damper stuck tests and two sets of airflow sensor tests, each including a fault-free scenario and scenarios with varying levels of faults, over one day. The faults were implemented in Oak Ridge National Laboratory’s two-story Flexible Research Platform building to generate a high-quality, well-controlled dataset covering fault-induced and fault-free scenarios. The test building, fault test scenarios, and data validation are described here. The open-source dataset includes 1 min intervals of weather and building data on the presence and absence of building faults. This dataset can be used to analyze the effects of HVAC system faults on system operation and indoor building conditions, and to develop or evaluate a fault detection and diagnosis algorithm.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Prediction of Distributed River Sediment Respiration Rates Using Community-Generated Data and Machine Learning

River sediment microbial respiration is a key indicator of ecosystem functioning and the biogeochemical fluxes across this critical zone link surface and subsurface waters. As such, there is tremendous interest in measuring and mapping these respiration rates. Respiration observations are expensive and labor intensive; there is limited data available to the community. An open science, collaborative initiative is collecting samples for respiration rate analysis and multi-scale metadata; this evolving data set is being used for making machine learning (ML) predictions at unsampled sites to help inform continued community engagement. However, it is a challenge to find an optimum configuration for ML models to work with this feature-rich (i.e., 100+ possible input variables) data set. Here, we present results from a two-tiered approach to managing the analysis of this complex data set: (a) a stacked ensemble of models that automatically optimizes hyperparameters and manages the training of many models and (b) feature permutation importance to detect the most important features in the models. The major elements of this workflow are modular, portable, open, and cloud-based thus making this implementation a potential template for other applications. The models developed here predict that sediment organic matter chemistry is one of the most important features for predicting sediment respiration rate. Other larger-scale, important features fall into the categories of climatic, ecological, geological, and fluvial settings. Leveraging these larger-scale features to generate data-driven estimates of river sediment respiration rates reveals spatially consistent but heterogeneous patterns across the river network of the Columbia River Basin.

54 ENVIRONMENTAL SCIENCES↗

mzPeak: Designing a Scalable, Interoperable, and Future-Ready Mass Spectrometry Data Format

Advances in mass spectrometry (MS) instrumentation, such as higher resolution, faster scan speeds, and improved sensitivity, have significantly increased the volume and complexity of data. The growing adoption of imaging and ion mobility further amplifies these challenges across MS-based omics fields, including proteomics, metabolomics, and lipidomics. While these technologies unlock new possibilities, they also present significant challenges in data management, storage, and accessibility. Existing open formats, such as the XML-based community standards mzML and imzML, struggle to meet the demands of modern MS workflows due to their large file sizes, slow data access, and limited metadata support. Vendor-specific formats, while optimized for proprietary instruments, lack interoperability, comprehensive metadata support and long-term archival reliability. This white paper lays the groundwork for mzPeak, a next-generation community data format designed to address these challenges and support high-throughput, multi-dimensional MS workflows. By adopting a hybrid model that combines efficient binary storage for numerical data and both human and machine-readable metadata storage, mzPeak will reduce file sizes, accelerate data access, and offer a scalable, adaptable solution for evolving MS technologies. For researchers, mzPeak will enable enhanced interoperability across platforms, seamless support for complex workflows including ion mobility and MS imaging, and faster data access compared to existing community formats such as mzML. Its design will ensure data is managed in compliance with regulatory standards, essential for applications such as precision medicine and chemical safety, where long-term data integrity and accessibility are critical. For vendors, mzPeak provides a streamlined, open alternative to proprietary formats, reducing the burden of regulatory compliance while aligning with the industry's push for transparency and standardization. By offering a high-performance, interoperable solution, mzPeak positions vendors to meet customer demands for sustainable data management tools which will be able to handle emerging and future data types and workflows. mzPeak aspires to become the cornerstone of MS data management, empowering researchers, vendors, and developers to innovate and collaborate more effectively.

data formats↗

BEAST DB: Grand-Canonical Database of Electrocatalyst Properties

We present BEAST DB, an open-source database comprised of ab initio electrochemical data computed using grand-canonical density functional theory in implicit solvent at consistent calculation parameters. The database contains over 20,000 surface calculations and covers a broad set of heterogeneous catalyst materials and electrochemical reactions. Calculations were performed at self-consistent fixed potential as well as constant charge to facilitate comparisons to the computational hydrogen electrode. This article presents common use cases of the database to rationalize trends in catalyst activity, screen catalyst material spaces, understand elementary mechanistic steps, analyze the electronic structure, and train machine learning models to predict higher fidelity properties. Users can interact graphically with the database by querying for individual calculations to gain a granular understanding of reaction steps or by querying for an entire reaction pathway on a given material using an interactive reaction pathway tool. BEAST DB will be periodically updated, with planned future updates to include advanced electronic structure data, surface speciation studies, and greater reaction coverage.

database↗

AmeriFlux FLUXNET-1F US-Dk1 Duke Forest-open field

This is the AmeriFlux Management Project (AMP) created FLUXNET-1F version of the carbon flux data for the site US-Dk1 Duke Forest-open field. This is the FLUXNET version of the carbon flux data for the site US-Dk1 Duke Forest-open field produced by applying the standard ONEFlux (1F) software. Site Description - The Duke Forest grass field is approximately 480×305 m, dominated by the C3 grass Festuca arundinacea Shreb. (tall fescue) includes minor components of C3 herbs and the C4 grass Schizachyrium scoparium (Michx.) Nash, not considered here. The site was burned in 1979 and is mowed annually during the summer for hay according to local practices. Lai, C.T. and G.G. Katul, 2000, "The dynamic role of root-water uptake in coupling potential to actual transpiration" , Advances in Water Resources, 23, 427-439; Novick , K.A., P. C. Stoy, G. G. Katul, D. S. Ellsworth, M. B. S. Siqueira, J. Juang, R. Oren, 2004, Carbon dioxide and water vapor exchange in a warm temperate grassland, Oecologia, 138, 259-274; Stoy PC, Katul GG, Siqueira MBS, Juang J-Y, McCarthy HR, Oishi AC, Uebelherr JM, Kim H-S, Oren R (2006). Separating the effects of climate and vegetation on evapotranspiration along a successional chronosequence in the southeastern U.S. Global Change Biology 12:2115-2135

Oishi, Chris [USDA Forest Service]↗

Aligning NASA Earth Science Data Stewardship with FAIR Principles: Outcomes, Recommendations, and Future Directions

The FAIR Principles—Findable, Accessible, Interoperable, and Reusable—offer a widely accepted framework for improving the sharing and reuse of digital scientific data by both human and machine users. Following these principles is critical for effective scientific data stewardship, broader scientific collaboration, and compliance with federal and agency data policies. This paper, based on the work of NASA’s Open, Free, and FAIR Working Group (O’FAIR WG) under the Earth Science Data Systems Program, presents an overview of how FAIR is being applied within NASA’s Earth science data landscape. It highlights ongoing progress and challenges, identifies FAIR-enabling resources, and offers recommendations and strategic actions to enhance the FAIRness of NASA-funded open and free Earth science data products. The FAIR-enabling resources identified underscore the vital role of NASA's existing enterprise processes, standards, tools, and infrastructures in supporting FAIR implementation. Our findings show strong performance in making NASA Earth science data more findable and accessible. However, further work is needed—especially in enhancing interoperability, so that different systems and tools can better understand and exchange data. This is especially important for enabling machine-driven discovery and analysis. We emphasize the importance of a balanced strategy that combines a centralized, top-down approach—focused on building enterprise-level capabilities and processes—with a decentralized, bottom-up approach driven by discipline-specific needs and community practices. We advocate for coordinated efforts to enhance (meta)data interoperability to facilitate seamless data and information sharing and exchange of Earth science data both within NASA and across other agencies managing Earth science data.

Data Product↗