Search NASA⌕ Search

SEARCH · Search NASA

Results for “Open Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Public Data Set: Impurity Dynamics and Radiative Losses During Local Helicity Injection Startup in the Pegasus-III Spherical Tokamak

This public dataset contains openly-documented, machine readable digital research data corresponding to figures published in C. Rodriguez Sanchez et al., “ Impurity Dynamics and Radiative Losses During Local Helicity Injection Startup in the Pegasus-III Spherical Tokamak,” accepted for publication in Physics of Plasmas .

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Announcing the Biomedical Data Translator: Initial Public Release

ABSTRACT The growing availability of biomedical data offers vast potential to improve human health, but the complexity and lack of integration of these datasets often limit their utility. To address this, the Biomedical Data Translator Consortium has developed an open‐source knowledge graph–based system—Translator—designed to integrate, harmonize, and make inferences over diverse biomedical data sources. We announce here Translator's initial public release and provide an overview of its architecture, standards, user interface, and core features. Translator employs a scalable, federated, knowledge graph framework for the integration of clinical, genomic, pharmacological, and other biomedical knowledge sources, enabling query retrieval, inference, and hypothesis generation. Translator's user interface is designed to support the exploration of knowledge relationships and the generation of insights, without requiring deep technical expertise and gradually revealing more detailed evidence, provenance, and confidence information, as needed by a given user. To demonstrate Translator's application and impact, we highlight features of the user interface in the context of three real‐world use cases: suggesting potential therapeutics for patients with rare disease; explaining the mechanism of action of a pipeline drug; and screening and validating drug candidates in a model organism. We discuss strengths and limitations of reasoning within a largely federated system and the need for rich concept modeling and deep provenance tracking. Finally, we outline future directions for enhancing Translator's functionality and expanding its data sources. Translator represents a significant step forward in making complex biomedical knowledge more accessible and actionable, aiming to accelerate translational research and improve patient care.

Research & Experimental Medicine↗

Development of Real-Time High-Density Pulsar Data Transmission and Processing for Grid Synchronization

Taking advantage of the extreme stability of the pulsar period, it can serve as the timing source for grid synchronization to compensate for the timing drift instigated by the loss of GPS signal. Nevertheless, the real-time transmission and processing of the pulsar data suffer from its high-frequency data rate, varying from megahertz to gigahertz, resulting in reduced computing speed and increased time delay. To mitigate this issue, the hardware and software frameworks are implemented for the high-density pulsar data transmission and processing for grid synchronization in this research. Initially, the high-density pulsar data is transferred using open-source software. The complementary duty cycle timing module is designed to coordinate the operation of the dual-channel high-speed interface and software. Subsequently, the multiple-threading is applied to the receiving, parsing, and splicing pulsar data. Next, the pulsar signal extraction method is implemented based on the polyphase filterbank and time of arrival estimation. Ultimately, real-time performance verification experiments are carried out for different components under two hardware platforms. Finally, the results demonstrate that only 0.482 s is required for processing 4 Gigabyte data through multiple-threading, which is 3.8 times faster than the single thread. The pulsar signal extraction can also be executed within 707 ms for 4.8 seconds of data, thereby indicating that real-time requirements can be met.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Data from: "Towards CONUS-Wide ML-Augmented Conceptually-Interpretable Modeling of Catchment-Scale Precipitation-Storage-Runoff Dynamics"

This data package was generated to support the manuscript “Towards CONUS-Wide Machine Learning-Augmented Conceptually Interpretable Modeling of Catchment-Scale Precipitation-Storage-Runoff Dynamics.” It provides input files, model outputs, plotting data, scripts, notebooks, and documentation used to develop, evaluate, and reproduce Mass-Conserving Perceptron (MCP)-based hydrologic modeling experiments across 513 selected Catchment Attributes and Meteorology for Large-sample Studies in the United States (CAMELS-US) basins. The files are organized by modeling component and analysis purpose, including rainfall–runoff experiments, snow module experiments, coupled hydrologic-snow experiments, Long Short-Term Memory (LSTM) benchmark results, model skill metrics, initialization and epoch records, cell-state normalization files, Akaike Information Criterion (AIC)-based model comparison files, and data used to generate manuscript figures. Tabular files can be opened using standard spreadsheet software or Python/R data-analysis tools. Python scripts, Jupyter notebooks, and selected MATLAB scripts are included for model execution, postprocessing, plotting, and statistical analysis. Quality assurance and quality control were conducted through the source-data selection and modeling workflow. Meteorological forcing, streamflow, and static catchment attributes were derived from the CAMELS-US dataset, and snow water equivalent data were derived from the University of Arizona (UA) Snow Water Equivalent dataset. Selected basins and time periods were screened during the associated research workflow to avoid missing observations or poor-quality cases. Static geospatial features were processed primarily using Quantum Geographic Information System (QGIS) and Geospatial Data Abstraction Library (GDAL) workflows. Additional details are provided in the associated manuscript and documentation.

ESS-DIVE CSV File Formatting Guidelines Reporting ↗

A Science Gateway for the Repeatable Analysis of Machine Learning Predicted Gravity Anomalies

In recent years, deep learning has become an increasingly popular alternative for modeling in geoscience applications due to its scalability and efficiency. However, the interpretability, compute, data volume, and hyperparameter tuning requirements of deep learning models make development and monitoring difficult. Furthermore, model explainability and communicating results obtained by these models to users or domain experts is a challenge, as domain experts in geoscience also need to have a deep understanding of how those models function in order to support their scientific works. Here, we describe a science gateway and machine learning pipeline for predicting gravity anomalies from geophysical data. The gateway, built on open-source technologies, provides a holistic view of the pipeline through interactive visualizations aimed at enabling efficient exploratory data analysis. The repeatability, reproducibility, and monitoring capabilities of this overall system allow us to iterate and analyze at scale. Using this pipeline and gateway, we can repeatedly produce accurate high-resolution gravity anomaly datasets. By describing the underlying technologies, implementation, and results, here we provide a foundation for the broader adoption of science gateways into cross-cutting geoscience and machine learning research projects as a means to improve the scientific discovery and collaboration in the geophysics and computational sciences community.

58 GEOSCIENCES↗

Cloverleaf Data Artifacts for ArtIMis LDRD

This report summarizes the use of the open-source CloverLeaf/CloverLeaf3D mini-apps to generate synthetic data sets to train foundation models for the ArtIMis LDRD DI. These data artifacts are intended to be used by LANL collaborators and shared externally with our university and institutional partners. Note that CloverLeaf/CloverLeaf3D is not a LANL simulation code.

97 MATHEMATICS AND COMPUTING↗

CatTestHub: A benchmarking database of experimental heterogeneous catalysis for evaluating advanced materials

The ability to quantitatively compare newly evolving catalytic materials and technologies is hindered by the widespread availability of catalytic data collected in a consistent manner. While certain catalytic chemistries have been widely studied across decades of scientific research, quantitative comparisons based on literature information is hindered by variability in reaction conditions, types of reported data, and reporting procedures. Here, we present CatTestHub, an open-access database dedicated to benchmarking experimental heterogeneous catalysis data. Combining systematically reported catalytic activity data for selected probe chemistries, with relevant material characterization and reactor configuration information, the database provides a collection of catalytic benchmarks for distinct classes of active site functionality. Through key choices in data access, availability, and traceability, CatTestHub seeks to balance the fundamental information needs of chemical catalysis and the FAIR data design principles. Details of the database architecture and the means through which to navigate it are presented, highlighting examples of catalytic insights readily drawn from the available benchmarking data. In its current iteration, CatTestHub spans over 250 unique experimental data points, collected over 24 solid catalysts, that facilitated the turnover of 3 distinct catalytic chemistries. Here, a roadmap is presented through which to expand the open-access platform that serves as a community wide benchmark, primarily through continuous addition of kinetic information on select catalytic systems by members of the heterogeneous catalysis community at large.

Benchmark↗

Impact of Vegetation Assimilation on Flash Drought Characteristics across the Continental United States

Predicting and managing the impacts of flash droughts is difficult owing to their rapid onset and intensification. Flash drought monitoring often relies on assessing changes in root-zone soil moisture. However, the lack of widespread soil moisture measurements means that flash drought assessments often use process-based model data like that from the North American Land Data Assimilation System (NLDAS). Such reliance opens flash drought assessment to model biases, particularly from vegetation processes. Here, we examine the influence of vegetation on NLDAS-simulated flash drought characteristics by comparing two experiments covering 1981–2017: open loop (OL), which uses NLDAS surface meteorological forcing to drive a land surface model using prognostic vegetation, and data assimilation (DA), which instead assimilates near-real-time satellite-derived leaf area index (LAI) into the land surface model. The OL simulation consistently underestimates LAI across the United States, causing relatively high soil moisture values. Both experiments produce similar geographic patterns of flash droughts, but OL produces shorter duration events and regional trends in flash drought occurrence that are sometimes opposite to those in DA. Across the Midwest and Southern United States, flash droughts are 4 weeks (about 70%) longer on average in DA than OL. Moreover, across much of the Great Plains, flash drought occurrence has trended upward according to the DA experiment, opposite to the trend in OL. This sensitivity of flash drought to the representation of vegetation suggests that representing plants with greater fidelity could aid in monitoring flash droughts and improve the prediction of flash drought transitions to more persistent and damaging long-term droughts.

Drought↗

PV Degradation Modeling: Applying Geospatial Workflows with "PVDeg"

Accurate degradation modeling is essential for predicting photovoltaic (PV) module performance, estimating longevity and informing design decisions. With degradation rates varying significantly by location, geospatial analysis is critical for PV and broader applications, such as agrivoltaics, weathering and environmental data analysis. This work presents PVDeg, an open-source tool designed for geospatial degradation analysis. PVDeg integrates meteorological data from global sources, including the National Solar Radiation Database (NSRDB) and Photovoltaic Geographical Information System (PVGIS), with degradation models. The toolkit enables users to customize geospatial workflows by integrating weather data, material parameters, and user-defined Python functions. It facilitates accelerated downloads of NSRDB and PVGIS datasets and optimizes geospatial point selection to preserve data density in regions of interest. Additionally, PVDeg provides a local database for storage and spatial queries, supporting large-scale analyses without the need for high-performance computing (HPC) resources. PVDeg provides a foundational workflow that extends its utility beyond PV applications, enabling researchers to analyze geospatial processes across discipline.

14 SOLAR ENERGY↗

A Performant, Scalable Processing Pipeline for High‐Quality and FAIR Environmental Sensor Data

High-resolution environmental monitoring is necessary to record, understand, and predict biogeochemical and ecological changes particularly in coastal systems but brings significant challenges in processing and making rapidly available the resulting data. The COMPASS-FME project established a network of coastal observational sites across the Chesapeake Bay and western Lake Erie regions extensively instrumented with soil, vegetation, and weather sensors logging data every 15 min. Our data processing framework, written in R and completely open source, prioritizes rapid model-experiment iteration and makes biogeochemical data rapidly available for quality assurance/quality control, analysis, and model ingestion. This pipeline is distinguished by a standardized and modular approach to data curation, extensive metadata and documentation, and its high performance. These attributes combine to make biogeochemical data rapidly accessible across COMPASS-FME and the broader community. Flexible, powerful, and reproducible approaches to handling high-volume environmental data are crucial for accelerating biogeosciences research.

Pennington, Stephanie C. [Pacific Northwest Nation↗

A Survey of Open-Source Tools for Transmission and Distribution Systems Research

This work presents a review of open-source electric power transmission and distribution systems analysis tools suitable for use by industry professionals and academic researchers. Due to the high complexity of the electric grid, there exist numerous tools and extensive research pertaining to nearly every aspect of the design, operation, and control of transmission and distribution networks. In addition to the commercial tools, a wide range of free, open-source tools, models, and data usable by the scientific community for related research have been developed by different organizations, including both international and US universities and national laboratories. However, due to the absence of a catalog of available tools, models and data, researchers often lack a knowledge of existing capabilities and may develop duplicative software and tools. Increasing awareness of these available resources seeks to accelerate their broader use, leading to more efficient and standardized grid analysis. This review paper (which is part of a larger survey effort that studied over 400 tools in the transmission, distribution, buildings, and electric vehicles space) outlines selected open-source resources that have been developed in power transmission and distribution systems research. It is anticipated that this work can serve as a guide for industry and academic researchers alike, ensuring that research efforts are well-channeled.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Novel Results Visualization for Dynamic PSA and New Modeling Features in EMRALD

The Event Modeling Risk Assessment Linked Diagram (EMRALD) tool, developed at the Idaho National Laboratory (INL), was designed to simplify the creation of dynamic models and support various research projects. One of the primary goals of EMRALD was to provide visual methods for modeling. EMRALD consists of two main components: a web-based user interface for model development and a solve engine for running model simulations. Over time, it has evolved to meet the diverse needs of its users. Initially, EMRALD's results were simple text outputs with final key state percentages and uncertainty bounds. However, because EMRALD utilizes a three-phase discrete event simulation and tracks the paths of each simulation run leading to a key state, there is significant potential to analyze large sets of path results data, including state paths, events, and timing. Visualizing this data meaningfully posed a challenge. To address this, a novel time-based Sankey diagram was developed. EMRALD exports results data in a format that can be opened in this Sankey viewer, allowing users to visualize paths, occurrences, events, and probability data for the entire simulation run in a single diagram. Moreover, when EMRALD was first created, there were limited tools capable of meeting its graphical requirements, many of which are no longer supported. In 2024, a new web-based interface was developed using modern graphing tools, enabling additional modeling features. This paper discusses the new dynamic PSA results visualization capability and the enhanced modeling tools available in EMRALD.

97 - MATHEMATICS AND COMPUTING↗

A Proxy Method to Bridge LCA Data Gaps Using Automated Material Classification and Probabilistic Under-Specification

Life cycle assessments (LCAs) are essential for understanding the environmental impacts of material production. However, gaps in life cycle inventory (LCI) data for material and chemical inputs present a key challenge for LCA practitioners, especially in the early design stages. Strategies for filling in these gaps require additional time and expertise, which can hinder the LCA’s completion. This study combined automatic material classification and probabilistic under-specification to create a time-efficient method to fill material LCI data gaps. To illustrate the proposed method, proxy environmental impact distributions were generated using publicly available material LCI data classified into the ChemOnt chemical taxonomy using the open-source chemical classification software ClassyFire. Input materials with data gaps were then classified into the same taxonomy, where proxy environmental impact values could be selected from the available distributions to quickly fill in any data gaps. Although these methods were applied to classify material production processes available in the Federal LCA Commons and Ecoinvent databases, they can be applied to any LCA database. This study shows that classifying materials by their chemical structure produces taxonomies with increased granularity relative to industrial classification, improving the ability of under-specified proxy data to be used for differentiating the environmental impacts of competing designs.

biological databases↗

Machine learning tools for epigenetics

The software provides machine learning analysis and visualization to detect patterns in epigenetic data, including conventional machine learning and statistical methods, and open-source packages like pyBigWig (https://github.com/deeptools/pyBigWig) for data processing. The software is written in python, it uses some python libraries.

Kim, Anastasiia↗

BASIN-3D Data Integration for Selected ARM Data Field Campaign Report

The purpose of this data services request was to demonstrate integration of the Atmospheric Radiation Measurement (ARM) User Facility’s “met” datastreams with time series data from other earth science data sources using the BASIN-3D data synthesis software tool. BASIN-3D is an open-source Python library that enables researchers to integrate data across configured public and private data sources. It provides a common query language for researchers to request measurement locations and time series data based on specified locations, variables, time period, statistics, aggregation, and data quality. BASIN-3D acquires the data that match the query from each configured data source and translates the results into harmonized vocabularies, thus reducing researchers' data-wrangling effort. In addition, because the queries are executed on demand, researchers can easily regenerate their synthesized data sets as new data and/or data updates become available, eliminating one-off data products. BASIN-3D can output data using a variety of different data structures for end-user applications including Python pandas data frames and hdf5 output formats.

54 ENVIRONMENTAL SCIENCES↗

Harnessing large language models’ zero-shot and few-shot learning capabilities for regulatory research

Abstract Large language models (LLMs) are sophisticated AI-driven models trained on vast sources of natural language data. They are adept at generating responses that closely mimic human conversational patterns. One of the most notable examples is OpenAI's ChatGPT, which has been extensively used across diverse sectors. Despite their flexibility, a significant challenge arises as most users must transmit their data to the servers of companies operating these models. Utilizing ChatGPT or similar models online may inadvertently expose sensitive information to the risk of data breaches. Therefore, implementing LLMs that are open source and smaller in scale within a secure local network becomes a crucial step for organizations where ensuring data privacy and protection has the highest priority, such as regulatory agencies. As a feasibility evaluation, we implemented a series of open-source LLMs within a regulatory agency’s local network and assessed their performance on specific tasks involving extracting relevant clinical pharmacology information from regulatory drug labels. Our research shows that some models work well in the context of few- or zero-shot learning, achieving performance comparable, or even better than, neural network models that needed thousands of training samples. One of the models was selected to address a real-world issue of finding intrinsic factors that affect drugs' clinical exposure without any training or fine-tuning. In a dataset of over 700 000 sentences, the model showed a 78.5% accuracy rate. Our work pointed to the possibility of implementing open-source LLMs within a secure local network and using these models to perform various natural language processing tasks when large numbers of training examples are unavailable.

Biochemistry & Molecular Biology↗

The DECADE cosmic shear project IV: cosmological constraints from 107 million galaxies across 5,400 deg$^2$ of the sky

We present cosmological constraints from the Dark Energy Camera All Data Everywhere (DECADE) cosmic shear analysis. This work uses shape measurements for 107 million galaxies measured through Dark Energy Camera (DECam) imaging of $5,\!412$ deg$^2$ of sky that is outside the Dark Energy Survey (DES) footprint. We derive constraints on the cosmological parameters $S_8 = 0.791^{+0.027}_{-0.032}$ and $Ω_{\rm m} =0.269^{+0.034}_{-0.050}$ for the $Λ$CDM model, which are consistent with those from other weak lensing surveys and from the cosmic microwave background. We combine our results with cosmic shear results from DES Y3 at the likelihood level, since the two datasets span independent areas on the sky. The combined measurements, which cover $\approx\! 10,\!000$ deg$^2$, prefer $S_8 = 0.791 \pm 0.023$ and $Ω_{\rm m} = 0.277^{+0.034}_{-0.046}$ under the $Λ$CDM model. These results are the culmination of a series of rigorous studies that characterize and validate the DECADE dataset and the associated analysis methodologies (Anbajagane et. al 2025a,b,c). Overall, the DECADE project demonstrates that the cosmic shear analysis methods employed in Stage-III weak lensing surveys can provide robust cosmological constraints for fairly inhomogeneous datasets. This opens the possibility of using data that have been previously categorized as ``unusable'' for cosmic shear analyses, thereby increasing the statistical power of upcoming weak lensing surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Drone Flight Data Logs

This dataset represents the open-air tests for the drones when testing different flight scenarios. For some flights we created and tested with a set of onboard sensors. For others we used the native logs for the drones. We recorded relevant conditions for each of the flights to examine environmental issues and weight impacts. We also looked at segmentations of flights to investigate the energy used in each type of flight.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗