Search NASA⌕ Search

SEARCH · Search NASA

Results for “open datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Bragg Spot Finder (BSF): a new machine-learning-aided approach to deal with spot finding for rapidly filtering diffraction pattern images

Macromolecular crystallography contributes significantly to understanding diseases and, more importantly, how to treat them by providing atomic resolution 3D structures of proteins. This is achieved by collecting X-ray diffraction images of protein crystals from important biological pathways. Spotfinders are used to detect the presence of crystals with usable data, and the spots from such crystals are the primary data used to solve the relevant structures. Having fast and accurate spot finding is essential, but recent advances in synchrotron beamlines used to generate X-ray diffraction images have brought us to the limits of what the best existing spotfinders can do. This bottleneck must be removed so spotfinder software can keep pace with the X-ray beamline hardware improvements and be able to see the weak or diffuse spots required to solve the most challenging problems encountered when working with diffraction images. In this paper, we first present Bragg Spot Detection (BSD), a large benchmark Bragg spot image dataset that contains 304 images with more than 66 000 spots. We then discuss the open source extensible U-Net-based spotfinder Bragg Spot Finder (BSF), with image pre-processing, a U-Net segmentation backbone, and post-processing that includes artifact removal and watershed segmentation. Finally, we perform experiments on the BSD benchmark and obtain results that are (in terms of accuracy) comparable to or better than those obtained with two popular spotfinder software packages ( Dozor and DIALS ), demonstrating that this is an appropriate framework to support future extensions and improvements.

36 MATERIALS SCIENCE↗

Restructuring Big Data to Improve Data Access and Performance in Analytic Services Making Research More Efficient for the Study of Extreme Weather Events and Application User Communities

By developing and enhancing various services and tools, the GES DISC provides users with the capability to access and visualize data, and to make comparisons of data from multiple sensor and models via a number of cross-discipline projects. Discovering Data via Faceted Web Interface Web interface to data products and services Search and Download mechanisms Dataset Landing Pages Accessing Data through Interoperable Services: GDS – GrADS Data Server OPeNDAP - Open-source Project for a Network Data Access Protocol WMS – OGC service GIS connector – allowing IS tools to access data easier (coming soon) HTTPS -- direct online access Downloading Data Basics: Subset and egridding Service – Parameter, Spatial, Time, Vertical, Mean averaging, format conversion, and regridding for L3/L4 gridded data Swath Data Subsetter – Parameter, spatial subset of L2 /L1 data. Visualizing Data Online: Giovanni –Visualization and Analysis L3/L4 gridded data AIRS NRT Viewer – AIRS near-real-time DQVis – L2 data quality visualization

data cube↗

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge, reduce risk, and support safe, productive human space missions. Through the powerful emerging approaches of artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in biomedical science and engineered astronaut health systems, to enable Earth independence and autonomy of mission operations. Here we present an overview of AI/ML architecture to support deep space mission goals, developed with leaders in the field. First, we focus on the fundamental biological research that supports our understanding of physiological responses to spaceflight, and we describe current efforts to support AI/ML research including data standardization and data engineering through maximally open and FAIR (findable, accessible, interoperable, reusable) databases and the generation of AI-ready datasets for reuse and analysis. We also discuss remote data management frameworks for research data as well as environmental and health data that are generated during deep space missions. We highlight several research projects that leverage data standardization and management for fundamental biological discovery to uncover the complex effects of space travel on living systems. Next, we provide an overview of cutting-edge AI/ML approaches that can be integrated to support remote monitoring and analysis during deep space missions, including generative models and large language models to learn the underlying biomedical patterns and predict outcomes or answer questions during off world medical scenarios. We also describe current AI/ML methods to support this research and monitoring through automated cloud-based labs which enable limited human intervention and closed-loop experimentation in remote settings. These labs could support mission autonomy by analyzing environmental data streams, and would be facilitated through in situ analytics capabilities to avoid sending large raw data files through low bandwidth communications. Finally, in the context of deep space missions with limited communications or access to medical advice from Earth, we describe a solution for integrated, real-time mission biomonitoring across hierarchical levels from continuous environmental monitoring, to wearables and point-of-care devices, to molecular and physiological monitoring. We introduce a precision space health system that will ensure that the future of space health is predictive, preventative, participatory and personalized.

artificial intelligence↗

Re-Organizing Earth Observation Data Storage to Support Temporal Analysis of Big Data

The Earth Observing System Data and Information System archives many datasets that are critical to understanding long-term variations in Earth science properties. Thus, some of these are large, multi-decadal datasets. Yet the challenge in long time series analysis comes less from the sheer volume than the data organization, which is typically one (or a small number of) time steps per file. The overhead of opening and inventorying complex, API-driven data formats such as Hierarchical Data Format introduces a small latency at each time step, which nonetheless adds up for datasets with O(10^6) single-timestep files. Several approaches to reorganizing the data can mitigate this overhead by an order of magnitude: pre-aggregating data along the time axis (time-chunking); storing the data in a highly distributed file system; or storing data in distributed columnar databases. Storing a second copy of the data incurs extra costs, so some selection criteria must be employed, which would be driven by expected or actual usage by the end user community, balanced against the extra cost.

data storage↗

Empowering Geothermal Research: The Geothermal Data Repository's New AI Research Assistant: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has integrated a Large Language Model (LLM) with the metadata and supporting documents associated with GDR datasets to create an Artificially Intelligent (AI) research assistant. By leveraging work done to make GDR metadata machine-readable and an open-source LLM integration model called the Energy Language Model, developed by the National Renewable Energy Laboratory, AskGDR serves as a virtual research assistant to GDR users. It provides answers to a variety of user-provided questions using natural language processing and generative machine learning. Users can get answers to questions about specific datasets, including inquiries about the equipment, assumptions and methodologies used in the origination of the data; or more abstract questions, such as the applicability of data to specific research fields. AskGDR improves the discoverability of geothermal data by helping guide users to datasets beyond simple keyword searches. It enables users to find data based on properties of the data, discover information contained within supporting documents, and explore data from projects related to their research objectives.

access↗

PyCMG-based Simulation of Volumetric Concrete Microstructure

Concrete is a complex, heterogeneous material with a microstructure composed of aggregates, cement paste, and pores spanning multiple length scales. Understanding this microstructure is critical for advancing the performance, durability, and modeling of concrete-based systems. While experimental imaging such as X-ray computed tomography (XCT) provides valuable insights, generating large datasets with detailed ground truth annotations is both costly and labor-intensive due to challenges in segmenting similar phases, such as aggregates and cement paste, that often share similar attenuation properties. To address this, we developed a pipeline to simulate realistic 3D concrete microstructures using the open-source Python package PyCMG. This simulation effort focuses on generating high-fidelity, annotated microstructures that can serve as training or benchmarking datasets for image analysis, segmentation algorithms, and machine learning models, particularly in scenarios where experimental data is scarce.

Ziabari, Amir [Oak Ridge National Laboratory; ORNL↗

Open Energy Data Initiative (OEDI) FY22-24 (Final Technical Report)

Final technical report for the Open Energy Data Initiative (OEDI) project covering fiscal years FY22 through FY24. The DOE Open Energy Data Initiative (OEDI) is a partnership between the National Renewable Energy Laboratory (NREL), the U.S. Department of Energy (DOE), and major cloud providers including Amazon, Microsoft, and Google to provide universal access to big data in the cloud. At the heart of OEDI is a centralized repository of high-value energy research datasets aggregated from the U.S. Department of Energy's Program Offices, National Laboratories and other collaborators. It aggregates smaller, domain-specific repositories, allows direct data submissions, and includes support for big data through its energy data lakes. OEDI's data lakes make high-value data universally accessible and help researchers, collaborators and the general public overcome many of the obstacles to accessing and using big data.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Empowering Geothermal Research: The Geothermal Data Repository's New AI Research Assistant

The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has integrated a Large Language Model (LLM) with the metadata and supporting documents associated with GDR datasets to create an Artificially Intelligent (AI) research assistant. By leveraging work done to make GDR metadata machine-readable and an open-source LLM integration model called the Energy Language Model, developed by the National Renewable Energy Laboratory, AskGDR serves as a virtual research assistant to GDR users. It provides answers to a variety of user-provided questions using natural language processing and generative machine learning. Users can get answers to questions about specific datasets, including inquiries about the equipment, assumptions and methodologies used in the origination of the data; or more abstract questions, such as the applicability of data to specific research fields. AskGDR improves the discoverability of geothermal data by helping guide users to datasets beyond simple keyword searches. It enables users to find data based on properties of the data, discover information contained within supporting documents, and explore data from projects related to their research objectives. This paper will outline the development, integration, output, and efficacy of the AskGDR LLM, including adherence to scientific rigor through improvements designed to increase the accuracy of generated answers, avoid speculation, and provide proper references for all resources used.

access↗

CO2 and CH4 leaf-level fluxes and soil porewater concentrations from common vegetation patches in Louisiana’s coastal wetlands

This dataset contains leaf-level flux and soil porewater concentration measurements of carbon dioxide (CO2) and methane (CH4 ) in plots in the footprint of Ameriflux sites US-LA2 and US-LA3. Leaf fluxes in US-LA2 were measured on patches dominated by Sagittaria lancifolia and co-dominated by Sagittaria lancifolia and Typha latifolia. In US-LA3, fluxes were measured from distinct Juncus roemerianus and Spartina alterniflora patches. The porewater concentrations were collected across a vertical profile (~50 cm depth) at centric locations within 25 m2 plots where we measured the leaf fluxes. US-LA3 included an additional set of measurements in open water spots. We aimed to evaluate differences in leaf fluxes and porewater pools of CO2 and CH4 of representative ecohydrological patches across a salinity gradient. We also used this dataset to help develop ELM-Wet, a more realistic representation of wetland carbon biogeochemical processes within the U.S. Department of Energy’s Energy Exascale Earth System Model (E3SM) Land Model version 1 (ELM v.1). The files can be opened with regular text editors or spreadsheet programs. Version 2.0 (8/26/2025): This is the latest version of this dataset. The update includes additional samples of soil porewater CH4/CO2 concentrations from June-2021 to November-2022, as well as minor adjustments made to V1 samples via changing Henry's solubility to account for porewater salinity. Additionally leaf-level measurments of spectral indices, PSRI, NDVI, and PRI have been added to complement Leaf-level flux measurements. All V1 data sets have been integrated into V2 sheets, ensuring data from the previous version is contained with the additional samples and consistent with V2 metadata.

54 ENVIRONMENTAL SCIENCES↗

Announcing the Biomedical Data Translator: Initial Public Release

ABSTRACT The growing availability of biomedical data offers vast potential to improve human health, but the complexity and lack of integration of these datasets often limit their utility. To address this, the Biomedical Data Translator Consortium has developed an open‐source knowledge graph–based system—Translator—designed to integrate, harmonize, and make inferences over diverse biomedical data sources. We announce here Translator's initial public release and provide an overview of its architecture, standards, user interface, and core features. Translator employs a scalable, federated, knowledge graph framework for the integration of clinical, genomic, pharmacological, and other biomedical knowledge sources, enabling query retrieval, inference, and hypothesis generation. Translator's user interface is designed to support the exploration of knowledge relationships and the generation of insights, without requiring deep technical expertise and gradually revealing more detailed evidence, provenance, and confidence information, as needed by a given user. To demonstrate Translator's application and impact, we highlight features of the user interface in the context of three real‐world use cases: suggesting potential therapeutics for patients with rare disease; explaining the mechanism of action of a pipeline drug; and screening and validating drug candidates in a model organism. We discuss strengths and limitations of reasoning within a largely federated system and the need for rich concept modeling and deep provenance tracking. Finally, we outline future directions for enhancing Translator's functionality and expanding its data sources. Translator represents a significant step forward in making complex biomedical knowledge more accessible and actionable, aiming to accelerate translational research and improve patient care.

Research & Experimental Medicine↗

Distribution Substation Planning Toolkit (dsp-toolkit) v1.0

The Distribution Substation Planning Toolkit (DSP Toolkit) is a software suite designed to streamline the planning and optimization of distribution substations. This toolkit offers a comprehensive set of tools and APIs for data curation, short-term electric load forecasting, and weather-sensitive load adjustment, making it an essential resource for utility companies, engineers, and researchers. Features • Data Preprocessing and Curation: Efficiently manage and preprocess large datasets to ensure high-quality input for analysis. • Short-Term Load Forecasting: Utilize data-driven models to predict short-term electric loads accurately. • Weather-Sensitive Modeling: Automatically adjust load forecasts based on weather data to predict future peak demands more precisely. Uses The DSP Toolkit is ideal for planning and optimizing distribution substations, providing a user-friendly interface and comprehensive documentation. It is suitable for both novice and experienced users, facilitating efficient and accurate planning processes. Advantages • Efficiency: Automates complex planning tasks, reducing manual effort and minimizing errors. • Scalability: Handles large datasets and complex models, making it suitable for large-scale projects. • Community and Support: Open-source with active community contributions, ensuring continuous improvement and support. • Extensibility: Easily extendable with custom modules and plugins, allowing users to tailor the toolkit to their specific needs. The DSP Toolkit stands out by offering a robust, flexible, and user-friendly solution for distribution substation planning. Public Abstract

Li, Han [Lawrence Berkeley National Laboratory (LB↗

The Mock LISA Data Challenges: History, Status, Prospects

This slide presentation reviews the importance for the Mock LISA Data Challenges (MLDC). Laser Interferometer Space Antenna (LISA) is a gravitational wave (GW) observatory that will return data such that data analysis is integral to the measurement concept. Further rationale of the MLDC are to kickstart the development of a LISA data-analysis computational infrastructure, and to encourage, track, and compare progress in LISA data-analysis development in the open community. The MLDCs is a coordinated, voluntary effort in GW community, that will periodically issue datasets with synthetic noise and GW signals from sources of undisclosed parameters; increasing difficulty. The challenge participants return parameter estimates and descriptions of search methods. Some of the challenges and the resultant entries are reviewed. The aim is to show that LISA data analysis is possible, and to develop new techniques, using multiple international teams for the development of LISA core analysis tools

data analysis↗

DAQ: Software Architecture for Data Acquisition in Sounding Rockets

A multithreaded software application was developed by Jet Propulsion Lab (JPL) to collect a set of correlated imagery, Inertial Measurement Unit (IMU) and GPS data for a Wallops Flight Facility (WFF) sounding rocket flight. The data set will be used to advance Terrain Relative Navigation (TRN) technology algorithms being researched at JPL. This paper describes the software architecture and the tests used to meet the timing and data rate requirements for the software used to collect the dataset. Also discussed are the challenges of using commercial off the shelf (COTS) flight hardware and open source software. This includes multiple Camera Link (C-link) based cameras, a Pentium-M based computer, and Linux Fedora 11 operating system. Additionally, the paper talks about the history of the software architecture's usage in other JPL projects and its applicability for future missions, such as cubesats, UAVs, and research planes/balloons. Also talked about will be the human aspect of project especially JPL's Phaeton program and the results of the launch.

Ahmad, Mohammad↗

ImageLabler: Labeling and Managing Image Data for Machine Learning in the Earth Sciences

While machine learning techniques for image classification have been around for a long time, storing and managing the vast number of images required as training data is still a problem for scientists. This is especially true for the field of Earth science, where only recently have experts begun using machine learning techniques for image-based phenomena classification. Image Labeler, a fast and scalable cloud-based tagging platform for Earth science images, seeks to improve upon existing methods of managing images and associated metadata, such as maintaining categorized folders of images on a local machine, a process that can be cumbersome and difficult to scale. The platform facilitates rapid development of image-based Earth science phenomena training datasets by allowing scientists to upload their existing imagery as well as extract new samples from open satellite imagery services made available through NASA’s Global Imagery Browse Service (GIBS). Image Labeler also supports GeoTIFF data, with capabilities such as displaying GeoTIFFs on an interactive map, drawing shapefiles over them, and tagging them with additional metadata. This allows scientists to perform spatiotemporal subsetting with geographic information and develop training data more quickly. Built using modern web technologies, Image Labeler includes additional capabilities such as team collaboration for large-scale image tagging projects. Users can download their data in a machine-learning-ready format, allowing scientists to spend time on experimentation rather than on the collection of training data. In this presentation, we demonstrate how Image Labeler seeks to become a one-stop image data management solution for machine learning applications in Earth science.

Ashish Acharya↗

Using CALIPSO's New Ocean Derived Column Optical Depths

CALIPSO’s Version 4.51 Level 2 data release introduces an all-new group of science data sets containing estimates of total column two-way transmittances and effective optical depths derived from CALIOP ocean surface backscatter measurements and MERRA-2 reanalysis wind speed data. These estimates use data from the standard CALIOP lidar signal but in a passive sensor-like way, thus creating a unique link to passive instrument measurements. These new retrievals are provided for the entire mission, day and night, at 532nm, and are reported at single shot, 1km, and 5km resolutions for all profiles in which a valid lidar ocean surface return is detected. The addition of a total column optical depth constraint on subsequent retrievals of cloud and aerosol optical properties opens the door for many exciting new ways to leverage the already rich and versatile CALIPSO dataset. Following a brief review of the retrieval technique, this talk will focus on quality assurance assessments, estimated uncertainties, and data usage scenarios. We will conclude with examples highlighting some of the exciting work already being done using this new addition to CALIPSO’s already rich data record.

R Ryan↗

Improved Soil Moisture Estimation and Detection of Irrigation Signal By Incorporating SMAP Soil Moisture Into the Indian Land Data Assimilation System (ILDAS)

Land surface models have facilitated the estimation of soil moisture over a range of spatiotemporal scales. However, limitations in model parameterization and under-representation of anthropogenic processes restrict their ability to estimate local-scale soil moisture variability, especially over irrigated areas. Assimilation of satellite-based soil moisture retrievals into land surface models can be a viable approach to overcome these constraints, specially over highly irrigated countries such as India, where such applications are rare. Additionally, large-scale validation of modeled soil moisture has been limited over India till now due to lack of a representative station network. By assimilating Soil Moisture Active Passive (SMAP)-based estimates into the state-of-the-art Indian Land Data Assimilation System (ILDAS) and combining with a new soil moisture station network of more than 200 stations, this study demonstrates improved soil moisture estimations and capture of irrigation signals over the region. The Noah-MP land surface model is forced by multiple local and global meteorological datasets and Ensemble Kalman Filter (EnKF) is used for assimilation of soil moisture. Comparison of open-loop and data assimilated soil moisture against station soil moisture data shows relative spatial mean improvement of 0.0178 in correlation and 0.0029 m3/m3 in RMSE. Further statistical comparison with in-situ data has also shown better results over most of the stations, as evident from improved correlations and reduced unbiased RMSE after assimilation. Finally, the climatology of soil moisture over the different irrigation fractions reveals that data assimilated outputs over irrigated grid cells tend to have higher soil moisture during dry winter season, demonstrating the ability to capture irrigation signals. These findings quantify the value of data assimilation in improving soil moisture estimates and the ability to capture unmodeled processes such as irrigation, which lays the science groundwork for upcoming space missions such as NASA ISRO Synthetic Aperture Radar (NISAR).

Soil Moisture↗

High School Citizen Scientists Use AI/ML to Predict Intra-Ocular Pressure From Gene Expression Data for Spaceflown Mice

Artificial Intelligence (AI) and Machine Learning (ML) have increasingly become pivotal in biological and biomedical research, largely due to the culture of open data sharing and its associated benefits. The methodologies inherent in AI/ML are particularly adept at identifying and forecasting biological phenotypes from the vast amounts of data generated by next-generation sequencing technologies. These techniques offer substantial promise for advancing research in space biosciences and for the development of automated systems for monitoring space health. Nevertheless, there are crucial aspects to consider when training, validating, and testing machine learning models in both biological research and clinical contexts. It is essential that Open Science principles, including data sharing and the availability of open-source code, are complemented by high-quality, publicly accessible training resources. These resources should focus on best practices and include modules based on real-world scientific cases and data to ensure that future AI/ML practitioners gain practical experience with genuine problems. Addressing this knowledge gap, we have designed, developed, and delivered both interactive and self-paced training programs for citizen scientists worldwide, enabling them to utilize AI/ML for space biology research. This initiative was made possible through generous funding from a Transformation to Open Science Training grant. The interactive training sessions, conducted this summer, utilized AI/ML techniques to analyze data from the Open Science Data Repository, specifically targeting the effects of spaceflight on ocular structure and function. The dataset OSD-583, from the Rodent Research 9 mission, provides experimental data detailing the ocular responses of mice subjected to a 35-day spaceflight, compared with ground control counterparts. Using OSD-583 as observational data, our summer training participants applied AI/ML methods to predict intraocular pressure from RNA-seq data and identify the genes most predictive of the observed responses. Further analysis through pathway enrichment and gene set enrichment revealed that these genes are involved in molecular and cellular processes contributing to retinal degeneration.

James Casaletto↗