Search NASA⌕ Search

SEARCH · Search NASA

Results for “Open Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29

NASA's Earth Science Data Systems Standards Process Experiences

NASA has impaneled several internal working groups to provide recommendations to NASA management on ways to evolve and improve Earth Science Data Systems. One of these working groups is the Standards Process Group (SPC). The SPG is drawn from NASA-funded Earth Science Data Systems stakeholders, and it directs a process of community review and evaluation of proposed NASA standards. The working group's goal is to promote interoperability and interuse of NASA Earth Science data through broader use of standards that have proven implementation and operational benefit to NASA Earth science by facilitating the NASA management endorsement of proposed standards. The SPC now has two years of experience with this approach to identification of standards. We will discuss real examples of the different types of candidate standards that have been proposed to NASA's Standards Process Group such as OPeNDAP's Data Access Protocol, the Hierarchical Data Format, and Open Geospatial Consortium's Web Map Server. Each of the three types of proposals requires a different sort of criteria for understanding the broad concepts of "proven implementation" and "operational benefit" in the context of NASA Earth Science data systems. We will discuss how our Standards Process has evolved with our experiences with the three candidate standards.

Ullman, Richard E.↗

Empirically-calibrated H100 node power models for accurate AI training energy estimation

Accurately quantifying the energy use of artificial intelligence (AI) training is critical for infrastructure planning, carbon accounting, and sustainable data center operation, but few studies have directly measured the power consumption of production workloads on contemporary hardware. By combining empirical measurements from Brookhaven National Laboratory during AI training on 8-graphics-processing-unit H100 systems with open-source benchmarking data, we develop statistical models relating computational intensity to node-level power consumption. We measure the gap between manufacturer-rated thermal design power (TDP) and actual power demand during AI training. Our analysis reveals that even computationally intensive workloads operate at only 76% of the 10.2 kW TDP rating. Our architecture-specific model, calibrated to floating-point operations, predicts energy consumption with 11.4% mean absolute percentage error, significantly outperforming TDP-based approaches (27%–37% error). We identified distinct power signatures between transformer and convolutional neural network architectures, with transformers showing characteristic fluctuations that may impact grid stability. These results provide a measurement-grounded basis for improving AI training energy estimates, enabling more reliable infrastructure sizing, cost projections, and environmental impact assessments.

Newkirk, Alex C↗

BRE‐X Emissions Database for End‐of‐Life Scenarios of Selective Building Construction Materials to Enable Circular Economy in Construction

In the United States, construction and demolition debris predominately end up in landfills with minimal end‐of‐life Re‐X (recover, recycle, reuse, etc.) scenarios, resulting in large environmental impacts and lost opportunities for material recovery. Except for concrete and metals, which seem to have a few well‐defined end‐of‐life pathways, there seems to be a lack of well‐documented end‐of‐life scenarios for other construction materials, let alone their emissions data. Hence, there is a need for documented end‐of‐life Re‐X scenarios and end‐of‐life data of more building materials to motivate widespread use of Re‐X strategies in building design. This paper outlines the efforts of the National Renewable Energy Laboratory, Carbon Leadership Forum, Building Transparency, and Skidmore, Owings & Merrill to (a) create an open‐access BRE‐X (Building Re‐X) end‐of‐life emissions database consisting of greenhouse gas emissions data associated with various end‐of‐life scenarios for a select list of high‐impact building construction materials, and (b) integrate the BRE‐X end‐of‐life emissions database with CAD/BIM/LCA tools for evaluating various end‐of‐life scenarios. The paper also presents a few existing life cycle inventory databases that contain sparse amounts of end‐of‐life data for a few construction materials and their limitations in terms of scaling and data consolidation. Finally, a sample of how the collected data can be ingested into whole‐building LCA tools using open data formats and a public access link to the BRE‐X end‐of‐life emissions database is also included.

36 MATERIALS SCIENCE↗

Dataset of Generative AI Workload Power Profiles

This dataset provides a collection of high-resolution (5/10 Hz or every 0.2/0.1 seconds) power consumption profiles for generative artificial intelligence (GenAI) workloads executed on NLR's High Performance Computing (HPC) platform Kestrel. The dataset also includes examples of representative whole-facility power profiles generated using a bottom-up, event-driven, data center energy model . This dataset is designed to support research in energy modeling, infrastructure planning, energy system integration, and sustainability analysis for AI-driven computing systems. The dataset captures time-resolved electrical power measurements across a diverse set of configurations, including variations in job type (inference vs. training), workload (LLM vs. image generation), datasets, and number of compute nodes. Power traces are provided in a standardized format and include both raw/instantaneous and aggregated files. Each profile is accompanied by metadata describing workload parameters, enabling reproducibility and cross-study comparison. The dataset is intended for use in applications such as data center infrastructure planning, energy modeling, demand response and grid impact studies, and development and validation of system-level simulation tools. By making these workload-specific power profiles publicly available, this dataset aims to address the current lack of open, empirical energy data for generative AI systems and to facilitate transparent, reproducible research on the energy and environmental impacts of large-scale AI deployment. If you use this dataset, please cite the associated publication: Vercellino et al., “Measurement of Generative AI Workload Power Profiles for Whole-Facility Data Center Infrastructure Planning,” arXiv:2604.07345 (2026).

97 MATHEMATICS AND COMPUTING↗

VEDA: Visualization, Exploration, & Data Analysis

NASA's Visualization, Exploration, and Data Analysis (VEDA) project is an open-source science cyberinfrastructure for data processing, visualization, exploration, and geographic information systems (GIS) capabilities. Developed collaboratively and mostly reusing existing open-source components, VEDA consolidates GIS delivery mechanisms, processing platforms, analysis services, and visualization tools and provides an ecosystem of open tools for addressing Earth science research and application needs through the public-facing VEDA Dashboard. In this presentation, Dr. Freitag will provide an overview of VEDA and how it can potentially serve the AOS community.

Brian Freitag↗

Data Requirements for Oceanic Processes in the Open Ocean, Coastal Zone, and Cryosphere

The type of information system that is needed to meet the requirements of ocean, coastal, and polar region users was examined. The requisite qualities of the system are: (1) availability, (2) accessibility, (3) responsiveness, (4) utility, (5) continuity, and (6) NASA participation. The system would not displace existing capabilities, but would have to integrate and expand the capabilities of existing systems and resolve the deficiencies that currently exist in producer-to-user information delivery options.

Nagler, R. G.↗

Fusion Approach for Remotely-Sensed Mapping of Agriculture (FARMA): A Scalable Open Source Method for Land Cover Monitoring Using Data Fusion

The increasing availability of very-high resolution (VHR; <2 m) imagery has the potential to enable agricultural monitoring at increased resolution and cadence, particularly when used in combination with widely available moderate-resolution imagery. However, scaling limitations exist at the regional level due to big data volumes and processing constraints. Here, we demonstrate the Fusion Approach for Remotely-Sensed Mapping of Agriculture (FARMA), using a suite of open source software capable of efficiently characterizing time-series field-scale statistics across large geographical areas at VHR resolution. We provide distinct implementation examples in Vietnam and Senegal to demonstrate the approach using WorldView VHR optical, Sentinel-1 Synthetic Aperture Radar, and Sentinel-2 and Sentinel-3 optical imagery. This distributed software is open source and entirely scalable, enabling large area mapping even with modest computing power. FARMA provides the ability to extract and monitor sub-hectare fields with multisensor raster signals, which previously could only be achieved at scale with large computational resources. Implementing FARMA could enhance predictive yield models by delineating boundaries and tracking productivity of smallholder fields, enabling more precise food security observations in low and lower-middle income countries.

fusion↗

The Cumulus Ecosystem: Open Source and Beyond to Foster Collaboration

NASA’s Earth Observing System Data and Information System (EOSDIS) open source Cumulus software is designed as a common set of code and services that can be used to create a pipeline to deliver and manage earth science data in the cloud. Cumulus strives to create an ecosystem on the foundation of open source that unites those with shared problems and goals by encouraging users to contribute solutions back to the platform. Large parts of ingesting and managing data are common and much of what is created can be used by others. Our goal is to maximize collaboration and code reuse while allowing users to design a custom solution that meets their needs without having to take on extraneous functionality. In this talk we will describe how the Cumulus ecosystem works beyond just open source software. We will review the technology, the successes and challenges, and the evolution and future of Cumulus as an ecosystem.

Cumulus↗

Atmospheric Measurements by the 2002 Geoscience Laser Altimeter System Mission

The NASA Earth Observing System (EOS) program is a multiple platform NASA initiative for the study of global change. As part of the EOS project, the Geoscience Laser Altimeter System (GLAS) was selected as a laser sensor filling complementary requirements for several earth science disciplines including atmospheric and surface applications. Late in 2002, the GaAs instrument is to be launched for a three to five year observational mission. For the atmosphere, the instrument is designed to full fill comprehensive requirements for profiling of radiatively significant cloud and aerosol. Algorithms have been developed to process the cloud and aerosol data and provide standard data products. After launch there will be a three-month project to analyze and understand the system performance and accuracy of the data products. As an EOS mission, the GaAs measurements and data products will be openly available to all investigators. An overview of the instrument, data products and evaluation plan is given.

Spinhirne, James D.↗

Acquisition of and Access to Research Omics Data

Omics data are essential for understanding the myriad and complex effects of space environments on humans. To assure maximum benefit from these kinds of data, the NASA Human Research Program Data Management Plan stipulates that human omics data should be archived within and accessed through the NASA Life Sciences Portal (NLSP). The NLSP has the capability to acquire and provision access to omics (and other kinds of) research results for individual and ad-hoc groups of subjects at the direction of institutional review boards, or other authorizing bodies or individuals, per institutional, program and investigation-specific policies and procedures. However, because some single-subject omics data, like CT scans and other kinds of large, complex biomedical data, could be used to identify heretofore unknown risks to the subject’s health, or, in certain cases, be used to identify a subject, NASA Policy Directive 7170.1 describes various policies regarding the management of and access to “research genetic testing” data, which includes many kinds of omics data. For example, NPD 7170.1 prohibits access to human research genetic data by NASA personnel who make employment decisions for the subjects from whom the data were obtained. To meet the objective of acquiring research omics data for NLSP in compliance with the policies in NPD 7170.1 and other applicable NASA policies, we designed NOMADS (the NLSP Omics Multimodal Acquisition of Data System), a new component that supports the transfer of large research data files, including research genetic testing data, using one of several different transfer mechanisms. The choice of mechanism is made by the submitter of the data, with guiding information from the system, and is likely to often be determined in large part by the nature and source location of the data. For example, for small files where the source data files are not already stored in a cloud storage system, users are likely to prefer to transfer their data to the NLSP via a web browser. Conversely, for large sets of files already organized and stored in a cloud storage system, users may opt for NOMAD’s cloud-to-cloud transfer method. All omics datasets targeted for the NASA Life Sciences Data Archive must pass a variety of quality checks to ensure data integrity and adherence to the standards defined by the LSDA Data Submission Guidelines (DSG) (see https://nlsp.nasa.gov/explore/lsdahome/datasubmit). These include requirements that data are consistent with open standards established by the omics community. Non-compliant data will not be accepted however archivists are available to advise submitters on how to revise data submissions and re-submit until compliance is achieved. Following compliance with the LSDA DSG, omics data next undergo a variety of additional quality checks to ensure the data meet omics community standards. Domain specific Omics data quality control tools and techniques are continually evolving and linked to the advancements in omics assays utilized and thus, the tools and techniques utilized by the LSDA for data quality control and validation will need to be sustained accordingly. All human omics data will be access controlled according to the policies described above, and requiring IRB approval for any additional access grants once the data are acquired (including access for analysis using the NLSP workspace tools).

Omics↗

Acquisition of and Access to Research Omics Data

Omics data are essential for understanding the myriad and complex effects of space environments on humans. To assure maximum benefit from these kinds of data, the NASA Human Research Program Data Management Plan stipulates that human omics data should be archived within and accessed through the NASA Life Sciences Portal (NLSP). The NLSP has the capability to acquire and provision access to omics (and other kinds of) research results for individual and ad-hoc groups of subjects at the direction of institutional review boards, or other authorizing bodies or individuals, per institutional, program and investigation-specific policies and procedures. However, because some single-subject omics data, like CT scans and other kinds of large, complex biomedical data, could be used to identify heretofore unknown risks to the subject’s health, or, in certain cases, be used to identify a subject, NASA Policy Directive 7170.1 describes various policies regarding the management of and access to “research genetic testing” data, which includes many kinds of omics data. For example, NPD 7170.1 prohibits access to human research genetic data by NASA personnel who make employment decisions for the subjects from whom the data were obtained. To meet the objective of acquiring research omics data for NLSP in compliance with the policies in NPD 7170.1 and other applicable NASA policies, we designed NOMADS (the NLSP Omics Multimodal Acquisition of Data System), a new component that supports the transfer of large research data files, including research genetic testing data, using one of several different transfer mechanisms. The choice of mechanism is made by the submitter of the data, with guiding information from the system, and is likely to often be determined in large part by the nature and source location of the data. For example, for small files where the source data files are not already stored in a cloud storage system, users are likely to prefer to transfer their data to the NLSP via a web browser. Conversely, for large sets of files already organized and stored in a cloud storage system, users may opt for NOMAD’s cloud-to-cloud transfer method. All omics datasets targeted for the NASA Life Sciences Data Archive must pass a variety of quality checks to ensure data integrity and adherence to the standards defined by the LSDA Data Submission Guidelines (DSG) (see https://nlsp.nasa.gov/explore/lsdahome/datasubmit). These include requirements that data are consistent with open standards established by the omics community. Non-compliant data will not be accepted however archivists are available to advise submitters on how to revise data submissions and re-submit until compliance is achieved. Following compliance with the LSDA DSG, omics data next undergo a variety of additional quality checks to ensure the data meet omics community standards. Domain specific Omics data quality control tools and techniques are continually evolving and linked to the advancements in omics assays utilized and thus, the tools and techniques utilized by the LSDA for data quality control and validation will need to be sustained accordingly. All human omics data will be access controlled according to the policies described above, and requiring IRB approval for any additional access grants once the data are acquired (including access for analysis using the NLSP workspace tools).

Omics↗

Assessing Spatial Representativeness of Global Flux Tower Eddy-Covariance Measurements Using Data from FLUXNET2015

Large datasets of carbon dioxide, energy, and water fluxes were measured with the eddy-covariance (EC) technique, such as FLUXNET2015. These datasets are widely used to validate remote-sensing products and benchmark models. One of the major challenges in utilizing EC-flux data is determining the spatial extent to which measurements taken at individual EC towers reflect model-grid or remote sensing pixels. To minimize the potential biases caused by the footprint-to-target area mismatch, it is important to use flux datasets with awareness of the footprint. This study analyze the spatial representativeness of global EC measurements based on the open-source FLUXNET2015 data, using the published flux footprint model (SAFE-f). The calculated annual cumulative footprint climatology (ACFC) was overlaid on land cover and vegetation index maps to create a spatial representativeness dataset of global flux towers. The dataset includes the following components: (1) the ACFC contour (ACFCC) data and areas representing 50%, 60%, 70%, and 80% ACFCC of each site, (2) the proportion of each land cover type weighted by the 80% ACFC (ACFCW), (3) the semivariogram calculated using Normalized Difference Vegetation Index (NDVI) considering the 80% ACFCW, and (4) the sensor location bias (SLB) between the 80% ACFCW and designated areas (e.g. 80% ACFCC and window sizes) proxied by NDVI. Finally, we conducted a comprehensive evaluation of the representativeness of each site from three aspects: (1) the underlying surface cover, (2) the semivariogram, and (3) the SLB between 80% ACFCW and 80% ACFCC, and categorized them into 3 levels. The goal of creating this dataset is to provide data quality guidance for international researchers to effectively utilize the FLUXNET2015 dataset in the future.

54 ENVIRONMENTAL SCIENCES↗

Improved Soil Moisture Estimation and Detection of Irrigation Signal By Incorporating SMAP Soil Moisture Into the Indian Land Data Assimilation System (ILDAS)

Land surface models have facilitated the estimation of soil moisture over a range of spatiotemporal scales. However, limitations in model parameterization and under-representation of anthropogenic processes restrict their ability to estimate local-scale soil moisture variability, especially over irrigated areas. Assimilation of satellite-based soil moisture retrievals into land surface models can be a viable approach to overcome these constraints, specially over highly irrigated countries such as India, where such applications are rare. Additionally, large-scale validation of modeled soil moisture has been limited over India till now due to lack of a representative station network. By assimilating Soil Moisture Active Passive (SMAP)-based estimates into the state-of-the-art Indian Land Data Assimilation System (ILDAS) and combining with a new soil moisture station network of more than 200 stations, this study demonstrates improved soil moisture estimations and capture of irrigation signals over the region. The Noah-MP land surface model is forced by multiple local and global meteorological datasets and Ensemble Kalman Filter (EnKF) is used for assimilation of soil moisture. Comparison of open-loop and data assimilated soil moisture against station soil moisture data shows relative spatial mean improvement of 0.0178 in correlation and 0.0029 m3/m3 in RMSE. Further statistical comparison with in-situ data has also shown better results over most of the stations, as evident from improved correlations and reduced unbiased RMSE after assimilation. Finally, the climatology of soil moisture over the different irrigation fractions reveals that data assimilated outputs over irrigated grid cells tend to have higher soil moisture during dry winter season, demonstrating the ability to capture irrigation signals. These findings quantify the value of data assimilation in improving soil moisture estimates and the ability to capture unmodeled processes such as irrigation, which lays the science groundwork for upcoming space missions such as NASA ISRO Synthetic Aperture Radar (NISAR).

Soil Moisture↗

A Dynamic Landslide Hazard Monitoring Framework for the Lower Mekong Region

The Lower Mekong region is one of the most landslide-prone areas of the world. Despite the need for dynamic characterization of landslide hazard zones within the region, it is largely understudied for several reasons. Dynamic and integrated understanding of landslide processes requires landslide inventories across the region, which have not been available previously. Computational limitations also hamper regional landslide hazard assessment, including accessing and processing remotely sensed information. Finally, open-source software and modelling packages are required to address regional landslide hazard analysis. Leveraging an open-source data-driven global Landslide Hazard Assessment for Situational Awareness model framework, this study develops a region-specific dynamic landslide hazard system leveraging satellite-based Earth observation data to assess landslide hazards across the lower Mekong region. A set of landslide inventories were prepared from high-resolution optical imagery using advanced image-processing techniques. Several static and dynamic explanatory variables (i.e., rainfall, soil moisture, slope, relief, distance to roads, distance to faults, distance to rivers) were considered during the model development phase. An extreme gradient boosting decision tree model was trained for the monsoon period of 2015–2019 and the model was evaluated with independent inventory information for the 2020 monsoon period. The model performance demonstrated considerable skill using receiver operating characteristic curve statistics, with Area Under the Curve values exceeding 0.95. The model architecture was designed to use near-real-time data, and it can be implemented in a cloud computing environment (i.e., Google Cloud Platform) for the routine assessment of landslide hazards in the Lower Mekong region. This work was developed in collaboration with scientists at the Asian Disaster Preparedness Center as part of the NASA SERVIR Program’s Mekong hub. The goal of this work is to develop a suite of tools and services on accessible open-source platforms that support and enable stakeholder communities to better assess landslide hazard and exposure at local to regional scales for decision making and planning.

Nishan Kumar Biswas↗