Search NASA⌕ Search

SEARCH · Search NASA

Results for “Research data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Hierarchical Estimation For Planetary Protection

The software uses Bayes' theorem to describe the probability of an event based on prior knowledge of conditions that might be related to the event. The purpose of Bayesian analysis is to determine posterior probabilities based on prior probabilities where new information can be used in the decision-making process as additional data is gathered. The software will be used in Probabilistic Risk Assessments (PRAs) related to the Europa Clipper mission, which is one of NASA’s top priorities. Ultimately, the mission entails sending the Europa Clipper spacecraft to Jupiter’s Europa moon to orbit the planet and collect data for research and development. Europa is the smallest of the four Galilean moons orbiting Jupiter and is believed by researchers to be the most promising place to look for present-day environments suitable for life. Europa is thought to have an iron core, a rocky mantle, and a salt-water ocean covered by an ice-layered surface.

Gribok, Andrei [Idaho National Laboratory (INL), I↗

2010-2012 Minneapolis - St. Paul Travel Behavior Inventory

The 2010-2012 Travel Behavior Inventory (TBI) provided Minnesota policymakers and researchers with data about travel in the Minneapolis - St. Paul region. It also updated the region's travel demand forecasting, including transit ridership for major transportation projects. The Metropolitan Council in the Minneapolis-St. Paul area conducted the survey. The TBI consists of a paper-based survey and a wearable global positioning system survey. The data collection process for these two surveys was independent, and the results are not intended to function together.

1Hz data↗

'Omics and Big Data in Harmful Algal Bloom Research

Phytoplankton, a group including eukaryotic microalgae and cyanobacteria, play a crucial climate role converting CO 2 into organic carbon through global primary production. They support a wide range of life, both freshwater and marine, from zooplankton to fish and mammals. While they are essential in nutrient cycles, certain phytoplankton species can proliferate excessively under favorable conditions, leading to harmful algal blooms (HABs) that pose significant threats to human and ecosystem health through the toxins they produce.

59 BASIC BIOLOGICAL SCIENCES↗

Using probability distribution function as a scaling approach to incorporate soil heterogeneity into biogeochemical models for greenhouse gas predictions (Final Technical Report)

The project investigated biogeochemical processes at terrestrial-aquatic interfaces (TAIs), focusing on soil microsite heterogeneity and its impact on greenhouse gas (GHG) fluxes. Using laboratory experiments, modeling, and data integration, researchers explored redox-driven microbial processes under fluctuating hydrological conditions. Key advancements included modifying the DAMM-GHG model to incorporateelectron acceptor availability and enhancing the AquaMEND model for improved microbial metabolism representation. Results highlighted microsite redox variability as a key driver of GHG fluxes, informing Earth system models. The project fostered interdisciplinary collaborations, student training, and the development of novel modeling frameworks to improve Earth'senergy budget.

54 ENVIRONMENTAL SCIENCES↗

Data Cards for Standardized Metadata Across DOE-Aligned Data Initiatives: Toward Transparent, Interoperable, and Governed Dataset Documentation

As data-intensive research, advanced computing, and artificial intelligence become increasingly central to scientific and operational workflows, the need for consistent, transparent, and machine-actionable documentation has grown correspondingly. Multiple DOE-aligned communities—including Office of Science, Genesis Mission, American Science Cloud (AmSC), National Nuclear Security Administration (NNSA) stewardship and governance, and related cross-laboratory collaborations—have independently developed metadata practices to support discovery, access, reuse, repository deposit, and compliance.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Report of the 2026 Workshop on Next-Generation Ecosystems for Scientific Computing: Harnessing Community, Software, and AI for Cross-Disciplinary Team Science

Scientific computing is undergoing rapid transformation as advances in artificial intelligence, heterogeneous computing, automation, and data-intensive research reshape not only computational tools but also the institutions, workforce models, and collaborative practices that support scientific discovery. This report synthesizes insights from the 2026 Workshop on Next-Generation Ecosystems for Scientific Computing, the second in a three-year series focused on strengthening scientific computing ecosystems through socio-technical co-design. Workshop discussions identified four interdependent strategic themes: software ecosystems for AI-enabled scientific discovery; trust, validation, and traceability; human-AI teaming and paradigm shifts; and workforce, pedagogy, and governance. The report translates these themes into eight priorities for community action spanning shared research infrastructure, trust and traceability, user experience, human-AI teaming, workforce development, cross-sector coordination, stewardship and sustainability, and evaluation of scientific value. Together, these priorities outline directions for building scientific computing ecosystems that remain trustworthy, sustainable, innovative, and resilient as AI assumes a growing role in scientific work.

AI↗

Using automated machine learning for the upscaling of gross primary productivity

Estimating gross primary productivity (GPP) over space and time is fundamental for understanding the response of the terrestrial biosphere to climate change. Eddy covariance flux towers provide in situ estimates of GPP at the ecosystem scale, but their sparse geographical distribution limits larger-scale inference. Machine learning (ML) techniques have been used to address this problem by extrapolating local GPP measurements over space using satellite remote sensing data. However, the accuracy of the regression model can be affected by uncertainties introduced by model selection, parameterization, and choice of explanatory features, among others. Recent advances in automated ML (AutoML) provide a novel automated way to select and synthesize different ML models. In this work, we explore the potential of AutoML by training three major AutoML frameworks on eddy covariance measurements of GPP at 243 globally distributed sites. We compared their ability to predict GPP and its spatial and temporal variability based on different sets of remote sensing explanatory variables. Explanatory variables from only Moderate Resolution Imaging Spectroradiometer (MODIS) surface reflectance data and photosynthetically active radiation explained over 70 % of the monthly variability in GPP, while satellite-derived proxies for canopy structure, photosynthetic activity, environmental stressors, and meteorological variables from reanalysis (ERA5-Land) further improved the frameworks' predictive ability. We found that the AutoML framework Auto-sklearn consistently outperformed other AutoML frameworks as well as a classical random forest regressor in predicting GPP but with small performance differences, reaching an r 2 of up to 0.75. We deployed the best-performing framework to generate global wall-to-wall maps highlighting GPP patterns in good agreement with satellite-derived reference data. This research benchmarks the application of AutoML in GPP estimation and assesses its potential and limitations in quantifying global photosynthetic activity.

54 ENVIRONMENTAL SCIENCES↗

Merged Observatory Data Files (MODFs): an integrated observational data product supporting process-oriented investigations and diagnostics

A large and ever-growing body of geophysical information is measured in campaigns and at specialized observatories as a part of scientific expeditions and experiments. These collections of observed data include many essential climate variables (as defined by the Global Climate Observing System) but are often distinguished by a wide range of additional non-routine measurements that are designed to not only document the state of the environment but also the drivers that contribute to that state. These field data are used not only to further understand environmental processes through observation-based studies but also to provide baseline data to test model performance and to codify understanding to improve predictive capabilities. To address the considerable barriers and difficulty in utilizing these diverse and complex data for observation–model research, the Merged Observatory Data File (MODF) concept has been developed. A MODF combines measurements from multiple instruments into a single file that complies with well-established data format and metadata practices and has been designed to parallel the development of corresponding Merged Model Data Files (MMDFs). Using the MODF and MMDF protocols will facilitate the evolution of model intercomparison projects into model intercomparison and improvement projects by putting observation and model data “on the same page” in a timely manner. The MODF concept was developed especially for weather forecast model studies in the Arctic. The surprisingly complex process of implementing MODFs in that context refined the concept itself. Thus, this article explains the concept of MODFs by providing details on the issues that were revealed and resolved during that first specific implementation. Detailed instructions are provided on how to make MODFs, and this article can be considered a MODF creation manual.

54 ENVIRONMENTAL SCIENCES↗

Mechanical and durability properties of ultra-high-performance concrete of spent nuclear fuel dry storage systems: a review

Dry storage systems are used for interim storage of spent nuclear fuel (SNF). However, with the growing need to extend the operational periods of these systems, there are concerns about the degradation of their concrete overpacks, which could compromise the system's structural integrity and safety during hazardous events. Traditional concrete mixtures used in SNF dry storage systems have remained largely unchanged since their inception and often use conventional ingredients. These materials are susceptible to degradation mechanisms such as chemical attacks, alkali-silica reactions (ASR), and freeze–thaw cycles, which can lead to a loss of strength and durability over time. To address these challenges, this paper reviews the application of ultra-high-performance concrete (UHPC) as a promising alternative for spent nuclear fuel dry storage system overpacks. UHPC offers superior mechanical properties, exceptional durability, and reduced susceptibility to degradation mechanisms compared to conventional concrete. This paper focuses on the role of supplementary cementitious materials (SCMs) such as silica fume, fly ash, and metakaolin in enhancing UHPC performance for SNF storage applications. These SCMs have been shown to significantly improve the material’s microstructure, strength, and resistance to environmental stressors typically encountered in SNF storage environments. Moreover, incorporating SCMs supports sustainable construction by reducing cement consumption and associated carbon emissions. The review brings together existing research and experimental data, providing insights for engineers and researchers on developing UHPC mixtures that meet the rigorous demands of spent nuclear fuel dry storage systems, extending their service life and minimizing inspection intervals.

36 - MATERIALS SCIENCE↗

Sustainable Port Operations: Powered by NREL

Seaports are vital economic hubs that allow the United States to compete on a global scale. But the heavy vehicles and cargo equipment that enable their operations also emit harmful air pollutants and greenhouse gas emissions. For nearly two decades, National Renewable Energy Laboratory (NREL) researchers have worked toward comprehensive seaport decarbonization. They fuse world-class analysis with deep vehicle and transportation systems knowledge to guide strategic deployment of low- and zero-emissions vehicles, charging and refueling infrastructure, and grid improvements. Together, these capabilities can enable sustainable port operations. This fact sheet outlines major seaport and airport decarbonization capabilities across the laboratory, including: fleet research, energy data, and insights for decarbonization; comprehensive hydrogen infrastructure deployment; optimized charging through grid integration; strategic blueprinting for clean, optimized technology deployment; and integrating diversity, equity, inclusion, and accessibility considerations into decarbonization efforts.

ADVANCED PROPULSION SYSTEMS,ENERGY CONSERVATION, C↗

Geospatial Data Platform for All

Spatiotemporal data has evolved in scale due to augmented use in cross-domain applications. Simultaneously, there is substantial growth in the availability of Geographic Information Systems (GIS) data provided by the United States Geological Survey (USGS) along with other federal, state, county, or local agencies through open-data portals and public access APIs. However, data availability does not equate with accessibility. Large-scale analyses and applications require robust, performant data management with co-location of data storage and computing. The insufficiency of data management infrastructure compels researchers to adopt ad hoc project- specific GIS data storage solutions (e.g., copying data to High-Performance computer file systems). As an ad hoc storage strategy does not scale, it hampers cross-domain analyses causing difficulty in data reuse and utilizing existing code bases. Furthermore, GIS data is complex and requires expertise to analyze and manipulate due to its intricate data structures and data-specific projection transformations. Despite the challenges, we recognize that derived GIS data products, e.g., satellite or LIDAR-based images, can be used in downstream applications such as AI by domain, but non-GIS experts. To address the data needs and overcome the challenges, we are working towards a GIS Data Platform focused on efficient data storage, data discovery and access, and an API to enable common workflows. We propose a knowledge-graph (KG) approach for data discovery, whereby datasets are semantically linked to higher- level constructs such as projects and research areas. The semantic data links enable researchers to explore datasets in a top-down approach by specifying relevant and meaningful terms (assists in finding hidden data). An advantage is that the nodes and edges in a knowledge graph create built-in semantic documentation. Deeper spatiotemporal connections between data sources can be encoded via Graph Neural Networks (GNN) (Zhang et al., 2021). The KG approach can be extended to integrate the data itself in a Virtual KG (VKG). Our work will derive inspiration from large-scale VKG efforts that have been undertaken or are currently underway as part of the OpenStreetMap project (Ding et al., 2021). For DOE Data Days, we share the proposed geospatial data platform hybrid (cloud/on-prem) architecture, our work-to-date on storing, retrieving, and transforming LiDAR and raster data relevant to two important NREL use-cases, including the Renewable Energy Potential (reV) Model, and present our proposal for a KG based data discovery engine.

data platform↗

32 examples of LLM applications in materials science and chemistry: towards automation, assistants, agents, and accelerated scientific discovery

Abstract Large language models (LLMs) are reshaping many aspects of materials science and chemistry research, enabling advances in molecular property prediction, materials design, scientific automation, knowledge extraction, and more. Recent developments demonstrate that the latest class of models are able to integrate structured and unstructured data, assist in hypothesis generation, and streamline research workflows. To explore the frontier of LLM capabilities across the research lifecycle, we review applications of LLMs through 32 total projects developed during the second annual LLM hackathon for applications in materials science and chemistry, a global hybrid event. These projects spanned seven key research areas: (1) molecular and material property prediction, (2) molecular and material design, (3) automation and novel interfaces, (4) scientific communication and education, (5) research data management and automation, (6) hypothesis generation and evaluation, and (7) knowledge extraction and reasoning from the scientific literature. Collectively, these applications illustrate how LLMs serve as versatile predictive models, platforms for rapid prototyping of domain-specific tools, and much more. In particular, improvements in both open source and proprietary LLM performance through the addition of reasoning, additional training data, and new techniques have expanded effectiveness, particularly in low-data environments and interdisciplinary research. As LLMs continue to improve, their integration into scientific workflows presents both new opportunities and new challenges, requiring ongoing exploration, continued refinement, and further research to address reliability, interpretability, and reproducibility.

Computer Science↗

Energy Systems Integration Facility (ESIF): World-Class Systems Integration Capabilities and Research

The Energy Systems Integration Facility (ESIF), located at the National Renewable Energy Laboratory (NREL) South Table Mountain campus, is a world-renowned user facility for research and development of modern, advanced, and clean energy technologies. ESIF is distinguished by its continuously evolving, highly integrated systems that span throughout the building, connecting research capabilities across multiple laboratories and test areas. The primary ESIF research systems include: [1] data, cyber, and control networks, [2] research electrical distribution buses (REDB), [3] thermal integration infrastructure, and [4] hydrogen systems. The data, cyber, and control networks provide monitoring, control, communication, automation, visualization, and time series data storage and tagging capabilities for research projects and ESIF systems, including facility safety functions. The REDB system consists of four dedicated AC and DC electrical power networks that can connect devices located across the facility through versatile, automatic circuit configuration to support complex power electronics experiments up to the megawatt-scale. The thermal integration infrastructure consists of three temperature-conditioned water loops that provide heating and cooling interfaces and capabilities for thermal energy research. The hydrogen systems provide megawatt-scale hydrogen production, drying, compression, high-pressure storage, and delivery to laboratory end uses, including hydrogen fuel cell vehicle fueling. The ESIF research systems interconnect and extend throughout the various lab areas of the facility to create elaborate networks composed of diverse technologies for cutting-edge research. The ESIF capabilities are operated and stewarded by the ESIF Research Operations group, who also actively upgrade and advance the systems to ensure they remain ahead of anticipated research - enabling the success of many pioneering energy integration projects. The poster, created by members of the ESIF Research Operations team, highlights and summarizes the four core integrated systems at ESIF. The poster was first presented at the internal NREL Energize Forum on May 13th, 2024, and received the "Best Poster" award.

capabilities↗

Trust Not Verify? The Critical Need for Data Curation Standards in Materials Informatics

The importance of data curation has been recognized in multiple areas of research; however, the discussion of this important issue is only beginning to emerge in materials science. In this Perspective, we highlight the benefits of using the standardized data curation protocols in materials science and discuss current gaps in accurate and reproducible data reporting using case studies drawn from high-impact materials science papers and well-known databases such as the Crystallography Open Database (COD) and the Cambridge Structural Database (CSD). We argue that both experimental and computational materials scientists need to embrace a culture of rigorous data curation as part of modern research data management. We propose a sample data curation pipeline for materials chemistry and illustrate its use by creating two new materials chemistry databases. Here, we hope that this perspective will serve to catalyze further discussion and promote the continuous development of rigorous data curation practices within the materials science research community. We posit that adherence to best practices of data curation will promote and enhance the reliability, reproducibility, and integrity of materials research and enable the development of reliable AI and machine learning models that critically depend on the use of quality data.

Chemical structure↗

A systematic review of machine learning in groundwater monitoring

With increasing concerns about water scarcity, groundwater has become crucial since this resource provides most of the freshwater needs. However, various human and natural activities often contaminate the groundwater, making it unsuitable for use. Over the years, scientists and engineers have used many methods to predict and track groundwater contamination as part of environmental monitoring. Consequently, there is an urgent need for improved methods, particularly in the face of increasing contamination. Machine learning has sometimes been used to monitor groundwater, air quality, and climate. Traditional methods must be improved due to the complexity and large amount of environmental data. This includes using hybrid models that combine traditional and new techniques. Despite the use of machine learning in many scientific areas, there is a lack of comprehensive reviews focusing on its use in environmental monitoring, especially groundwater monitoring. We aim to fill this gap by exploring machine-learning applications in groundwater monitoring. We discuss relevant methods, their limitations, and future potential. We summarize research on automating data processing and model training using groundwater sensor data. Our research underscores the transformative potential of machine learning to revolutionize long-term groundwater monitoring and contamination detection, providing valuable insights for future research and practical applications.

AI/ML↗

FunM2C: A Filter for Uncertainty Visualization of Multivariate Data on Multi-Core Devices

Uncertainty visualization is an emerging research topic in data visualization because neglecting uncertainty in visualization can lead to inaccurate assessments. In this paper, we study the propagation of multivariate data uncertainty in visualization. Although there have been a few advancements in probabilistic uncertainty visualization of multivariate data, three critical challenges remain to be addressed. First, the state-of-the-art probabilistic uncertainty visualization framework is limited to bivariate data (two variables). Second, existing uncertainty visualization algorithms use computationally intensive techniques and lack support for cross-platform portability. Third, as a consequence of the computational expense, integration into production visualization tools is impractical. In this work, we address all three issues and make a threefold contribution. First, we take a step to generalize the state-of-the-art probabilistic framework for bivariate data to multivariate data with an arbitrary number of variables. Second, through utilization of VTK-m’s shared-memory parallelism and cross-platform compatibility features, we demonstrate acceleration of multivariate uncertainty visualization on different many-core architectures, including OpenMP and AMD GPUs. Third, we demonstrate the integration of our algorithms with the ParaView software. We demonstrate the utility of our algorithms through experiments on multivariate simulation data with three and four variables.

Hari, Gautam↗

Dataset: Breaking the barrier of human-annotated training data for machine-learning-aided plant research using aerial imagery

This dataset supports the implementation described in the manuscript "Breaking the Barrier of Human-Annotated Training Data for Machine-Learning-Aided Biological Research Using Aerial Imagery." It comprises UAV aerial imagery used to execute the code available at https://github.com/pixelvar79/GAN-Flowering-Detection-paper. For detailed information on dataset usage and instructions for implementing the code to reproduce the study, please refer to the GitHub repository.

generative and adversarial learning↗

Developing Data-Driven Synthetic Infrastructure Models for Resilience Analysis

Research on infrastructure resilience has produced promising methods to simulate and optimize complex networks to improve performance. However, restrictions on sharing infrastructure models and the steep cost of developing and maintaining infrastructure models presents a roadblock to adoption. To overcome this limitation, this research focuses on methods to create data-driven infrastructure models that will help improve infrastructure resilience and security. The analysis couples incomplete utility data, geospatial data, machine learning, and synthetic network generation methods to rapidly develop and update infrastructure models. The methods are validated using realistic utility models and site-specific data, with a focus on Puerto Rico due to its unique infrastructure challenges and available data. This research highlights promising opportunities for the use of synthetic network generation and machine learning to create infrastructure models when very little data is available. Results demonstrate that hybrid methods, which combine sparse utility data with synthetic models, can enhance model accuracy, and machine learning can predict model attributes using training data from other models. However, the complexity of infrastructure systems means that even minor changes in network connectivity can significantly impact simulation results. Resilience analysis using synthetic infrastructure models shows that while some system behaviors are preserved, the magnitude of disruptions may not be accurately represented, indicating the need for more research and validation before using synthetic models for critical infrastructure investment decisions. The framework outlined in this report represents a significant advance to infrastructure model development and could be applied to additional domains and sites. Future research will continue to streamline and validate methods to help reduce roadblocks to resilience analysis.

24 POWER TRANSMISSION AND DISTRIBUTION↗