Search NASA⌕ Search

SEARCH · Search NASA

Results for “reproducible research”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Not just for programmers: How GitHub can accelerate collaborative and reproducible research in ecology and evolution

Abstract Researchers in ecology and evolutionary biology are increasingly dependent on computational code to conduct research. Hence, the use of efficient methods to share, reproduce, and collaborate on code as well as document research is fundamental. GitHub is an online, cloud‐based service that can help researchers track, organize, discuss, share, and collaborate on software and other materials related to research production, including data, code for analyses, and protocols. Despite these benefits, the use of GitHub in ecology and evolution is not widespread. To help researchers in ecology and evolution adopt useful features from GitHub to improve their research workflows, we review 12 practical ways to use the platform. We outline features ranging from low to high technical difficulty, including storing code, managing projects, coding collaboratively, conducting peer review, writing a manuscript, and using automated and continuous integration to streamline analyses. Given that members of a research team may have different technical skills and responsibilities, we describe how the optimal use of GitHub features may vary among members of a research collaboration. As more ecologists and evolutionary biologists establish their workflows using GitHub, the field can continue to push the boundaries of collaborative, transparent, and open research.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

The need for reproducible research in soft robotics

In recent years we have witnessed the rise of commercialization efforts for soft robotic technology, including soft grippers (Soft Robotics, Inc.), stretchable sensors (StretchSense, Inc. and Lightlace, Inc.), and platforms for human-robot interaction (Festo Inc., Meta Reality Labs, Toyota Research Institute, and Disney Research). However, commercialization as a whole lags the trends enjoyed by other robotic technology at equivalent points in their respective lifecycles.

43 PARTICLE ACCELERATORS↗

Membrane Pretreatment and Cell Conditioning for Proton Exchange Membrane Water Electrolysis

As the research community studying proton exchange membrane water electrolysis (PEMWE) grows, it is important to develop methods that achieve transparent, reproducible research. Reproducibility of performance and durability is a challenge facing the PEMWE field, as the published literature includes a wide spread of results obtained with nominally similar materials. Prior round-robin performance benchmarking efforts [1] have identified inadequate cell conditioning as a major source of variation in apparent cell performance. Inadequate pre-treatment and conditioning can lead to instability in initial performance and adds ambiguity to durability measurements. However, excessively prolonged conditioning procedures limit the throughput of testing and the pace of research. Pretreatment and conditioning methods vary significantly across the research literature, but little systematic investigation is available into the mechanisms of these procedures or how procedural differences may impact results. This presentation will discuss investigations into the effects of membrane pre-treatment and operating procedures during cell conditioning on initial performance and catalyst-specific accelerated stress tests, with the aim of recommending procedures to enable clear, reproducible, and high-throughput research. Methods investigated include the use of hydrogen peroxide, acids, and hydration at elevated temperature for membrane pretreatment, and conditioning procedures such as current or voltage holds and cycling. The investigations cover both the impacts of these procedures on cell performance and stability as well as underlying mechanisms and processes taking place in the cell materials.

cell↗

SANSMIC v.1.0

SAND2024-01036O The updated Sandia Solution Mining Code (SANSMIC) modeling software models the dissolution of underground salt cavern walls due to raw-water injection (leaching), assuming an oil blanket. A previous version of this software modeled the impacts of leaching at the Strategic Petroleum Reserve and has been used in research for several decades, but the source code was not available. The new version will enhance current research as additional subsurface storage of energy in solution-mined caverns increases. It will also allow research reproducibility by making the software available to the wider research community. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Hart, David↗

How should reproducibility be approached in plastic recycling?

With the growing importance of developing new and improved methodologies for plastic recycling, conducting reproducible research and ensuring that results are transferable across labs are increasingly important. This Voices article reflects on how academia and industry view the path forward for strengthening reproducibility to advance science and enable a circular plastics economy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

PubChemLite Plus Collision Cross Section (CCS) Values for Enhanced Interpretation of Nontarget Environmental Data

Finding relevant chemicals in the vast (known) chemical space is a major challenge for environmental and exposomics studies leveraging nontarget high resolution mass spectrometry (NT-HRMS) methods. Chemical databases now contain hundreds of millions of chemicals, yet many are not relevant. This article details an extensive collaborative, open science effort to provide a dynamic collection of chemicals for environmental, metabolomics, and exposomics research, along with supporting information about their relevance to assist researchers in the interpretation of candidate hits. The PubChemLite for Exposomics collection is compiled from ten annotation categories within PubChem, enhanced with patent, literature and annotation counts, predicted partition coefficient (logP) values, as well as predicted collision cross section (CCS) values using CCSbase. Monthly versions are archived on Zenodo under a CC-BY license, supporting reproducible research, and a new interface has been developed, including historical trends of patent and literature data, for researchers to browse the collection. This article details how PubChemLite can support researchers in environmental and exposomics studies, describes efforts to increase the availability of experimental CCS values, and explores known limitations and potential for future developments. The data and code behind these efforts are openly available.

PubChem↗

Dataset of Generative AI Workload Power Profiles

This dataset provides a collection of high-resolution (5/10 Hz or every 0.2/0.1 seconds) power consumption profiles for generative artificial intelligence (GenAI) workloads executed on NLR's High Performance Computing (HPC) platform Kestrel. The dataset also includes examples of representative whole-facility power profiles generated using a bottom-up, event-driven, data center energy model . This dataset is designed to support research in energy modeling, infrastructure planning, energy system integration, and sustainability analysis for AI-driven computing systems. The dataset captures time-resolved electrical power measurements across a diverse set of configurations, including variations in job type (inference vs. training), workload (LLM vs. image generation), datasets, and number of compute nodes. Power traces are provided in a standardized format and include both raw/instantaneous and aggregated files. Each profile is accompanied by metadata describing workload parameters, enabling reproducibility and cross-study comparison. The dataset is intended for use in applications such as data center infrastructure planning, energy modeling, demand response and grid impact studies, and development and validation of system-level simulation tools. By making these workload-specific power profiles publicly available, this dataset aims to address the current lack of open, empirical energy data for generative AI systems and to facilitate transparent, reproducible research on the energy and environmental impacts of large-scale AI deployment. If you use this dataset, please cite the associated publication: Vercellino et al., “Measurement of Generative AI Workload Power Profiles for Whole-Facility Data Center Infrastructure Planning,” arXiv:2604.07345 (2026).

97 MATHEMATICS AND COMPUTING↗

Vehicular Re-Identification from Uncontrolled Multiple Views

Vehicle re-identification (re-ID) across disparate sensing modalities remains a fundamental challenge for transportation research. In this work, we introduce a deep multi-view vehicle re-ID framework that leverages Siamese networks to compare pairs of vehicle images and produce matching scores, enabling robust association across drastically different viewpoints such as those from UAVs, surveillance cameras, and ground sensors. The model exploits convolutional neural networks to learn features that remain discriminative under changes in angle, distance, and illumination, supporting more generalizable re-ID performance. As part of this effort, we also developed an automated pipeline to synchronize roadside and UAV video streams, producing a multi-perspective dataset that complements preexisting real collections and a synthetic dataset generated in this study. Together, these contributions advance the capability to re-identify vehicles across wide viewing baselines; establish a foundation for scalable, reproducible research in vehicle re-ID; and open pathways for future applications, such as inferring routine behaviors, movement patterns, and daily habits of the individual associated with the vehicle.

convolutional neural networks↗

duqling

The duqling R package is a tool for facilitating reproducible research in UQ and ML by providing easy-to-use and consistent access to a wide variety of popular (public) test functions

Rumsey, Kelin↗

User Guide: A Curated Dataset of Regional Meteor Events with Simultaneous Optical and Infrasound Observations

This user guide supports a curated dataset of 71 meteor events recorded between 2006 and 2011 in Southwestern Ontario, Canada. Each event was simultaneously observed by ground-based optical cameras and an infrasound array, providing a rare opportunity to examine meteor trajectories and acoustic signals from the same atmospheric entry events. The dataset includes raw and processed optical data, meteor trajectories, photometric light curves, infrasound waveforms, and atmospheric specifications relevant for acoustic modeling. The archive is structured to support reproducible research in meteor physics, atmospheric acoustics, and shock wave analysis. It is organized following transparent file naming conventions and structured folders to facilitate scientific reuse, comparison, and integration across research domains. The dataset is freely available on Zenodo, doi: 10.5281/zenodo.15868512.

54 ENVIRONMENTAL SCIENCES↗

Persistence Control of Engineered Functions in Complex Soil Microbiomes (PerCon SFA), Secure Biosystems Design Project Data Catalog at PNNL DataHub

The Persistence Control of Engineered Functions in Complex Soil Microbiomes Project (PerCon SFA) at Pacific Northwest National Laboratory (PNNL) is a Genomic Sciences Program Biosystems Design, Science Focus Area research project consortium. Collaborating across highly integrated institutions, PerCon SFA scientists are exploring how environmental niches can be sculpted using the mechanisms of genome reduction and metabolic addiction to drive secure rhizosphere community design for robust biomass cropping in challenging environments. The PerCon SFA DataHub project repository contains publication-relevant digital dataset and metadata DOI packages, enabling exploration and download of integrated experimental dataset catalogs publicly available to a global scientific community. and metadata repository allow for exploring and downloading integrated experimental biodesign omics dataset catalogs, including experimental protocols and/or workflows, raw and/or processed data, as required by the repository, and other relevant supporting materials and/or metadata required for research reproducibility and reporting.

59 BASIC BIOLOGICAL SCIENCES↗

Data for reproducing the figures of the paper Multimodal Super-Resolution: Discovering hidden physics and its application to fusion plasmas

This deposit contains the raw data for reproducing research results of the paper Multimodal Super-Resolution: Discovering hidden physics and its application to fusion plasmas. The main contribution of this work is to utilize machine learning techniques to reconstruct and enhance the resolution of a diagnostic measurement from other available diagnostics in a system. The proposed techniques is called Diag2Diag.

diag2diag↗

PSY (PowerSystems.jl) [SWR-23-105]

PowerSystems.jl is a package to organize and manipulate data for the study of energy systems with diverse modeling requirements. This software serves two main purposes: to reduce the burden of large power system data set development, and to promote reproducible research and simulation. PowerSystems.jl implements an abstract hierarchy to represent and customize power systems data and includes data containers for quasi-static and dynamic simulation applications. Key features include efficient management of large quantities of time series data, optimized serialization, and comprehensive validation capabilities. For further information regarding the Sienna modeling framework at NREL, see: https://www.nrel.gov/analysis/sienna.html

Lara Aguilar, Jose Daniel↗

Reproducibility, Replicability, and Research Quality in Homogeneous Catalysis

Researchers from all sectors of homogeneous catalysis convened in response to concerns regarding reproducibility in science to analyze the issue and provide recommendations. In addition to an in- person workshop, the group engaged the broader homogeneous catalysis community through a webinar series and virtually during the workshop. Results of the project affirm that homogeneous catalysis is not in a reproducibility crisis, as evidenced by the field’s current and past contributions to society that have led to economic growth and advances in a range of industries from agriculture to consumer goods to human health. However, it is not uncommon for researchers to encounter obstacles related to reproducibility. Ensuring reproducibility remains the responsibility of the community, both in current work and in training future researchers. This report is intended to engage key stakeholders, including disciplinary societies, publishers, employers, research leaders, and researchers, in practices that maximize reproducible homogeneous catalysis and ensure continued innovation and translatable discoveries. Recommendations made herein are also framed to be applicable beyond homogeneous catalysis, empowering the broader chemical if not scientific community.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Managing Software Provenance to Enhance Reproducibility in Computational Research

Scientific processes rely on software as an important tool for data acquisition, analysis, and discovery. Over the years, sustainable software development practices have made progress in being considered as an integral component of research. However, management of computation-based scientific studies is often left to individual researchers who design their computational experiments based on personal preferences and the nature of the study. Here, we believe that the quality, efficiency, and reproducibility of computation-based scientific research can be improved by explicitly creating an execution environment that allows researchers to provide a clear record of traceability. This is particularly relevant to complex computational studies in high-performance computing (HPC) environments. In this article, we review the documentation required to maintain a comprehensive record of HPC computational experiments for reproducibility. We also provide an overview of tools and practices that we have developed to perform such studies around Flash-X, a multiphysics scientific software.

97 MATHEMATICS AND COMPUTING↗

Good practices for documenting AI-based studies on energy and buildings

Artificial intelligence has transformed building science research over the past decade, with applications spanning energy modeling, energy prediction, HVAC optimization and controls, fault detection, and occupancy modeling. However, many studies lack adequate documentation of datasets, algorithms, training procedures, and validation methods. Building science research faces additional challenges including inconsistent evaluation metrics, limited generalizability across building types, climates, and significant gaps between experimental studies and deployed systems. This communication provides practical guidance for good practices in documenting and publishing AI-based research following established standards from the computer science and machine learning communities. By adopting frameworks such as Datasheets for Datasets, Model Cards, and standardized reproducibility checklists, researchers can ensure their work meets the rigorous documentation standards necessary for reproducible, comparable, and impactful building science research.

Hong, Tianzhen [Lawrence Berkeley National Laborat↗

Efforts to enhance reproducibility in a human performance research project

Background: Ensuring the validity of results from funded programs is a critical concern for agencies that sponsor biological research. In recent years, the open science movement has sought to promote reproducibility by encouraging sharing not only of finished manuscripts but also of data and code supporting their findings. While these innovations have lent support to third-party efforts to replicate calculations underlying key results in the scientific literature, fields of inquiry where privacy considerations or other sensitivities preclude the broad distribution of raw data or analysis may require a more targeted approach to promote the quality of research output. Methods: We describe efforts oriented toward this goal that were implemented in one human performance research program, Measuring Biological Aptitude, organized by the Defense Advanced Research Project Agency's Biological Technologies Office. Our team implemented a four-pronged independent verification and validation (IV&V) strategy including 1) a centralized data storage and exchange platform, 2) quality assurance and quality control (QA/QC) of data collection, 3) test and evaluation of performer models, and 4) an archival software and data repository. Results: Our IV&V plan was carried out with assistance from both the funding agency and participating teams of researchers. QA/QC of data acquisition aided in process improvement and the flagging of experimental errors. Holdout validation set tests provided an independent gauge of model performance. Conclusions: In circumstances that do not support a fully open approach to scientific criticism, standing up independent teams to cross-check and validate the results generated by primary investigators can be an important tool to promote reproducibility of results.

59 BASIC BIOLOGICAL SCIENCES↗

Making digital objects FAIR in high energy physics: An implementation for Universal FeynRules Output (UFO) models

Research in the data-intensive discipline of high energy physics (HEP) often relies on domain-specific digital contents. Reproducibility of research relies on proper preservation of these digital objects. This paper reflects on the interpretation of principles of Findability, Accessibility, Interoperability, and Reusability (FAIR) in such context and demonstrates its implementation by describing the development of an end-to-end support infrastructure for preserving and accessing Universal FeynRules Output (UFO) models guided by the FAIR principles. UFO models are custom-made python libraries used by the HEP community for Monte Carlo simulation of collider physics events. Our framework provides simple but robust tools to preserve and access the UFO models and corresponding metadata in accordance with the FAIR principles.

Neubauer, Mark S.↗