Search NASA⌕ Search

SEARCH · Search NASA

Results for “scientific reproducibility”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Science Capsule: Towards Sharing and Reproducibility of Scientific Workflows

Workflows are increasingly processing large volumes of data from scientific instruments, experiments and sensors. These workflows often consist of complex data processing and analysis steps that might include a diverse ecosystem of tools and also often involve human-in-the-loop steps. Sharing and reproducing these workflows with collaborators and the larger community is critical but hard to do without the entire context of the workflow including user notes and execution environment. In this paper, we describe Science Capsule, which is a framework to capture, share, and reproduce scientific workflows. Science Capsule captures, manages and represents both computational and human elements of a workflow. It automatically captures and processes events associated with the execution and data life cycle of workflows, and lets users add other types and forms of scientific artifacts. Science Capsule also allows users to create `workflow snapshots' that keep track of the different versions of a workflow and their lineage, allowing scientists to incrementally share and extend workflows between users. Our results show that Science Capsule is capable of processing and organizing events in near real-time for high-throughput experimental and data analysis workflows without incurring any significant performance overheads.

Ghoshal, Devarshi↗

Software Quality Assurance for High Performance Computing Containers

Software containers are a key channel for delivering portable and reproducible scientific software in high performance computing (HPC) environments. HPC environments are different from other types of computing environments primarily due to usage of the message passing interface (MPI) and drivers for specialized hard- ware to enable distributed computing capabilities. This distinction directly impacts how software containers are built for HPC applications and can complicate software quality assurance efforts including portability and performance. This work introduces a strategy for building containers for HPC applications that adopts layering as a mechanism for software quality assurance. The strategy is demonstrated across three different HPC systems, two of them petaflops scale with entirely different interconnect technologies and/or processor chipsets but running the same container. Performance consequences of the containerization strategy are found to be less than 5-14% while still achieving portable and reproducible containers for HPC systems.

97 MATHEMATICS AND COMPUTING↗

Characterizing Families of Spectral Similarity Scores and Their Use Cases for Gas Chromatography–Mass Spectrometry Small Molecule Identification

Metabolomics provides a unique snapshot into the world of small molecules and the complex biological processes that govern the human, animal, plant, and environmental ecosystems encapsulated by the One Health modeling framework. However, this “molecular snapshot” is only as informative as the number of metabolites confidently identified within it. The spectral similarity (SS) score is traditionally used to identify compound(s) in mass spectrometry approaches to metabolomics, where spectra are matched to reference libraries of candidate spectra. Unfortunately, there is little consensus on which of the dozens of available SS metrics should be used. This lack of standard SS score creates analytic uncertainty and potentially leads to issues in reproducibility, especially as these data are integrated across other domains. In this work, we use metabolomic spectral similarity as a case study to showcase the challenges in consistency within just one piece of the One Health framework that must be addressed to enable data science approaches for One Health problems. Here, using a large cohort of datasets comprising both standard and complex datasets with expert-verified truth annotations, we evaluated the effectiveness of 66 similarity metrics to delineate between correct matches (true positives) and incorrect matches (true negatives). We additionally characterize the families of these metrics to make informed recommendations for their use. Our results indicate that specific families of metrics (the Inner Product, Correlative, and Intersection families of scores) tend to perform better than others, with no single similarity metric performing optimally for all queried spectra. This work and its findings provide an empirically-based resource for researchers to use in their selection of similarity metrics for GC-MS identification, increasing scientific reproducibility through taking steps towards standardizing identification workflows.

59 BASIC BIOLOGICAL SCIENCES↗

Research Software Engineering Efforts for DataFed: FY2021 Developments

DataFed is a scientific data management system for big data providing simple and uniform data access, organization, discovery, and sharing within and across scientific facilities - with the goals of enhancing productivity and scientific reproducibility. We hope to aid in communicating development contributions to DataFed over the last year and any planning we can provide for the future developments in the next year

42 ENGINEERING↗

Science Capsule - Capturing the Data Life Cycle

The data generated from scientific workflows often become unusable due to the lack or incompleteness of information required for processing and analyzing the data. Reproducibility of scientific data and workflows facilitates efficient processing and analyses. A key to enabling reproducibility is to capture the end-to-end workflow life cycle, and any contextual metadata and provenance.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Standards, dissemination, and best practices in systems biology

In this study, the reproducibility of scientific research is crucial to the success of the scientific method. Here, we review the current best practices when publishing mechanistic models in systems biology. We recommend, where possible, to use software engineering strategies such as testing, verification, validation, documentation, versioning, iterative development, and continuous integration. In addition, adhering to the Findable, Accessible, Interoperable, and Reusable modeling principles allows other scientists to collaborate and build off of each other’s work. Existing standards such as Systems Biology Markup Language, CellML, or Simulation Experiment Description Markup Language can greatly improve the likelihood that a published model is reproducible, especially if such models are deposited in well-established model repositories. Where models are published in executable programming languages, the source code and their data should be published as open-source in public code repositories together with any documentation and testing code. For complex models, we recommend container-based solutions where any software dependencies and the run-time context can be easily replicated.

59 BASIC BIOLOGICAL SCIENCES↗

Universal Workflow Language and Software Enable Geometric Learning and FAIR Scientific Protocol Reporting

Written language and conventional data structures for representing scientific procedures suffer from low process detail, often fail to accurately represent protocols, and lack universality. New strategies for the handling of experimental data are needed to provide viable process information for both humans and machines. In this work, we present the universal workflow language (UWL) and interface (UWLi). UWL is a findable, accessible, interoperable, and reusable (FAIR)-compatible, graph-based data architecture that can capture arbitrary scientific procedures through workflow representation, and UWLi is an accompanying software package for building, manipulating, and interpreting UWL entries. The UWL format was found to be highly effective in identifying deficiencies in the reported process details of high-impact, peer-reviewed scientific journals, and in simulated scenarios, the graph format was shown to be more effective than conventional methods in predictively modeling the outcome of diverse scientific protocols. Implementation of UWL could enable more accurate scientific communication and more impactful process datasets.

14 SOLAR ENERGY↗

SurFE-XD (Surface curvature-driven Finite Elements model for Diffusion under eXtreme conditions)

SurFE-XD is a mesoscale finite element framework to model surface diffusion under mutliphysics environments. The code uses legacy C++ library dolphin wrapped with python in a FEniCS driven unified form language and just-in-time (JIT) compilation setting. The purpose of the release is to attract wide-ranging usage of the code along with publication supplementation to support reproducibility of scientific data. SurFE-XD has been originally conceived under the LDRD-DR funding for “High-Gradient (C-BAND) Breakdown tolerant accelerator materials project. Currently SrFE-XD support electrostatics and Thermo-elasticity driven surface diffusion kernels. Releasing the code will also enable to include contributions from other physical regimes e.g., plasticity and electrodynamics etc as well portability to GPU-based platforms.

Bagchi, Soumendu↗

Code for BALDR Study 07.04

SAND2024-11256O The Code for BALDR Study 07.04 software reproduces results from the BALDR study concerning "Multilabel Proportion Prediction and Out-of-Distribution Detection on Gamma Spectra of Short-Lived Fission Products." This code can reproduce a scientific study following the step numbers present in the file names. The scientific study uses synthetic and measured data to find the best model for the radioisotope proportion estimation task of interest and generates results. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Morrow, Tyler↗

Lost in transduction: Critical considerations when using viral vectors

The application of retroviral vectors in the laboratory requires considerations that often go overlooked but are often easy to circumvent. Here, we discuss the relationship between the observed transduction efficiency of a cell population and per-cell viral insertions—and describe how differential cell-type susceptibilities can confound results. We consider the math underlying this problem and review an alternative approach to the commonly used “multiplicity of infection” (MOI) method of titering and using viral vectors in the biomedical research laboratory.

59 BASIC BIOLOGICAL SCIENCES↗

Code for BALDR Study 07.05 v.1.0.0

SAND2024-01293O The Code for BALDR Study 07.05 software reproduces results from the BALDR study concerning "A Semi-Supervised Learning Method to Produce Explainable Radioisotope Proportion Estimates for NaI-based Synthetic and Measured Gamma Spectra." The code for BALDR Study 07.05 can reproduce a scientific study following the step numbers within the file names. Using synthetic and measured data to find the best model for the radioisotope proportion estimation task of interest, the study generated results for inclusion in a paper. The high-level methods involved are neural networks, semi-supervised learning, and out-of-distribution detection. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Morrow, Tyler↗

HPC-driven computational reproducibility in numerical relativity codes: a use case study with IllinoisGRMHD

Abstract Reproducibility of results is a cornerstone of the scientific method. Scientific computing encounters two challenges when aiming for this goal. Firstly, reproducibility should not depend on details of the runtime environment, such as the compiler version or computing environment, so results are verifiable by third-parties. Secondly, different versions of software code executed in the same runtime environment should produceconsistent numerical results for physical quantities. In this manuscript, we test the feasibility of reproducing scientific results obtained using theIllinoisGRMHDcode that is part of an open-source community software for simulation in relativistic astrophysics, theEinstein Toolkit. We verify that numerical results of simulating a single isolated neutron star withIllinoisGRMHDcan be reproduced, and compare them to results reported by the code authors in 2015. We use two different supercomputers: Expanse at SDSC, and Stampede2 at TACC. By compiling the source code archived along with the paper on both Expanse and Stampede2, we find thatIllinoisGRMHDreproduces results published in its announcement paper up to errors comparable to round-off level changes in initial data parameters. We also verify that a current version ofIllinoisGRMHDreproduces these results once we account for bug fixes which have occurred since the original publication.

Astronomy & Astrophysics↗

Managing Software Provenance to Enhance Reproducibility in Computational Research

Scientific processes rely on software as an important tool for data acquisition, analysis, and discovery. Over the years, sustainable software development practices have made progress in being considered as an integral component of research. However, management of computation-based scientific studies is often left to individual researchers who design their computational experiments based on personal preferences and the nature of the study. Here, we believe that the quality, efficiency, and reproducibility of computation-based scientific research can be improved by explicitly creating an execution environment that allows researchers to provide a clear record of traceability. This is particularly relevant to complex computational studies in high-performance computing (HPC) environments. In this article, we review the documentation required to maintain a comprehensive record of HPC computational experiments for reproducibility. We also provide an overview of tools and practices that we have developed to perform such studies around Flash-X, a multiphysics scientific software.

97 MATHEMATICS AND COMPUTING↗

AI Benchmark Democratization and Carpentry

Benchmarks are a cornerstone of modern machine learning, enabling reproducibility, comparison, and scientific progress. However, AI benchmarks are increasingly complex, requiring dynamic, AI-focused workflows. Rapid evolution in model architectures, scale, datasets, and deployment contexts makes evaluation a moving target. Large language models often memorize static benchmarks, causing a gap between benchmark results and real-world performance. Beyond traditional static benchmarks, continuous adaptive benchmarking frameworks are needed to align scientific assessment with deployment risks. This calls for skills and education in AI Benchmark Carpentry. From our experience with MLCommons, educational initiatives, and programs like the DOE's Trillion Parameter Consortium, key barriers include high resource demands, limited access to specialized hardware, lack of benchmark design expertise, and uncertainty in relating results to application domains. Current benchmarks often emphasize peak performance on top-tier hardware, offering limited guidance for diverse, real-world scenarios. Benchmarking must become dynamic, incorporating evolving models, updated data, and heterogeneous platforms while maintaining transparency, reproducibility, and interpretability. Democratization requires both technical innovation and systematic education across levels, building sustained expertise in benchmark design and use. Benchmarks should support application-relevant comparisons, enabling informed, context-sensitive decisions. Dynamic, inclusive benchmarking will ensure evaluation keeps pace with AI evolution and supports responsible, reproducible, and accessible AI deployment. Community efforts can provide a foundation for AI Benchmark Carpentry.

von Laszewski, Gregor [Virginia U.]↗

MLCommons Science Benchmarks

Benchmarks are a cornerstone of modern machine learning practice, providing standardized eval- uations that enable reproducibility, comparison, and scientific progress. Yet, as AI systems particularly deep learning models become increasingly dynamic, traditional static benchmarking approaches are losing their relevance. Models rapidly evolve in architecture, scale, and capability; datasets shift; and deployment contexts continuously change, creating a moving target for evaluation. Without adaptive benchmarking frame- works, both scientific assessment and real-world de- ployment risk becoming misaligned with actual system behavior. Drawing on our experience from MLCommons, educa- tional initiatives, and government programs such as the DOE s Million Parameter Consortium, we identify key barriers that hinder the broader adoption and utility of benchmarking in AI. These include substantial resource demands, limited access to specialized hardware, lack of expertise in benchmark design, and uncertainty among practitioners about how to relate benchmark results to their own application domains. Moreover, current benchmarks often emphasize peak performance on leadership-class hardware, offering limited guidance for more diverse, real-world deployment scenarios. We argue that benchmarking itself must become dy- namic in order to incorporate evolving models, updated data, and heterogeneous computational platforms while maintaining transparency, reproducibility, and inter- pretability. Democratizing this process requires not only technical innovation, but also systematic educational efforts spanning undergraduate to professional levels to develop sustained expertise in benchmark design and use. Finally, benchmarks should be framed and com- municated to support application-relevant comparisons, enabling both developers and users to make informed, context-sensitive decisions. Advancing dynamic and inclusive benchmarking practices will be essential to ensure that evaluation keeps pace with the evolving AI landscape and supports responsible, reproducible, and accessible AI deployment.

Hawks, Benjamin G. [Fermilab]↗

Evaluating integration and performance of containerized climate applications on a Hewlett Packard Enterprise Cray system

Containers have taken over large swaths of cloud computing as the most convenient way of packaging and deploying applications. The features that containers offer for packaging and deploying applications translate to high performance computing (HPC) as well. At The National Oceanic and Atmospheric Administration, containers provide an easy way to build and distribute complex HPC applications, allowing faster collaboration, portability, and experiment computer environment reproducibility amongst the scientific community. The challenge arises when applications rely on message passing interface (MPI). This necessitates investigation into how to properly run these applications with their own unique requirements and produce performance on par with native runs. We investigate the MPI performance for benchmarks and containerized climate models for various containers covering selection of compiler and MPI library combinations from the Cray provided programming environments on the Cray XC supercomputer GAEA. Performance from the benchmarks and the climate models shows that for the most part containerized applications perform on par with the natively built applications when the system optimized Cray MPICH libraries are bound into the container, and the hybrid model containers have poor performance in comparison. We also describe several challenges and our solutions in running these containers, particularly challenges with heterogeneous jobs for the containerized model runs.

Abraham, Subil↗

A guideline to document occupant behavior models for advanced building controls

The availability of computational power, and a wealth of data from sensors have boosted the development of model-based predictive control for smart and effective control of advanced buildings in the last decade. More recently occupant-behavior models have been developed for including people in the building control loops. However, while important objectives of scientific research are reproducibility and replicability of results, not all information is available from published documents. Therefore, the aim of this paper is to propose a guideline for a thorough and standardized occupant-behavior model documentation. For that purpose, the literature screening for the existing occupant behavior models in building control was conducted, and the occupant behavior modeling processes were studied to extract practices and gaps for each of the following phases: problem statement, data collection, and preprocessing, model development, model evaluation, and model implementation. Here, the literature screening pointed out that the current state-of-the-art on model documentation shows little unification, which poses a particular burden for the model application and replication in field studies. In addition to the standardized model documentation, this work presented a model-evaluation schema that enabled benchmarking of different models in field settings as well as the recommendations on how OB models are integrated with the building system.

Building control↗

A reporting format for field measurements of soil respiration

Field observations of the soil-to-atmosphere CO2 flux–soil respiration, RS–are a prime example of ‘long tail’ data that historically have had neither centralized databases nor an agreed-upon reporting format. This has hindered scientific transparency, analytical reproducibility, and novel syntheses with respect to this globally-important component of the carbon cycle. Here we propose a new data and metadata reporting format for RS data, based on engagement with a wide range of researchers in the field as well as expert advisory panels. Our goal was a reporting format that would be relevant and useful for synthesis activities, and optimizing data discoverability and usability while not placing an undue burden on data contributors. We describe previous RS data collection efforts, lessons learned from related databases and data-oriented networks (e.g. FLUXNET) in earth and ecological sciences, and the process of community consultation. The proposed reporting format focuses on chamber-level data and metadata, specifying measurement conditions and, for a given measurement period defined by beginning and ending timestamps, a mean RS flux (or CO2 concentration) and associated ancillary measurements. Fundamentally, this format aims to enable findable, accessible, interoperable, and reusable data, while providing ‘future-proofing’ capabilities to support reanalyses using as yet unknown algorithms or approaches. Finally, this proposed RS reporting format is available online, and is intended to be a dynamic document, subject to further community feedback and/or change in the future.

Bond-Lamberty, Benjamin↗