Search NASA⌕ Search

SEARCH · Search NASA

Results for “workflow development”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Arrayed in vivo barcoding for multiplexed sequence verification of plasmid DNA and demultiplexing of pooled libraries

Sequence verification of plasmid DNA is critical for many cloning and molecular biology workflows. To leverage high-throughput sequencing, several methods have been developed that add a unique DNA barcode to individual samples prior to pooling and sequencing. However, these methods require an individual plasmid extraction and/or in vitro barcoding reaction for each sample processed, limiting throughput and adding cost. Here, we develop an arrayed in vivo plasmid barcoding platform that enables pooled plasmid extraction and library preparation for Oxford Nanopore sequencing. This method has a high accuracy and recovery rate, and greatly increases throughput and reduces cost relative to other plasmid barcoding methods or Sanger sequencing. We use in vivo barcoding to sequence verify >45 000 plasmids and show that the method can be used to transform error-containing dispersed plasmid pools into sequence-perfect arrays or well-balanced pools. In vivo barcoding does not require any specialized equipment beyond a low-overhead Oxford Nanopore sequencer, enabling most labs to flexibly process hundreds to thousands of plasmids in parallel.

59 BASIC BIOLOGICAL SCIENCES↗

CVEVOLVE

CVEvolve is an agentic AI system for autonomous algorithm discovery for scientific data processing. It creates workflows where large language model agents freely set up and configure development environments and evaluation harnesses, develop and improve data processing algorithms with designed exploration-exploitation balancing mechanisms, log history and findings in a structured database, and run holdout testing to ensure algorithm generalizability. CVEvolve offers a zero-code interface and does not require users to provide structured data and evaluation scripts.

Cherukara, MatthewJoseph [Argonne National Laborat↗

Bottom-up design of actinide materials from molecular clusters: Demonstration of a general-purpose simulation capability leveraging machine-learned atomic potentials

Actinide thin-film coatings such as uranium dioxide (UO 2 ) play an important role in nuclear reactors and other mission-relevant applications, but realization of their potential requires a deep fundamental understanding of the chemical vapor deposition (CVD) processes used for their growth. The slow experimental progress can be attributed, in part, to the standard safety guidelines associated with handling uranium byproducts, which are often corrosive, toxic, and radioactive. Accurate simulation techniques, when used in concert with experiment, can improve laboratory safety, material durability, and deliverable timeframes. However, state-of-the-art computational methods are either insufficiently accurate or intractably expensive. To remedy this situation, in this project we suggested a machine-learning (ML) accelerated workflow for simulating molecular clustering toward deposition. As a benchmark test case, we considered molecular clustering in steam and assessed independent components of our workflow by comparing with measured thermodynamic properties of water. After analyzing each component individually and finding no fundamental barrier to realization of the workflow, we attempted to integrate the ML component, a Sandia-developed tool called FitSNAP. As this was the first application of FitSNAP to atoms and molecules in the gas phase at Sandia, the method required more fitting data than was originally anticipated. Systematic improvements were made by including in the fit data diatomic potentials, molecular single-bond-breaking curves, and symmetry-constrained intermolecular potentials. We concluded that our strategy provides a feasible pathway toward modeling CVD and related processes, but that extensive training data must be generated before it can be of practical use.

36 MATERIALS SCIENCE↗

PCOR partnership initiative to accelerate CCCUS deployment

The Plains CO 2 Reduction (PCOR) Partnership Initiative, led by the Energy & Environmental Research Center (EERC), was a 5-year collaborative effort under the U.S. Department of Energy’s Regional Initiative to Accelerate CCUS (carbon capture, utilization, and storage). With support from the University of Wyoming School of Energy Resources (UW SER), the University of Alaska Fairbanks Institute of Northern Engineering (UAF INE), and a broad network of public and private stakeholders, the program focused on overcoming regional challenges to CCUS across ten U.S. states and four Canadian provinces. This report summarizes the notable findings of the EERC and its collaborators through the 5-year PCOR Partnership Initiative phase. The EERC developed and demonstrated a risk-based area of review (AOR) workflow for U.S. Environmental Protection Agency Class VI carbon dioxide (CO 2 ) storage permitting.

01 COAL, LIGNITE, AND PEAT↗

Low-Temperature Geothermal Play Fairway Analysis for the Denver Basin: Preprint

This project is part of a national initiative to showcase the benefits of incorporating low-temperature geothermal resource assessment into the deployment of geothermal heating and cooling (GHC), combined heat and power (CHP), and geothermal direct use (GDU) technologies. The initiative was established to accelerate the country's decarbonization efforts by identifying potential for low-temperature geothermal resource utilization (< 150 Degrees Celsius, i.e., GHC, CHP, and GDU) in selected sedimentary basins with numerous population centers. The Play Fairway Analysis (PFA) methodologies in this study were adapted from previous PFA investigations of sedimentary basin geothermal play types (SBGPTs) that evaluated the potential for low-temperature resources (< 150 Degrees Celsius). Workflows, relevant datasets, python code, common and composite geological criteria maps are utilized to develop low-temperature geothermal resource favorability maps for the Denver Basin, a sedimentary basin spanning Colorado, Nebraska, and Wyoming. The replication of these methodologies in other SBGPTs can evaluate potential for low-temperature resources. To facilitate future assessment of low-temperature geothermal resources in SBGPTs, this project provides PFA workflows, data, tools, and favorability maps that will ultimately support the utilization of low-temperature geothermal resources in sedimentary basins.

Denver Basin↗

Numerical Modeling and Optimization of the iProTech Pitching Inertial Pump (PIP) Wave Energy Converter (WEC) (Cooperative Research and Development Final Report, CRADA Number: CRD-22-22968)

This work generated a first-of-its-kind automated workflow to couple time-domain simulations of wave energy converters written in one software language with a set of design generation and evaluation scripts written in another software language. This automated workflow used an existing optimization package to analyze the sensitivity of different design parameters on the power output of a specific WEC, iProTech’s Pitching Inertial Pump (PIP). Geometric, inertial, and power take-off variables were all varied and optimized to find values that produced the highest amount of power generated over varying wave conditions. The findings on these parameter sensitivity studies are used to inform future design iterations of the PIP WEC. Including more design variables in the optimizations will only increase computational run time and further software development is needed to analyze a larger optimization.

16 TIDAL AND WAVE POWER↗

DOE Repository Metadata Profile (DRMP): A Metadata Framework for Advancing Interoperability and AI Readiness Across Scientific Repositories

The Department of Energy (DOE) funds a diverse and distributed ecosystem of repositories that steward scientific data, publications, and software across its research programs, user facilities, and national laboratories. While significant progress has been made in standardizing dataset-level metadata, the metadata describing repositories themselves (their identity, governance, access interfaces, policies, and technical capabilities) remains inconsistent and fragmented across DOE-funded systems. This variability limits discoverability, interoperability, automated validation, and AI-driven analysis, all of which are increasingly essential for modern scientific workflows. To address this gap, the DOE Data Curation Working Group (DCWG) developed the DOE Repository Metadata Profile (DRMP). The DRMP is a practical, community-driven framework that defines how repositories can describe themselves in a consistent, machine-actionable, and scalable manner. The DRMP is not a new metadata schema. Instead, it is a mapping profile and structured element set capturing the essential characteristics of DOE repositories. It harmonizes repository-level metadata across six widely adopted community schemas: RE3Data; DCAT-US v3; Schema.org; Dublin Core; DataCite 4.6; and PREMIS 3.0. This harmonization eliminates reinvention and enables interoperability within DOE and across the broader scientific ecosystem. A core objective of the DRMP is to reduce burden on repositories by allowing them to reuse their existing metadata through a Rosetta-style crosswalk rather than redesigning local implementations. The profile introduces a three-level conformance model that supports incremental adoption: • Level 1 – Minimum Viable Record (MVR): foundational identification elements required for workflows, project registration, and basic repository presence. • Level 2 – Interoperable: structured metadata enabling alignment with national and international discovery systems. • Level 3 – AI-Ready: enhanced provenance, policy transparency, fixity, semantic context, and capabilities that support automated reasoning, model training governance, and machine-assisted curation. To support implementation, the DRMP includes JSON Schema definitions, OpenAPI patterns, and MCP templates that allow repositories to publish machine-readable metadata directly within existing platforms. These resources are modular and lightweight, enabling adoption without major architectural change. Adopting the DRMP enables repositories to: • Enhance discoverability and interoperability by aligning identifiers, classifications, and descriptive elements across widely used schema standards. • Support federated discovery and cross-registration across DOE systems, Data.gov, and international catalogs. • Enable AI agents and workflow orchestration systems to interpret repository-level metadata within the American Science Cloud (AmSC) through Model Context Protocol (MCP)-based context publication. • Demonstrate alignment with DOE’s open science, stewardship, and FAIR data priorities. This guidance represents a community-driven step forward. Through voluntary adoption and continued feedback, the DRMP advances a cohesive, machine-actionable description of DOE repositories that supports FAIR data practices, preparing the infrastructure for AI-enabled research, and strengthening the discoverability and reuse of DOE’s scientific outputs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Exocortex Network for AI-Augmented Human-Led Scientific Expedition

AI advances in science can be viewed along two main directions with a fluid boundary: enhancing efficiency through automation and smart tools to accelerate tasks that humans can already perform; and enabling exploration into uncharted territories and potentially toward AGI. These advances manifest in the AI cognitive core through the development and explainability of foundation models; in the physical embodiment of instruments and facilities; and in the integrated agency of AI workflows exemplified by the science exocortex. To address the role of humans in this evolving landscape, in this Perspective, we suggest a third direction: the development of personalized agents that form human-centered networks, supporting both efficiency and exploration while ensuring that AI remains aligned with human vision.

97 MATHEMATICS AND COMPUTING↗

Developing multi-gene CRISPRa/i programs to accelerate DBTL cycles in ABF hosts engineered for chemical production

This project developed and implemented a modular CRISPR activation and interference (CRISPRa/i) platform to accelerate strain optimization and pathway development for industrially relevant microbial hosts. By integrating multiplexed transcriptional perturbation tools with data-driven Design–Build–Test–Learn (DBTL) workflows, the team achieved reductions in cycle time and enhanced production of industrial aromatics, particularly 4-aminocinnamic acid (4-ACA), in Pseudomonas putida. Key accomplishments included: ● Development of a robust, tunable CRISPRa/i system in P. putida that enabled efficient multi-target gene regulation via guide RNA (gRNA) programs ● Completion of two full DBTL cycles, guided by machine learning (ML) models trained on transcriptomic and performance data, reducing engineering time by over 30% ● Optimization of multi-gene regulatory programs to balance expression of host and pathway modules, improve 4-ACA titers, and resolve metabolic bottlenecks ● Demonstration of system portability through a limited proof-of-concept extension in Acinetobacter baylyi, underscoring the generalizability of the approach ● Evaluation of strain performance on lignocellulosic biomass-derived substrates, demonstrating the feasibility of converting renewable carbon into aromatic building blocks These results illustrate the feasibility of applying ML-guided CRISPRa/i perturbation strategies to accelerate strain development in complex microbial systems. The resulting tools and datasets contribute to DOE objectives by improving platform predictability, reducing development costs, and enabling broader access to sustainable, economically viable bioproduction technologies.

09 BIOMASS FUELS↗

Thermo-Fluid Modeling Framework for Supercomputing Digital Twins: Part 2, Automated Cooling Models

The development of digital twins for the purpose of improving the energy efficiency of supercomputing facilities is a non-trivial endeavor that is complicated by the difficulty of creating physics-based thermo-fluid cooling system models (CSMs). Within ExaDigit---an open-source framework for liquid-cooled supercomputing digital twins---a thermo-fluid modeling framework is being developed. This effort has been segmented into two with two companion papers describing each portion of the overall effort. Part 1 focuses on the development of a cooling system library in Dymola for the Frontier supercomputer at Oak Ridge National Laboratory {\cite{Kumar2024}. Part 2, this paper, describes an effort to create a template-based auto-generation methodology for CSMs, called \textit{AutoCSM}. In this paper, an overview of the initial AutoCSM architecture and workflow is provided, along with a practical example using the Oak Ridge Leadership Computing Facility's (OLCF) Frontier supercomputer CSM. AutoCSM will (1) improve ExaDigiT's user accessibility by providing a flexible workflow for modularizing the creation of the CSM system and control logic, (2) decrease the development time of CSMs, and (3) standardize the method for incorporating CSMs into the ExaDigiT framework.

Greenwood, Scott↗

MFiX Development Updates

This presentation discusses recent developments of the Multiphase Flow with Interphase eXchanges (MFiX) software. A brief review of all modeling approaches along with their cost versus accuracy is provided to guide users when selecting a model for a given application. The main new features of the past five official releases of MFiX are described. Improvement in chemistry management allow for easier simulation setup, and faster simulation speed. Progress in the implementation of a thin-wall boundary condition is presented. The major new model development over the past year is the release of the Glued Sphere Particle model (GSP), where component spheres are combined together to represent non-spherical particles. The integration of Machine Learning (ML) workflow in the CFD process is discussed with two applications: a surrogate model for stiff chemistry and the development of a PIC stress model using ML.

Dietiker, Jeff↗

Reward Driven Workflows for Unsupervised Explainable Analysis of Phases and Ferroic Variants From Atomically Resolved Imaging Data

Rapid progress in aberration corrected electron microscopy necessitates development of robust methods for the identification of phases, ferroic variants, and other pertinent aspects of materials structure from imaging data. While unsupervised methods for clustering and classification are widely used for these tasks, their performance can be sensitive to hyperparameter selection in the analysis workflow. In this study, the effects of descriptors and hyperparameters are explored on the capability of unsupervised ML methods to distill local structural information, exemplified by the discovery of polarization and lattice distortion in Sm − dopped BiFeO 3 (BFO) thin films. It is demonstrated that a reward-driven approach can be used to optimize these key hyperparameters across the full workflow, where rewards are designed to reflect domain wall continuity and straightness, ensuring that the analysis aligns with the material's physical behavior. This approach allows the discovery of local descriptors that are best aligned with the specific physical behavior, providing insight into the fundamental physics of materials. The reward driven workflow is further extended to disentangle structural factors of variation via an optimized variational autoencoder (VAE). Lastly, the importance of well-defined rewards is explored as a quantifiable measure of the success of the workflow.

Barakati, Kamyar [University of Tennessee, Knoxvil↗

Evaluating FRI3D for Cost Savings in Fire Hazard Analysis at DOE Sites

A fire hazard analysis, required for many U.S. Department of Energy (DOE) facilities, is a complex, cumbersome, and costly process. Fire hazard analyses may be viewed as a checkbox, but ideally and in spirit with the DOE-STD-1066, the fire hazard analysis (FHA) should be a part of the workflow and used to help in modifications, maintenance, and improving operational safety. With current FHA development processes, it is both time and cost prohibitive for true integration. A tool called Fire Risk Investigation in 3D or FRI3D was developed under the DOE Light Water Reactor Sustainability program to simplify and automate many aspects of a fire probabilistic risk analysis for existing nuclear power plants. The FRI3D tool automates fire scenarios by combining approved fire simulation codes, U.S. Nuclear Regulatory Commission fire calculations methods, 3D modeling and visualization, and probabilistic risk analysis models into a single workflow supported with a user interface. FRI3D was initially designed for used in combination with a PRA, this case study, evaluated using FRI3D for a plant modification, determined the benefits that detailed fire modeling can have for U.S. Department of Energy facilities with or without a PRA model. It also looked at what tasks from DOE requirements could be reduced using the tool and what is needed to integrate fire hazard analysis into site workflow.

97 - MATHEMATICS AND COMPUTING↗

HARD: A performance portable radiation hydrodynamics code based on FleCSI framework

Hydrodynamics And Radiation Diffusion (HARD) is an open-source application for high-performance simulations of compressible hydrodynamics with radiation-diffusion coupling. Built on the FleCSI (Bergen et al., 2021 [1]) (Flexible Computational Science Infrastructure) framework, HARD expresses its computational units as tasks whose execution can be orchestrated by multiple back-end runtimes, including Legion (Bauer et al., 2012 [2]), MPI (Forum, 1994 [3]), and HPX (Kaiser et al., 2020 [4]). Node-level parallelism is handled through Kokkos (Edwards et al., 2014 [5]), providing a single-source, portable code base that runs efficiently on laptops, small homogeneous clusters, and the largest heterogeneous supercomputers currently available. To ensure scientific reliability, HARD includes a regression test suite that automatically reproduces canonical verification problems such as the Sod and LeBlanc shock tubes, and the Sedov blast wave, comparing numerical solutions against known analytical results. The project is distributed under an OSI-approved license, hosted on GitHub, and accompanied by reproducible build scripts and continuous integration workflows. This combination of performance portability, verification infrastructure, and community-focused development makes HARD a sustainable platform for advancing radiation hydrodynamics research across multiple domains.

97 MATHEMATICS AND COMPUTING↗

Data for Yield from Iowa’s first commercial miscanthus fields: implications of spatial variability for productivity and sustainability beyond research plots

This dataset contains biomass yield measurements and associated vegetation index data collected from commercial Miscanthus × giganteus fields in eastern Iowa during the 2022–2023 growing seasons. The data support the analyses presented in the article: “Yield From Iowa's First Commercial Miscanthus Fields: Implications of Spatial Variability for Productivity and Sustainability Beyond Research Plots.” We collected 105 ground-truth biomass samples from four mature commercial fields (>4 years old) covering 92.81 ha. Samples were taken from 3 m² quadrats that were hand-harvested in alignment with commercial harvest timing. Stem biomass (excluding leaves) was weighed, moisture-corrected, and converted to dry-matter yield expressed in Mg DM ha⁻¹. Sampling locations were selected to capture spatial variability visible in aerial imagery and were recorded using RTK GPS. Each biomass observation was paired with vegetation indices derived from high-resolution PlanetScope satellite imagery (3 m resolution). Images were acquired throughout the growing season, and indices were calculated to evaluate their ability to predict end-of-season biomass yield. Statistical and machine learning approaches were used to identify key predictors, and a linear regression model based on end-of-July Green Normalized Difference Vegetation Index (GNDVI) was developed and evaluated. This repository includes the data used in that modeling workflow. Management practices, economic data, full imagery time series, and additional methodological details are described in the associated publication and are not included here. The dataset consists of three comma-separated value (CSV) files: 1. Combine_Groundtruth_Yield_VI_22_23.csv This file contains ground-truth biomass yield measurements and associated key vegetation index values collected during the 2022 and 2023 growing seasons. Rows: 105 observations Columns: Year — Year of observation (2022 or 2023) Field — Field location identifier Sample_number — Unique sample identifier GNDVI_End_Jul — Green Normalized Difference Vegetation Index calculated at end of July GNDVI_End_Aug — Green Normalized Difference Vegetation Index calculated at end of August NDRE_End_Aug — Normalized Difference Red Edge index calculated at end of August Biomass_Stem_Yield_MgDM/ha — Measured stem biomass yield (megagrams dry matter per hectare) 2. trainData_GNDVI.csv This file contains the subset of observations used to train the predictive relationship between July GNDVI and biomass yield. Rows: 76 observations Columns: Unnamed: 0 — Row index retained from the original data processing workflow GNDVI_End_Jul — GNDVI at end of July Stem_Yield_MgDM/ha — Observed stem biomass yield (Mg DM ha⁻¹) 3. testData_GNDVI.csv This file contains the test dataset used to evaluate model performance. Rows: 29 observations Columns: Unnamed: 0 — Row index retained from the original data processing workflow GNDVI_End_Jul — GNDVI at end of July Predicted_Yield_MgDM/ha — Model-predicted stem biomass yield (Mg DM ha⁻¹) Observed_Yield_MgDM/ha — Measured stem biomass yield (Mg DM ha⁻¹)

Potential yield, yield gap, in-field management, y↗

AI-Ready Data Pilot Project Report

The proliferation of artificial intelligence in scientific research has created an urgent need to define "AI-ready data" for researchers and, more importantly, provide resources to help them produce AI-ready data. At Pacific Northwest National Laboratory, we conducted a pilot study with three data scientists evaluating three CSV datasets from different scientific domains, followed by semi-structured interviews capturing assessment practices. Our findings reveal that AI-readiness evaluation is intuition-based, with practitioners asking "How fast can I go from raw data to my machine learning pipeline?" Data scientists consistently prioritized workflow efficiency, human interpretability, and quality stewardship signals. From these insights, we developed a practical evaluation framework comprising data requirements, metadata standards, and validation tests that provides actionable criteria for producing and curating AI-ready datasets, addressing the gap between theoretical understanding and practical implementation.

97 MATHEMATICS AND COMPUTING↗

plexosdb: A Modular Library for Programmatic PLEXOS Model Construction

plexosdb is a lightweight Python library for constructing PLEXOS models using a SQLite-backed data structure. It provides a clear, modular interface that maps relational data directly to model components. By leveraging SQLite and idiomatic Python, it enables fast iteration and reproducible workflows. The result is a performant, composable foundation for scalable PLEXOS model development.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Scenario Generation for Built Environment Decision Support under Uncertainty: Case Studies of Airflow Modeling and Climate-Resilient Infrastructure System Design

When confronted with unforeseen challenges, practicing informed decision making is crucial for enhancing resilience in the built environment. While scan-to-building information modeling (BIM) is a well-established approach for creating detailed digital representations of physical assets, its application in assessing and improving infrastructure resilience remains underexplored. This study addresses this gap by proposing a novel application of scan-to-BIM, namely, scan-to-BIM-to-digital twin (S-BIM-DT) workflow. By integrating reality capture and digital twin technologies, this workflow creates continuously updated and accurate digital representations of physical assets, enabling the generation of various scenarios. Unlike traditional methods, the S BIM-DT workflow facilitates continuous model refinement, supporting informed resilience strategies. By combining these technologies into a cohesive process, the workflow facilitates decision making under uncertainty, enabling stakeholders to evaluate and respond to various scenarios effectively. We demonstrate the implementation of the S-BIM-DT workflow through two use cases that highlight its capability to enhance resilience at different scales. The first use case involves the Combined Transportation, Emergency, and Communications Center (CTECC) in Austin, Texas. BIM-enriched computational fluid dynamics (CFD) modeling simulates airflow and develops alternative scenarios for optimizing the heating, ventilation, and air conditioning (HVAC) systems. This approach enhances resilience against airborne health threats in a postCOVID context. The second use case focuses on designated areas within Beaumont, Texas, as part of the Southeast Texas Urban Integrated Field Laboratory (SETx-UIFL) research. By developing inundation maps to assess extreme weather events, this modeling aids in preparedness efforts and informs the development of climate-resilient infrastructure in vulnerable neighborhoods. Results indicate that the S-BIM-DT workflow effectively generates scenarios that enhance resilience in the built environment by facilitating informed decision making. Furthermore, this study serves as a bridge between advanced scan-to-BIM methodologies and the practical strategies needed to improve built infrastructure resilience.

Built environment↗