Search NASASearch

SEARCH · Search NASA

Results for “workflows”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Integrating Ultra-Coarse-Grained Protein Models into Accessible Workflows for Multiscale Molecular Dynamics

To capture protein conformational transitions using molecular dynamics (MD), several simulation resolutions covering different spatial and temporal scales are typically needed. All-atom (AA) simulations provide fine resolution, but are computationally infeasible for large systems over longer durations. Coarse-grained (CG) and ultra-coarse-grained (UCG) models have a lower resolution and computational cost while still being able to conserve essential protein features. Prior work on a Multiscale Machinelearned Modeling Infrastructure (MuMMI) combined both AA and CG simulations to study RAS-RAF protein interactions, leveraging CG models for longer time scales and using AA to investigate unusual conformations in greater detail. However, MuMMI is still resource-intensive, and this study aims to maximize exploration of the protein conformational space while reducing computational cost. In this paper, we build on prior work that integrates UCG models based on heterogeneous elastic network modeling (hENM) into the MuMMI workflow. We demonstrate that UCG models enable accurate sampling of protein conformations, focusing on simulating RAS-RAF protein interactions. Using higher-resolution CG Martini simulation data, we can automatically refine intramolecular interactions in UCG models. We present a scalable Python package that uses fluctuations observed in higher-resolution CG Martini simulations to estimate bond coefficients of the UCG model. We built novel machine learning-based backmapping methods to recover more detailed CG Martini structures from UCG structures, using diffusion models to learn the mapping between scales. Finally, we present UCG-mini-MuMMI, an accessible and less compute-intensive version of MuMMI as a resource for the scientific community. Incorporating UCG models into MD studies is applicable to a broad range of systems and proteins, and our study offers insights into the advantages and limitations of these methods.

Chemical structure

Machine Learning‐Assisted Microearthquake Location Workflow for Monitoring the Newberry Enhanced Geothermal System

Abstract Enhanced geothermal systems (EGS) offer a sustainable energy source but face challenges in accurately locating microearthquakes induced during reservoir stimulation. Locating these microearthquakes provides reliable feedback on the stimulation progress. Current deep learning methods for locating earthquakes require extensive data sets for training, which is problematic as detected microearthquakes are often limited. To address the scarcity of training data, we propose a practical workflow using probabilistic multilayer perceptron (PMLP) which predicts microearthquake locations from cross‐correlation time lags in waveforms. Utilizing a 3D velocity model of Newberry site derived from ambient noise interferometry, we generate numerous synthetic microearthquakes and 3D acoustic waveforms for PMLP training. Accurate synthetic tests prompt us to apply the trained network to the 2012 and 2014 stimulation field waveforms. To enhance the accuracy of source localization, we carefully handpick the P‐arrival times. Predictions on the 2012 stimulation data set show major microseismic activity at depths of 0.5–1.2 km, correlating with a known casing leakage scenario. In the 2014 data set, the majority of predictions concentrate at 2.0–2.9 km depths, consistent with results obtained from conventional physics‐based inversion, and align with the presence of natural fractures from 2.0 to 2.7 km. We validate our findings by comparing the synthetic and field picks, demonstrating a satisfactory match for the first arrivals. By combining the benefits of quick inference speeds and accurate location predictions, we demonstrate the feasibility of using realistic synthetic data set to locate microseismicity for EGS monitoring.

15 GEOTHERMAL ENERGY

Developing a complete AI-accelerated workflow for superconductor discovery

The quest to identify new superconducting materials with enhanced properties is hindered by the prohibitive cost of computing electron-phonon spectral functions, severely limiting the materials space that can be explored. Here, we introduce a Bootstrapped Ensemble of Equivariant Graph Neural Networks (BEE-NET), a machine-learning model trained to predict the Eliashberg spectral function and superconducting critical temperature with a mean-absolute-error of 0.87 K relative to DFT-based Allen-Dynes calculations. Intriguingly, BEE-NET achieves a true-negative-rate of 99.4%, enabling highly efficient screening for the rare property of superconductivity. Integrated into a multi-stage, AI-accelerated discovery pipeline that incorporates elemental-substitution strategies and machine-learned interatomic potentials, our workflow reduced over 1.3 million candidate structures to 741 dynamically and thermodynamically stable compounds with DFT-confirmed T c > 5 K. We report the successful synthesis and experimental confirmation of superconductivity in two of these previously unreported compounds. This study establishes a data-driven framework that integrates machine learning, quantum calculations, and experiments to systematically accelerate superconductor discovery.

Gibson, Jason B. [Quantum Formatics, Cambridge, MA

Building workflows for an interactive human-in-the-loop automated experiment (hAE) in STEM-EELS

Exploring the structural, chemical, and physical properties of matter on the nano- and atomic scales has become possible with the recent advances in aberration-corrected electron energy-loss spectroscopy (EELS) in scanning transmission electron microscopy (STEM). However, the current paradigm of STEM-EELS relies on the classical rectangular grid sampling, in which all surface regions are assumed to be of equal a priori interest. However, this is typically not the case for real-world scenarios, where phenomena of interest are concentrated in a small number of spatial locations, such as interfaces, structural and topological defects, and multi-phase inclusions. One of the foundational problems is the discovery of nanometer- or atomic-scale structures having specific signatures in EELS spectra. Herein, we systematically explore the hyperparameters controlling deep kernel learning (DKL) discovery workflows for STEM-EELS and identify the role of the local structural descriptors and acquisition functions in experiment progression. In agreement with the actual experiment, we observe that for certain parameter combinations the experiment path can be trapped in the local minima. We demonstrate the approaches for monitoring the automated experiment in the real and feature space of the system and knowledge acquisition of the DKL model. Based on these, we construct intervention strategies defining the human-in-the-loop automated experiment (hAE). This approach can be further extended to other techniques including 4D STEM and other forms of spectroscopic imaging. The hAE library is available on Github at https://github.com/utkarshp1161/hAE/tree/main/hAE.

Pratiush, Utkarsh [Univ. of Tennessee, Knoxville,

Agentic workflow enables the recovery of critical materials from complex feedstocks via selective precipitation

We present a multi-agentic workflow for critical materials recovery that deploys a series of AI agents and automated instruments to recover critical materials from produced water and magnet leachates. This approach achieves selective precipitation from real-world feedstocks using simple chemicals, accelerating the development of efficient, adaptable, and scalable separations to a timeline of days, rather than months and years.

Ritchhart, Andrew J.

Unsupervised Segmentation and Clustering Workflow for Efficient Processing of 4D-STEM and 5D-STEM Data

Four-dimensional scanning transmission electron microscopy (4D-STEM) enables mapping of diffraction information with nanometer-scale spatial resolution, offering detailed insight into local structure, orientation, and strain. However, as data dimensionality and sampling density increase, particularly for in situ scanning diffraction experiments (5D-STEM), robust segmentation of structurally consistent behavior across sequential measurements becomes essential for efficient and physically meaningful analysis. Here, we introduce a clustering framework that identifies crystallographically distinct domains from 4D-STEM datasets. By using local diffraction-pattern similarity as a metric, the method extracts closed contours delineating spatially contiguous regions. This approach produces cluster-averaged diffraction patterns that improve signal quality while reducing data volume by orders of magnitude, enabling rapid and accurate orientation, phase, and strain mapping. We demonstrate the applicability of this approach to in situ liquid-cell 4D-STEM data of gold nanoparticle growth. Our method provides a scalable and generalizable route for spatially coherent segmentation, data compression, and quantitative structure–strain mapping across diverse 4D-STEM modalities. The full analysis code and example workflows are publicly available to support reproducibility and reuse.

4D-STEM

Improving I/O-aware Workflow Scheduling via Data Flow Characterization and trade-off Analysis

The scientific computing paradigm has transitioned from compute-intensive to I/O-intensive and memory-intensive in the past decade, especially when data-driven science has become common practice. Numerous empirical I/O-aware scheduling optimizations have been developed by incorporating I/O capacity and bandwidth as constraints into scheduling. Unfortunately, there is a lack of data flow (I/O) characterization tool and an understanding of trade-offs between concurrency, locality, and I/O bandwidth. To bridge the gap, this work 1) presents a set of descriptors to characterize, organize, and visualize I/O profiles, including flow size, I/O bandwidth, and operation count, which group data flows by I/O types, tasks, and files; 2) proposes an I/O Roofline model-based trade-off analysis to find the optimal trade-off between flow operational intensity, concurrency, and flow performance. The I/O descriptors generate useful insights into complicated I/O behaviors, suggesting distinct concurrency, storage, and scheduling to be used by types, tasks, and files. The proposed trade-off analysis guides scheduling decisions that generate resource assignment with the best flow parallelism. We evaluate our I/O-aware scheduling methodology on a highly I/O-intensive workflow–1000 Genomes. The experimental results demonstrate speedups of up to 2.4× compared to the state-of-the- art methods.

Guo, Luanzheng [BATTELLE (PACIFIC NW LAB)]

Towards FAIR Workflows for Federated Experimental Sciences

A de-centralized, peer-to-peer AI metadata framework is demonstrated which can enable end-to-end metadata & lineage tracking for distributed Machine Learning pipelines spanning edge, High Performance Computing, and cloud environments. With a specific example of end-to-end microscopy algorithm and datasets, the proposed method shows how to enable reproducibility, audit trail, provenance of metadata artifacts. The emerging needs of automation in experimental sciences, ML-centric workflows, and FAIR metadata management across federated compute environments is addressed.

machine learning

ABLE Workflow Copier

A copier template for generating a snakemake workflow with an associated python package for implementing dataset transformation, feature extraction, and modeling.

Pathak, Maharshi [Northeastern Univ., Boston, MA (

OpenStudio® HPXML workflow [SWR-25-13]

OpenStudio-HPXML allows running residential EnergyPlus™ simulations using an HPXML file for the building description. It is intended to be used by user interfaces or other automated software workflows that automatically produce the HPXML file. OpenStudio-HPXML can accommodate a wide range of different building technologies and geometries. End-to-end simulations typically run in 3-10 seconds, depending on complexity, computer platform and speed, etc. For more information on running simulations, generating HPXML files with the appropriate inputs to run EnergyPlus, etc., please visit the documentation linked below. https://openstudio-hpxml.readthedocs.io/en/latest

Horowitz, Scott

torc (Torc Workflow Management System) [SWR-24-127]

This software package orchestrates execution of a workflow of jobs on distributed computing resources. It is optimized for use on HPCs with Slurm, but also can be used in the cloud and on local computers. Please refer to the documentation at https://nrel.github.io/torc

Thom, Daniel [National Renewable Energy Laboratory

Software-Defined Data Center Network Architecture using VXLAN-based BGP EVPN for Dynamic Workflows in a Supercomputing Environment (VXLAN-based BGP EVPN Fabric for HPC) v1

This software repository automates the deployment of a multi-vendor VXLAN-based BGP EVPN architecture, leveraging Containerlab to instantiate a stretched CLOS topology. It integrates Linux, Nokia SR Linux, and Arista cEOS, using BGP for underlay, overlay, and topology extension. The software enables rapid prototyping and testing of advanced network configurations. Its key advantage lies in providing a dynamic, programmable environment for research and development of critical technologies supporting dynamic workflows within supercomputing environments, surpassing the limitations of static, vendor-locked alternatives by fostering interoperability and agility.

Kumar, Ronal [Lawrence Berkeley National Laborator

A standardized workflow for kinetic metabolic model curation and dissemination

Kinetic metabolic models provide invaluable insights into cellular metabolism, supporting applications in synthetic biology, metabolic engineering, and systems biology. However, reproducibility and utility of these models hinge on clear and rigorous documentation, standardized annotation, and accessible visualization. This paper presents a workflow for building, annotating, visualizing, and sharing kinetic metabolic models. Our method integrates community standards and open-source tools to ensure reproducibility, interoperability, and user accessibility. This procedure enables researchers to produce reusable and well-documented kinetic models, advancing their role as powerful tools in metabolic research.

Cook, Margaret [Univ. of Washington, Seattle, WA (

FIRM image analysis: A machine learning workflow for quantifying extracellular matrix components from electron microscopy images

The extracellular matrix (ECM) is a complex network of biomolecules that plays an integral role in the structure, processes, and signaling mechanisms of cells and tissues. Identifying and quantifying changes in these matrix components provides insight into the mechanisms behind specific tissue remodeling processes; however, quantifying these changes is challenging due to difficult imaging conditions, complexity of the ECM, and the subtlety of these changes. Current imaging techniques allow us to visualize these critical remodeling events and developments in image analysis have employed a combination of analysis software and machine learning techniques to improve the efficiency and accuracy with which features are measured. Although image analysis has seen much improvement in recent years, there has been no technique developed to address ambiguity in feature edges in electron microscopy images. Presented here is a new machine learning-based workflow for the analysis of microscopy images named FIRM (Feature Identification from Raw Microscopy) that uses a random forest classifier to identify ECM features of interest and generate binary segmentation masks for quantification with ImageJ-FIJI. FIRM performed with an F1 score of 0.794 and greater than 80% accuracy for number and size of features detected. FIRM had similar deviation from the ground truth in the number of identified fibrils, fibril size, and size distributions when compared to human analyses. The results suggest that FIRM performs as well as manual analysis and requires a fraction of the time. This analysis technique is more efficient, eliminates user bias, and can be easily optimized to identify a variety of features, making it useful for any discipline requiring image analysis.

Science & Technology - Other Topics

A Workflow for Characterizing Legacy Wells as Potential Leakage Pathways for Integration to NRAP-Open-IAM

Carbon capture and storage is a crucial component of climate change mitigation strategies, involving the capture of carbon dioxide (CO2) from point sources and its injection into permeable subsurface formation. Many suitable CO2 storage sites coincide with legacy wells since the conditions that kept hydrocarbons in-situ for thousands of years are also ideal for storage of carbon dioxide. To protect underground sources of drinking water (USDW) during greenhouse gas injection, the Environmental Protection Agency (EPA) mandates area of review evaluations. These evaluations ensure that drinking water sources would not be contaminated by injected fluids. They include identification of legacy wellbores, integrity assessments, and implementing any necessary corrective action. Previous assessment approaches of legacy wells include high-level scoring of regional data and well construction and abandonment evaluation. This work describes a novel methodology that evaluates well construction and abandonment, ranks them based on complexity, and performs a risk assessment with NRAP-Open-IAM. A workflow of the methodology is presented, highlighting its capabilities and limitations.

Wise, Jarrett

Modularization of EDGE Workflows Using Nextflow: Improving the Efficiency and Maintainability of Bioinformatics Software

EDGE is a bioinformatics platform developed in 2016 by researchers at Los Alamos National Laboratory (LANL) to facilitate the analysis of next-generation sequencing data by researchers with varying levels of experience in bioinformatics (Li et al., 2017). Users with single-end, paired-end or long-read sequencing data can provide their reads as input to EDGE and select the combination of workflows to run that are most useful for their research (e.g., quality control of reads, genome assembly, or the taxonomic classification of input reads). Table 1 summarizes the modules available in EDGE. EDGE is available as a web platform at https://edgebioinformatics.org, as installable source code maintained on GitHub under a GPLv3 license, and as a publicly hosted Docker image.

59 BASIC BIOLOGICAL SCIENCES