Search NASA⌕ Search

SEARCH · Search NASA

Results for “workflow development”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

Enhancing occupant behavior representation for interoperability between building information modeling and building energy modeling

Building Performance Simulation (BPS) has been adopted as an essential tool for designing, operating, and retrofitting buildings to optimize energy efficiency throughout the building life cycle. The Green Building XML (gbXML) schema facilitates seamless data exchange between Building Information Modeling (BIM) and Building Energy Modeling (BEM) software tools. However, limited occupant behavior (OB) representation in BIM often leads to inconsistent and inaccurate energy simulation in BEM software. This paper presents 154 systematic enhancements to the existing occupant behavior XML (obXML) schema v1.3.4, initially developed for standardizing OB representation for BEM, to address existing limitations and improve interoperability with BIM models. The enhancements encompass improved integration with BIM models through extended building representations and system operations, expanded support for advanced OB models with additional environmental parameters and mathematical capabilities, and implementation of a standardized model documentation framework. To facilitate seamless data transformation between gbXML and obXML schemas, we developed a publicly available gb-obXML Schema Converter. Three case studies demonstrate the enhanced schema’s capabilities: representation of building information using a two-story office building model, documentation of a window operation behavior model, and validation of the schema converter’s functionality. The enhanced obXML schema v1.4 enables sophisticated modeling of occupant-building interactions while maintaining consistency with industry-standard BIM schemas. The standardized documentation framework facilitates reproducibility and knowledge sharing in the OB research community, while the schema converter automates the integration of building information into OB simulation workflows. These enhancements establish a foundation for more accurate building performance simulation by supporting sophisticated representation of occupant behavior within the BIM-to-BEM simulation workflows.

Chung, Jihoon↗

An Integral Activity-Based Protein Profiling Method for Higher Throughput Determination of Protein Target Sensitivity to Small Molecules

Activity-based protein profiling (ABPP) is a chemoproteomic technique that uses small molecule probes to label active enzymes selectively and covalently in complex proteomes. Competitive ABPP, which involves treatment of the active proteome with an analyte of interest, is especially powerful for profiling how small molecules impact specific protein activities. Advances in higher throughput workflows have made it possible to generate extensive competitive ABPP data across diverse biological samples, making this approach highly appealing for characterizing shared and unique proteins affected by perturbations such as drug or chemical exposures. To use the competitive ABPP approach effectively to understand potential adverse effects of chemicals of concern (CoC), a wide range of concentrations may be needed, particularly for chemicals that lack potency or toxicity data. In this work, we present an integral competitive ABPP method that enables target sensitivity determination for different organophosphate (OP) pesticides as model toxicants. Using previously developed OP-ABPs, we optimized conditions for tandem mass tag (TMT) multiplexing of ABPP samples and compared conventional competitive ABPP involving samples at discrete paraoxon concentrations to pooled samples across that same concentration range. We then expanded our approach to compare protein target sensitivities toward two additional OP pesticides, chlorpyrifos oxon and malaoxon. The results showed that differences in integral intensities for the pooled competition sample can be used to evaluate the relative sensitivity of specific proteins without increasing the overall number of samples. For 8 CoC concentrations of interest, this strategy reduced the number of TMT plexes and the corresponding number of LC–MS/MS analyses 3-fold. In conclusion, we envision the integral ABPP (IABPP) method will provide a means to screen diverse chemicals more rapidly to identify both high and low sensitivity protein targets.

activity-based probes↗

The high level trigger and express data production at STAR

To meet the demands of the Beam Energy Scan phase-II (BES-II) program, the STAR experiment at the Relativistic Heavy Ion Collider (RHIC) developed a dual real-time framework consisting of a High Level Trigger (HLT) and an Express Data Production system (xProduction). The HLT operates online within the Data Acquisition (DAQ) chain on a dedicated multi-core CPU cluster with the option to offload compute-intensive kernels to Xeon Phi coprocessors. It uses parallelized algorithms, such as the Cellular Automaton (CA) Track Finder, to perform rapid tracking, vertexing, and event filtering. This allows it to select events of interest in real time and provide immediate feedback on detector and beam conditions. In contrast, the xProduction workflow runs concurrently and independently of the DAQ loop. It applies near offline-quality calibration and reconstruction within hours of data collection. The xProduction input is the express data stream, whose content can be enriched by HLT trigger/priority selections under DAQ/HLT resource constraints, and it uses the STAR calibration/conditions framework, incorporating online calibration/QA information when available. This enables early preliminary physics analysis, including the reconstruction of rare signals, such as hyperons and hypernuclei. It also provides collaboration-wide access to analysis-ready datasets. Together, the HLT and xProduction systems form a complementary architecture: the HLT performs online event selection while the xProduction chain delivers high-quality results within a short amount of time. This integrated framework has enabled the prompt reconstruction of the $^5_Λ$ He hypernucleus with high statistical significance and the efficient processing of hundreds of millions of heavy-ion collision events. In conclusion, its demonstrated scalability and robustness establish a model for future high-luminosity experiments requiring both online event filtering and rapid access to analysis-quality data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A combinatorially complete epistatic fitness landscape in an enzyme active site

Protein engineering often targets binding pockets or active sites which are enriched in epistasis—nonadditive interactions between amino acid substitutions—and where the combined effects of multiple single substitutions are difficult to predict. Few existing sequence-fitness datasets capture epistasis at large scale, especially for enzyme catalysis, limiting the development and assessment of model-guided enzyme engineering approaches. We present here a combinatorially complete, 160,000-variant fitness landscape across four residues in the active site of an enzyme. Assaying the native reaction of a thermostable β-subunit of tryptophan synthase (TrpB) in a nonnative environment yielded a landscape characterized by significant epistasis and many local optima. These effects prevent simulated directed evolution approaches from efficiently reaching the global optimum. There is nonetheless wide variability in the effectiveness of different directed evolution approaches, which together provide experimental benchmarks for computational and machine learning workflows. The most-fit TrpB variants contain a substitution that is nearly absent in natural TrpB sequences—a result that conservation-based predictions would not capture. Thus, although fitness prediction using evolutionary data can enrich in more-active variants, these approaches struggle to identify and differentiate among the most-active variants, even for this near-native function. Overall, this work presents a large-scale testing ground for model-guided enzyme engineering and suggests that efficient navigation of epistatic fitness landscapes can be improved by advances in both machine learning and physical modeling.

biocatalysis↗

Visualization techniques for the gyrokinetic tokamak simulation code

Gyrokinetic simulations of plasma microturbulence in tokamaks are challenging to visualize because the compute grid follows the magnetic field lines that spiral around the torus. We have overcome this challenge by developing three new approaches that improve visualization of gyrokinetics. Our techniques work directly with the topology of magnetic flux surfaces where the simulation stores variables in concentric rings on poloidal planes (vertical cross sections of the torus). Our visualization preview step triangulates each consecutive pair of rings to display the data on a poloidal plane. The second visualization technique follows spiral field lines around the torus and constructs polygons to visualize a flux surface. Third, the poloidal triangles are connected between planes to form prisms that compose a 3-D model of the entire torus. The visualization workflow produces detailed geometry that matches the high resolution, irregular compute grid for every time step. The surface and solid models are displayed in scientific visualization programs to effectively explore and communicate the results, including fluctuation of electron density, ion temperature, and electrostatic potential. Highly detailed renderings verify plasma behavior along magnetic field lines over time.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

WorkflowHub: a registry for computational workflows

The rising popularity of computational workflows is driven by the need for repetitive and scalable data processing, sharing of processing know-how, and transparent methods. As both combined records of analysis and descriptions of processing steps, workflows should be reproducible, reusable, adaptable, and available. Workflow sharing presents opportunities to reduce unnecessary reinvention, promote reuse, increase access to best practice analyses for non-experts, and increase productivity. In reality, workflows are scattered and difficult to find, in part due to the diversity of available workflow engines and ecosystems, and because workflow sharing is not yet part of research practice. WorkflowHub provides a unified registry for all computational workflows that links to community repositories, and supports both the workflow lifecycle and making workflows findable, accessible, interoperable, and reusable (FAIR). By interoperating with diverse platforms, services, and external registries, WorkflowHub adds value by supporting workflow sharing, explicitly assigning credit, enhancing FAIRness, and promoting workflows as scholarly artefacts. The registry has a global reach, with hundreds of research organisations involved, and more than 800 workflows registered.

97 MATHEMATICS AND COMPUTING↗

ANS Winter 2024 Summary: MCCAFE: The Monte Carlo Constructor for ATR Fuel Elements

The Irradiation Experiment Neutronics Analysis Department at Idaho National Laboratory (INL) has implemented a new analysis workflow for experiments in the Advanced Test Reactor (ATR). One key piece of this workflow is the Monte Carlo Constructor for ATR Fuel Elements, or MCCAFE. For each ATR operating cycle, the Reactor and Nuclear Safety Engineering (RNSE) Department first solves the core in eigenvalue mode and depletes the driver fuel materials. In a separate calculation, neutronics analysts model and deplete the materials of one or more irradiation experiments, usually in a series of fixed-source Monte Carlo N-Particle (MCNP) models of the ATR for neutron transport calculations. It was desirable to use the results of the former calculations to inform the models of the latter. MCCAFE is a Python program developed using American Society of Mechanical Engineers Nuclear Quality Assurance-1 procedures at INL. Its purpose is to take the calculated results from the RNSE depletion solutions and the measured or projected operating parameters from the Nuclear Data Management and Analysis System (NDMAS) to generate fixed-source models of the ATR core at given points in time across one or more cycles.

99 - GENERAL AND MISCELLANEOUS↗

Synthetic Biology PacBio/JAWS QC Analysis (PBJ) v3.0

This software was designed as a sequence validation tool for the assembly of synthetic constructs. It analyzes FASTQ files against a list of reference sequences, combining the results from eight sequencing libraries to generate a summary, and the files needed to view the results in the Integrative Genomics Viewer (IGV) application for manual verification. This was developed for FASTQ files generated by PacBio sequencing, but could be used on any FASTQ files that do not have paired end reads. It can be used to analyze one - eight libraries at a time, and assumes that each construct sequence in the reference will be in each pool, however, this is not a requirement. This is used to identify which libraries of pooled sequences contains a perfect match, or fixable match to the reference file. This pipeline uses many freely available open source libraries, the value added is that in our application the steps of the pipeline are defined in Workflow Description Language (WDL) and run through the Cromwell workflow engine in Docker containers, for easy distribution and set up, as well as the user friendly html summary that is generated.

Simirenko, Lisa↗

Data-flow parallelism for high-energy and nuclear physics frameworks

The processing tasks of an event-processing workflow in high-energy and nuclear physics (HENP) can typically be represented as a directed acyclic graph formed according to the data flow—i.e. the data dependencies among algorithms executed as part of the workflow. With this representation, an HENP framework can optimally execute a workflow, exploiting the parallelism inherent among independent tasks. Despite such a natural description of a workflow, most HENP frameworks do not make use of technologies that provide concurrent execution of graph-based tasking structures. In this talk, we describe Fermilab efforts to adopt a graph-based technology (specifically Intel’s oneTBB flow graph) for meeting the framework needs of its experiments, notably DUNE. Building on the Meld project as presented at CHEP2023, we demonstrate that all common processing idioms supported by current frameworks can naturally be supported by oneTBB’s data-flow technology, optimally leveraging the concurrent capabilities of the machine. In addition, we discuss collaborative efforts between Fermilab and the Intel oneTBB development team, who is considering improvements to the flow-graph technology to better support HENP use cases.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Review and Outlook on Experimental Advances and Innovations in Geological CO 2 Storage: Insights from Depleted Gas Reservoirs and Saline Aquifers

Geological storage of carbon dioxide (CO 2 ) in depleted gas reservoirs and deep saline aquifers is a key part of global decarbonization efforts. As carbon capture and storage advances toward commercial-scale deployment, the credibility and scalability of laboratory experiments are increasingly vital for guiding safe and effective field implementation. This review offers a comprehensive, cross-scale evaluation of experimental methodologies, including core flooding, high-pressure, high-temperature systems, microfluidic visualization, and emerging systems such as multilayer commingled/compartmentalized core flooding, 3D-printed micromodels, and AI-powered digital twins. These innovations are demonstrated to enhance representativeness, reproducibility, and real-time insight, thereby addressing the limitations of conventional workflows. A critical analysis of methodological gaps, such as inconsistent pressure–temperature conditions, oversimplified brine chemistry, and a lack of standardization, reveals experimental sources of scale translation errors and performance uncertainty. By comparing the unique challenges of depleted gas reservoirs (such as low water saturation and legacy well leakage) to those of saline aquifers (including pressure buildup and caprock integrity), this review identifies formation-specific priorities for experimental design. Novel contributions include a synthesis of best practices, integration strategies for model calibration, and recommendations for standardizing core handling, saturation procedures, and reporting protocols. Furthermore, this work serves as a guide for developing robust, field-relevant experimental strategies that can increase the deployment and regulatory acceptance of CO 2 storage technologies at scale.

58 GEOSCIENCES↗

Identifying Climate Patterns Using Clustering Autoencoder Techniques

Abstract The complexity of growing spatiotemporal resolution of climate simulations produces a variety of climate patterns under different projection scenarios. This paper proposes a new data-driven climate classification workflow via an unsupervised deep learning technique that can dimensionally reduce the vast volume of spatiotemporal numerical climate projection data into a compact representation. We aim to identify distinct zones that capture multiple climate variables as well as their future changes under different climate change scenarios. Our approach leverages convolutional autoencoders combined with k -means clustering (standard autoencoder) and online clustering based on the Sinkhorn–Knopp algorithm (clustering autoencoder) across the conterminous United States (CONUS) to capture unique climate patterns in a data-driven fashion from the Geophysical Fluid Dynamics Laboratory Earth System Model with GOLD component (GFDL-ESM2G). The developed approach compresses 70 years of GFDL-ESM2G simulation at 0.125° spatial resolution across the CONUS under multiple warming scenarios to a lower-dimensional space by a factor of 660 000 and then tested on 150 years of GFDL-ESM2G simulation data. The results show that five climate clusters capture physically reasonable and spatially stable climatological patterns matched to known climate classes defined by human experts. Results also show that using a clustering autoencoder can reduce the computational time for clustering by up to 9.2 times when compared to using a standard autoencoder. Our five unique climate patterns resulting from the deep learning–based clustering of the lower-dimensional space thereby enable us to provide insights on hydrometeorology and its spatial heterogeneity across the conterminous United States immediately without downloading large climate datasets. Significance Statement This paper presents a data-driven climate classification approach using unsupervised deep learning to dimensionally reduce climate model outputs and to identify distinct climate regions for their future changes. Our approach compresses climate information for 70 years of Geophysical Fluid Dynamics Laboratory Earth System Model data across the conterminous United States (CONUS) at 0.125° spatial resolution. The results reveal that five climate clusters capture reasonable and stable climatological patterns matched to known climate patterns. The embedded clustering process in deep learning provides ×9.2 times faster execution than the k -means clustering technique. These results give us insight about climate spatial patterns and heterogeneity of hydrological patterns across the conterminous United States without downloading large climate datasets.

Kurihana, Takuya↗

Data-flow parallelism for high-energy and nuclear physics computing frameworks

The processing tasks of a scientific workflow in high-energy and nuclear physics (HENP) can typically be represented as a directed acyclic graph formed according to the data flow—i.e. the data dependencies among algorithms executed as part of the workflow. With this representation, an HENP computing framework can optimally execute a workflow, exploiting the parallelism inherent among independent tasks. Despite such a natural description of a workflow, most HENP frameworks do not make use of technologies that provide concurrent execution of graph-based tasking structures. In this session, we describe Fermilab efforts to adopt a graph-based technology (specifically Intel’s oneTBB flow graph) for meeting the framework needs of its experiments, notably DUNE. After introducing the physics DUNE intends to explore, we will show that all common processing idioms supported by current HENP frameworks can naturally be supported by oneTBB’s data-flow technology, optimally leveraging the concurrent capabilities of the machine. In addition, we discuss collaborative efforts between Fermilab and the Intel oneTBB development team, who is considering improvements to the flow-graph technology to better support HENP use cases.

43 PARTICLE ACCELERATORS↗

Computationally evaluating high-yield metabolites for sustainable aviation fuel (SAF) using machine learning

The computational tool described in this report helps identify promising biological pathways that produce SAF platform molecules (either a drop-in SAF, or a precursor that can be easily converted to a drop-in SAF). The workflow the computational tool follows first identifies possible biological pathways from a user-defined metabolite. These pathways may, or may not lead to a SAF platform molecule, thus the second step involves insilico testing of the end product of each pathway to assess whether it is, or is not, a SAF platform molecule. The identification of biological pathways performed in the first step is facilitated by linking the metabolite to a biological reaction database. Pathways are found by identifying pathways in the reaction database that include the metabolite. The computational tool includes an alternative way to find pathways. The alternative way develops a Flux Balanced Analysis (FBA), and modifying the FBA to include reactions that transform the metabolite. These modifications serve as a basis for understanding, in a semi-quantitative way, if there is an increase in the flux to desirable products. The second step, in silico testing of the end-products, is accomplished by estimating key physical properties relevant to SAF. When good models are available, we have integrated those models into the computational tool. In a few instances, we have developed our own models. In all instances, we have validated the models against available measured data. Finally, we have evaluated the effectiveness of our computational tool by genetically engineering Rhodosporidium toruloides. Validation occurred without the use of a FBA, and further validation is required.

09 BIOMASS FUELS↗

Reimagining metal-organic framework discovery: Integrating experiment, computation, and artificial intelligence

The traditional development of novel metal–organic frameworks (MOFs) is often hindered by challenges such as synthetic accessibility and time- and resource-intensive experimentation. High-throughput, automated experimental and computational techniques have enabled rapid chemical space exploration and theoretical MOF design. When combined with artificial intelligence (AI), these methods can be used to lead autonomous laboratories to new frontiers for MOF discovery, where these materials can be designed for a specific application, efficiently synthesized, characterized, and evaluated. Here, this perspective highlights the role of AI in advancing automated MOF synthesis and characterization, computational MOF design and screening, and the integration of these approaches within autonomous workflows to ultimately enable the MOF laboratories of the future.

Gaidimas, Madeleine A. [Northwestern University, E↗

Leveraging data mining, active learning, and domain adaptation for efficient discovery of advanced oxygen evolution electrocatalysts

Developing advanced catalysts for acidic oxygen evolution reaction (OER) is crucial for sustainable hydrogen production. This study presents a multistage machine learning (ML) approach to streamline the discovery and optimization of complex multimetallic catalysts. Our method integrates data mining, active learning, and domain adaptation throughout the materials discovery process. Unlike traditional trial-and-error methods, this approach systematically narrows the exploration space using domain knowledge with minimized reliance on subjective intuition. Then, the active learning module efficiently refines element composition and synthesis conditions through iterative experimental feedback. The process culminated in the discovery of a promising Ru-Mn-Ca-Pr oxide catalyst. Our workflow also enhances theoretical simulations with domain adaptation strategy, providing deeper mechanistic insights aligned with experimental findings. By leveraging diverse data sources and multiple ML strategies, we demonstrate an efficient pathway for electrocatalyst discovery and optimization. This comprehensive, data-driven approach represents a paradigm shift and potentially benchmark in electrocatalysts research.

Science & Technology - Other Topics↗

Building MCP-native hierarchical AI scientist ecosystems: a perspective on scaling multi-agent scientific discovery

Large language models (LLMs) are evolving from chatbots with limited tool-using capabilities to agentic AI systems that can perform deep research, assist in proposing hypotheses, help design experiments, automate data analysis, and draft scientific reports. However, there are currently two bottlenecks limiting LLMs' real-world impact on the broader scientific research community beyond academic demonstrations: lack of interoperability (repetitive manual tool-integration is required across scenarios) and the need for scalable coordination (unstructured communication and memory become brittle as the number of agents grows). In this Perspective, we argue that the next phase of agentic scientific discovery requires the development of an ecosystem of protocol-native agents and tools organized through hierarchies inspired by human society, beyond the current paradigm of a single monolithic “AI scientist”. We use Model Context Protocol (MCP) as a concrete example of an emerging interoperability layer for scientific tool and context exchange, and we propose three complementary pathways to increase the scaling capabilities of an MCP-native scientific ecosystem by addressing the composability issues: (1) MCP servers for high-value scientific tools maintained by domain experts, (2) automated transformation of existing code repositories into MCP services, and (3) autonomous invention and evolution of new agents and workflows. Finally, we provide a practical roadmap for scaling AI-driven scientific discovery by expanding tool supply and coordination in MCP-native scientific ecosystems.

97 MATHEMATICS AND COMPUTING↗

Multibody for Everybody (M4E) - A Linearization Approach to Enable Frequency Domain Analysis, Time Integration and Control Co-Design

1.1 Background/Objectives: Marine energy represents a promising yet underexploited source of power. To increase the harvested power, significant efforts have been made to improve wave energy converter (WEC) modeling capabilities and optimize power take-off (PTO) performance; however, these efforts have often treated WEC dynamics, PTO design, and controller development sequentially. In contrast, control co-design (CCD) is emerging as a promising strategy to address these issues directly, creating a growing need for fast analysis tools suitable for repeated simulation and parametric studies [1]. To support this need, this work presents the Multibody for Everybody (M4E) [2] linearization module, which employs a symbolic toolbox to provide deeper insight of WEC design parameters. The objective is to demonstrate that a minimal-coordinate linearization of articulated WEC dynamics can provide accurate wave response predictions and substantial computational savings relative to nonlinear time-domain simulation, while preserving compatibility with broader wave-energy analysis workflows, enabling CCD. 1.2 Approach/Activities: The proposed approach linearizes the equations of motion, generated by M4E, in minimal coordinates about a selected operating point and combines the resulting system with frequencydomain hydrodynamic terms to incorporate the reduced mass, damping, stiffness, and forcing operators. The linearized model is used for both impedance-based response amplitude operator (RAO) prediction and rapid regular-wave time integration. The methodology is demonstrated on a single-flap device and a FOSWEC configuration, with linearized M4E responses compared against the corresponding nonlinear M4E simulations and WEC-Sim results. Regular-wave time histories, RAO trends, and runtime differences are assessed. The framework is also compatible with broader wave-energy workflows, including coupling to WecOptTool, although that capability is not the focus of this work [3]. 1.3 Results/Lessons: The linearized M4E model reproduces key regularwave response characteristics such as integration and Response Amplitude over multiple frequencies. This module matches nonlinear M4E and WEC-Sim results while substantially reducing integration cost. Thus, the proposed framework can serve as a rapid analysis layer for articulated WEC design, parameter studies, and controls-oriented workflows. The analysis is most appropriate in the near-equilibrium regime, about the linearization point.

16 TIDAL AND WAVE POWER↗

17 O NMR Spectroscopy Reveals CO 2 Speciation and Dynamics in Hydroxide-Based Carbon Capture Materials

Carbon dioxide capture technologies are set to play a vital role in mitigating the current climate crisis. Solid-state 17 O NMR spectroscopy can provide key mechanistic insights that are crucial to effective sorbent development. In this work, we present the fundamental aspects and complexities for the study of hydroxide-based CO 2 capture systems by 17 O NMR. We perform static density functional theory (DFT) NMR calculations to assign peaks for general hydroxide CO 2 capture products, finding that 17 O NMR can readily distinguish bicarbonate, carbonate and water species. However, in application to CO 2 binding in two test case hydroxide-functionalised metal-organic frameworks (MOFs) – MFU-4l and KHCO 3 -cyclodextrin-MOF, we find that a dynamic treatment is necessary to obtain agreement between computational and experimental spectra. We therefore introduce a workflow that leverages machine-learning force fields to capture dynamics across multiple chemical exchange regimes, providing a significant improvement on static DFT predictions. In MFU-4l, we parameterise a two-component dynamic motion of the bicarbonate motif involving a rapid carbonyl seesaw motion and intermediate hydroxyl proton hopping. For KHCO 3 -CD-MOF, we combined experimental and modelling approaches to propose a new mixed carbonate-bicarbonate binding mechanism and thus, we open new avenues for the study and modelling of hydroxide-based CO 2 capture materials by 17 O NMR.

NMR spectroscopy↗