Search NASA⌕ Search

SEARCH · Search NASA

Results for “workflow development”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Correlation function metrology for warm dense matter: Recent developments and practical guidelines

X-ray Thomson scattering (XRTS) has emerged as a valuable diagnostic for matter under extreme conditions, as it captures the intricate many-body physics of the probed sample. Recent advances, such as the model-free temperature diagnostic of Dornheim et al. [Nat. Commun. 13 , 7911 (2022)], have demonstrated how much information can be extracted directly within the imaginary-time formalism. However, since the imaginary-time formalism is a concept often difficult to grasp, we provide here a systematic overview of its theoretical foundations and explicitly demonstrate its practical applications to temperature inference, including relevant subtleties. Furthermore, we present recent developments that enable the determination of the absolute normalization, Rayleigh weight, and density from XRTS measurements without reliance on uncontrolled model assumptions. Finally, we outline a unified workflow that guides the extraction of these key observables, offering a practical framework for applying the method to interpret experimental measurements.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Doppler Backscattering Data Analysis and Integrated Modeling with OMFIT

One Modeling Framework for Integrated Tasks (OMFIT) is a widely used software tool in the magnetic fusion research community. OMFIT provides magnetic fusion energy researchers with a framework for the development of special-purpose physics modules. This paper describes an OMFIT physics module pertaining to the Doppler Backscattering (DBS) fusion plasma diagnostic. DBS measures density fluctuations and flow velocity through plasma scattering of electromagnetic waves. The OMFIT DBS module was developed to analyze experimental DBS data and facilitate modeling of DBS systems installed on multiple tokamak devices. The OMFIT DBS module is designed to support several analysis workflows: detailed analysis of experimental data, experimental planning, and theory-based synthetic diagnostic modeling. The DBS module uses integrated modeling by leveraging other OMFIT physics modules to perform tasks related to DBS, e.g. ray/beam–tracing simulations, edge-localized mode–synchronized data analysis, magnetic equilibrium reconstruction, and fitting kinetic profile data. Furthermore, this paper describes several supported workflows and serves a reference for the OMFIT DBS module.

Doppler backscattering↗

AutoCSM

A template system-of-systems modeling approach for automating the development, deployment, and integration of cooling system models (CSMs) for supercomputing facilities within the ExaDigiT framework. AutoCSM is a Python-based framework to assist in CSM developers in accelerating the creation and deployment of system-level thermal-hydraulic CSMs. The intention is for this tool specifically to help standardize digital twin workflows for ExaDigiT. However, this tool can be used independent of ExaDigiT (and even other systems besides CSMs).

Greenwood, MichaelScott [Oak Ridge National Labora↗

The need for standardization and improved open (meta)data practices in metaproteomics

Metaproteomics enables functional insight into microbial communities by identifying and quantifying proteins in complex samples. Yet, heterogeneous analytical workflows and the lack of standardization across experimental and bioinformatics stages hinder reproducibility and comparability, limiting integration with other omics data. We here present a community-developed reporting checklist tailored to the specific needs of metaproteomics. We also outline current efforts to enable structured and interoperable metadata capture, drawing on standards from proteomics and microbiome research wherever possible. By promoting transparent reporting and advancing metadata practices, our recommendations aim to align metaproteomics more closely with FAIR principles and support reproducible and interoperable research practices.

Armengaud, Jean [Universite Paris-Saclay, France]↗

gaia: An R package to estimate crop yield responses to temperature and precipitation

gaia is an open-source R package designed to estimate crop yield shocks in response to annual weather variations and CO 2 concentrations at the country scale for 17 major crops. This innovative tool streamlines the workflow from raw climate data processing to projections of annual shocks to crop yields at the country level, using the response surfaces from an empirical econometric model developed and documented in Waldhoff et al. (2020), which leverages historical weather, CO 2 , and crop yield data for robust empirical fitting for 17 crops. gaia uses these response surfaces with monthly temperature and precipitation projections (e.g., from the Coupled Model Intercomparison Project Phase 6 (CMIP6) (O’Neill et al., 2016) climate data bias-adjusted and statistically downscaled by the ISIMIP3BASD approach (Lange, 2019) in the Inter-Sectoral Impact Model Intercomparison Project (ISIMIP) (Warszawski et al., 2014)) to project yield shocks that can be applied to agricultural productivity changes at the country level for use in multisectoral economic models. The historical and future projections use gridded, country-and-crop specific monthly growing season precipitation and temperature data, aggregated to the national level, and weighted by cropland area derived from the global Monthly Irrigated and Rainfed Crop Areas around the year 2000 (MIRCA2000) dataset (Portmann et al., 2010). These annual, country, and crop-specific yield shocks can be aggregated to different definitions of regions, crop commodities, and time periods, as needed by specific multisectoral economic models. gaia serves as a lightweight, powerful tool that can aid exploration of crop yield responses under a broad range of future climate projections, enhancing human-Earth system analysis capabilities.

60 APPLIED LIFE SCIENCES↗

Scalable workflow for evaluating and optimizing large language models

This work describes the improved workflow for evaluating open-source large language models (LLMs) for trustworthiness. The workflow facilitates the acquisition of LLMs, the generation of LLM responses, and the evaluation of the responses for their trustworthiness. As a use case, the workflow is employed to evaluate dense, quantized, and pruned Meta Llama3.1 LLMs for their truthfulness. The outcome of the project could set the stage for understanding and developing trustworthy models in the future projects.

97 MATHEMATICS AND COMPUTING↗

Improving I/O-aware Workflow Scheduling via Data Flow Characterization and trade-off Analysis

The scientific computing paradigm has transitioned from compute-intensive to I/O-intensive and memory-intensive in the past decade, especially when data-driven science has become common practice. Numerous empirical I/O-aware scheduling optimizations have been developed by incorporating I/O capacity and bandwidth as constraints into scheduling. Unfortunately, there is a lack of data flow (I/O) characterization tool and an understanding of trade-offs between concurrency, locality, and I/O bandwidth. To bridge the gap, this work 1) presents a set of descriptors to characterize, organize, and visualize I/O profiles, including flow size, I/O bandwidth, and operation count, which group data flows by I/O types, tasks, and files; 2) proposes an I/O Roofline model-based trade-off analysis to find the optimal trade-off between flow operational intensity, concurrency, and flow performance. The I/O descriptors generate useful insights into complicated I/O behaviors, suggesting distinct concurrency, storage, and scheduling to be used by types, tasks, and files. The proposed trade-off analysis guides scheduling decisions that generate resource assignment with the best flow parallelism. We evaluate our I/O-aware scheduling methodology on a highly I/O-intensive workflow–1000 Genomes. The experimental results demonstrate speedups of up to 2.4× compared to the state-of-the- art methods.

Guo, Luanzheng [BATTELLE (PACIFIC NW LAB)]↗

The microbiologist's guide to metaproteomics

Metaproteomics is an emerging approach for studying microbiomes, offering the ability to characterize proteins that underpin microbial functionality within diverse ecosystems. As the primary catalytic and structural components of microbiomes, proteins provide unique insights into the active processes and ecological roles of microbial communities. By integrating metaproteomics with other omics disciplines, researchers can gain a comprehensive understanding of microbial ecology, interactions, and functional dynamics. This review, developed by the Metaproteomics Initiative (www.metaproteomics.org), serves as a practical guide for both microbiome and proteomics researchers, presenting key principles, state-of-the-art methodologies, and analytical workflows essential to metaproteomics. Topics covered include experimental design, sample preparation, mass spectrometry techniques, data analysis strategies, and statistical approaches.

bioinformatics↗

Omics-driven onboarding of the carotenoid producing red yeast Xanthophyllomyces dendrorhous CBS 6938

Transcriptomics is a powerful approach for functional genomics and systems biology, yet it can also be used for genetic part discovery. Here, we derive constitutive and light-regulated promoters directly from transcriptomics data of the basidiomycete red yeast Xanthophyllomyces dendrorhous CBS 6938 (anamorph Phaffia rhodozyma) and use these promoters with other genetic elements to create a modular synthetic biology parts collection for this organism. X. dendrorhous is currently the sole biotechnologically relevant yeast in the Tremellomycete class-it produces large amounts of astaxanthin, especially under oxidative stress and exposure to light. Thus, we performed transcriptomics on X. dendrorhous under different wavelengths of light (red, green, blue, and ultraviolet) and oxidative stress. Differential gene expression analysis (DGE) revealed that terpenoid biosynthesis was primarily upregulated by light through crtI, while oxidative stress upregulated several genes in the pathway. Further gene ontology (GO) analysis revealed a complex survival response to ultraviolet (UV) where X. dendrorhous upregulates aromatic amino acid and tetraterpenoid biosynthesis and downregulates central carbon metabolism and respiration. The DGE data was also used to identify 26 constitutive and regulated genes, and then, putative promoters for each of the 26 genes were derived from the genome. Simultaneously, a modular cloning system for X. dendrorhous was developed, including integration sites, terminators, selection markers, and reporters. Each of the 26 putative promoters were integrated into the genome and characterized by luciferase assay in the dark and under UV light. The putative constitutive promoters were constitutive in the synthetic genetic context, but so were many of the putative regulated promoters. Notably, one putative promoter, derived from a hypothetical gene, showed ninefold activation upon UV exposure. Thus, this study reveals metabolic pathway regulation and develops a genetic parts collection for X. dendrorhous from transcriptomic data. Therefore, this study demonstrates that combining systems biology and synthetic biology into an omics-to-parts workflow can simultaneously provide useful biological insight and genetic tools for nonconventional microbes, particularly those without a related model organism. This approach can enhance current efforts to engineer diverse microbes.

60 APPLIED LIFE SCIENCES↗

MVP: a modular viromics pipeline to identify, filter, cluster, annotate, and bin viruses from metagenomes

While numerous computational frameworks and workflows are available for recovering prokaryote and eukaryote genomes from metagenome data, only a limited number of pipelines are designed specifically for viromics analysis. With many viromics tools developed in the last few years alone, it can be challenging for scientists with limited bioinformatics experience to easily recover, evaluate quality, annotate genes, dereplicate, assign taxonomy, and calculate relative abundance and coverage of viral genomes using state-of-the-art methods and standards. Here, we describe Modular Viromics Pipeline (MVP) v.1.0, a user-friendly pipeline written in Python and providing a simple framework to perform standard viromics analyses. MVP combines multiple tools to enable viral genome identification, characterization of genome quality, filtering, clustering, taxonomic and functional annotation, genome binning, and comprehensive summaries of results that can be used for downstream ecological analyses. Overall, MVP provides a standardized and reproducible pipeline for both extensive and robust characterization of viruses from large-scale sequencing data including metagenomes, metatranscriptomes, viromes, and isolate genomes. As a typical use case, we show how the entire MVP pipeline can be applied to a set of 20 metagenomes from wetland sediments using only 10 modules executed via command lines, leading to the identification of 11,656 viral contigs and 8,145 viral operational taxonomic units (vOTUs) displaying a clear beta-diversity pattern. Further, acting as a dynamic wrapper, MVP is designed to continuously incorporate updates and integrate new tools, ensuring its ongoing relevance in the rapidly evolving field of viromics. MVP is available at https://gitlab.com/ccoclet/mvp and as versioned packages in PyPi and Conda.

59 BASIC BIOLOGICAL SCIENCES↗

Accelerating lattice gauge theory studies with Agentic AI

Lattice gauge theory research, with its computationally intensive simulations and complex multi‑stage workflows, is well positioned to benefit from agentic AI systems. We demonstrate how such tools can support key components of lattice gauge theory research, including novel simulation code development using standard LQCD frameworks, HPC job orchestration, simulation data analysis, and expert‑guided tuning of algorithmic parameters such as Hasenbusch mass preconditioning and multigrid solvers. Our results show that agentic AI can reduce manual effort, improve productivity, and accelerate the research cycle while maintaining essential human oversight.

Ayyar, Venkitesh [Fermilab]↗

Sustainable Enablers of Knowledge Management Strategies in a Higher Education Institution

By facilitating the capture, organization, and dissemination of knowledge within and beyond the institution, knowledge management (KM) in higher education institutions (HEIs) fuels innovation, enhances research impact, and strengthens collaboration, ultimately leading to the creation of new knowledge and its valuable exchange. However, there is still much to explore in terms of the enablers of knowledge creation, sharing, and transfer. Therefore, this paper aims to identify the enablers of effective KM in the Polytechnique University of Leiria, which serves as a benchmark for other higher education institutions due to its leadership role in RUN-EU, a consortium of European universities. To achieve this, a narrative analysis based on information from SCOPUS and the institute’s website, focusing on innovation, research, and development strategies, is proposed. The findings suggest that for KM initiatives to be successful, they need to be strategically designed, culturally supported, technologically enabled, and integrated into existing workflows.

Santos, Eleonora (ORCID:0000000346930804)↗

Automated Strain Construction for Biosynthetic Pathway Screening in Yeast

Automation accelerates the Design-Build-Test-Learn (DBTL) cycle for synthetic biology; however, most strain construction pipelines lack robotic integration. Here, in this study, we present the workflow design and source code for a modular, integrated protocol that automates the Build step in Saccharomyces cerevisiae. We programmed the Hamilton Microlab VANTAGE to integrate off-deck hardware via its central robotic arm, enabling automated steps that increased throughput to 2,000 transformations per week. We developed a user interface with the Hamilton VENUS software to support on-demand parameter customization. As a proof of concept, we screened a gene library in an engineered yeast strain producing verazine, a key intermediate in the biosynthesis of steroidal alkaloids. Our pipeline rapidly identified pathway bottlenecks and genes that enhanced verazine production by 2.0- to 5-fold. This technical note provides resources for synthetic biologists designing yeast workflows for biofoundries to screen libraries for pathway discovery/optimization, combinatorial biosynthesis, and protein engineering.

automation↗

HDF5 in the exascale era: Delivering efficient and scalable parallel I/O for exascale applications

Accurately modeling real-world systems requires scientific applications at exascale to generate massive amounts of data and manage data storage efficiently. However, parallel input and output (I/O) faces challenges due to new application workflows and the state-of-the-art memory, interconnect, and storage architectures considered in exascale designs. The storage hierarchy has expanded with node-local persistent memory, solid-state storage, and traditional disk and tape-based storage, thus requiring efficiency at each layer and much more efficient data movement among these layers. This paper discusses how the ExaHDF5 project improved the I/O performance and data management for exascale architectures by enhancing HDF5, a widely used parallel I/O library. The team developed an Asynchronous I/O Virtual Object Layer (VOL) connector that allowed overlapping I/O with computation. They also created a Cache VOL to complement asynchronous I/O by incorporating fast storage layers, such as burst buffer and node-local storage, into the parallel I/O workflow through caching and staging data. Additionally, the team enabled data aggregation and I/O at the node level by using a Subfiling Virtual File Driver (VFD). To demonstrate superior I/O performance with HDF5 at exascale, the ExaHDF5 team collaborated with several exascale applications. In this paper, we show I/O performance improvements for three applications: Cabana (a particle-based simulation library), EQSIM (a regional earthquake simulation software), and E3SM (a climate system modeling library).

Asynchronous I/Ol↗

Polarized Deep-Inelastic Scattering with Spin Correlations in Herwig 7

This repository is the research software and reproducibility companion for the HerwigPol polarized deep-inelastic scattering implementation developed for Herwig 7. It brings together the modified Herwig and ThePEG source snapshots, the curated POLDIS fixed-order reference code, the custom Rivet analyses, the DIS validation workflow, and the paper source in a single formal repository layout. The repository is intended to preserve the source-level ingredients needed to rebuild and re-run the validated DIS studies. It therefore tracks code, input cards, workflow drivers, and technical notes, while intentionally excluding generated artifacts such as build products, campaign outputs, merged YODA files, plots, and rendered paper outputs.

Papaefstathioou, Andreas [Kennesaw State Universit↗

Deployment of inference as a service at the US CMS Tier-2 data centers

Coprocessors, especially GPUs, will be a vital ingredient of data production workflows at the HL-LHC. At CMS, the GPU-as-a-service approach for production workflows is implemented by the SONIC project (Services for Optimized Network Inference on Coprocessors). SONIC provides a mechanism for outsourcing computationally demanding algorithms, such as neural network inference, to remote servers, where requests from multiple clients are intelligently distributed across multiple GPUs by a load-balancing service. This talk highlights the recent progress in deploying SONIC at selected U.S. CMS Tier-2 data centers. Using realistic CMS Run3 data processing workflows, such as those containing transformer-based algorithms, we demonstrate how SONIC is integrated into the production-like environment to enable accelerated inference offloading. We will present developments from both the client and server sides, including production job and data center configurations for NVIDIA and AMD GPUs. We will also present performance scaling benchmarks and discuss the challenges of operating SONIC in CMS production, such as server discovery, GPU saturation, fallback server logic, etc.

Holzman, Burt↗

Automated Immunoprecipitation Workflow for Comprehensive Acetylome Analysis

Immunoprecipitation is one of the most effective methods for enrichment of lysine-acetylated peptides for comprehensive acetylome analysis using mass spectrometry. Manual acetyl peptide enrichment method using non-conjugated antibodies and agarose beads has been developed and applied in various studies. However, it is time consuming, and can introduce contaminants and variability that leads to potential sample loss and decreased sensitivity and robustness of the analysis. Here we describe a fast, automated enrichment protocol that enables reproducible and comprehensive acetylome analysis using a magnetic bead-based immunoprecipitation reagent.

Lysine acetylation, Acetylome, Acetyl peptide enri↗

Preliminary Results on Process Modeling Tools for Determining Variability in Additively Manufactured Stainless Steel 316 Parts

The Advanced Materials and Manufacturing Technologies program aims to accelerate the development, qualification, demonstration, and deployment of advanced materials and manufacturing technologies to enable reliable and economical nuclear energy. However, the distinct characteristics of additive manufacturing (AM) materials, stemming from their unique processing history, microstructure, and properties, pose significant challenges for the qualification and certification of nuclear components. These challenges primarily arise from component-scale variations in microstructure and properties influenced by local process conditions and geometry, which affect thermal history, melt pool dynamics, and microstructure evolution. Computational modeling tools can play a crucial role in predicting and controlling this variability. This report presents preliminary results on process modeling tools designed to predict microstructure variability in additively manufactured stainless steel 316 parts. It details the software packages and physical modeling approaches employed to simulate an AM component within an automated process modeling workflow. Initial results are demonstrated through comparisons between predicted microstructures and experimental measurements across various representative processing conditions. The report concludes by discussing the challenges inherent in process modeling of AM components and outlines a plan for future development needs.

36 MATERIALS SCIENCE↗