Search NASA⌕ Search

SEARCH · Search NASA

Results for “workflow development”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

2018 NISAR Applications Workshop: Agriculture and Soil Moisture

Agricultural lands cover the globe and play an essential role in not only sustaining a growing global population, but can have significant implications on the Earth system through land use change (e.g., deforestation, grazing, etc.). As such, countries around the world have dedicated programs for managing these lands. Accurate and timely information concerning the status of agricultural crops (soil moisture, crop health, crop type, etc.) is essential to those nations’ anthropogenic and ecological health as well as economy. The joint NASA/US Department of Agriculture Agricultural Research Service (USDA-ARS) workshop focused on advancing agriculture and soil moisture applications by using remote sensing data from the NASA-ISRO Synthetic Aperture Radar (NISAR) mission (expected launch 2022). Participants included representatives from the international agriculture community that are key players in facilitating integration of Earth Observations into decision support workflows including US Federal Agencies, nonprofits, and private sector. They included scientists, technicians, and program managers with a responsibility for data acquisition and exploitation such as product development, delivery, and use, as well as capacity building. Discussions were held over two and a half days to convey the broader agriculture and soil moisture community information needs, the mission and procedures for various representative participants and programs involved in the delivery of geospatial products, and the capabilities and status of the NISAR mission. Case studies were presented to demonstrate the current state of practice in the use of SAR remote sensing for applications of direct importance for the agriculture and soil moisture communities. Eleven organizations presented their information requirements in response to a set of questions provided by the NASA team, then the NASA team responded by describing the degree to which NISAR could meet these requirements. Discussion ensued about needed data product specifications to increase utility (e.g., projection, latency, etc.), tools and capacity building.

Stavros, Natasha↗

LaRC SmartLab Apps For Instrument Control And Data Processing: Optical Micrometer Data Visualizer

The LaRC Smart Lab applications are a series of software tools to greatly enhance researcher efficiency by streamlining and automating workflows. Python scripts and applications are increasingly being used in scientific workflows, including for instrument control and data processing. Interactive Python scripting environments such as Jupyter Lab provide powerful tools for using Python. In some use cases, the development of standalone applications with dedicated graphical user interfaces (GUIs) can enhance the utility of the code and open it up to more users, including non-programmers. Here, we describe a GUI based optical micrometer data visualization application developed as part of the LaRC SmartLab project. We highlight its use in visualizing experimental data and briefly discuss its implementation to give pointers to programmers who wish develop work based on this application's or similar co de.

LaRC SmartLab↗

Improving I/O-aware Workflow Scheduling via Data Flow Characterization and trade-off Analysis

The scientific computing paradigm has transitioned from compute-intensive to I/O-intensive and memory-intensive in the past decade, especially when data-driven science has become common practice. Numerous empirical I/O-aware scheduling optimizations have been developed by incorporating I/O capacity and bandwidth as constraints into scheduling. Unfortunately, there is a lack of data flow (I/O) characterization tool and an understanding of trade-offs between concurrency, locality, and I/O bandwidth. To bridge the gap, this work 1) presents a set of descriptors to characterize, organize, and visualize I/O profiles, including flow size, I/O bandwidth, and operation count, which group data flows by I/O types, tasks, and files; 2) proposes an I/O Roofline model-based trade-off analysis to find the optimal trade-off between flow operational intensity, concurrency, and flow performance. The I/O descriptors generate useful insights into complicated I/O behaviors, suggesting distinct concurrency, storage, and scheduling to be used by types, tasks, and files. The proposed trade-off analysis guides scheduling decisions that generate resource assignment with the best flow parallelism. We evaluate our I/O-aware scheduling methodology on a highly I/O-intensive workflow–1000 Genomes. The experimental results demonstrate speedups of up to 2.4× compared to the state-of-the- art methods.

Guo, Luanzheng [BATTELLE (PACIFIC NW LAB)]↗

The microbiologist's guide to metaproteomics

Metaproteomics is an emerging approach for studying microbiomes, offering the ability to characterize proteins that underpin microbial functionality within diverse ecosystems. As the primary catalytic and structural components of microbiomes, proteins provide unique insights into the active processes and ecological roles of microbial communities. By integrating metaproteomics with other omics disciplines, researchers can gain a comprehensive understanding of microbial ecology, interactions, and functional dynamics. This review, developed by the Metaproteomics Initiative (www.metaproteomics.org), serves as a practical guide for both microbiome and proteomics researchers, presenting key principles, state-of-the-art methodologies, and analytical workflows essential to metaproteomics. Topics covered include experimental design, sample preparation, mass spectrometry techniques, data analysis strategies, and statistical approaches.

bioinformatics↗

Omics-driven onboarding of the carotenoid producing red yeast Xanthophyllomyces dendrorhous CBS 6938

Transcriptomics is a powerful approach for functional genomics and systems biology, yet it can also be used for genetic part discovery. Here, we derive constitutive and light-regulated promoters directly from transcriptomics data of the basidiomycete red yeast Xanthophyllomyces dendrorhous CBS 6938 (anamorph Phaffia rhodozyma) and use these promoters with other genetic elements to create a modular synthetic biology parts collection for this organism. X. dendrorhous is currently the sole biotechnologically relevant yeast in the Tremellomycete class-it produces large amounts of astaxanthin, especially under oxidative stress and exposure to light. Thus, we performed transcriptomics on X. dendrorhous under different wavelengths of light (red, green, blue, and ultraviolet) and oxidative stress. Differential gene expression analysis (DGE) revealed that terpenoid biosynthesis was primarily upregulated by light through crtI, while oxidative stress upregulated several genes in the pathway. Further gene ontology (GO) analysis revealed a complex survival response to ultraviolet (UV) where X. dendrorhous upregulates aromatic amino acid and tetraterpenoid biosynthesis and downregulates central carbon metabolism and respiration. The DGE data was also used to identify 26 constitutive and regulated genes, and then, putative promoters for each of the 26 genes were derived from the genome. Simultaneously, a modular cloning system for X. dendrorhous was developed, including integration sites, terminators, selection markers, and reporters. Each of the 26 putative promoters were integrated into the genome and characterized by luciferase assay in the dark and under UV light. The putative constitutive promoters were constitutive in the synthetic genetic context, but so were many of the putative regulated promoters. Notably, one putative promoter, derived from a hypothetical gene, showed ninefold activation upon UV exposure. Thus, this study reveals metabolic pathway regulation and develops a genetic parts collection for X. dendrorhous from transcriptomic data. Therefore, this study demonstrates that combining systems biology and synthetic biology into an omics-to-parts workflow can simultaneously provide useful biological insight and genetic tools for nonconventional microbes, particularly those without a related model organism. This approach can enhance current efforts to engineer diverse microbes.

60 APPLIED LIFE SCIENCES↗

MVP: a modular viromics pipeline to identify, filter, cluster, annotate, and bin viruses from metagenomes

While numerous computational frameworks and workflows are available for recovering prokaryote and eukaryote genomes from metagenome data, only a limited number of pipelines are designed specifically for viromics analysis. With many viromics tools developed in the last few years alone, it can be challenging for scientists with limited bioinformatics experience to easily recover, evaluate quality, annotate genes, dereplicate, assign taxonomy, and calculate relative abundance and coverage of viral genomes using state-of-the-art methods and standards. Here, we describe Modular Viromics Pipeline (MVP) v.1.0, a user-friendly pipeline written in Python and providing a simple framework to perform standard viromics analyses. MVP combines multiple tools to enable viral genome identification, characterization of genome quality, filtering, clustering, taxonomic and functional annotation, genome binning, and comprehensive summaries of results that can be used for downstream ecological analyses. Overall, MVP provides a standardized and reproducible pipeline for both extensive and robust characterization of viruses from large-scale sequencing data including metagenomes, metatranscriptomes, viromes, and isolate genomes. As a typical use case, we show how the entire MVP pipeline can be applied to a set of 20 metagenomes from wetland sediments using only 10 modules executed via command lines, leading to the identification of 11,656 viral contigs and 8,145 viral operational taxonomic units (vOTUs) displaying a clear beta-diversity pattern. Further, acting as a dynamic wrapper, MVP is designed to continuously incorporate updates and integrate new tools, ensuring its ongoing relevance in the rapidly evolving field of viromics. MVP is available at https://gitlab.com/ccoclet/mvp and as versioned packages in PyPi and Conda.

59 BASIC BIOLOGICAL SCIENCES↗

Accelerating lattice gauge theory studies with Agentic AI

Lattice gauge theory research, with its computationally intensive simulations and complex multi‑stage workflows, is well positioned to benefit from agentic AI systems. We demonstrate how such tools can support key components of lattice gauge theory research, including novel simulation code development using standard LQCD frameworks, HPC job orchestration, simulation data analysis, and expert‑guided tuning of algorithmic parameters such as Hasenbusch mass preconditioning and multigrid solvers. Our results show that agentic AI can reduce manual effort, improve productivity, and accelerate the research cycle while maintaining essential human oversight.

Ayyar, Venkitesh [Fermilab]↗

Sustainable Enablers of Knowledge Management Strategies in a Higher Education Institution

By facilitating the capture, organization, and dissemination of knowledge within and beyond the institution, knowledge management (KM) in higher education institutions (HEIs) fuels innovation, enhances research impact, and strengthens collaboration, ultimately leading to the creation of new knowledge and its valuable exchange. However, there is still much to explore in terms of the enablers of knowledge creation, sharing, and transfer. Therefore, this paper aims to identify the enablers of effective KM in the Polytechnique University of Leiria, which serves as a benchmark for other higher education institutions due to its leadership role in RUN-EU, a consortium of European universities. To achieve this, a narrative analysis based on information from SCOPUS and the institute’s website, focusing on innovation, research, and development strategies, is proposed. The findings suggest that for KM initiatives to be successful, they need to be strategically designed, culturally supported, technologically enabled, and integrated into existing workflows.

Santos, Eleonora (ORCID:0000000346930804)↗

Lunar Impact Flash Locations from NASA's Lunar Impact Monitoring Program

Meteoroids are small, natural bodies traveling through space, fragments from comets, asteroids, and impact debris from planets. Unlike the Earth, which has an atmosphere that slows, ablates, and disintegrates most meteoroids before they reach the ground, the Moon has little-to-no atmosphere to prevent meteoroids from impacting the lunar surface. Upon impact, the meteoroid's kinetic energy is partitioned into crater excavation, seismic wave production, and the generation of a debris plume. A flash of light associated with the plume is detectable by instruments on Earth. Following the initial observation of a probable Taurid impact flash on the Moon in November 2005,1 the NASA Meteoroid Environment Office (MEO) began a routine monitoring program to observe the Moon for meteoroid impact flashes in early 2006, resulting in the observation of over 330 impacts to date. The main objective of the MEO is to characterize the meteoroid environment for application to spacecraft engineering and operations. The Lunar Impact Monitoring Program provides information about the meteoroid flux in near-Earth space in a size range-tens of grams to a few kilograms-difficult to measure with statistical significance by other means. A bright impact flash detected by the program in March 2013 brought into focus the importance of determining the impact flash location. Prior to this time, the location was estimated to the nearest half-degree by visually comparing the impact imagery to maps of the Moon. Better accuracy was not needed because meteoroid flux calculations did not require high-accuracy impact locations. But such a bright event was thought to have produced a fresh crater detectable from lunar orbit by the NASA spacecraft Lunar Reconnaissance Orbiter (LRO). The idea of linking the observation of an impact flash with its crater was an appealing one, as it would validate NASA photometric calculations and crater scaling laws developed from hypervelocity gun testing. This idea was dependent upon LRO finding a fresh impact crater associated with one of the impact flashes recorded by Earth-based instruments, either the bright event of March 2013 or any other in the database of impact observations. To find the crater, LRO needed an accurate area to search. This Technical Memorandum (TM) describes the geolocation technique developed to accurately determine the impact flash location, and by association, the location of the crater, thought to lie directly beneath the brightest portion of the flash. The workflow and software tools used to geolocate the impact flashes are described in detail, along with sources of error and uncertainty and a case study applying the workflow to the bright impact flash in March 2013. Following the successful geolocation of the March 2013 flash, the technique was applied to all impact flashes detected by the MEO between November 7, 2005, and January 3, 2014.

Moser, D. E.↗

Streamling the Change Management with Business Rules

Will discuss how their organization is trying to streamline workflows and the change management process with business rules. In looking for ways to make things more efficient and save money one way is to reduce the work the workflow task approvers have to do when reviewing affected items. Will share the technical details of the business rules, how to implement them, how to speed up the development process by using the API to demonstrate the rules in action.

Basic Rules↗

Automated Strain Construction for Biosynthetic Pathway Screening in Yeast

Automation accelerates the Design-Build-Test-Learn (DBTL) cycle for synthetic biology; however, most strain construction pipelines lack robotic integration. Here, in this study, we present the workflow design and source code for a modular, integrated protocol that automates the Build step in Saccharomyces cerevisiae. We programmed the Hamilton Microlab VANTAGE to integrate off-deck hardware via its central robotic arm, enabling automated steps that increased throughput to 2,000 transformations per week. We developed a user interface with the Hamilton VENUS software to support on-demand parameter customization. As a proof of concept, we screened a gene library in an engineered yeast strain producing verazine, a key intermediate in the biosynthesis of steroidal alkaloids. Our pipeline rapidly identified pathway bottlenecks and genes that enhanced verazine production by 2.0- to 5-fold. This technical note provides resources for synthetic biologists designing yeast workflows for biofoundries to screen libraries for pathway discovery/optimization, combinatorial biosynthesis, and protein engineering.

automation↗

HDF5 in the exascale era: Delivering efficient and scalable parallel I/O for exascale applications

Accurately modeling real-world systems requires scientific applications at exascale to generate massive amounts of data and manage data storage efficiently. However, parallel input and output (I/O) faces challenges due to new application workflows and the state-of-the-art memory, interconnect, and storage architectures considered in exascale designs. The storage hierarchy has expanded with node-local persistent memory, solid-state storage, and traditional disk and tape-based storage, thus requiring efficiency at each layer and much more efficient data movement among these layers. This paper discusses how the ExaHDF5 project improved the I/O performance and data management for exascale architectures by enhancing HDF5, a widely used parallel I/O library. The team developed an Asynchronous I/O Virtual Object Layer (VOL) connector that allowed overlapping I/O with computation. They also created a Cache VOL to complement asynchronous I/O by incorporating fast storage layers, such as burst buffer and node-local storage, into the parallel I/O workflow through caching and staging data. Additionally, the team enabled data aggregation and I/O at the node level by using a Subfiling Virtual File Driver (VFD). To demonstrate superior I/O performance with HDF5 at exascale, the ExaHDF5 team collaborated with several exascale applications. In this paper, we show I/O performance improvements for three applications: Cabana (a particle-based simulation library), EQSIM (a regional earthquake simulation software), and E3SM (a climate system modeling library).

Asynchronous I/Ol↗

Polarized Deep-Inelastic Scattering with Spin Correlations in Herwig 7

This repository is the research software and reproducibility companion for the HerwigPol polarized deep-inelastic scattering implementation developed for Herwig 7. It brings together the modified Herwig and ThePEG source snapshots, the curated POLDIS fixed-order reference code, the custom Rivet analyses, the DIS validation workflow, and the paper source in a single formal repository layout. The repository is intended to preserve the source-level ingredients needed to rebuild and re-run the validated DIS studies. It therefore tracks code, input cards, workflow drivers, and technical notes, while intentionally excluding generated artifacts such as build products, campaign outputs, merged YODA files, plots, and rendered paper outputs.

Papaefstathioou, Andreas [Kennesaw State Universit↗

Deployment of inference as a service at the US CMS Tier-2 data centers

Coprocessors, especially GPUs, will be a vital ingredient of data production workflows at the HL-LHC. At CMS, the GPU-as-a-service approach for production workflows is implemented by the SONIC project (Services for Optimized Network Inference on Coprocessors). SONIC provides a mechanism for outsourcing computationally demanding algorithms, such as neural network inference, to remote servers, where requests from multiple clients are intelligently distributed across multiple GPUs by a load-balancing service. This talk highlights the recent progress in deploying SONIC at selected U.S. CMS Tier-2 data centers. Using realistic CMS Run3 data processing workflows, such as those containing transformer-based algorithms, we demonstrate how SONIC is integrated into the production-like environment to enable accelerated inference offloading. We will present developments from both the client and server sides, including production job and data center configurations for NVIDIA and AMD GPUs. We will also present performance scaling benchmarks and discuss the challenges of operating SONIC in CMS production, such as server discovery, GPU saturation, fallback server logic, etc.

Holzman, Burt↗

Automated Immunoprecipitation Workflow for Comprehensive Acetylome Analysis

Immunoprecipitation is one of the most effective methods for enrichment of lysine-acetylated peptides for comprehensive acetylome analysis using mass spectrometry. Manual acetyl peptide enrichment method using non-conjugated antibodies and agarose beads has been developed and applied in various studies. However, it is time consuming, and can introduce contaminants and variability that leads to potential sample loss and decreased sensitivity and robustness of the analysis. Here we describe a fast, automated enrichment protocol that enables reproducible and comprehensive acetylome analysis using a magnetic bead-based immunoprecipitation reagent.

Lysine acetylation, Acetylome, Acetyl peptide enri↗

Preliminary Results on Process Modeling Tools for Determining Variability in Additively Manufactured Stainless Steel 316 Parts

The Advanced Materials and Manufacturing Technologies program aims to accelerate the development, qualification, demonstration, and deployment of advanced materials and manufacturing technologies to enable reliable and economical nuclear energy. However, the distinct characteristics of additive manufacturing (AM) materials, stemming from their unique processing history, microstructure, and properties, pose significant challenges for the qualification and certification of nuclear components. These challenges primarily arise from component-scale variations in microstructure and properties influenced by local process conditions and geometry, which affect thermal history, melt pool dynamics, and microstructure evolution. Computational modeling tools can play a crucial role in predicting and controlling this variability. This report presents preliminary results on process modeling tools designed to predict microstructure variability in additively manufactured stainless steel 316 parts. It details the software packages and physical modeling approaches employed to simulate an AM component within an automated process modeling workflow. Initial results are demonstrated through comparisons between predicted microstructures and experimental measurements across various representative processing conditions. The report concludes by discussing the challenges inherent in process modeling of AM components and outlines a plan for future development needs.

36 MATERIALS SCIENCE↗

Integrating ORNL’s HPC and Neutron Facilities with a Performance-Portable CPU/GPU Ecosystem

We explore the development of a performance-portable CPU/GPU ecosystem to integrate two of the US Department of Energy’s (DOE’s) largest scientific instruments, the Oak Ridge Leadership Computing facility and the Spallation Neutron Source (SNS), both of which are housed at Oak Ridge National Laboratory. We select a relevant data reduction workflow use-case to obtain the differential scattering cross-section from data collected by SNS’s CORELLI and TOPAZ instruments. We compare the current CPU-only production implementation using the Garnet Python multiprocess package based on the Mantid C++ framework against our proposed CPU/GPU implementation that uses the LLVM-based, just-in-time Julia scientific language and the JACC.jl performance-portable package. Two proxy apps were developed: (i) an app for extracting relevant Mantid kernels (MDNorm) in C++ and (ii) the Julia MiniVATES.jl miniapp. We present performance results for NVIDIA A100 and AMD MI100 GPUs and AMD EPYC 7513 and 7662 CPUs. The results provide insights for future generations of data reduction software that can embrace performance portability for an integrated research infrastructure across DOE’s experimental and computational facilities.

Hahn, Steven↗

A high-throughput workflow to analyze sequence-conformation relationships and explore hydrophobic patterning in disordered peptoids

Understanding how a macromolecule’s primary sequence governs its conformational landscape is crucial for elucidating its function, yet these design principles are still emerging for macromolecules with intrinsic disorder. Herein, we introduce a high-throughput workflow that implements a practical colorimetric conformational assay, introduces a semi-automated sequencing protocol using matrix-assisted laser desorption/ionization and tandem mass spectrometry (MALDI-MS/MS), and develops a generalizable sequence-structure algorithm. Using a model system of 20mer peptidomimetics containing polar glycine and hydrophobic N-butylglycine residues, we identified nine classifications of conformational disorder and isolated 122 unique sequences across varied compositions and conformations. Conformational distributions of three compositionally identical library sequences were corroborated through atomistic simulations and ion mobility spectrometry coupled with liquid chromatography. A data-driven strategy was developed using existing sequence variables and data-derived “motifs” to inform a machine-learning algorithm toward conformation prediction. Here, this multifaceted approach enhances our understanding of sequence-conformation relationships and offers a powerful tool for accelerating the discovery of materials with conformational control.

data-driven analysis↗