Search NASA⌕ Search

SEARCH · Search NASA

Results for “analysis orchestration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Curifactory: A research experiment manager

Curifactory is a command line tool and framework for organizing Python experiment code, configuration parameters, and results. It is an opinionated and lightweight approach to workflow management infrastructure and is primarily intended to support researchers conducting experiments on one machine. This software was developed to support the reproducibility of results for several data science projects in the Nuclear Nonproliferation Division at Oak Ridge National Laboratory. Curifactory is intended to be a general framework and is not specific to machine learning or data science. It can aid in any field in which experiments are primarily computation-based studies and can be implemented in Python (e.g., high-energy physics, astronomy, computational chemistry). Here, the design emphasizes the automated caching of intermediate data analysis artifacts to speed up development involving computationally intensive tasks. It also allows for data provenance and experiment reproduction. Individual experiment runs are tracked through logs and their output reports, and entire copies of a run with all cached data and metadata can be exported for others to run using Curifactory on another machine. Curifactory experiments can either be integrated into a project from the beginning or can be written on top of an existing codebase without needing significant modification. A few important views of the Curifactory library can be seen in Figure 1.

97 MATHEMATICS AND COMPUTING↗

Lightfall v0.0.1

Lightfall is a desktop application for synchrotron beamline instrument control, data acquisition, and live analysis at the Advanced Light Source (ALS). Built on Python and Qt, it provides a native graphical interface for operating beamline hardware, configuring and executing experimental scans, and visualizing results in real time. Key features include direct integration with EPICS control systems, a built-in electronic logbook, remote beamline access over secure tunnels, and an interprocess communication (IPC) architecture that coordinates with external analysis applications via ZMQ and EPICS process variables. This IPC approach allows Lightfall to orchestrate specialized analysis tools—including GPU-accelerated streaming correlators—without embedding them, avoiding the dependency conflicts common in monolithic scientific software platforms. Compared to prior approaches such as Xi-CAM's plugin-based architecture, Lightfall's design cleanly separates instrument control from domain-specific analysis, enabling feedback-driven acquisition where live analysis results can adjust scan parameters during an experiment. Its native Qt interface provides responsive performance for real-time data visualization that web-based alternatives struggle to match. Lightfall is designed for use by beamline scientists and staff operating synchrotron instruments at national user facilities.

Pandolfi, Ronald [Lawrence Berkeley National Labor↗

Toward Unified Autonomous Scattering Experiments: A Cross-Facility Case Study at ALS and PETRA III

Autonomous experiments rely on the integration of control, data acquisition, analysis, and decision-making frameworks. While such systems have been demonstrated at individual facilities, adapting them to additional instruments remains challenging due to differences in local infrastructure. We present a modular workflow that connects existing open-source tools for data access (Tiled), workflow orchestration (Prefect), analysis and visualization (pyFAI, Plotly Dash), and Gaussian-process-based adaptive sampling (gpCAM) into a unified framework for autonomous scattering experiments. The same configuration operates across two synchrotron beamlines (ALS 7.3.3 and PETRA III P03) with only minimal facility-specific adjustments, as shown in proof-of-concept demonstrations. This validates that a consistent design emphasizing modularity and shared interfaces can ease deployment across diverse experimental environments. The resulting framework provides a flexible foundation for extending autonomous control and analysis capabilities beyond a single beamline or instrument.

47 OTHER INSTRUMENTATION↗

ellora-spack-gen

The project contains software to analyze the behavior of large language models at code generation tasks. Publicly available information is used to generate Spack package recipes for HPC developers. The software contains orchestration tooling, analysis, and plotting functionality.

Melone, CaetanoN↗

Accelerating lattice gauge theory studies with Agentic AI

Lattice gauge theory research, with its computationally intensive simulations and complex multi‑stage workflows, is well positioned to benefit from agentic AI systems. We demonstrate how such tools can support key components of lattice gauge theory research, including novel simulation code development using standard LQCD frameworks, HPC job orchestration, simulation data analysis, and expert‑guided tuning of algorithmic parameters such as Hasenbusch mass preconditioning and multigrid solvers. Our results show that agentic AI can reduce manual effort, improve productivity, and accelerate the research cycle while maintaining essential human oversight.

Ayyar, Venkitesh [Fermilab]↗

Preparation of the Multi-Site Data Processing at the Vera C. Rubin Observatory

The Vera C. Rubin Observatory’s Legacy Survey of Space and Time (LSST) Camera is scheduled to start taking data in the summer of 2025. The Data Release Production will run the LSST Science Pipe software at data facilities in the US, France and the UK. The LSST Science Pipeline consists of complex directed acyclic graphs (DAGs) of tasks. Rubin will use the Production and Distributed Analysis (PanDA) workflow and workload management system to orchestrate this complex workflow and the distribution of workloads to the data facilities. When run end-to-end by a team of data production staff, this processing (the Science Pipelines, distributed by the workflow and workload management system) is referred to as a 'campaign'. This paper describes the central services and data facility specific services that support this multi-site data process model, including the service deployment infrastructure, the workload and workflow system, the Campaign Management tools, and connection to Rubin Data Management. This paper will also mention the experience of processing the Rubin Commissioning Camera data. All these are part of the effort to scale up the processing capabilities for the expected very large data volume from the LSST Camera.

Yang, Wei [SLAC]↗

Machine Learning Decoding of Full Duplex Signals

A full duplex signal is when two endpoints (server one and server two) transmit on a single conductor pair simultaneously and with the same frequency. This results in the waveforms created from each server to be merged with one another when observed at any point along the transmission line making physical analysis of the wave unobtainable. This project was orchestrated to find the means to separate the merged signal into two separate signals which represent the signals originally sent from each server without an active tap.

97 - MATHEMATICS AND COMPUTING↗

Geospatial Data Workflow Orchestration and Architecture

In an era characterized by explosive growth in geospatial data, the selection of appropriate technologies for data storage, processing, and orchestration is critical for organizations aiming to maintain competitive advantages. This white paper provides a comprehensive analysis of how Oak Ridge National Laboratory (ORNL) has effectively employed various cloud technologies, including containerized applications, container orchestrators, and workflow orchestrators, to develop robust geospatial data processing solutions. We explore the fundamental concepts behind these technologies and compare multiple deployment models tailored to diverse use cases. Our findings conclude that while Kubernetes has emerged as the preferred platform for truly scalable and fault-tolerant production workflows, the choice of workflow orchestration tool requires careful consideration of team needs, pipeline complexity, and deployment environments. This paper aims to serve as a strategic guide for organizations leveraging geospatial data, articulating the balance between technology choices and practical implementation to enhance workflow efficacy and scalability.

97 MATHEMATICS AND COMPUTING↗

Mapping of flumioxazin tolerance in a snap bean diversity panel leads to the discovery of a master genomic region controlling multiple stress resistance genes

Effective weed management tools are crucial for maintaining the profitable production of snap bean (Phaseolus vulgaris L.). Preemergence herbicides help the crop to gain a size advantage over the weeds, but the few preemergence herbicides registered in snap bean have poor waterhemp (Amaranthus tuberculatus) control, a major pest in snap bean production. Waterhemp and other difficult-to-control weeds can be managed by flumioxazin, an herbicide that inhibits protoporphyrinogen oxidase (PPO). However, there is limited knowledge about crop tolerance to this herbicide. We aimed to quantify the degree of snap bean tolerance to flumioxazin and explore the underlying mechanisms. We investigated the genetic basis of herbicide tolerance using genome-wide association mapping approach utilizing field-collected data from a snap bean diversity panel, combined with gene expression data of cultivars with contrasting response. The response to a preemergence application of flumioxazin was measured by assessing plant population density and shoot biomass variables. Snap bean tolerance to flumioxazin is associated with a single genomic location in chromosome 02. Tolerance is influenced by several factors, including those that are indirectly affected by seed size/weight and those that directly impact the herbicide's metabolism and protect the cell from reactive oxygen species-induced damage. Transcriptional profiling and co-expression network analysis identified biological pathways likely involved in flumioxazin tolerance, including oxidoreductase processes and programmed cell death. Transcriptional regulation of genes involved in those processes is possibly orchestrated by a transcription factor located in the region identified in the GWAS analysis. Several entries belonging to the Romano class, including Bush Romano 350, Roma II, and Romano Purpiat presented high levels of tolerance in this study. The alleles identified in the diversity panel that condition snap bean tolerance to flumioxazin shed light on a novel mechanism of herbicide tolerance and can be used in crop improvement.

60 APPLIED LIFE SCIENCES↗

Gene network centrality analysis identifies key regulators coordinating day-night metabolic transitions in Synechococcus elongatus PCC 7942 despite limited accuracy in predicting direct regulator-gene interactions

Synechococcus elongatus PCC 7942 is a model organism for studying circadian regulation and bioproduction, where precise temporal control of metabolism significantly impacts photosynthetic efficiency and CO 2 -to-bioproduct conversion. Despite extensive research on core clock components, our understanding of the broader regulatory network orchestrating genome-wide metabolic transitions remains incomplete. We address this gap by applying machine learning tools and network analysis to investigate the transcriptional architecture governing circadian-controlled gene expression. While our approach showed moderate accuracy in predicting individual transcription factor-gene interactions - a common challenge with real expression data - network-level topological analysis successfully revealed the organizational principles of circadian regulation. Our analysis identified distinct regulatory modules coordinating day-night metabolic transitions, with photosynthesis and carbon/nitrogen metabolism controlled by day-phase regulators, while nighttime modules orchestrate glycogen mobilization and redox metabolism. Through network centrality analysis, we identified potentially significant but previously understudied transcriptional regulators: HimA as a putative DNA architecture regulator, and TetR and SrrB as potential coordinators of nighttime metabolism, working alongside established global regulators RpaA and RpaB. This work demonstrates how network-level analysis can extract biologically meaningful insights despite limitations in predicting direct regulatory interactions. The regulatory principles uncovered here advance our understanding of how cyanobacteria coordinate complex metabolic transitions and may inform metabolic engineering strategies for enhanced photosynthetic bioproduction from CO 2 .

59 BASIC BIOLOGICAL SCIENCES↗

Analysis of 26S Proteasome Activity across Arabidopsis Tissues

Plants utilize the ubiquitin proteasome system (UPS) to orchestrate numerous essential cellular processes, including the rapid responses required to cope with abiotic and biotic stresses. The 26S proteasome serves as the central catalytic component of the UPS that allows for the proteolytic degradation of ubiquitin-conjugated proteins in a highly specific manner. Despite the increasing number of studies employing cell-free degradation assays to dissect the pathways and target substrates of the UPS, the precise extraction methods of highly potent tissues remain unexplored. Here, we utilize a fluorogenic reporting assay using two extraction methods to survey proteasomal activity in different Arabidopsis thaliana tissues. This study provides new insights into the enrichment of activity and varied presence of proteasomes in specific plant tissues.

26S proteasome↗

IRIS-MEMFLOW: Data Flow-Enabled Portable Memory Orchestration in IRIS Runtime for Diverse Heterogeneity

Task-based programming models and execution paradigms provide a means to decompose a computation by expressing it as a graph in which each node represents a specific computation operating on memory objects and the edges define the dependencies in the execution flow. In this execution model, independent nodes in the graph can be executed concurrently in different computing devices, making it suitable for heterogeneous systems in which computing devices with different architectures coexist. However, careful memory orchestration across heterogeneous devices is needed because copies of the same memory object may reside in multiple devices during execution. Manually ensuring such an orchestration is quite challenging. Not only must an application developer guard against race conditions, but they must also optimize data movement between the host and devices because unnecessary data movement significantly impacts performance. To mitigate these challenges, we enhance the IRIS heterogeneous runtime and introduce IRIS-MEMFLOW–a data flow–enabled portable memory abstraction for seamlessly orchestrating memory in diverse heterogeneous computing environments. By using data-flow analysis, IRIS-MEMFLOW guards against race conditions while multiple heterogeneous devices access memory objects. IRIS-MEMFLOW also optimizes data movement between the host and devices without manual intervention. As a result, IRIS provides improved programming productivity, performance, and portability for multidevice heterogeneous executions in high-performance computing and cloud systems that run diverse architectures from different vendors. The efficacy of IRIS-MEMFLOW is evaluated through experiments that show its capability in terms of programming productivity, multidevice heterogeneity, portability, and low overhead versus the state of the art.

Monil, M. A. H. [ORNL] (ORCID:0000000334194037)↗

LeWRON: Agentic Analysis of Electroweak Phase Transitions

The electroweak phase transition (EWPT) is a central topic in particle physics and cosmology, connecting collider phenomenology, baryogenesis, and gravitational-wave observatories. Its analysis requires a technically demanding, convention-sensitive, and model-dependent pipeline, from constructing the finite-temperature effective potential to tracking thermal histories, computing bubble nucleation rates, and predicting gravitational-wave spectra. We present LeWRON (Learning ElectroWeak phase tRansitiON), an agentic framework that orchestrates this pipeline starting from an input Lagrangian. LeWRON combines audited toolbox construction with an Explorer module that uses the generated model-specific code for further analysis, including scans and plots. Intermediate analytic outputs are checked by auditor agents and stored as structured artifacts, enabling reproducible human inspection and downstream use through both a command-line interface and a public Python API. The framework supports a reproduction mode, which infers conventions from the literature and reproduces published results, and a discovery mode, which guides users through structured checkpoints for new models. We demonstrate LeWRON across representative beyond-the-Standard-Model scenarios and release the code on GitHub.

Wang, Isaac R. [Fermilab] (ORCID:000000030789218X)↗

VA EDH Advanced Software Pipeline Framework Report: Enhancing Automation and Scalability

The VA Environmental Determinants of Health (EDH) Advanced Software Pipeline Framework is designed to enhance the efficiency, scalability, and security of geospatial data processing workflows. This framework integrates modern data orchestration and containerization technologies, including Prefect for workflow automation, Docker for containerization, and PostgreSQL/PostGIS for geospatial data storage and analysis. It ensures standardized, reproducible, and automated data processing, supporting VA objectives related to substance use risk assessment and recovery research. The pipeline addresses key scalability and performance challenges through horizontal and vertical scaling, high-performance computing (HPC) integration, parallel processing, task caching, and dynamic resource allocation. These optimizations improve throughput and reduce latency, allowing the system to efficiently manage large and complex datasets. Additionally, security and compliance measures—such as data encryption (SSL), Role-Based Access Control (RBAC), and adherence to GDPR and HIPAA standards—safeguard sensitive information throughout data transmission and storage. A key implementation of this framework includes the automation of shelter list geolocation workflows, ensuring that up-to-date data is readily available for VA decision-making. Lessons learned from this project include the transition from in-memory processing to incremental storage writes, improving resource management and reliability. Future enhancements aim to expand automation, integrate AI-driven anomaly detection, and incorporate high-performance computing resources. This framework provides a scalable, secure, and adaptable solution for managing geospatial datasets, reinforcing the VA’s ability to support clinical and strategic initiatives through data-driven decision-making.

97 MATHEMATICS AND COMPUTING↗

DeepLynx Ecosystem 2025

Poor data integration and governance continue to plague complex engineering projects, resulting in missed cost, schedule, and performance targets. Departments operate in isolated systems with manual data exchange, creating fragmented information that compounds errors and leads to significant delays and cost overruns. The DeepLynx ecosystem addresses these challenges through an open-source, modular data management platform that transforms fragmented project data into an integrated digital thread. Built on a federated microservice architecture, the ecosystem comprises seven specialized tools centered around DeepLynx Nexus, a unified data catalog with hierarchical organization and graph-based navigation capabilities. The ecosystem includes: DeepLynx Stream for real-time timeseries data ingestion from industrial sources; DeepLynx Ingest for governed data uploads with formal review workflows; DeepLynx Lattice for ontology-based entity and relationship extraction; DeepLynx Run for workflow orchestration and secure AI/ML compute; DeepLynx Visualize for 3D digital twin visualization; and DeepLynx Insight for AI-assisted document analysis with traceable, grounded responses. Deployable in cloud, on-premise, or hybrid environments using containerized Docker applications and Helm charts, the DeepLynx ecosystem provides flexible infrastructure that adapts to organizational requirements. By consolidating project data into a unified data lake with role-based access controls and OAuth2 authentication, DeepLynx enables digital thread and digital twin capabilities that improve decision-making, reduce risk, and support complex engineering workflows throughout the project lifecycle.

42 - ENGINEERING↗

NGPINT V3: a containerized orchestration Python software for discovery of next-generation protein–protein interactions

Abstract Summary Batch yeast two-hybrid (Y2H) assays, leveraged with next-generation sequencing, have afforded successful innovations for the analysis of protein–protein interactions. NGPINT is a Conda-based software designed to process the millions of raw sequencing reads resulting from Y2H–next-generation interaction screens. Over time, increasing compatibility and dependency issues have prevented clean NGPINT installation and operation. A system-wide update was essential to continue effective use with its companion software, Y2H-SCORES. We present NGPINT V3, a containerized implementation built with both Singularity and Docker, allowing accessibility across virtually any operating system and computing environment. Availability and implementation This update includes streamlined dependencies and container images hosted on Sylabs (https://cloud.sylabs.io/library/schuyler/ngpint/ngpint) and Dockerhub (https://hub.docker.com/r/schuylerds/ngpint), facilitating easier adoption and integration into high-throughput and cloud-computing workflows. Full instructions and software can be also found in the GitHub repository https://github.com/Wiselab2/NGPINT_V3 and Zenodo https://doi.org/10.5281/zenodo.15256036.

Biochemistry & Molecular Biology↗

From sensing to acclimation: The role of membrane lipid remodeling in plant responses to low temperatures

Abstract Low temperatures pose a dramatic challenge to plant viability. Chilling and freezing disrupt cellular processes, forcing metabolic adaptations reflected in alterations to membrane compositions. Understanding the mechanisms of plant cold tolerance is increasingly important due to anticipated increases in the frequency, severity, and duration of cold events. This review synthesizes current knowledge on the adaptive changes of membrane glycerolipids, sphingolipids, and phytosterols in response to cold stress. We delve into key mechanisms of low-temperature membrane remodeling, including acyl editing and headgroup exchange, lipase activity, and phytosterol abundance changes, focusing on their impact at the subcellular level. Furthermore, we tabulate and analyze current gycerolipidomic data from cold treatments of Arabidopsis, maize, and sorghum. This analysis highlights congruencies of lipid abundance changes in response to varying degrees of cold stress. Ultimately, this review aids in rationalizing observed lipid fluctuations and pinpoints key gaps in our current capacity to fully understand how plants orchestrate these membrane responses to cold stress.

Plant Sciences↗