Search NASA⌕ Search

SEARCH · Search NASA

Results for “workflows”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Integrated Simulation of Weld Residual Stress Evolution and Crack Propagation Using XFEM

Nuclear power plant components operate in environments that promote multiple degradation mecha- nisms, several of which involve crack initiation and growth. An ongoing effort in the U.S. Department of Energy’s Nuclear Energy Advanced Modeling and Simulation (NEAMS) program is developing a general capability within the Multiphysics Object Oriented Simulation Environment (MOOSE) framework for simulating three-dimensional crack growth under a range of driving conditions, including fatigue, stress corrosion cracking (SCC), brittle fracture, and stress-relaxation cracking. This report demonstrates an end-to-end workflow that uses this capability to model weld-residual-stress-driven SCC in the J-groove weld of a pressurized-water reactor control rod drive mechanism penetration in the vessel head. The workflow consists of a thermomechanical welding simulation with temperature-dependent plasticity, followed by cooldown to ambient conditions, and a restart of the simulation using the MOOSE extended finite element method (XFEM) module to propagate a three-dimensional crack through the residual stress field. New welding capabilities were developed to properly initialize newly activated elements in the weld region, and robustness improvements were made to the mesh-based algorithm for defining cutting planes in the 3D XFEM algorithm, allowing it to handle complex crack fronts and stress fields. Together these advances allowed the simulated SCC crack to grow from an initial elliptical flaw in the weld, across the weld, through the tube wall, and almost to the triple point (where the weld, tube, and reactor pressure vessel head intersect) over roughly 36 years of simulated service. These results demonstrate a workflow that can be extended to fully three-dimensional welding simulations and more complex crack interaction problems.

42 - ENGINEERING↗

Development and implementation of high-throughput proteomic and metabolomics assays by using advanced chromatographic and mass spectrometric systems (CRADA Final Report)

The mission of this CRADA with Agilent was to couple powerful MS platforms (QQQ, IM-QTOFMS) with Agilent’s novel Ultra-High-Performance Liquid Chromatography (UHPLC) fast metabolomic workflows and perform ABF Machine Learning (ML) to generated datasets. Agilent transferred UHPLC methods to PNNL and LBNL and methods were implemented and demonstrated in both labs, achieving total acquisition times of < 10 min. Metabolites analyzed using Agilent’s shared methods included metabolites from central carbon metabolism, common across hosts, and metabolites unique to engineered strains. Standards were acquired in an UHPLC-Drift Tube Ion Mobility Mass Spectrometer (DTIMS) system for the first time within the context of ABF and methods were optimized based on Agilent’s protocols. Samples from ABF hosts Pseudomonas putida, Aspergillus pseudoterreus, Aspergillus niger and Rhodosporidium toruloides were analyzed using the UHPLC-DTIMS platform for a total of 276 runs. A data analysis workflow compatible with the Experimental Data Depot (EDD) and completely shareable was developed for the acquired UHPLC-DTIMS data. Samples were analyzed using a Data Independent Acquisition Approach (DIA), which for most of the standards provided more transitions therefore increasing detection confidence. Using the data acquired by PNNL, LBNL, and Agilent’s specifications from previous ML projects, SNL applied an ensemble ML strategy to pick the best performing model for automated LC-method selection. Finally, with the contribution of the participant labs and Agilent, SNL developed an Automated Method Selection (AMS) software tool to predict the best liquid chromatography method for analysis of any new molecules of interest. Samples with novel pathways and new metabolite targets of interest are generated at a high pace in the ABF. Overall, the project advanced rapid metabolomics by combining liquid chromatography, ion mobility spectrometry, and data-independent mass spectrometry with machine learning. This multidimensional approach uses retention time, collision cross-section, precursor mass, and fragment-ion information to distinguish chemically similar metabolites that can be difficult to resolve using conventional liquid- or gas-chromatography methods. The resulting workflow also provided automated metabolite-identification error estimates, addressing a recognized need for statistical confidence measures in metabolomics.

Petzold, Christopher [Lawrence Berkeley National L↗

Machine Learning-Based Anomaly Detection for PMT Data Quality Monitoring in the SBN and DUNE

Maintaining high-quality detector data is essential for achieving the scientific objectives of the Short-Baseline Neutrino (SBN) Program at Fermilab. Current data quality monitoring (DQM) procedures rely primarily on threshold-based metrics and manual inspection of detector monitoring plots, making the detection of subtle or gradually developing anomalies both time-consuming and dependent on expert interpretation. This project developed and evaluated a machine-learning workflow for automatically identifying anomalous photomultiplier tube (PMT) channels in the Short-Baseline Near Detector (SBND) using optical-hit amplitude data. A Python-based analysis program was developed to process ROOT files, extract statistical features describing individual PMT amplitude distributions, and generate feature vectors for anomaly detection. These features were used to train an Isolation Forest model using data representing normal detector operation. The trained model was subsequently applied to independent detector runs to identify channels exhibiting statistically unusual behavior relative to the learned reference response. To support expert interpretation, the workflow generated complementary diagnostic products, including anomaly score distributions, normalized amplitude comparisons, decision-tree visualizations, and principal component analysis (PCA) projections. This project demonstrated the feasibility of integrating unsupervised machine learning into detector data-quality monitoring and developed a complete workflow for automated PMT performance assessment to aid expert-driven review. Beyond its technical contributions, the VFP appointment fostered a research collaboration between Aurora University and Fermilab and provided direct workforce development benefits by training the visiting faculty member in detector-scale machine-learning methods that are now being incorporated into undergraduate coursework and research. The methodology developed here provides a foundation for future applications to ProtoDUNE and other liquid argon time projection chamber (LArTPC) detectors, contributing to ongoing efforts to improve detector reliability, reduce manual monitoring requirements, and enable scalable data quality monitoring for future large-scale neutrino experiments, including the Deep Underground Neutrino Experiment (DUNE).

Colón Santana, Juan A. [Unlisted, US, IL]↗

FY26 Progress on Demonstration of a Multiphysics Steady State Capability for Modeling Core Radial Expansion in SFRs

Under the U.S. Department of Energy Office of Nuclear Energy Advanced Modeling and Simulation (NEAMS) Program, an integrated multiphysics approach is being developed to model the core bowing phenomena important to liquid metal-cooled fast reactors. Core bowing is an important passive safety mechanism in liquid metal-cooled fast reactors and involves multiphysics effects including radiation transport, fluid flow, heat transfer, and mechanical response to temperature and flux gradients. This report summarizes recent progress on developing a multiphysics, MOOSE-based workflow to predict core bowing and associated reactivity feedback. Significant new capabilities in the reactor physics code Griffin - sodium backfill and pin power reconstruction for deformed geometries - were applied in this effort. This year’s work included verification, code comparisons, sensitivity studies, and coupled demonstrations that advance the state of MOOSE-based core bowing workflow. Griffin’s sodium backfill capability was verified by demonstrating that its automated treatment of geometry expansion and material-density updates reproduces manual calculations exactly, confirming solid mass conservation and proper coolant backfilling in expanded geometries. Reconstructed pin powers were compared for Griffin’s ductheterogeneous and ring-heterogeneous treatments in single-, seven-, and nineteen-assembly cases, with best agreement observed in lower-leakage configurations and the duct-heterogeneous approach offering substantially lower computational cost. Thermal-hydraulic sensitivity sensitivities showed that MOOSE SCM, SAM, and CFD are expected to produce similar deformation predictions despite variances in their temperature predictions, and that explicit treatment of inter-assembly flow becomes increasingly important as gap flow rate increases. Finally, coupled demonstrations on small multi-assembly configurations using Griffin, MOOSE Solid Mechanics, MOOSE SCM, and Heat Conduction produced physically consistent reactivity feedback from thermal expansion and bowing. The coupled demonstrations simulated grid plate expansion as well as resultant core bowing at full power conditions. Simplifications were made in current workflow, namely the assumption of instantaneous full power conditions following hot zero power, and pre-expanding the Griffin geometry axially due to lack of an axial fuel pin expansion model and temperature feedback to Griffin.

Wozniak, Nicholas↗

Toward a persistent event-streaming system for high-performance computing applications

High-performance computing (HPC) applications have traditionally relied on parallel file systems and file transfer services to manage data movement and storage. Alternative approaches have been proposed that use direct communications between application components, trading persistence and fault tolerance for speed. Event-driven architectures, as popularized in enterprise contexts, present a compelling middle ground, avoiding the performance cost and API constraints of parallel file systems while retaining persistence and offering impedance matching between application components. However, adapting streaming frameworks to HPC workloads requires addressing challenges unique to HPC systems. This paper investigates the potential for a streaming framework designed for HPC infrastructures and use cases. We introduce Mofka, a persistent event-streaming framework designed specifically for HPC environments. Mofka combines the capabilities of a traditional streaming service with optimizations tailored to the HPC context, such as support for massively multicore nodes, efficient scaling for large producer-consumer workflows, RDMA-enabled high-performance network communications, specialized network fabrics with multiple links per node, and efficient handling of large scientific data payloads. Built using the Mochi suite of HPC data service components, Mofka provides a lightweight, modular, and high-performance solution for persistent streaming in HPC systems. We present the architecture of Mofka and evaluate its performance against Kafka and Redpanda using benchmarks on diverse platforms, including Argonne's Polaris and Oak Ridge's Frontier supercomputers, showing up to 8× improvement in throughput in some scenarios. We then demonstrate its utility in several real-world applications: a tomographic reconstruction pipeline, a workflow for the discovery of metal-organic frameworks for carbon capture, and the instrumentation of Dask workflows for provenance tracking and performance analysis.

HPC↗

Myna

The additive manufacturing (AM) community has been developing digital factory tools over the past decade to better leverage the multi-modal process data coming out of the advanced manufacturing process. As a result, numerous databases of additive manufacturing process data exist in the literature and in the archival storage of disparate research groups. While some efforts have been made to create a standard ontology for storing and sharing AM data, in practice a variety of data structures are used to store AM build data, even within a single institution. This causes many problems for maintainability and extensibility when attempting to integrate computational modeling tools with experimental data to either validate models or to provide further insight into results and trends. Myna is a Python-based framework that aims to decrease the effort needed to connect individual computational models to the variety of AM process data that exist in different research groups and institutions. This type of software is sometimes referred to as "middleware" or “glueware,” in that it connects disparate databases and applications into a single computational ecosystem. Instead of maintaining unique interfaces between each application and each database, developers can create a single interface from each application to Myna and thereby gain access to the implemented database connections. Similarly, developing a database connection in Myna provides access to the developed simulation applications. This framework greatly simplifies the maintainability of model applications that rely on experimental data. Using external simulation tools, users will also be able to run pre-configured workflows using the built-in workflow manager. Several examples of input files are provided with Myna for different workflows, including melt pool geometry predictions and detailed melt pool and solidification microstructure predictions.

Knapp, GerryL. [Oak Ridge National Laboratory (ORN↗

PVDeg: Enhancing Usability and AI-Driven Multi-Mechanism Degradation Modeling

PVDeg version 0.7.0, released in December 2025, introduced major enhancements to improve usability and performance. This update reorganized tutorials and tool notebooks to create a more intuitive experience, enabling users to easily follow and adapt workflows for their specific analyses. In addition to structural improvements, both the notebooks and core logic underwent significant optimization for efficiency, robustness, and style. These refinements were supported by new testing frameworks built on nbval and pytest, adherence to PEP8 standards, and extensive code refactoring, which collectively simplify onboarding for new developers. Looking ahead, version 0.8.0 will deliver advanced AI-driven capabilities. The primary focus is to further develop and automate the degradation workflow, designed to analyze PV module degradation across diverse locations and system configurations. By integrating large language models (LLMs) to scan literature and compile a comprehensive database of materials and degradation rates, this feature will enable modeling of multiple materials and mechanisms within a single, streamlined workflow. Users will be able to evaluate degradation impacts on different system architectures under varying environmental conditions, facilitating informed decisions on bill-of-materials optimization for specific deployment scenarios. These advancements position PVDeg as a powerful, user-friendly tool for accelerating PV reliability research and system design.

14 SOLAR ENERGY↗

Multibody for Everybody (M4E) - A Linearization Approach to Enable Frequency Domain Analysis, Time Integration and Control Co-Design

1.1 Background/Objectives: Marine energy represents a promising yet underexploited source of power. To increase the harvested power, significant efforts have been made to improve wave energy converter (WEC) modeling capabilities and optimize power take-off (PTO) performance; however, these efforts have often treated WEC dynamics, PTO design, and controller development sequentially. In contrast, control co-design (CCD) is emerging as a promising strategy to address these issues directly, creating a growing need for fast analysis tools suitable for repeated simulation and parametric studies [1]. To support this need, this work presents the Multibody for Everybody (M4E) [2] linearization module, which employs a symbolic toolbox to provide deeper insight of WEC design parameters. The objective is to demonstrate that a minimal-coordinate linearization of articulated WEC dynamics can provide accurate wave response predictions and substantial computational savings relative to nonlinear time-domain simulation, while preserving compatibility with broader wave-energy analysis workflows, enabling CCD. 1.2 Approach/Activities: The proposed approach linearizes the equations of motion, generated by M4E, in minimal coordinates about a selected operating point and combines the resulting system with frequencydomain hydrodynamic terms to incorporate the reduced mass, damping, stiffness, and forcing operators. The linearized model is used for both impedance-based response amplitude operator (RAO) prediction and rapid regular-wave time integration. The methodology is demonstrated on a single-flap device and a FOSWEC configuration, with linearized M4E responses compared against the corresponding nonlinear M4E simulations and WEC-Sim results. Regular-wave time histories, RAO trends, and runtime differences are assessed. The framework is also compatible with broader wave-energy workflows, including coupling to WecOptTool, although that capability is not the focus of this work [3]. 1.3 Results/Lessons: The linearized M4E model reproduces key regularwave response characteristics such as integration and Response Amplitude over multiple frequencies. This module matches nonlinear M4E and WEC-Sim results while substantially reducing integration cost. Thus, the proposed framework can serve as a rapid analysis layer for articulated WEC design, parameter studies, and controls-oriented workflows. The analysis is most appropriate in the near-equilibrium regime, about the linearization point.

16 TIDAL AND WAVE POWER↗

Using WorldView-2 Imagery to Track Flooding in Thailand in a Multi-Asset Sensorweb

For the flooding seasons of 2011-2012 multiple space assets were used in a "sensorweb" to track major flooding in Thailand. Worldview-2 multispectral data was used in this effort and provided extremely high spatial resolution (2m / pixel) multispectral (8 bands at 0.45-1.05 micrometer spectra) data from which mostly automated workflows derived surface water extent and volumetric water information for use by a range of NGO and national authorities. We first describe how Worldview-2 and its data was integrated into the overall flood tracking sensorweb. We next describe the use of Support Vector Machine learning techniques that were used to derive surface water extent classifiers. Then we describe the fusion of surface water extent and digital elevation map (DEM) data to derive volumetric water calculations. Finally we discuss key future work such as speeding up the workflows and automating the data registration process (the only portion of the workflow requiring human input).

surface water extent↗

A Dose of Reality: Radiation Analysis for Realistic Human Spacecraft

INTRODUCTION As with most computational analyses, a tradeoff exists between problem complexity, resource availability and response accuracy when modeling radiation transport from the source to a detector. The largest amount of analyst time for setting up an analysis is often spent ensuring that any simplifications made have minimal impact on the results. The vehicle shield geometry of interest is typically simplified from the original CAD design in order to reduce computation time, but this simplification requires the analyst to "re-draw" the geometry with a limited set of volumes in order to accommodate a specific radiation transport software package. The resulting low-fidelity geometry model cannot be shared with or compared to other radiation transport software packages, and the process can be error prone with increased model complexity. The work presented here demonstrates the use of the DAGMC (Direct Accelerated Geometry for Monte Carlo) Toolkit from the University of Wisconsin, to model the impacts of several space radiation sources on a CAD drawing of the US Lab module. METHODS The DAGMC toolkit workflow begins with the export of an existing CAD geometry from the native CAD to the ACIS format. The ACIS format file is then cleaned using SpaceClaim to remove small holes and component overlaps. Metadata is then assigned to the cleaned geometry file using CUBIT/Trelis from csimsoft (Registered Trademark). The DAGMC plugin script removes duplicate shared surfaces, facets the geometry to a specified tolerance, and ensures that the faceted geometry is water tight. This step also writes the material and scoring information to a standard input file format that the analyst can alter as desired prior to running the radiation transport program. The scoring results can be transformed, via python script, into a 3D format that is viewable in a standard graphics program. RESULTS The CAD model of the US Lab module of the International Space Station, inclusive of all the racks and components, was simplified to remove holes and volume overlaps. Problematic features within the drawing were also removed or repaired to prevent runtime issues. The cleaned drawing was then run through the DAGMC workflow to prepare for analysis. Pilot tests modeling transport of 1GeV proton and 800MeV/A oxygen sources show that reasonable results are converged upon in an acceptable amount of overall computation time from drawing preparation to data analysis. The FLUKA radiation transport code will next be used to model both a GCR and a trapped radiation source. These results will then be compared with measurements that have been made by the radiation instrumentation deployed inside the US Lab module. DISCUSSION Early analyses have indicated that the DAGMC workflow is a promising toolkit for running vehicle geometries of interest to NASA through multiple radiation transport codes. In addition, recent work has shown that a realistic human phantom, provided via a subcontract with the University of Florida, can be placed inside any vehicle geometry for a combinatorial analysis. This added functionality gives the user the ability to score various parameters at the organ level, and the results can then be used as input for cancer risk models.

Barzilla, J. E.↗

Toward a Common Earth Data Publication Framework

Data publication is an essential activity for all data archives. Each of NASA's twelve Distributed Active Archive Centers (DAACs) have established publication workflows which account for the heterogeneous suite of missions, instruments, data providers, and datasets managed within the Earth Observation System Data and Information System (EOSDIS) program. Some aspects of data publication vary across DAACs: workflows range from manual to automatic, terms used to describe publication elements differ, and systems used to publish and manage data vary. Despite these differences, the DAAC data publication processes are generally the same: obtain the data and related information from data providers, describe the data with metadata and documentation, and release the data for access by the user community. In order to improve consistency and reduce the time required to publish data, we have developed a cross-DAAC initiative called the Common Earthdata Publication Framework (Earthdata Pub). Earthdata Pub seeks to: standardize communications and interactions with data providers; identify and standardize common workflows and steps in the data publication process; and design/implement a front-end system with features that include a common web interface, email & status tracking, and common application programming interfaces (APIs) to communicate with various DAAC-specific software components (services and applications) on the back-end. We will present the latest updates on this effort's progress and future plans.

data publication↗

A Model-Based Systems Engineering Journey to Developing a Concept of Operations

Starting in 2017, NASA’s Human Research Program (HRP) Exploration Medical Capability (ExMC) element began a systems engineering transition from traditional, document-centric development to model-centric development when defining its foundation medical systems. These foundation medical systems define a Concept of Operations (ConOps) and identify the generic requirements for a medical system based on assumptions about a generic crew and mission environments and guidance from NASA standards (e.g., Medical “Levels of Care”). By making the transition, ExMC intends to improve communication among stakeholders about foundation medical system requirements and content. In addition, this transition will enable ExMC to lower both development and crew treatment risks for future, mission-specific medical systems. ExMC followed a Model Based Systems Engineering (MBSE) paradigm when developing the foundation medical systems. A model-based approach provides several advantages over a traditional, document-centric approach. First, when Systems Engineers (SE) develop diagrams in a model using a standard modeling language, they produce information dense pictures that facilitate understanding much more efficiently with less room for misinterpretation than text. Second, due to the evolving nature of projects, documentation becomes out of date the minute it is published. This can result in people making decisions based on information that is no longer current, especially if they are referencing a locally-stored copy of a document. A model, on the other hand, is always up to date with the latest approved changes and information. It serves as a single point of truth. Third, a model-centric approach centralizes all important information in one place. Rather than having to flip through separate ConOps documents, design specifications, requirements specifications, and the like to coordinate information, a model captures the content in one, integrated spot. This integration makes tracing information from end-to-end easier with greater reliability. The ExMC Systems Engineering Lifecycle follows a well-defined process. ExMC Systems Engineers perform all major steps of the process, regardless of the development methodology. One of the first steps in the process is developing the ConOps that describes the operation of the system from the point of view of the users. It includes a list of the users and their needs, the goals of the medical system, key assumptions about the system, and definitions of the medical system’s operational environments. For this development effort, ExMC chose to replace the traditional text-based ConOps document with a model. While the decision to change the development workflow was not difficult, implementing the structural and organizational workflows were. It required showing ExMC’s users, most of whom are not Systems Engineers, how the information they require would be presented in the model and to gain their acceptance of this approach. This paper documents key lessons learned during the ConOps transformation by focusing on how the model represents information, the agile workflow used by SEs when developing the model and how it integrates into a project plan, how leadership influenced key users to accept the transformation, and how the users interact with the model information.

Jeffrey Robert Cohen↗

Cross-Validation of Computational and Experimental Distributed Surface Pressures on the Space Launch System

This paper presents a new workflow for comparing experimental pressure-sensitive paint (PSP) data to computational fluid dynamic (CFD) simulations by way of mapping data from corresponding grids utilizing interpolation methods. In addition to generating quantitative and qualitative point-to-point comparisons between PSP and CFD data, this workflow extracts sectional loading data from both grids and generates lineload comparison charts for corresponding PSP and CFD runs. Experimental PSP data presented in this paper were taken from a 2016 NASA Ames Research Center Unitary Plan Wind Tunnel 11- by 11-Foot Transonic WindTunnel Facility test of the NASA Space Launch System. CFD simulation data for comparison purposes were generated using the FUN3D code. Overall, interpolation onto PSP grids versus CFD grids yields comparable surface pressure fields. However, lineload comparisons are easier to make on the CFD grid-mapped data due to the grid topology and the current capabilities of the lineload analysis tools at NASA Langley Research Center. This workflow is written using contemporary software (Python, Tecplot, PyTecplot), is compatible with existing tools at NASA Langley, and is developed to be adaptable depending on the situation.

SLS↗

Broadening JPL’s Mission Formulation Paradigm with Human Centered Design

Within NASA’s highly competitive environment for funding, Human Centered Design (HCD) and cybernetics could provide advantages to proposers during the mission formulation phase. Opportunities are limited when it comes to funding new science missions. Proposers are challenged to make a compelling case about the scientific desirability, technical feasibility, and resource viability of their concepts. Organizations follow established processes for proposal development using teams that typically include scientists, engineers, and managers. These team members are highly experienced subject matter experts (SME) in their own disciplines, and can respond to requirements from the solicitation. However, they are typically not trained as designers and communicators. Their approach is rooted within NASA’s science and technology paradigm. How can we improve the proposal development process, refine workflow between team members, and deliver clear and appealing offerings to the stakeholders and evaluators? These questions have been addressed by today’s most innovative companies (e.g., Apple, Google, 3M, Dyson), where the design process is not limited simply to engineering and management, but involves an all-encompassing approach drawing from fields such as social sciences, design, and the arts. Like these commercial enterprises, NASA currently employs systems thinking and integrated design, but can benefit further by moving beyond its current practices, which are mostly driven by rigid engineering, technology, science, and project management considerations. At JPL’s Innovation Foundry and through the Solar System Mission Formulation Office, we broadened this paradigm by including HCD in the mission formulation workflow. Our goal was to create a proposal with improved clarity and appeal, thus helping our team to communicate its message and aid evaluators with their work. In this paper we provide examples and lessons learned from our recent proposal development effort using HCD. We discuss touch points where we infused non-linear designerly approaches and cybernetic circularity into the workflow. Implemented design topics include operational design for team building; process design throughout distinct phases of the proposal development and writing process; communication design for streamlined exchange of information within the team and to stakeholders; interaction design; graphic design; and creating boundary objects. While these approaches may feel new or foreign to SMEs and managers in the aerospace community, they produced significant benefits in this mission formulation effort. We will describe how such approaches can be used to broaden NASA’s technology-driven paradigm through design, thus creating an environment which fosters innovation, improved communication, and strategic advantage for proposers and their organizations.

Turner, Neal↗

NASA GeneLab RNASeq Consensus Pipeline: A Nextflow Implementation

The NASA GeneLab project (genelab.nasa.gov) seeks to accelerate space biology research through cataloging and democratizing omics data. Since raw omics data is largely inaccessible to non-bioinformaticians, GeneLab works with the scientific community to develop standard processing pipelines to generate and publish processed data. Unlike raw data, processed data has greater immediate value to a wide range of users with varying technical backgrounds and computational capabilities. Standardizing processing workflows is essential to match the pace of raw data generation, ensure reproducibility, and enable standardized processed data for comparison across datasets. Previously, GeneLab developed a standardized pipeline for processing RNAseq data, referred to as the ‘GeneLab RNAseq Consensus Pipeline (RCP)’, in collaboration with GeneLab’s Analysis Working Groups. The work presented here is a Nextflow implementation of GeneLab’s RCP that automates and accelerates data processing of RNASeq datasets hosted on GeneLab. In addition to the core data processing, the workflow also includes staging of GeneLab raw data and a robust verification and validation (V&V) program that runs after each processing step to identify errors in real-time, stop additional downstream computation, and preserve computational resources. The workflow, including the staging and V&V functionality, is open source for others to reuse and modify at https://github.com/nasa/GeneLab_Data_Processing/tree/master/RNAseq.

Jonathan Dejesus Oribello↗

NASA GeneLab RNASeq Consensus Pipeline: A Nextflow Implementation

The NASA GeneLab project (genelab.nasa.gov) seeks to accelerate space biology research through cataloging and democratizing omics data. Since raw omics data is largely inaccessible to non-bioinformaticians, GeneLab works with the scientific community to develop standard processing pipelines to generate and publish processed data. Unlike raw data, processed data has greater immediate value to a wide range of users with varying technical backgrounds and computational capabilities. Standardizing processing workflows is essential to match the pace of raw data generation, ensure reproducibility, and enable standardized processed data for comparison across datasets. Previously, GeneLab developed a standardized pipeline for processing RNAseq data, referred to as the ‘GeneLab RNAseq Consensus Pipeline (RCP)’, in collaboration with GeneLab’s Analysis Working Groups. The work presented here is a Nextflow implementation of GeneLab’s RCP that automates and accelerates data processing of RNASeq datasets hosted on GeneLab. In addition to the core data processing, the workflow also includes staging of GeneLab raw data and a robust verification and validation (V&V) program that runs after each processing step to identify errors in real-time, stop additional downstream computation, and preserve computational resources. The workflow, including the staging and V&V functionality, is open source for others to reuse and modify at https://github.com/nasa/GeneLab_Data_Processing/tree/master/RNAseq.

Jonathan D Oribello↗

Developing a Vision for Maturing the Heliophysics Infrastructure towards Open Science: The DIARieS Analysis Ecosystem

In the dawn of open science and the upcoming requirements, we speak about the existing state of Heliophysics infrastructure and detail the evolution required to address capability or interconnection shortcomings. Such a daunting barrier calls for an analysis ecosystem with multi-faceted capability. We propose such an ecosystem, called DIARieS, to be built upon five conceptual pillars: Discovery, Implementation, Analysis, Reproducibility, and Sharing of results. The combination of these concepts in a single platform will enable users to more intuitively combine recent advances in technology to create ‘DIARieS’ of their workflows, which can be easily made open to others in the community. The DIARieS ecosystem will also increase our efficiency by streamlining our various workflow processes, including automatic incorporation of the impending requirements of open science. The various components of the ecosystem will simplify software installation and data implementation, including automatically generated citation lists based on the components included. Automatic containerization and version control of the ecosystem will make the custom workflows easily reproducible. Employing widget technology will ease the difficulty of producing publication and commercial quality visualizations and applying common analyses techniques. Incorporating multiple technologies will streamline the various sharing methods common in our work environments today. Overall, the totality of capabilities to be offered by this analysis ecosystem will drastically simplify the application of open science principles to our work in addition to improving our efficiency and ease of collaboration. This talk summarizes a vision of the proposed ecosystem, which is described in more detail in Ringuette et al. (2022: https://doi.org/10.1016/j.asr.2022.05.012).

infrastructure↗

Kamodo’s Satellite Constellation Mission Planning Tool

Kamodo provides a functional model-agnostic interface to a growing collection of Heliophysics model outputs. The CCMC, in collaboration with the Geospace Dynamics Constellation Science Team, has recently developed Kamodo’s satellite constellation mission planning tool to perform reconstructions in any pair of dimensions, including time. The ‘reconstruction’ tool enables users to fly any 4-dimensional grid of satellites through a given model data set, reconstructing what the given constellation would observe during the mission. This capability facilitates determination of what satellite configuration is best for a given science question, even allowing comparison across multiple models. This tool, written in Python, is built upon Kamodo’s flythrough tool, which in turn depends on a growing network of model-specific interfaces. Since each model interface is designed with model-agnostic syntax, the flythrough tool and the satellite constellation mission planning tool also feature model-agnostic syntax. In this work, we will describe the basic analysis choices available in the tool and provide a variety of sample workflows. The tool is freely available at https://github.com/nasa/Kamodo for the public. We invite the community to use the reconstruction tool and adapt the provided workflows for their mission planning, and to contribute their own workflows to share with others.

software↗