Search NASA⌕ Search

SEARCH · Search NASA

Results for “Software pipelines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

PSTN-019: The LSST Science Pipelines Software: Optical Survey Pipeline Reduction and Analysis Environment

The NSF-DOE Vera C. Rubin Observatory is executing the Legacy Survey of Space and Time (LSST) as its prime mission, producing a series of data releases over the ten-year survey. The LSST Science Pipelines Software will be used to create these data releases and to perform the nightly prompt processing and alert production. This paper provides an overview of the LSST Science Pipelines Software, describing the components and their integration into pipelines that generate science-ready data products.

79 ASTRONOMY AND ASTROPHYSICS↗

Disclosure of the XRD pipeline software

The XRD pipeline software is a program to reduce 2D X-ray diffraction data from area detectors with advanced algorithms on automatic masking and image segmentation, which help characterization of multiple phases recorded in the data and facilitate data analysis.SF-25-114

XU, WENQIAN [Argonne National Laboratory (ANL), Ar↗

VA EDH Advanced Software Pipeline Framework Report: Enhancing Automation and Scalability

The VA Environmental Determinants of Health (EDH) Advanced Software Pipeline Framework is designed to enhance the efficiency, scalability, and security of geospatial data processing workflows. This framework integrates modern data orchestration and containerization technologies, including Prefect for workflow automation, Docker for containerization, and PostgreSQL/PostGIS for geospatial data storage and analysis. It ensures standardized, reproducible, and automated data processing, supporting VA objectives related to substance use risk assessment and recovery research. The pipeline addresses key scalability and performance challenges through horizontal and vertical scaling, high-performance computing (HPC) integration, parallel processing, task caching, and dynamic resource allocation. These optimizations improve throughput and reduce latency, allowing the system to efficiently manage large and complex datasets. Additionally, security and compliance measures—such as data encryption (SSL), Role-Based Access Control (RBAC), and adherence to GDPR and HIPAA standards—safeguard sensitive information throughout data transmission and storage. A key implementation of this framework includes the automation of shelter list geolocation workflows, ensuring that up-to-date data is readily available for VA decision-making. Lessons learned from this project include the transition from in-memory processing to incremental storage writes, improving resource management and reliability. Future enhancements aim to expand automation, integrate AI-driven anomaly detection, and incorporate high-performance computing resources. This framework provides a scalable, secure, and adaptable solution for managing geospatial datasets, reinforcing the VA’s ability to support clinical and strategic initiatives through data-driven decision-making.

97 MATHEMATICS AND COMPUTING↗

An Algorithmic and Software Pipeline for Very Large Scale Scientific Data Compression with Error Guarantees

Efficient data compression is becoming increasingly critical for storing scientific data because many scientific applications produce vast amounts of data. This paper presents an end-to-end algorithmic and software pipeline for data compression that guarantees both error bounds on primary data (PD) and derived data, known as Quantities of Interest (QoI).We demonstrate the effectiveness of the pipeline by compressing fusion data generated by a large-scale fusion code, XGC, which produces tens of petabytes of data in a single day. We demonstrate that the compression is conducted by setting aside computational resources known as staging nodes, and does not impact the simulation performance. For efficient parallel I/O, the pipeline uses ADIOS2, which many codes such as XGC already use for their parallel I/O. We show that our approach can compress the data by two orders of magnitude while guaranteeing high accuracy on both the PD and the QoIs. Further, the amount of resources required by compression is a few percent of the resources required by simulation while ensuring that the compression time for each stage is less than the corresponding simulation time.This pipeline consists of three main steps. The first step decomposes the data using domain decomposition into small subdomains. Each subdomain is then compressed independently to achieve a high level of parallelism. The second step uses existing techniques that guarantee error bounds on the primary data for each subdomain. The third step uses a post-processing optimization technique based on Lagrange multipliers to reduce the QoI errors for data corresponding to each subdomain. The Lagrange multipliers generated can be further quantized or truncated to increase the compression level. All of the above characteristics of our approach make it highly practical to apply on-the-fly compression while guaranteeing errors on QoIs that are critical to the scientists.

Banerjee, Tania↗

VA-Environmental Determinants of Health’s Software Pipeline Framework

In order to characterize US communities in terms of social and environmental determinants of health (EDH) that affect veterans, the Social and Environmental Determinants of Health (SEDH) project, a collaboration between the Veterans Affairs (VA) Social Department and the Oak Ridge National Laboratory (ORNL), aims to identify, extract, and prepare datasets that include community variables known to be associated with health outcomes in veterans. The EDH project's responsibilities include investigating SEDH data, providing technical documentation for SEDH datasets, and automating the download, extraction, and preparation of the data through the creation of a software pipeline framework.

54 ENVIRONMENTAL SCIENCES↗

PyPop: a mature open-source software pipeline for population genomics

Python for Population Genomics (PyPop) is a software package that processes genotype and allele data and performs large-scale population genetic analyses on highly polymorphic multi-locus genotype data. In particular, PyPop tests data conformity to Hardy-Weinberg equilibrium expectations, performs Ewens-Watterson tests for selection, estimates haplotype frequencies, measures linkage disequilibrium, and tests significance. Standardized means of performing these tests is key for contemporary studies of evolutionary biology and population genetics, and these tests are central to genetic studies of disease association as well. Here, we present PyPop 1.0.0, a new major release of the package, which implements new features using the more robust infrastructure of GitHub, and is distributed via the industry-standard Python Package Index. New features include implementation of the asymmetric linkage disequilibrium measures and, of particular interest to the immunogenetics research communities, support for modern nomenclature, including colon-delimited allele names, and improvements to meta-analysis features for aggregating outputs for multiple populations.

59 BASIC BIOLOGICAL SCIENCES↗

Toward an AI-Powered Software Pipeline for Real-Time Tracking and Analysis of Wildfire and Smoke

Real-time tracking of wildfires and smoke is crucial for effective response, minimizing damage, protecting lives, and efficiently managing resources during fire emergencies. We develop a web-based AI-powered pipeline that detects wildfires in aerial video and estimates deployment-relevant behavior metrics, including cumulative burned area, burned-area growth rate, fire spread direction, and smoke dispersion. The system combines a YOLO-based detector with YCbCr-based fire segmentation, HSV-based smoke segmentation, Farneback optical flow, and centroid-based spatiotemporal tracking. Using ground sampling distance (GSD), pixel-level fire masks are converted to physical burned-area measurements by correlating fire pixel counts with camera altitude and tilt angle. We benchmark YOLO variants and non-YOLO baselines (GoogLeNet, CNN, DBN, Autoencoder, U-Net, and AlexNet) on the IEEE FLAME dataset and a newly created aerial frame dataset, Wildfire-DB. Cross-dataset evaluation uses a strict threshold-transfer protocol: decision thresholds are selected on FLAME validation and transferred unchanged to Wildfire-DB to quantify generalization under domain shift. YOLOv6 achieves the strongest cross-dataset frame-level fire detection on Wildfire-DB (ROC-AUC 0.8200, PR-AUC 0.8044, and transferred-threshold F1 0.7596). For tracking-oriented deployment requiring oriented localization, YOLO11-OBB provides the most reliable cross-dataset behavior among OBB-capable models while remaining computationally feasible. To analyze the feasibility of UAV deployment, we further measure inference efficiency using synchronized GPU and CPU power logs on a fixed workload of 1569 frames. YOLO-family models process the video in 5.73–12.47 seconds with net energy of 1247.28–1775.39 J, substantially lower latency and energy than heavier classification and reconstruction baselines. Overall, model optimality depends on operational objectives: YOLOv6 is best for cross-dataset detection robustness, whereas YOL...

Color segmentation↗

ADEPT: a domain independent sequence alignment strategy for gpu architectures

Bioinformatic workflows frequently make use of automated genome assembly and protein clustering tools. At the core of most of these tools, a significant portion of execution time is spent in determining optimal local alignment between two sequences. This task is performed with the Smith-Waterman algorithm, which is a dynamic programming based method. With the advent of modern sequencing technologies and increasing size of both genome and protein databases, a need for faster Smith-Waterman implementations has emerged. Multiple SIMD strategies for the Smith-Waterman algorithm are available for CPUs. However, with the move of HPC facilities towards accelerator based architectures, a need for an efficient GPU accelerated strategy has emerged. Existing GPU based strategies have either been optimized for a specific type of characters (Nucleotides or Amino Acids) or for only a handful of application use-cases. In this paper, we present ADEPT, a new sequence alignment strategy for GPU architectures that is domain independent, supporting alignment of sequences from both genomes and proteins. Our proposed strategy uses GPU specific optimizations that do not rely on the nature of sequence. We demonstrate the feasibility of this strategy by implementing the Smith-Waterman algorithm and comparing it to similar CPU strategies as well as the fastest known GPU methods for each domain. ADEPT’s driver enables it to scale across multiple GPUs and allows easy integration into software pipelines which utilize large scale computational systems. We have shown that the ADEPT based Smith-Waterman algorithm demonstrates a peak performance of 360 GCUPS and 497 GCUPs for protein based and DNA based datasets respectively on a single GPU node (8 GPUs) of the Cori Supercomputer. Overall ADEPT shows 10x faster performance in a node-to-node comparison against a corresponding SIMD CPU implementation. ADEPT demonstrates a performance that is either comparable or better than existing GPU strategies. We demonstrated the efficacy of ADEPT in supporting existing bionformatics software pipelines by integrating ADEPT in MetaHipMer a high-performance denovo metagenome assembler and PASTIS a high-performance protein similarity graph construction pipeline. Our results show 10% and 30% boost of performance in MetaHipMer and PASTIS respectively.

59 BASIC BIOLOGICAL SCIENCES↗

Automating the Analysis of Large Language Models Responses through Zero-Shot Question Answering

Recent advancements in Large Language Models (LLMs) have shown significant potential in various applications, yet their evaluation, particularly in zero-shot question answering scenarios, remains a challenging task. In this study, our objective was to explore precision metrics for Large Language Models (LLM) and design and implement a software pipeline to automatically evaluate LLMs' outputs under zero-shot question answering. Zero-shot question answering involves a model providing answers to questions about topics it hasn't seen during training. It leverages the principles of zero-shot learning by relying on semantic understanding and generalization from related knowledge. The data used was metadata from medical databases on congenital heart disease. We explored eleven LLM metrics and selected three for our evaluation: BLEU, BERTScore, and MoverScore. BLEU calculates a score based on the overlap of n-grams (contiguous sequences of n items, typically words) between the machine-generated translation and the reference translations. Higher BLEU scores indicate better correspondence between the machine-generated and human-generated translations. BERTScore is a metric used to evaluate the quality of machine-generated text by measuring the similarity of token embeddings produced by BERT (Bidirectional Encoder Representations from Transformers) between the generated text and reference text. MoverScore is a metric that quantifies the dissimilarity between the distributions of word embeddings from machine-generated text and reference text, emphasizing semantic similarity over exact token overlap. We also introduced HBKI, a composite metric summarizing these approaches. We tested five models —GPT-3, Llama-2, Gemini 1.5 Pro, Solar 10.7B, and Mixtral-8x7b. Our software pipeline, designed and implemented using Object-Oriented Programming principles, allows users to customize the selection and extraction of features for topics of interest in their own research. Our results show that MoverScore delivered the most precise evaluation of the LLM's outputs, while Mixtral-8x7b achieved the best overall performance in extracting metadata from the databases.

97 MATHEMATICS AND COMPUTING↗

System and methods for hardware-software cooperative pipeline error detection

An error reporting system utilizes a parity checker to receive data results from execution of an original instruction and a parity bit for the data. A decoder receives an error correcting code (ECC) for data resulting from execution of a shadow instruction of the original instruction, and data error correction is initiated on the original instruction result on condition of a mismatch between the parity bit and the original instruction result, and the decoder asserting a correctable error in the original instruction result.

97 MATHEMATICS AND COMPUTING↗

Numerical Codes for the DESC-LSST Analysis Pipeline: Core Cosmology Library Standard Modules and Beyond wCDM Modules (Final Technical Report)

The overall objective of the project is to investigate and develop specific software modules and analysis components for the software pipeline of LSST Dark Energy Science Collaboration (DESC). Following the key projects of DESC Science Roadmap (SRM), we will write and test computer codes for the Core Cosmology Library (CCL) in order to complete its modules, functionalities, and interface to work with the analysis pipelines from the five science probes of DESC (parts of SRM deliverables CX4.2TJP, CX6.2CS). We will also code modules for CCL to test models beyond w-Cold-Dark-Matter (wCDM) and modification to gravity (MG). Interfaces for MG models will also be developed for the TJPCOSMO software which is the main pipeline of the Theory and Joint Probe (TJP) working group (deliverable TJP2.3). In order to use the full power of LSST data to constrain MG models, we will also work on constraints from nonlinear regime by running and analyzing MG N-Body simulations using a Parameterized-Post-Friedmann framework into the Gadget-2 simulation package (deliverable TJP2.2). Preliminary results for the simulations were obtained in the past. In collaboration with other DESC groups, we plan to make these simulations feedable to cosmic emulators that are practical for likelihood analyzes (deliverable TJP2.2, parts of CX6.2CS). We will also modify and integrate our current codes for consistency tests between data sets and probes into the pipeline (parts of deliverables CX8.2TJP, TJP2.3). We understand that other groups will contribute to some of these objectives but our team will focus and collaborate with others on the particular part of testing MG and models beyond wCDM and refine the DESC pipeline for this purpose. PI has been coordinating his work with the TJP and CS working groups and the DESC management team. PI is a full member of DESC since June 2013. He and his students have been contributing to LSST-DESC activities and work including TJP telecons, collaboration meetings, hack-weeks, and workshops. PI chaired or co-chaired sessions at collaboration meetings and hack-weeks about testing gravity and models beyond wCDM using LSST. He is coordinating the TJP2 projects for testing models beyond wCDM including the writing of DESC-research-note, development of code for pipeline, and N-Body simulations for MG and beyond wCDM models testable with LSST analyses. As stressed in the DESC white paper, SRM, and P5 report, one of the important questions in understanding cosmic acceleration and dark energy is to be able to distinguish whether the acceleration is due to a dark energy component in the universe or a modification to gravity. Answering these questions will have a significant impact on the question of cosmic acceleration and dark energy. The methods that we will use include analytical work, numerical code, and N-Body simulations. A first approach that that we will use consists of using growth rate parameters that enter the perturbed dynamics equations. These parameters take distinctive values for distinct gravity theories and have potential to distinguish between Dark Energy and Modified Gravity. The second method is to look for inconsistencies in Dark Energy parameter spaces using specific combinations of cosmological data sets. Our investigation addresses the Dark Energy problem that is relevant to the mission of the HEP program to understand how our universe works at its most fundamental level. It will allow us to make progress on the HEP mission to explore the nature of Dark Energy and the basic nature of space and time using future surveys such as LSST. The investigation supports the DOE HEP program Cosmic Frontier as it will contribute to the study and understanding of dark energy and fundamental properties of the universe. The investigation contributes directly to LSST-DESC key projects and their deliverables as described in the Science Road-map document to build analysis pipeline and to test dark energy and beyond wCDM models using LSST.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

EPICS Deployment at Fermilab

Fermilab has traditionally not been an EPICS house, as such expertise in EPICS is limited and scattered. However, PIP-II will be using EPICS for its control system. Furthermore, when PIP-II is operating, it must to interface with the existing, though modernized (see ACORN) legacy control system. We have developed and deployed a software pipeline that addresses these needs and presents to developers a tested and robust software framework, including template IOCs from which new developers can quickly gain experience. In this presentation, we will discuss the motivation for this work, the implementation of a continuous integration/continuous deployment pipeline, testing, template IOCs, and the deployment of user applications. We will also discuss how this is used with the current PIP-II teststand and lessons learned.

43 PARTICLE ACCELERATORS↗

Laue-DIALS: Open-source software for polychromatic x-ray diffraction data

Most x-ray sources are inherently polychromatic. Polychromatic (“pink”) x-rays provide an efficient way to conduct diffraction experiments as many more photons can be used and large regions of reciprocal space can be probed without sample rotation during exposure—ideal conditions for time-resolved applications. Analysis of such data is complicated, however, causing most x-ray facilities to discard >99% of x-ray photons to obtain monochromatic data. Key challenges in analyzing polychromatic diffraction data include lattice searching, indexing and wavelength assignment, correction of measured intensities for wavelength-dependent effects, and deconvolution of harmonics. We recently described an algorithm, Careless, that can perform harmonic deconvolution and correct measured intensities for variation in wavelength when presented with integrated diffraction intensities and assigned wavelengths. Here, we present Laue-DIALS, an open-source software pipeline that indexes and integrates polychromatic diffraction data. Laue-DIALS is based on the dxtbx toolbox, which supports the DIALS software commonly used to process monochromatic data. As such, Laue-DIALS provides many of the same advantages: an open-source, modular, and extensible architecture, providing a robust basis for future development. We present benchmark results showing that Laue-DIALS, together with Careless, provides a suitable approach to the analysis of polychromatic diffraction data, including for time-resolved applications.

97 MATHEMATICS AND COMPUTING↗

Towards AI-assisted neutrino flavor theory design

Particle physics theories, such as those which explain neutrino flavor mixing, arise from a vast landscape of model-building possibilities. A model’s construction typically relies on the intuition of theorists. It also requires considerable effort to identify appropriate symmetry groups, assign field representations, and extract predictions for comparison with experimental data. We develop Autonomous Model Builder (AMBer), a framework in which a reinforcement learning agent interacts with a streamlined physics software pipeline to search these spaces efficiently. AMBer selects symmetry groups, particle content, and group representation assignments to construct models while minimizing the number of free parameters introduced. We validate our approach in well-studied regions of theory space and extend the exploration to a previously unexamined symmetry group. While demonstrated in the context of neutrino flavor theories, this approach of reinforcement learning with physics software feedback may be extended to other theoretical model-building problems in the future.

Baretz, Jason Benjamin↗

LDM-151: Data Management Science Pipelines Design

The LSST Science Requirements Document (the LSST SRD) specifies a set of data product guidelines, designed to support science goals envisioned to be enabled by the LSST observing program. Following these guidelines, the details of these data products have been described in the LSST Data Products Definition Document (DPDD), and captured in a formal flow-down from the SRD via the LSST System Requirements (LSR), Observatory System Specifications (OSS), to the Data Management System Requirements (DMSR). The LSST Data Management subsystem's responsibilities include the design, implementation, deployment and execution of software pipelines necessary to generate these data products. This document describes the design of the scientific aspects of those pipelines.

79 ASTRONOMY AND ASTROPHYSICS↗

Investigating Scientific Data Change with User Research Methods

Scientific datasets are continually expanding and changing due to fluctuations with instruments, quality assessment and quality control processes, and modifications to software pipelines. Datasets include minimal information about these changes or their effects requiring scientists manually assess modifications through a number of labor intensive and ad-hoc steps. The Deduce project is investigating data change to develop metrics, methods, and tools that will help scientists systematically identify and make decisions around data changes. Currently, there is a lack of understanding, and common practices, for identifying and evaluating changes in datasets since systematically measuring and managing data change is under explored in scientific work. We are conducting user research to address this need by exploring scientist's conceptualizations, behaviors, needs, and motivations when dealing with changing datasets. Our user research utilizes multiple methods to produce foundational, generative insights and evaluate research products produced by our team. In this paper, we detail our user research process and outline our findings about data change that emerge from our studies. Our work illustrates how scientific software teams can push beyond just usability testing user interfaces or tools to better probe the underlying ideas they are developing solutions to address.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

The Eighteenth Data Release of the Sloan Digital Sky Surveys: Targeting and First Spectra from SDSS-V

The eighteenth data release (DR18) of the Sloan Digital Sky Survey (SDSS) is the first one for SDSS-V, the fifth generation of the survey. SDSS-V comprises three primary scientific programs or “Mappers”: the Milky Way Mapper (MWM), the Black Hole Mapper (BHM), and the Local Volume Mapper. This data release contains extensive targeting information for the two multiobject spectroscopy programs (MWM and BHM), including input catalogs and selection functions for their numerous scientific objectives. We describe the production of the targeting databases and their calibration and scientifically focused components. DR18 also includes ~25,000 new SDSS spectra and supplemental information for X-ray sources identified by eROSITA in its eFEDS field. We present updates to some of the SDSS software pipelines and preview changes anticipated for DR19. We also describe three value-added catalogs (VACs) based on SDSS-IV data that have been published since DR17, and one VAC based on the SDSS-V data in the eFEDS field.

79 ASTRONOMY AND ASTROPHYSICS↗