Search NASA⌕ Search

SEARCH · Search NASA

Results for “data pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

AEPF: Attention-Enabled Point Fusion for 3D Object Detection

Current state-of-the-art (SOTA) LiDAR-only detectors perform well for 3D object detection tasks, but point cloud data are typically sparse and lacks semantic information. Detailed semantic information obtained from camera images can be added with existing LiDAR-based detectors to create a robust 3D detection pipeline. With two different data types, a major challenge in developing multi-modal sensor fusion networks is to achieve effective data fusion while managing computational resources. With separate 2D and 3D feature extraction backbones, feature fusion can become more challenging as these modes generate different gradients, leading to gradient conflicts and suboptimal convergence during network optimization. To this end, we propose a 3D object detection method, Attention-Enabled Point Fusion (AEPF). AEPF uses images and voxelized point cloud data as inputs and estimates the 3D bounding boxes of object locations as outputs. An attention mechanism is introduced to an existing feature fusion strategy to improve 3D detection accuracy and two variants are proposed. These two variants, AEPF-Small and AEPF-Large, address different needs. AEPF-Small, with a lightweight attention module and fewer parameters, offers fast inference. AEPF-Large, with a more complex attention module and increased parameters, provides higher accuracy than baseline models. Experimental results on the KITTI validation set show that AEPF-Small maintains SOTA 3D detection accuracy while inferencing at higher speeds. AEPF-Large achieves mean average precision scores of 91.13, 79.06, and 76.15 for the car class’s easy, medium, and hard targets, respectively, in the KITTI validation set. Results from ablation experiments are also presented to support the choice of model architecture.

Chemistry↗

Using active learning to improve quasar identification for the DESI spectra processing pipeline

The Dark Energy Spectroscopic Instrument (DESI) survey uses an automatic spectral classification pipeline to classify spectra. QuasarNET is a convolutional neural network used as part of this pipeline originally trained using data from the Baryon Oscillation Spectroscopic Survey (BOSS). In this paper we implement an active learning algorithm to optimally select spectra to use for training a new version of the QuasarNET weights file using only DESI data, with the goal of improving classification accuracy. This active learning algorithm includes a novel outlier rejection step using a Self-Organizing Map to ensure we label spectra representative of the larger quasar sample observed in DESI. We perform two iterations of the active learning pipeline, assembling a final dataset of 5600 labeled spectra, a small subset of the approximately 1.3 million quasar targets in DESI's Data Release 1. When splitting the spectra into training and validation subsets we achieve similar performance to the previously trained weights file in completeness and purity calculated on the validation dataset but do so with less than one tenth of the amount of training data. The new weights also more consistently classify objects in the same way when used on unlabeled data compared to the old weights file. In the process of improving QuasarNET's classification accuracy we discovered a systemic error in QuasarNET's redshift estimation and used our findings to improve our understanding of QuasarNET's redshifts.

Machine learning↗

MVP: a modular viromics pipeline to identify, filter, cluster, annotate, and bin viruses from metagenomes

While numerous computational frameworks and workflows are available for recovering prokaryote and eukaryote genomes from metagenome data, only a limited number of pipelines are designed specifically for viromics analysis. With many viromics tools developed in the last few years alone, it can be challenging for scientists with limited bioinformatics experience to easily recover, evaluate quality, annotate genes, dereplicate, assign taxonomy, and calculate relative abundance and coverage of viral genomes using state-of-the-art methods and standards. Here, we describe Modular Viromics Pipeline (MVP) v.1.0, a user-friendly pipeline written in Python and providing a simple framework to perform standard viromics analyses. MVP combines multiple tools to enable viral genome identification, characterization of genome quality, filtering, clustering, taxonomic and functional annotation, genome binning, and comprehensive summaries of results that can be used for downstream ecological analyses. Overall, MVP provides a standardized and reproducible pipeline for both extensive and robust characterization of viruses from large-scale sequencing data including metagenomes, metatranscriptomes, viromes, and isolate genomes. As a typical use case, we show how the entire MVP pipeline can be applied to a set of 20 metagenomes from wetland sediments using only 10 modules executed via command lines, leading to the identification of 11,656 viral contigs and 8,145 viral operational taxonomic units (vOTUs) displaying a clear beta-diversity pattern. Further, acting as a dynamic wrapper, MVP is designed to continuously incorporate updates and integrate new tools, ensuring its ongoing relevance in the rapidly evolving field of viromics. MVP is available at https://gitlab.com/ccoclet/mvp and as versioned packages in PyPi and Conda.

59 BASIC BIOLOGICAL SCIENCES↗

Data Agnostic Feature-Target Analysis & Ranking Machine Learning Pipeline (DAFTAR-ML) v0.1.0

DAFTAR-ML is a specialized machine-learning pipeline that identifies relevant features based on their relationship to a target variable. Many ML pipelines focus solely on prediction, and feature ranking is often absent or lacks robust statistical methods. DAFTAR-ML performs its tasks with this outcome in mind. Model training is robust, using nested cross-validation and hyperparameter tuning. Instead of relying on native feature-importance scores, it employs SHAP (SHapley Additive exPlanations) to quantify feature importance. The pipeline also produces comprehensive results, including publication-quality visualizations.

Melie, Tina [Lawrence Berkeley National Laboratory↗

Atacama Cosmology Telescope: DR6 gravitational lensing and SDSS BOSS cross-correlation measurement and constraints on gravity with the 𝐸 𝐺 statistic

We derive new constraints on the 𝐸 𝐺 statistic as a test of gravity, combining the cosmic microwave background (CMB) lensing map estimated from Data Release 6 (DR6) of the Atacama Cosmology Telescope with Sloan Digital Sky Survey III Baryon Oscillation Spectroscopic Survey (SDSS BOSS) CMASS and LOWZ galaxy data. We develop an analysis pipeline to measure the cross-correlation between CMB lensing maps and galaxy data, following a blinding policy and testing the approach through null and consistency checks. By testing the equivalence of the spatial and temporal gravitational potentials, the 𝐸 𝐺 statistic can distinguish Λ⁢ CDM from alternative models of gravity. We find 𝐸 𝐺 ⁡(𝑧 eff = 0.555) = 0.3⁢1$^{+0.06}_{−0.05}$ for Atacama Cosmology Telescope (ACT) and CMASS data at 68.28% confidence level, and 𝐸 𝐺 ⁡(𝑧 eff = 0.316) = 0.4⁢9$^{+0.14}_{−0.11}$ for the ACT and LOWZ. Systematic errors are estimated to be 3% and 4%, respectively. Including CMB lensing information from Planck PR4 results in 𝐸 𝐺 ⁡(𝑧 eff = 0.555) = 0.3⁢4$^{+0.05}_{−0.05}$ with CMASS and 𝐸 𝐺 ⁡(𝑧 eff = 0.316) = 0.4⁢3$^{+0.11}_{−0.09}$ with LOWZ. These are consistent with predictions for the Λ⁢ CDM model that best fits the Planck CMB anisotropy and SDSS BOSS baryon acoustic oscillations (BAO), where 𝐸$^{GR}_{𝐺⁡}$(𝑧 eff =0.555) =0.401 ± 0.005 for CMB lensing combined with CMASS and 𝐸$^{GR}_{𝐺}$⁡(𝑧 eff = 0.316) = 0.452 ± 0.005 combined with LOWZ. We also find 𝐸 𝐺 to be scale independent, with probability to exceed >5%, as predicted by general relativity. The methods developed in this work are also applicable to improved future analyses with upcoming spectroscopic galaxy samples and CMB lensing measurements.

79 ASTRONOMY AND ASTROPHYSICS↗

Proteomics Analysis of Human Contaminant Proteins

Complete characterization of unknowns via proteomics remains challenging. There exist regions of mass spectrometry-based proteomics data where empirical measurements are not attributed to peptides, and/or sequenced peptides from mass spectra are not attributed to any source. These uncharacterized regions are known as the “dark” proteome. Many proteomics tools rely on some a priori knowledge of sample composition; few tools allow for investigation of unknowns without relying on composition assumptions. Further, the potential low abundance of minor traces in these uncharacterized regions can make elucidation of the “dark” proteome challenging. Herein, we describe the development and evaluation of approaches to study the “dark” proteome and move towards an untargeted approach for more complete characterization, namely by studying minor human protein traces in non-human samples and combining that approach with non-human source organism identification without relying on assumptions. Human protein markers, in the form of genetically variant peptides, have been extensively examined in a variety of human matrices, including blood, plasma, and hair, but have yet to be investigated in non-human samples, such as cell cultures, as human contaminant traces. Genetically variant peptides are those that are found in proteins carrying single nucleotide polymorphisms. In this work, we aimed to (1) investigate the feasibility of detecting human contaminant genetically variant peptides (GVPs) in a diverse set of non-human organisms using public proteomics data and a computational pipeline, as well as to (2) develop a combined capability for untargeted source organism characterization and GVP detection. To our knowledge, this is the first report of applying these approaches towards a more complete proteomic characterization of unknowns. We successfully demonstrate the feasibility of broad human contaminant GVP detection in proteomics data, develop a better understanding of GVP detectability, characterize the sample-to-sample variability in GVP detection, and identify a core set of GVPs that can potentially be used as markers indicative of the human contaminant traces portion of the “dark” proteome. Further, we developed and evaluated a combined pipeline, MARLOWE-GVP, that enables both untargeted source organism characterization and GVP detection. We show high accuracy of correct source organism characterization and high degree of similarity of human contaminant GVP detection compared to the conventional approach. Success on both these efforts have allowed us to advance our understanding and characterization of the “dark” proteome.

59 BASIC BIOLOGICAL SCIENCES↗

The future of subsurface monitoring: AEC’s breakthroughs in CCS technology

Carbon capture and storage (CCS) has emerged as a key solution in the fight against climate change. However, for CCS to succeed, it is crucial to ensure that the sequestered CO2 stays safely trapped underground. The U.S. Department of Energy (DOE) has emphasized the need for advancements in subsurface monitoring, measurement, reporting, and verification. Aside from caprock integrity failure, the other primary failure points usually involve defective cement in the casing annulus of wellbores or plugged and abandoned wells. In addition, many energy producers (e.g., oil and gas, geothermal) and storage and disposal operators (e.g., H2 and water) must deal with the same issue. Poorly placed or degraded cement can create pathways for gas or fluid to escape from casing annuli and in plugged and abandoned or orphan wells, posing environmental risks. Yet, a reliable and cost-effective way to monitor cement and well integrity over multiple decades is still unavailable. Traditional geophysical methods like 4D seismic imaging and surface-based electromagnetic monitoring lack the resolution and accuracy for detecting these types of failures (Vasco et al., 2022; Fawad and Mondol, 2021). Wireline logging is expensive to run continuously and is obtrusive to the operation. While fiber optics can potentially be a solution, its bulkiness can significantly compromise the cement's integrity. To address these challenges, the Advanced Energy Consortium (AEC) at The University of Texas at Austin’s Bureau of Economic Geology (the Bureau) has been pioneering research in subsurface monitoring using its portfolio of distributed autonomous microfabricated sensors for harsh subsurface environments since 2008. A class of these microsensors [System on a Chip (SoC)] can be mixed in cement and permanently placed without compromising the cement column; the sensors would then communicate with each other or a data acquisition (DAQ) master node. Another class of the AEC microsensors can be fully autonomous, with rechargeable micro-batteries capable of exceeding 100°C, flash memory, and, currently, a pressure and temperature sensor. They are designed to circulate in mud, geothermal fluids, U-loops, or pipelines. They can log data into memory and are unobtrusive to operations. Our team has been working on a multi-year DOE-funded project (DE-FE0031856)—supported by $2.95M in federal funding and $0.75M in cost-matching from the AEC—to demonstrate SoC sensor utility for CO2 leakage monitoring in CCS applications. This multi-institutional collaboration developed a novel sensing architecture utilizing radiofrequency (RF) microsensors embedded within the cement sheath. These sensors detect CO2 migration and are interrogated via a Smart Casing Collar (SCC).

58 GEOSCIENCES↗

Optimizing inference of segmentation on high-resolution images in MLExchange

MLExchange is a machine learning (ML) operations platform providing web user-interfaces (UIs) for data visualization and analysis pipelines at synchrotron facilities. Among these UIs is the segmentation app which helps synchrotron users utilize ML algorithms to automatically segment high-resolution scientific images with minimal manual annotation effort. In this work, we share code optimizations that significantly speed up the segmentation inference workflow of large data in short time. By optimizing the sequence of CPU-GPU data transfers and introducing CPU parallelization to key operations, we improve the per-device, per-image frame computational efficiency and observe close to 3×$$\times$$ speedup over the original segmentation inference workflow run time when utilizing a single GPU. Further adaptations enabling multi-GPU inference yield more than 40×$$\times$$ speedup with 100 GPUs compared to the optimized single GPU inference workflow. This acceleration of the segmentation inference workflow will provide MLExchange users with easy access to segmentation results with little wait time.

Lu, Shizhao↗

Enabling Early Transient Discovery in LSST via Difference Imaging with DECam

We present SLIDE, a pipeline that enables transient discovery in data from the Vera C. Rubin Observatory’s Legacy Survey of Space and Time (LSST), using archival images from the Dark Energy Camera as templates for difference imaging. We apply this pipeline to the recently released Data Preview 1 (DP1; the first public release of Rubin commissioning data) and search for transients in the resulting difference images. The image subtraction, photometry extraction, and transient detection are all performed on the Rubin Science Platform. We demonstrate that SLIDE effectively extracts clean photometry by circumventing poor or missing LSST templates. We identified 29 previously unreported transients, 12 of which would not have been detected based on the DP1 DiaObject catalog. SLIDE will be especially useful for transient analysis in the early years of LSST, when template coverage will be largely incomplete or when templates may be contaminated by transients present at the time of acquisition. We present multiband light curves for a sample of known transients, along with new transient candidates identified through our search. Finally, we discuss the prospects of applying this pipeline during the main LSST survey. Our pipeline is broadly applicable and will support studies of all transients with slowly evolving phases.

Dong, Yize 一泽董 [Harvard-Smithsonian Center for Ast↗

The anomaly of the CMB power with the latest Planck data

Abstract The lack of power anomaly is an unexpected feature observed at large angular scales in the maps of Cosmic Microwave Background (CMB) produced by the COBE, WMAP andPlancksatellites. This signature, which consists in a missing of power with respect to that predicted by the ΛCDM model, might hint at a new cosmological phase before the standard inflationary era.The main point of this paper is taking into account the latestPlanckpolarisation data to investigate how the CMB polarisation improves the understanding of this feature. With this aim, we apply to the latestPlanckdata, both PR3 (2018) and PR4 (2020) releases, a new class of estimators capable of evaluating this anomaly by considering temperature and polarisation data both separately and in a jointly way. This is the first time that the PR4 dataset has been used to study this anomaly. To critically evaluate this feature, taking into account the residuals of known systematic effects present in thePlanckdatasets, we analyse the cleaned CMB maps using different combinations of sky masks, harmonic range and binning on the CMB multipoles.Our analysis shows that the estimator based only on temperature data confirms the presence of a lack of power with a lower-tail-probability (LTP), depending on the component separation method, ≤ 0.33% and ≤ 1.76% for PR3 and PR4, respectively. To our knowledge, the LTP≤ 0.33% for the PR3 dataset is the lowest one present in the literature obtained fromPlanck2018 data, considering thePlanckconfidence mask. We find significant differences between these two datasets when polarisation is taken into account most likely due to a different level of systematics. Especially, the analysis with PR3 data, unlike that with PR4, seems to point towards a lack of power at large scales also for polarisation.Moreover, we also show that for the PR3 dataset the inclusion of the subdominant polarisation information provides estimates that are less likely accepted in a ΛCDM cosmological model than the only-temperature analysis over the entire harmonic-range considered. In particular, at ℓ max = 26, we found that no simulation has a value as low as the data for all the pipelines.

Astronomy & Astrophysics↗

Cosmology with persistent homology: a Fisher forecast

Abstract Persistent homology naturally addresses the multi-scale topological characteristics of the large-scale structure as a distribution of clusters, loops, and voids. We apply this tool to the dark matter halo catalogs from theQuijotesimulations, and build a summary statistic for comparison with the joint power spectrum and bispectrum statistic regarding their information content on cosmological parameters and primordial non-Gaussianity. Through a Fisher analysis, we find that constraints from persistent homology are tighter for 8 out of the 10 parameters by margins of 13–50%. The complementarity of the two statistics breaks parameter degeneracies, allowing for a further gain in constraining power when combined. We run a series of consistency checks to consolidate our results, and conclude that our findings motivate incorporating persistent homology into inference pipelines for cosmological survey data.

Astronomy & Astrophysics↗

Effective Defect Detection Using Instance Segmentation for NDI

Ultrasonic testing is a common Non-Destructive Inspection (NDI) method used in aerospace manufacturing. However, the complexity and size of the ultrasonic scans make it challenging to identify defects through visual inspection or machine learning models. Using computer vision techniques to identify defects from ultrasonic scans is an evolving research area. In this study, we used instance segmentation to identify the presence of defects in the ultrasonic scan images of composite panels that are representative of real components manufactured in aerospace. We used two models based on Mask- RCNN (Detectron 2) and YOLO 11 respectively. Additionally, we implemented a simple statistical pre-processing technique that reduces the burden of requiring custom-tailored pre-processing techniques. Our study demonstrates the feasibility and effectiveness of using instance segmentation in the NDI pipeline by significantly reducing data pre-processing time, inspection time, and overall costs.

computer vision techniques↗

EV-ELM (Electric Vehicle Policies with the Energy Language Model) [SWR-25-156]

Electric Vehicle Policies with the Energy Language Model (EV-ELM) leverages previous work using Large Language Models (LLMs) to find, download, and parse policy information related to energy infrastructure. In this application, we use LLMs to find policy documents related to the permitting and installation of electric vehicle charging infrastructure. This software contains the code to find, download, and parse these documents, while a related data record in the Open Energy Data Initiative (OEDI) will include the resulting output dataset that can be used for downstream analysis. The EV-ELM repository contains code for the EV-ELM project, which focuses on retrieving and processing EV permitting processes using large language models. The project is composed of two pipelines: (1) a web scraping pipeline for discovering and downloading EV permitting documents, and (2) a document parsing and extraction pipeline that processes the downloaded files to produce structured data. The web scraping pipeline is designed to extract relevant information from various websites, while the document parsing pipeline processes and analyzes the extracted documents to derive meaningful insights. Both pipelines depend on the NLR elm repository, which provides essential tools and functionalities for handling and processing the data. The web scraping pipeline is a modified version of the ordinance_gpt example within the elm repository. It has been adapted to fit the specific requirements of the EV-ELM project, ensuring that it effectively captures and processes the necessary information related to EV permitting.

Olson, Reid [National Laboratory of the Rockies (N↗

LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning

Sequence-function data provides valuable information about the protein functional landscape but is rarely obtained during directed evolution campaigns. Here, we present Long-read every variant Sequencing (LevSeq), a pipeline that combines a dual barcoding strategy with nanopore sequencing to rapidly generate sequence-function data for entire protein-coding genes. LevSeq integrates into existing protein engineering workflows and comes with open-source software for data analysis and visualization. The pipeline facilitates data-driven protein engineering by consolidating sequence-function data to inform directed evolution and provide the requisite data for machine learning-guided protein engineering (MLPE). LevSeq enables quality control of mutagenesis libraries prior to screening, which reduces time and resource costs. Simulation studies demonstrate LevSeq’s ability to accurately detect variants under various experimental conditions. Lastly, we show LevSeq’s utility in engineering protoglobins for new-to-nature chemistry. Widespread adoption of LevSeq and sharing of the data will enhance our understanding of protein sequence-function landscapes and empower data-driven directed evolution.

59 BASIC BIOLOGICAL SCIENCES↗

Positron emission tomography harmonization in the Alzheimer's Disease Neuroimaging Initiative: A scalable and rigorous approach to multisite amyloid and tau quantification

Abstract INTRODUCTION A key goal of the Alzheimer's Disease NeuroImaging Initiative (ADNI) positron emission tomography (PET) Core is to harmonize quantification of β‐amyloid (Aβ) and tau PET image data across multiple scanners and tracers. METHODS We developed an analysis pipeline (Berkeley PET Imaging Pipeline, B‐PIP) for ADNI Aβ and tau PET images and applied it to PET data from other multisite studies. Steps include image pre‐processing, refacing, magnetic resonance imaging (MRI)/PET co‐registration, visual quality control (QC), quantification of tracer uptake, and standardization of Aβ and tau standardized uptake value ratios (SUVrs) across tracers. RESULTS Measurements from 10,105 cross‐sectional and longitudinal Aβ and tau PET scans acquired in several studies between 2010 and 2024 can be processed, harmonized, and directly merged across tracers and cohorts. DISCUSSION The B‐PIP developed in ADNI is a scalable image harmonization approach used in several observational studies and clinical trials that facilitates rigorous Aβ and tau PET quantification and data sharing. Highlights Quantitative results from ADNI Aβ and tau PET data are generated using a rigorous, scalable image processing pipeline This pipeline has been applied to PET data from several other large, multisite studies and trials Quantitative outcomes are harmonizable across studies and are shared with the scientific community

Neurosciences & Neurology↗

The 200 Gbps Challenge: Imagining HL-LHC analysis facilities

The IRIS-HEP software institute, as a contributor to the broader HEP Python ecosystem, is developing scalable analysis infrastructure and software tools to address the upcoming HL-LHC computing challenges with new approaches and paradigms, driven by our vision of what HL-LHC analysis will require. The institute uses a "Grand Challenge" format, constructing a series of increasingly large, complex, and realistic exercises to show the vision of HL-LHC analysis. Recently, the focus has been demonstrating the IRIS-HEP analysis infrastructure at scale and evaluating technology readiness for production. As a part of the Analysis Grand Challenge activities, the institute executed a "200 Gbps Challenge", aiming to show sustained data rates into the event processing of multiple analysis pipelines. The challenge integrated teams internal and external to the institute, including operations and facilities, analysis software tools, innovative data delivery and management services, and scalable analysis infrastructure. The challenge showcases the prototypes - including software, services, and facilities - built to process around 200 TB of data in both the CMS NanoAOD and ATLAS PHYSLITE data formats with test pipelines. The teams were able to sustain the 200 Gbps target across multiple pipelines. The pipelines focusing on event rate were able to process at over 30 MHz. These target rates are demanding; the activity revealed considerations for future testing at this scale and changes necessary for physicists to work at this scale in the future. The 200 Gbps Challenge has established a baseline on today's facilities, setting the stage for the next exercise at twice the scale.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

SITCOMTN-164: AnaCal Shear Profile of Abell 360 in LSSTComCam Data Preview 1

This technote presents the measurement of the weak-lensing shear profile for the Abell 360 galaxy cluster using LSSTComCam data processed with the AnaCal pipeline. We detail the procedures involved in bright-star masking, source selection, and the estimation of tangential and cross shear around the cluster center. The resulting shear profiles provide key insights into the mass distribution of Abell 360 and demonstrate the capabilities of AnaCal in processing early LSST data, with the tangential shear profile detected at 5σ significance.

79 ASTRONOMY AND ASTROPHYSICS↗

CONTROL AND DATA ACQUISITION IN A CYBER-PHYSICAL MIDSTREAM TESTBED

This thesis presents the development of a laboratory-scale cyber–physical midstream pipeline testbed designed to address this gap and support research in industrial control systems security. The platform integrates pumps, valves, sensors, programmable logic controllers (PLCs), and a human–machine interface (HMI) to emulate the monitoring and control architecture of real pipeline operations. The physical process is implemented as a closed-loop liquid circulation system designed to replicate flow behavior characteristic of midstream pipeline infrastructure. The testbed enables real-time data acquisition of key process variables, including flow rate and pressure facilitating the generation of datasets representative of normal pipeline operation. A threat model encompassing common ICS attack vectors was developed, including sensor spoofing, command injection, false data injection, denial-of-service attacks, and relay manipulation. Multiple attack scenarios were implemented and evaluated to demonstrate how cyber intrusions targeting sensors, actuators, networks, and software propagate into measurable physical consequences in pipeline flow and pressure. The developed platform serves as a practical, cost-effective environment for experimentation, education, and future cybersecurity research in midstream pipeline systems.

42 ENGINEERING↗