Search NASA⌕ Search

SEARCH · Search NASA

Results for “pipeline data processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

LCLS Big Data Handling – How I Learned to Stop Worrying and Love the Data Deluge

Advanced data and computing systems are vital to Linac Coherent Light Source (LCLS) operations, data interpretation and overall scientific productivity. The transition to MHz-era operation marks a fundamental change in scale that requires new infrastructure and architectures to link LCLS to the required scale of computing needed for scientific interpretation. The LCLS-II Data System meets big data challenges by implementing configurable data reduction that can adapt to multiple science areas, real-time analysis frameworks to provide visualization and fast feedback, and the ability to transfer data to local and remote computational facilities for near real time analysis at the appropriate scale. Feature extracted information generated in the data analysis pipeline - at the edge, local compute, or remote High-Performance Computing (HPC) resources - can be used to steer experiments and inform user decisions during beam time. Artificial Intelligence and Machine Learning (AI/ML) techniques present new opportunities to rapidly analyse large datasets and direct experiments, but create new challenges in scaling, adaptability, complexity, and trustworthiness. We describe how the LCLS-II Data System architecture addresses its data-driven challenges in the areas of data acquisition, data processing, data management, and workflow orchestration to decrease the overall time-to-science and provide a vision for future developments.

artificial intelligence↗

Machine Learning Automation Pipeline

Machine Learning Automation Pipeline (MLAP) is a package to perform machine learning (ML) analysis in a step by step manner, starting with data extraction until analysis and prediction. The scripts provide the users option to chose an action such as "Extract", "Prep", and "Train" and numerous cases can be launched with just a single command. The inputs for each case are provided using a JSON file. The simulation results of several cases can be assessed using an automated process and analyzed for various metrics pertinent to ML analysis.

Jha, Pankaj↗

Dark Energy Survey Year 6 results: cell-based coadds and METADETECTION weak lensing shape catalogue

We present the metadetection weak lensing galaxy shape catalogue from the 6-yr Dark Energy Survey (DES Y6) imaging data. This data set is the final release from DES, spanning 4422 deg 2 of the southern sky. We describe how the catalogue was constructed, including the two new major processing steps, cell-based image coaddition, and shear measurements with metadetection. The DES Y6 M etadetection weak lensing shape catalogue consists of 151 922 791 galaxies detected over riz bands, with an effective number density of n eff = 8.22 galaxies per arcmin 2 and shape noise of σ e = 0.29. We carry out a suite of validation tests on the catalogue, including testing for point spread function (PSF) leakage, testing for the impact of PSF modelling errors, and testing the correlation of the shear measurements with galaxy, PSF, and survey properties. In addition to demonstrating that our catalogue is robust for weak lensing science, we use the DES Y6 image simulation suite to estimate the overall multiplicative shear bias of our shear measurement pipeline. We find no detectable multiplicative bias at the roughly half-per cent level, with m = (3.4 ± 6.1) x 10 –3 , at 3σ uncertainty. This is the first time both cell-based coaddition and Metadetection algorithms are applied to observational data, paving the way to the Stage-IV weak lensing surveys.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

HARMONY: Large-Scale Architecture Search for Efficient Hybrid Language Models

As large language models scale to trillions of parameters, their computational and memory requirements present critical challenges for efficient training and deployment. While Mixture of Experts (MoE) architectures enable efficient scaling through sparse parameter activation, and state-space models like Mamba offer linear-time complexity, principled methods for combining these paradigms remain undeveloped. We introduce HARMONY (Hybrid Architecture Research for Mamba, Optimized with Neural efficiencY), a multi-objective evolutionary neural architecture search framework for discovering efficient hybrid language models that integrate Transformer attention mechanisms, Mixture-of-Experts routing, and Mamba state-space components. Through large-scale distributed search using 16,384 MI250X GPUs on the Frontier supercomputer, HARMONY explores a comprehensive design space encompassing six attention variants (MHA, MQA, GQA, MLA, SWA, and Mamba-2), variable MoE configurations with both routed and shared experts, and extensive Mamba hyperparameters. Our framework discovers heterogeneous architectures that balance training performance with computational efficiency through multi-objective optimization incorporating latency penalties and fitness-based selection. Analysis of discovered architectures reveals that optimal hybrid designs favor heterogeneous component mixing rather than homogeneous patterns, with Mamba-2 and Multi-Head Latent Attention (MLA) emerging as preferred mechanisms. Discovered architectures demonstrate superior training efficiency: our best configuration achieves a final perplexity of 1.0874 with 2.38B parameters while processing 4,320 tokens/second, outperforming significantly larger manually designed models. Full-scale evaluation shows HARMONY's top architectures achieve better loss trajectories than equivalently-sized models using state-of-the-art configurations including Mixtral, Jamba, and Samba. Additionally, we demonstrate 91% weak scaling efficiency when training discovered 36B-parameter models across 1,024 GPUs. HARMONY is released as an open framework with comprehensive tools for building and training hybrid models using expert-data-pipeline parallelism, democratizing access to automated architecture design for next-generation language models.

Herron, Emily [ORNL] (ORCID:0000000273008172)↗

Platform for Automated Anomaly Detection in the Mercury Process System at the Target System in the Spallation Neutron Source

The Spallation Neutron Source (SNS) at Oak Ridge National Laboratory accelerates proton beams, which are directed toward a mercury target to generate the world’s most intense neutron beams via spallation. The target system consists of several interconnected subsystems and accounts for a major share of the facility’s overall downtime. Early detection of anomalies in the target system response can thus provide the possibility of taking corrective actions to reduce downtime. Accelerator facilities have largely focused on the beam side for data-driven fault prognostics. On the target side, SNS relies on operational shift technicians (OSTs), who respond to alarms and manually flag anomalies onto the System Tracking and Reliability (STAR) platform. This paper presents one of the first studies of using machine learning (ML) to automate anomaly detection in the target system. The study focused on the mercury process system as the first use case and employed reconstruction-based anomaly detection on minutely sampled time series signals. The pipeline was integrated into the STAR platform to autonomously rank and flag anomalies every week. The STAR platform provides a user interface for the OSTs to evaluate the flagged anomalies, thereby incorporating human feedback.

Anomaly detection↗

Modeling Framework for the Assessment of a Sustainable Hydrogen Production and Supply Chain Network in California

The cost-effective and sustainable deployment of hydrogen supply and demand networks, especially in large economic regions like California, can be challenging considering the spatial-temporal availability and variability of the different actors across the network such as production processes, distribution modes, and end-users. In this presentation, we will provide an overview and demonstration of a modeling framework used to assess the environmental, economic, and human health impacts of plausible hydrogen production and supply chain networks in California. Scenarios focus on green hydrogen production pathways using water electrolysis and biomass gasification. End-use applications included in the model are transit, medium and heavy-duty trucking, port authorities, and power and aviation companies that currently consume natural gas, diesel, and aviation fuel for their day-to-day operation. Representative locations for hydrogen production and end-use are based on recent projections of the hydrogen economy in California. All mass and energy flows, as well as estimated emissions, are based on H2A process model designs and projections of technology performance, literature review, and LBNL process, economic and life cycle modeling, and not on company data for the sake of this presentation. Human health impacts are included following methodologies developed for the University of California Irvine HyDeal project. Life cycle phases associated with hydrogen production include feedstock preparation (water and biomass), energy production and consumption (renewable, grid, and combination of renewable and grid electricity), maintenance (chemical utilization in electrolysis and natural gas combustion in gasification), carbon sequestration, hydrogen storage (compression and liquefaction), and distribution (truck and pipeline). We apply the framework utilizing California specific emission factors, financial data, and human health damages and explore the impact of network characteristics on results. Example variations include: the inclusion of policy incentives or not, different representations of the electricity grid and source, electrolysis versus gasification versus combinations of both for production, liquefaction versus compression based on producer capacity cutoffs, transportation truck versus pipeline based on existing infrastructure, and ultimate end use. Comparison of these different scenarios can help inform future projects by demonstrating the trade-offs among environmental, economic, and human health impacts. This model, automated in R, is a starting platform upon which new analysis, modeling capabilities, locations, and emission factors can be rapidly tested and integrated.

Zaki, Mohammed Tamim↗

A Preferences Corpus and Annotation Scheme for Human-Guided Alignment of Time-Series GPTs

The process of time-series forecasting such as predicting trajectories of silicon content in blast furnaces is a difficult task. Most time-series approaches today focus on scalar-type MSE loss optimization. This optimization approach, while widely common, could benefit from the use of human expert or process-level preferences. In this paper, we introduce a novel alignment and fine-tuning approach that involves learning from a corpus of preferred and dis-preferred time-series prediction trajectories. Our contributions include (1) a preference annotation pipeline for time-series forecasts, (2) the application of Score-based Preference Optimization (SPO) to train decoder-only transformers from preferences, and (3) results showing improvements in forecast quality. The approach is validated on both proprietary blast furnace data and the UCI Appliances Energy dataset. The proposed preference corpus and training strategy offer a new option for fine-tuning sequence models in industrial settings.

DPO↗

PIPES (Pipeline for Integrated Projects in Energy Systems) [SWR-24-89]

The Pipeline for Integrated Projects in Energy Systems (PIPES) is a comprehensive project, data, and workflow management tool designed for integrated modeling teams. PIPES facilitates the management of data requirements, tasks, and progress tracking, serving as a higher-level integration layer that works across various data and modeling software. This tool integrates models, data, and tools to perform large-scale, integrated analysis work at scale. PIPES is designed to streamline integrated modeling projects, enhance collaboration, and ensure the quality and efficiency of data management and workflow processes. https://github.com/nrel-pipes/pipes-api https://github.com/nrel-pipes/pipes-web https://github.com/nrel-pipes/nrel-pipes

Gu, Jianli↗

Reinforcement Learning for In-Spill Optimization of the Mu2e Resonant Extraction: Compensating Non-Stationarity

We present design considerations and challenges for the fast machine learning component of a third-order resonant beam extraction regulation system being commissioned to deliver steady beam rates to the mu2e experiment at Fermilab. Dedicated quadrupoles drive the tune toward the 29/3 resonance each spill, extracting beam at kV multiwire septa. The overall Spill Regulation System consists of (1) a “slow” process using ~100-spill averages to adjust the base quad ramp infrequently, (2) a feedforward harmonic content compensator, and (3) the “fast” ML agent reacting during each ongoing spill with on-the-fly additive corrections to the sum of (1) and (2). We have demonstrated improved beam-rate steadying for a fast ML agent compared to a PID controller using a quasi-physical spill simulation, and demonstrated distillation of that simulation into a predictive surrogate model. Current work includes a data-and-training pipeline to generate data-aware surrogates with real-world dynamics, even as the dynamics shift unpredictably. The surrogates are to act as RL environments against which to train our fast ML control agents before deploying them on FPGA in the live system. Further current efforts focus on modeling and controlling beam loss around the storage ring, understanding additional available hardware inputs to the model, and the interplay of these with beam-steadying performance.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Optimal reconstruction of baryon acoustic oscillations for DESI 2024

Baryon acoustic oscillations (BAO) provide a robust standard ruler to measure the expansion history of the Universe through galaxy clustering. Density-field reconstruction is now a widely adopted procedure for increasing the precision and accuracy of the BAO detection. With the goal of finding the optimal reconstruction settings to be used in the DESI 2024 galaxy BAO analysis, we assess the sensitivity of the post-reconstruction BAO constraints to different choices in our analysis configuration, performing tests on blinded data from the first year of DESI observations (DR1), as well as on mocks that mimic the expected clustering and selection properties of the DESI DR1 target samples. Overall, we find that BAO constraints remain robust against multiple aspects in the reconstruction process, including the choice of smoothing scale, treatment of redshift-space distortions, fiber assignment incompleteness, and parameterizations of the BAO model. We also present a series of tests that DESI followed in order to assess the maturity of the end-to-end galaxy BAO pipeline before the unblinding of the large-scale structure catalogs.

79 ASTRONOMY AND ASTROPHYSICS↗

FiberFlex: Real-time FPGA-based Intelligent and Distributed Fiber Sensor System for Pedestrian Recognition

In recent years, security monitoring of public places and critical infrastructure has heavily relied on the widespread use of cameras, raising concerns about personal privacy violations. To balance the need for effective security monitoring with the protection of personal privacy, we explore the potential of optical fiber sensors for this application. This article proposes FiberFlex, an intelligent and distributed fiber sensor system. Ultizing Field Programmable Gate Arrays (FPGA) high-level synthesis (HLS) acceleration, FiberFlex offers real-time pedestrian detection by co-designing the entire pipeline of optical signal acquisition, processing, and recognition networks based on the principles of optical fiber sensing. As a promising alternative to traditional camera-based monitoring systems, FiberFlex achieves pedestrian detection by analyzing the vibration patterns caused by pedestrian footsteps, enabling security monitoring while preserving individual privacy. FiberFlex comprises three modules: First , fiber-optic sensing system: A fiber-optic distributed acoustic sensing (DAS) system is built and used to measure the ground vibration waves generated by people walking. Second , algorithms: We first collect the training data by measuring the ground vibration waves, label the data, and use the data to train the neural network models to perform pedestrian recognition. Third , hardware accelerators: We use HLS tools to design hardware modules on FPGA for data collection and pre-processing and integrate them with the downstream neural network accelerators to perform in-line real-time pedestrian detection. The final detection results are sent back from FPGA to the host CPU. We implement our system FiberFlex with the in-house built DAS system and AMD/Xilinx Kintex7 FPGA KC705 board and verify the whole system using the real-world collected data. We conduct recognition tests on five test subjects of varying ages, heights, and weights in a fixed sensing area. Each subject experienced 20 real-time recognition tests using their daily walking habits, and the subjects were given adequate rest between tests. After 100 tests on five test subjects, the overall real-time recognition accuracy exceeded \(88.0\%\) . The whole system uses 55 W of power, 33 W in the optical DAS system and 22 W in the FPGA. Relying on its end-to-end interdisciplinary design, FiberFlex seamlessly combines fiber-optic sensors with FPGA accelerators to enable low-power real-time security monitoring without compromising privacy, making it a valuable addition to the existing security monitoring network. According to FiberFlex, more valuable research can be conducted in the future, such as fall monitoring for the elderly, migration of identification networks between different application scenarios, and improvement of anti-interference performance in more complex environments. In future perception networks, where the “eyes” are not feasible, let’s use fiber optic touch instead.

Distributed↗

Towards Online Machine Learning in DUNE Data Acquisition

Processing the large volumes of data produced by liquid argon time projection chamber (LArTPC) experiments presents a significant challenge, especially those at the scale of DUNE. This is a particular challenge when aiming to trigger on low-energy neutrinos from core-collapse supernovae, which are typically buried in a high-rate radiological background. To enable real-time event selection suitable for such rare signals, we are developing machine learning based data filtering methods. In order to demonstrate the feasibility of this approach, we implemented such pipeline using the ICEBERG detector at Fermilab as a small-scale LArTPC, with a focus on identifying Michel electrons as a proxy for low-energy neutrino interactions. This poster will present the current status of integrating these machine learning models into the data acquisition (DAQ) system of this detector.

Dalager, Olivia [Fermilab]↗

Optimizing Cryo-Focused Pyrolysis GC/MS for Tracing Soil Organic Matter Across Diverse Ecosystems

The cycling of organic matter in terrestrial soils and sediments is central to a range of biogeochemical processes that regulate nutrient cycling, crop productivity, trace gas emissions, and contaminant transport. Pyrolysis-gas chromatography/mass spectrometry (py-GC/MS) is a powerful tool for characterizing bulk soil organic matter (SOM) at the molecular level. In this study, we used a cryo-focused py-GC/MS system to analyze soil samples from seven diverse ecosystems: vernal pool, prairie pothole, temperate forest, tropical forest, tundra, wildfire-affected boreal forest, and grassland. We addressed a key bottleneck in molecular-level SOM characterization by developing an automated data analysis pipeline to optimize py-GC/MS and complementary evolved gas analysis/mass spectrometry (EGA/MS) methods, incorporating advanced tools for peak deconvolution, developing a custom compound class library, and implementing fragmentation spectrum-based molecular networking for the first time. This improved workflow was applied to soil samples from all seven ecosystems, including multiple depths and density fractions. Our findings demonstrate that ecosystem type plays a dominant role in shaping compositional differences in SOM. We also identified trends in the source of SOM compounds (e.g., microbial vs plantderived) across soil depth and density fractions, which are critical for understanding persistence and turnover of SOM. Our molecular networking analysis indicated that although many compounds are widespread across ecosystems, others are restricted to specific environments, such as wetlands. This underscores the utility of molecular-level data in elucidating the complexity of SOM composition and the environmental drivers that shape it. Such molecular-level insights can deepen our knowledge of biogeochemical SOM cycles.

54 ENVIRONMENTAL SCIENCES↗

Force Field X: A computational microscope to study genetic variation and organic crystals using theory and experiment

Force Field X (FFX) is an open-source software package for atomic resolution modeling of genetic variants and organic crystals that leverages advanced potential energy functions and experimental data. FFX currently consists of nine modular packages with novel algorithms that include global optimization via a many-body expansion, acid–base chemistry using polarizable constant-pH molecular dynamics, estimation of free energy differences, generalized Kirkwood implicit solvent models, and many more. Applications of FFX focus on the use and development of a crystal structure prediction pipeline, biomolecular structure refinement against experimental datasets, and estimation of the thermodynamic effects of genetic variants on both proteins and nucleic acids. The use of Parallel Java and OpenMM combines to offer shared memory, message passing, and graphics processing unit parallelization for high performance simulations. Overall, the FFX platform serves as a computational microscope to study systems ranging from organic crystals to solvated biomolecular systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Beyond microbial abundance: metadata integration enhances disease prediction in human microbiome studies

Multiple studies have highlighted the interaction of the human microbiome with physiological systems such as the gut, immune, liver, and skin, via key axes. Advances in sequencing technologies and high-performance computing have enabled the analysis of large-scale metagenomic data, facilitating the use of machine learning to predict disease likelihood from microbiome profiles. However, challenges such as compositionality, high dimensionality, sparsity, and limited sample sizes have hindered the development of actionable models. One strategy to improve these models is by incorporating key metadata from both the human host and sample collection/processing protocols. This remains challenging due to sparsity and inconsistency in metadata annotation and availability. In this paper, we introduce a machine learning-based pipeline for predicting human disease states by integrating host and protocol metadata with microbiome abundance profiles from 68 different studies, processed through a consistent pipeline. Our findings indicate that metadata can enhance machine learning predictions, particularly at higher taxonomic ranks like Kingdom and Phylum, though this effect diminishes at lower ranks. Our study leverages a large collection of microbiome datasets comprising 11,208 samples, therefore enhancing the robustness and statistical confidence of our findings. This work is a critical step toward utilizing microbiome and metadata for predicting diseases such as gastrointestinal infections, diabetes, cancer, and neurological disorders.

Mathematics and Computing↗

Industrial Carbon Capture from a Cement Facility Using the Cryocap™ FG Process

The objective of the project was to execute and complete front-end engineering and design (FEED) studies for commercial-scale, carbon capture projects that separate 95% of the total CO 2 emissions at an industrial facility, producing at least 100,000 metric tonnes/year of CO 2 for sequestration. The industrial facility selected is the Holcim (US) Ste. Genevieve cement manufacturing facility (the largest single kiln line in the world), while the carbon capture system selected is Pressure Swing Adsorption system (PSA) assisted Cryocap™ technology developed by Air Liquide. The industrial host site emits approximately 3 million tonnes CO 2 /year based on plant data from 2020-2022. The captured CO 2 will meet the requirements of transport (Pipeline Grade) and geological storage, and the geological storage facilities within 80 miles of the CO 2 source. The impact of the project on Environmental Justice and the regional economy was also analyzed.

01 COAL, LIGNITE, AND PEAT↗

Storage Field Development Plan: One Earth Energy

This Storage Field Development plan presents the Storage Complex characterization results, construction, monitoring, and operational plans, and costs associated with the proposed One Earth Sequestration Carbon Capture and Storage (OES-CCS) site in McLean County, Illinois, near Gibson City. The proposed storage complex, known as the Mt. Simon Storage Complex, comprises the Cambrian Mt. Simon Sandstone reservoir and the primary seal, the Cambrian Eau Claire Formation. The lowermost Underground Source of Drinking Water (USDW) identified for the site is the Ordovician St. Peter Sandstone. Geologic characterization of the Mt. Simon Storage Complex at the OES-CCS site was performed by the Illinois Storage Corridor CarbonSAFE Phase III project, which also prepared and submitted three UIC Class VI applications to construct three injection wells; the permit applications were submitted and are in the federal EPA review process. A characterization well, OEE #1, was drilled to collect site-specific data. These data were analyzed and used to develop the UIC Class VI applications. The OEE #1 well will be converted to an in-zone monitoring (IZM) well for the injection phase. The proposed buildout for the OES-CCS site includes (1) three injection wells (OES #1, OES #2, and OES #3), (2) two IZM wells, (3) two above confining zone (ACZ) monitoring wells, one of which will be used to monitor the lowermost USDW, (4) capture and compression facilities, and (5) transportation facilities, i.e., pipelines. A pre-operational testing program was proposed in the Class VI permit application and will be employed at the site pending approval. Additional pre-injection (baseline), syn-injection, and post-injection monitoring and site care procedures will be followed by OES to ensure that injection activities are protective of human health and the environment. Injection is scheduled to begin in 2025, distributed across the three injection wells in accordance with the permit operating conditions. One Earth Sequestration intends to inject up to 90 million tonnes of CO 2 over a period of approximately 20 years. Injection will begin at approximately 0.5 million tonnes of CO 2 annually and ramp up to a maximum of 4.5 million tonnes annually. Daily injection rates are expected to range from 1,400 to 1,500 tonnes per day initially and reach a maximum of approximately 4,225 tonnes per day, depending on site geology and injectivity at each injection well location, and CO 2 availability. The costs associated with the OES-CCS project include pre-operational costs (e. g. additional seismic data acquisition and well drilling), capture and transportation facility and equipment costs, predicted field operating expenditures (OpEx), and decommissioning and post-injection site care (PISC) costs. The risks associated with project activities, such as site construction, injection operations, and verification of secure storage were evaluated, and mitigation strategies proposed to alleviate those risks.

09 BIOMASS FUELS↗

Machine learning-driven predictive resource management in complex science workflows

Here, the collaborative efforts of large communities in science experiments, often comprising thousands of global members, reflect a monumental commitment to exploration and discovery. Recently, advanced and complex data processing has gained increasing importance in science experiments. Data processing workflows typically consist of multiple intricate steps, and the precise specification of resource requirements is crucial for each step to allocate optimal resources for effective processing. Estimating resource requirements in advance is challenging due to a wide range of analysis scenarios, varying skill levels among community members, and the continuously increasing spectrum of computing options. One practical approach to mitigate these challenges involves initially processing a subset of each step to measure precise resource utilization from actual processing profiles before completing the entire step. While this two-staged approach enables processing on optimal resources for most of the workflow, it has drawbacks such as initial inaccuracies leading to potential failures and suboptimal resource usage, along with overhead from waiting for initial processing completion, which is critical for fast-turnaround analyses. In this context, our study introduces a novel pipeline of machine learning models within a comprehensive workflow management system, the Production and Distributed Analysis (PanDA) system. These models employ advanced machine learning techniques to predict key resource requirements, overcoming challenges posed by limited upfront knowledge of characteristics at each step. Accurate forecasts of resource requirements enable informed and proactive decision-making in workflow management, enhancing the efficiency of handling diverse, complex workflows across heterogeneous resources.

97 MATHEMATICS AND COMPUTING↗