Search NASA⌕ Search

SEARCH · Search NASA

Results for “pipeline data processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

LCLS Big Data Handling – How I Learned to Stop Worrying and Love the Data Deluge

Advanced data and computing systems are vital to Linac Coherent Light Source (LCLS) operations, data interpretation and overall scientific productivity. The transition to MHz-era operation marks a fundamental change in scale that requires new infrastructure and architectures to link LCLS to the required scale of computing needed for scientific interpretation. The LCLS-II Data System meets big data challenges by implementing configurable data reduction that can adapt to multiple science areas, real-time analysis frameworks to provide visualization and fast feedback, and the ability to transfer data to local and remote computational facilities for near real time analysis at the appropriate scale. Feature extracted information generated in the data analysis pipeline - at the edge, local compute, or remote High-Performance Computing (HPC) resources - can be used to steer experiments and inform user decisions during beam time. Artificial Intelligence and Machine Learning (AI/ML) techniques present new opportunities to rapidly analyse large datasets and direct experiments, but create new challenges in scaling, adaptability, complexity, and trustworthiness. We describe how the LCLS-II Data System architecture addresses its data-driven challenges in the areas of data acquisition, data processing, data management, and workflow orchestration to decrease the overall time-to-science and provide a vision for future developments.

artificial intelligence↗

Machine Learning Automation Pipeline

Machine Learning Automation Pipeline (MLAP) is a package to perform machine learning (ML) analysis in a step by step manner, starting with data extraction until analysis and prediction. The scripts provide the users option to chose an action such as "Extract", "Prep", and "Train" and numerous cases can be launched with just a single command. The inputs for each case are provided using a JSON file. The simulation results of several cases can be assessed using an automated process and analyzed for various metrics pertinent to ML analysis.

Jha, Pankaj↗

Dark Energy Survey Year 6 results: cell-based coadds and METADETECTION weak lensing shape catalogue

We present the metadetection weak lensing galaxy shape catalogue from the 6-yr Dark Energy Survey (DES Y6) imaging data. This data set is the final release from DES, spanning 4422 deg 2 of the southern sky. We describe how the catalogue was constructed, including the two new major processing steps, cell-based image coaddition, and shear measurements with metadetection. The DES Y6 M etadetection weak lensing shape catalogue consists of 151 922 791 galaxies detected over riz bands, with an effective number density of n eff = 8.22 galaxies per arcmin 2 and shape noise of σ e = 0.29. We carry out a suite of validation tests on the catalogue, including testing for point spread function (PSF) leakage, testing for the impact of PSF modelling errors, and testing the correlation of the shear measurements with galaxy, PSF, and survey properties. In addition to demonstrating that our catalogue is robust for weak lensing science, we use the DES Y6 image simulation suite to estimate the overall multiplicative shear bias of our shear measurement pipeline. We find no detectable multiplicative bias at the roughly half-per cent level, with m = (3.4 ± 6.1) x 10 –3 , at 3σ uncertainty. This is the first time both cell-based coaddition and Metadetection algorithms are applied to observational data, paving the way to the Stage-IV weak lensing surveys.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

HARMONY: Large-Scale Architecture Search for Efficient Hybrid Language Models

As large language models scale to trillions of parameters, their computational and memory requirements present critical challenges for efficient training and deployment. While Mixture of Experts (MoE) architectures enable efficient scaling through sparse parameter activation, and state-space models like Mamba offer linear-time complexity, principled methods for combining these paradigms remain undeveloped. We introduce HARMONY (Hybrid Architecture Research for Mamba, Optimized with Neural efficiencY), a multi-objective evolutionary neural architecture search framework for discovering efficient hybrid language models that integrate Transformer attention mechanisms, Mixture-of-Experts routing, and Mamba state-space components. Through large-scale distributed search using 16,384 MI250X GPUs on the Frontier supercomputer, HARMONY explores a comprehensive design space encompassing six attention variants (MHA, MQA, GQA, MLA, SWA, and Mamba-2), variable MoE configurations with both routed and shared experts, and extensive Mamba hyperparameters. Our framework discovers heterogeneous architectures that balance training performance with computational efficiency through multi-objective optimization incorporating latency penalties and fitness-based selection. Analysis of discovered architectures reveals that optimal hybrid designs favor heterogeneous component mixing rather than homogeneous patterns, with Mamba-2 and Multi-Head Latent Attention (MLA) emerging as preferred mechanisms. Discovered architectures demonstrate superior training efficiency: our best configuration achieves a final perplexity of 1.0874 with 2.38B parameters while processing 4,320 tokens/second, outperforming significantly larger manually designed models. Full-scale evaluation shows HARMONY's top architectures achieve better loss trajectories than equivalently-sized models using state-of-the-art configurations including Mixtral, Jamba, and Samba. Additionally, we demonstrate 91% weak scaling efficiency when training discovered 36B-parameter models across 1,024 GPUs. HARMONY is released as an open framework with comprehensive tools for building and training hybrid models using expert-data-pipeline parallelism, democratizing access to automated architecture design for next-generation language models.

Herron, Emily [ORNL] (ORCID:0000000273008172)↗

Platform for Automated Anomaly Detection in the Mercury Process System at the Target System in the Spallation Neutron Source

The Spallation Neutron Source (SNS) at Oak Ridge National Laboratory accelerates proton beams, which are directed toward a mercury target to generate the world’s most intense neutron beams via spallation. The target system consists of several interconnected subsystems and accounts for a major share of the facility’s overall downtime. Early detection of anomalies in the target system response can thus provide the possibility of taking corrective actions to reduce downtime. Accelerator facilities have largely focused on the beam side for data-driven fault prognostics. On the target side, SNS relies on operational shift technicians (OSTs), who respond to alarms and manually flag anomalies onto the System Tracking and Reliability (STAR) platform. This paper presents one of the first studies of using machine learning (ML) to automate anomaly detection in the target system. The study focused on the mercury process system as the first use case and employed reconstruction-based anomaly detection on minutely sampled time series signals. The pipeline was integrated into the STAR platform to autonomously rank and flag anomalies every week. The STAR platform provides a user interface for the OSTs to evaluate the flagged anomalies, thereby incorporating human feedback.

Anomaly detection↗

Reconfigurable Hardware for Compressing Hyperspectral Image Data

High-speed, low-power, reconfigurable electronic hardware has been developed to implement ICER-3D, an algorithm for compressing hyperspectral-image data. The algorithm and parts thereof have been the topics of several NASA Tech Briefs articles, including Context Modeler for Wavelet Compression of Hyperspectral Images (NPO-43239) and ICER-3D Hyperspectral Image Compression Software (NPO-43238), which appear elsewhere in this issue of NASA Tech Briefs. As described in more detail in those articles, the algorithm includes three main subalgorithms: one for computing wavelet transforms, one for context modeling, and one for entropy encoding. For the purpose of designing the hardware, these subalgorithms are treated as modules to be implemented efficiently in field-programmable gate arrays (FPGAs). The design takes advantage of industry- standard, commercially available FPGAs. The implementation targets the Xilinx Virtex II pro architecture, which has embedded PowerPC processor cores with flexible on-chip bus architecture. It incorporates an efficient parallel and pipelined architecture to compress the three-dimensional image data. The design provides for internal buffering to minimize intensive input/output operations while making efficient use of offchip memory. The design is scalable in that the subalgorithms are implemented as independent hardware modules that can be combined in parallel to increase throughput. The on-chip processor manages the overall operation of the compression system, including execution of the top-level control functions as well as scheduling, initiating, and monitoring processes. The design prototype has been demonstrated to be capable of compressing hyperspectral data at a rate of 4.5 megasamples per second at a conservative clock frequency of 50 MHz, with a potential for substantially greater throughput at a higher clock frequency. The power consumption of the prototype is less than 6.5 W. The reconfigurability (by means of reprogramming) of the FPGAs makes it possible to effectively alter the design to some extent to satisfy different requirements without adding hardware. The implementation could be easily propagated to future FPGA generations and/or to custom application-specific integrated circuits.

Aranki, Nazeeh↗

Modeling Framework for the Assessment of a Sustainable Hydrogen Production and Supply Chain Network in California

The cost-effective and sustainable deployment of hydrogen supply and demand networks, especially in large economic regions like California, can be challenging considering the spatial-temporal availability and variability of the different actors across the network such as production processes, distribution modes, and end-users. In this presentation, we will provide an overview and demonstration of a modeling framework used to assess the environmental, economic, and human health impacts of plausible hydrogen production and supply chain networks in California. Scenarios focus on green hydrogen production pathways using water electrolysis and biomass gasification. End-use applications included in the model are transit, medium and heavy-duty trucking, port authorities, and power and aviation companies that currently consume natural gas, diesel, and aviation fuel for their day-to-day operation. Representative locations for hydrogen production and end-use are based on recent projections of the hydrogen economy in California. All mass and energy flows, as well as estimated emissions, are based on H2A process model designs and projections of technology performance, literature review, and LBNL process, economic and life cycle modeling, and not on company data for the sake of this presentation. Human health impacts are included following methodologies developed for the University of California Irvine HyDeal project. Life cycle phases associated with hydrogen production include feedstock preparation (water and biomass), energy production and consumption (renewable, grid, and combination of renewable and grid electricity), maintenance (chemical utilization in electrolysis and natural gas combustion in gasification), carbon sequestration, hydrogen storage (compression and liquefaction), and distribution (truck and pipeline). We apply the framework utilizing California specific emission factors, financial data, and human health damages and explore the impact of network characteristics on results. Example variations include: the inclusion of policy incentives or not, different representations of the electricity grid and source, electrolysis versus gasification versus combinations of both for production, liquefaction versus compression based on producer capacity cutoffs, transportation truck versus pipeline based on existing infrastructure, and ultimate end use. Comparison of these different scenarios can help inform future projects by demonstrating the trade-offs among environmental, economic, and human health impacts. This model, automated in R, is a starting platform upon which new analysis, modeling capabilities, locations, and emission factors can be rapidly tested and integrated.

Zaki, Mohammed Tamim↗

A Preferences Corpus and Annotation Scheme for Human-Guided Alignment of Time-Series GPTs

The process of time-series forecasting such as predicting trajectories of silicon content in blast furnaces is a difficult task. Most time-series approaches today focus on scalar-type MSE loss optimization. This optimization approach, while widely common, could benefit from the use of human expert or process-level preferences. In this paper, we introduce a novel alignment and fine-tuning approach that involves learning from a corpus of preferred and dis-preferred time-series prediction trajectories. Our contributions include (1) a preference annotation pipeline for time-series forecasts, (2) the application of Score-based Preference Optimization (SPO) to train decoder-only transformers from preferences, and (3) results showing improvements in forecast quality. The approach is validated on both proprietary blast furnace data and the UCI Appliances Energy dataset. The proposed preference corpus and training strategy offer a new option for fine-tuning sequence models in industrial settings.

DPO↗

Lab Scale Demonstration of Pipeline Third-Party Damage Classification Using Convolutional Neural Networks

This research aims to propose a simple experiment for third party damage classification problem by generating a dataset of third-party damage events on a laboratory scale utilizing single mode-multi mode-single mode (SMS) fiber acoustic sensor. The sound samples representative of various third-party activities, such as vehicle movements, excavation, and digging, were sourced from open-source databases. These samples were then played through a speaker in proximity to an SMS sensor, and the resultant fiber acoustic vibration data were recorded for each event. This process yielded a collection of 200 samples across 13 distinct third-party events. Convolutional Neural Networks (CNNs) were employed to classify these samples into their respective categories, and an accuracy exceeding 97% was obtained from our results.

Bukka, Sandeep Reddy↗

Earth Science Data Processing With Nextflow

Earth science data processing tasks present many challenges. These tasks often process large input datasets and require scores of CPU-hours to generate results. All but the simplest tasks will be decomposed into a series of computational or data manipulation steps, also known as a scientific workflow. In order to reduce the burden of orchestrating and running the dependent processing steps, a workflow execution engine is required. This poster describes the lessons learned by the CLARREO Pathfinder (CPF) team while developing multiple scientific workflows and utilizing the open-source Nextflow engine to execute them in a cloud computing environment. The Nextflow engine is designed with the following stated goals: first, the engine does not dictate how individual steps in the task are implemented (i.e. it is language and interface agnostic); second, the engine supports easy configuration and modularity at the workflow level so that others can easily execute our workflows to reproduce results; lastly, the engine eases development by transparently scaling execution from local to remote environments. Nextflow was developed for the bioinformatics domain but is a good fit for other scientific workflows where the overall task is well-described by a dataflow diagram. The CPF team has developed Nextflow pipelines (i.e. scientific workflows) to simulate CLARREO radiance, generate large look-up tables for inter-calibration algorithms, and generate L4 intercalibration data products. These pipelines consume from single-digits to hundreds of thousands of CPU-hours. In the development and evolution of these pipelines we have discovered many design patterns, pitfalls, and solutions to common problems. Our goal is to demonstrate important aspects of how to design, implement, run, and ultimately share Nextflow pipelines in the domain of Earth science.

Aron D Bartle↗

PIPES (Pipeline for Integrated Projects in Energy Systems) [SWR-24-89]

The Pipeline for Integrated Projects in Energy Systems (PIPES) is a comprehensive project, data, and workflow management tool designed for integrated modeling teams. PIPES facilitates the management of data requirements, tasks, and progress tracking, serving as a higher-level integration layer that works across various data and modeling software. This tool integrates models, data, and tools to perform large-scale, integrated analysis work at scale. PIPES is designed to streamline integrated modeling projects, enhance collaboration, and ensure the quality and efficiency of data management and workflow processes. https://github.com/nrel-pipes/pipes-api https://github.com/nrel-pipes/pipes-web https://github.com/nrel-pipes/nrel-pipes

Gu, Jianli↗

Scalable Multiprocessor for High-Speed Computing in Space

A report discusses the continuing development of a scalable multiprocessor computing system for hard real-time applications aboard a spacecraft. "Hard realtime applications" signifies applications, like real-time radar signal processing, in which the data to be processed are generated at "hundreds" of pulses per second, each pulse "requiring" millions of arithmetic operations. In these applications, the digital processors must be tightly integrated with analog instrumentation (e.g., radar equipment), and data input/output must be synchronized with analog instrumentation, controlled to within fractions of a microsecond. The scalable multiprocessor is a cluster of identical commercial-off-the-shelf generic DSP (digital-signal-processing) computers plus generic interface circuits, including analog-to-digital converters, all controlled by software. The processors are computers interconnected by high-speed serial links. Performance can be increased by adding hardware modules and correspondingly modifying the software. Work is distributed among the processors in a parallel or pipeline fashion by means of a flexible master/slave control and timing scheme. Each processor operates under its own local clock; synchronization is achieved by broadcasting master time signals to all the processors, which compute offsets between the master clock and their local clocks.

Lux, James↗

Lab-Scale Demonstration of Pipeline Third-Party Damage Classification Using Convolutional Neural Networks

This research aims to mitigate the challenges of field tests for classification of third-party damages by generating a dataset of third-party damage events on a laboratory scale utilizing single mode-multi mode-single mode (SMS) fiber acoustic sensor. The sound samples representative of various third-party activities, such as vehicle movements, excavation, and digging, were sourced from open-source databases. These samples were then played through a speaker in proximity to an SMS sensor, and the resultant fiber acoustic vibration data were recorded for each event. This process yielded a collection of 200 samples across 13 distinct third-party events. Convolutional Neural Networks (CNNs) were employed to classify these samples into their respective categories, and an accuracy exceeding 97% was obtained from our results.

Bukka, Sandeep Reddy↗

A bit-serial VLSI array processing chip for image processing

An array processing chip integrating 128 bit-serial processing elements (PEs) on a single die is discussed. Each PE has a 16-function logic unit, a single-bit adder, a 32-b variable-length shift register, and 1 kb of local RAM. Logic in each PE provides the capability to mask PEs individually. A modified grid interconnection scheme allows each PE to communicate with each of its eight nearest neighbors. A 32-b bus is used to transfer data to and from the array in a single cycle. Instruction execution is pipelined, enabling all instructions to be executed in a single cycle. The 1-micron CMOS design contains over 1.1 x 10 to the 6th transistors on an 11.0 x 11.7-mm die.

Heaton, Robert↗

Telecommunications issues of intelligent database management for ground processing systems in the EOS era

Future NASA earth science missions, including the Earth Observing System (EOS), will be generating vast amounts of data that must be processed and stored at various locations around the world. Here we present a stepwise-refinement of the intelligent database management (IDM) of the distributed active archive center (DAAC - one of seven regionally-located EOSDIS archive sites) architecture, to showcase the telecommunications issues involved. We develop this architecture into a general overall design. We show that the current evolution of protocols is sufficient to support IDM at Gbps rates over large distances. We also show that network design can accommodate a flexible data ingestion storage pipeline and a user extraction and visualization engine, without interference between the two.

Touch, Joseph D.↗

Time Series Analysis in the Search for Other Worlds Through Transit Photometry

The Kepler Mission launched in June 2009 to commence NASA's first mission to search for potentially habitable, Earth-size planets orbiting Sun-like stars. Kepler discovered explanets via the transit method: searching for minute (100 ppm) drops in brightness lasting 1 - 13 hours corresponding to occasions where the planet crosses the face of its host star from Kepler's point of view. The exquisite precision required to carry out the Kepler mission (20 ppm in 6.5 hours) pushed astronomical time series analysis to the limits, and motivated the development of novel algorithmic approaches. Transit signatures of rocky planets are often dwarfed by the intrinsic stellar variability, which is not white noise, and often is non-stationary, and by instrumental systematic effects, which can include transients and electronic artifacts. Surmounting this challenging regime of weak, temporally compact, periodic signals in observation noise with strong systematics and other sources of variability motivated the development of 1) an overcomplete, non-decimated, wavelet-based matched filter to jointly estimate the properties of the non-stationary, non-white observation noise process, and 2) a multi-scale, maximum a posteriori (msMAP) approach to identifying and removing instrumental systematic effects. After over nine years of observations, the Kepler spacecraft finally ran out of fuel in November 2018, ending its data collection activities. Over 2300 planets were discovered by Kepler in its primary mission, and over 355 have been discovered by K2, the repurposed mission that followed Kepler's primary mission after the loss of a second reaction wheel in May 2013. We have ported the Kepler science pipeline for the Transiting Exoplanet Survey Satellite (TESS) Mission, which began science observations in July 2019, and report initial results and performance of the modified science pipeline.The Kepler and TESS Missions are supported by NASA's Science Mission Directorate.

transit surveys↗

Web Exploration Tools for a Fast Federated Optical Survey Database

We implemented several new web-based tools to improve the efficiency and versatility of access to the APS Catalog of the POSS I (Palomar Observatory-National Geographic Sky Survey) and its associated image database. The most important addition was a federated database system to link the APS Catalog and image database into one Internet-accessible database. With the FDBS, the queries and transactions on the integrated database are performed as if it were a single database. We installed Myriad the FDBS developed by Professor Jaideep Srivastava and members of his group in the University of Minnesota Computer Science Department. It is the first system to provide schema integration, query processing and optimization, and transaction management capabilities in a single framework. The attached figure illustrates the Myriad architecture. The FDBS permits horizontal access to the data, not just vertical. For example, for the APS, queries can be made not only by sky position, but also by any parameter present in either of the databases. APS users will be able to produce an image of all the blue galaxies and stellar sources for comparison with x-ray source error ellipses from AXAF (X Ray Astrophysics Facility) (Chandra) for example. The FDBS is now available as a beta release with the appropriate query forms at our web site. While much of our time was occupied with adapting Myriad to the APS environment, we also made major changes in Star Base, our DBMS for the Catalog, at the web interface to improve its efficiency for issuing and processing queries. Star Base is now three times faster for large queries. Improvements were also made at the web end of the image database for faster access; although work still needs to be done to the image database itself for more efficient return with the FDBS. During the past few years, we made several improvements to the database pipeline that creates the individual plate databases queries by StarBase. The changes include improved positions especially for galaxies, using a new median centroider and integrated magnitudes for galaxies with an improved density-to-intensity calibration with a "sky" background subtraction. In the original version of StarBase the object classification fainter than 19.5-20.0 mag., was an extrapolation of the networks trained on brighter objects. We have used a new catalog of galaxies at the NGP to train a neural network on objects fainter than 20th mag. This improved classification is used in the new version of StarBase. We have also added a FITS table option for the returned data from queries on the object catalog. The APS image database includes images in both colors so we have added a tool for querying the image database in both colors simultaneously. The images can be displayed in parallel or blinked for comparison.

Humphreys, Roberta M.↗

Reinforcement Learning for In-Spill Optimization of the Mu2e Resonant Extraction: Compensating Non-Stationarity

We present design considerations and challenges for the fast machine learning component of a third-order resonant beam extraction regulation system being commissioned to deliver steady beam rates to the mu2e experiment at Fermilab. Dedicated quadrupoles drive the tune toward the 29/3 resonance each spill, extracting beam at kV multiwire septa. The overall Spill Regulation System consists of (1) a “slow” process using ~100-spill averages to adjust the base quad ramp infrequently, (2) a feedforward harmonic content compensator, and (3) the “fast” ML agent reacting during each ongoing spill with on-the-fly additive corrections to the sum of (1) and (2). We have demonstrated improved beam-rate steadying for a fast ML agent compared to a PID controller using a quasi-physical spill simulation, and demonstrated distillation of that simulation into a predictive surrogate model. Current work includes a data-and-training pipeline to generate data-aware surrogates with real-world dynamics, even as the dynamics shift unpredictably. The surrogates are to act as RL environments against which to train our fast ML control agents before deploying them on FPGA in the live system. Further current efforts focus on modeling and controlling beam loss around the storage ring, understanding additional available hardware inputs to the model, and the interplay of these with beam-steadying performance.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗