Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data reduction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Quick Guide for SNS-NSE Reduction Software DrSpine (Data Reduction for Spin Echo)

This is a general guide to quick access the most common commands for DrSpine (Data Reduction for Spin Echo) data reduction software of the SNS-NSE spectrometer. Over the length of the manuscript the symbol “ $ “ represents the OS shell prompt, i.e., the beginning of a command line in a shell terminal, while the symbol “-->” represents the beginning of a command line in the DrSpine reduction software, after the activation of the reduction environment.

97 MATHEMATICS AND COMPUTING↗

Real-time data reduction at 100 Tbps: Challenge and opportunity for AI-based data reduction for next-generation large-scale nuclear physics collider experiment

The modern large-scale nuclear physics (NP) experiments in high-energy particle colliders utilize streaming-readout electronics to digitize detector response at O(100) Tbps bandwidth. Prominent examples at Brookhaven National Lab (BNL) include the sPHENIX experiment at Relativistic Heavy Ion Collider (RHIC), which is under construction, and the experiments proposed for the Electron-Ion Collider (EIC), planned for the 2030s . One of the main challenges for these streaming readout systems is to manage the data rate with sufficient data reduction in real time so the end-data fit persistent storage for offline analysis, which is typically at O(1000) times smaller and O(100) Gbps. Such data reduction traditionally is achieved via real-time high level triggers, which select and save a small subset of collisions of interest. Although triggering is applicable to high energy collider experiments such as those at the Large Hardron Collider at CERN, it is insufficient for these nuclear physics experiments which study diverse collision topologies. And traditional triggering approach is inefficient to preserve the max information harvested from the operation of colliders that costs O(100)M per year to DOE. Meanwhile, in recent years, ML-based high-throughput data reduction has emerged as a promising approach to efficiently preserve max information for a given space of persistent storage, e.g. via AI data compression, feature extraction, and noise filtering.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

HPDR: High-Performance Portable Scientific Data Reduction Framework

The rapid growth in scientific data generation is outpacing advancements in computing systems necessary for efficient storage, transfer, and analysis, particularly in the context of exascale computing. With the deployment of first-generation exascale computing systems and next-generation experimental facilities, this gap is widening and necessitates effective data reduction techniques to manage enormous data volumes. Over the past decade, various data reduction methods, including lossless compression, error-controlled lossy compression, and data refactoring, have been developed to accelerate I/O in scientific workflows. Despite significant reductions in data volume, these methods introduce considerable computational overhead, which can become the new bottleneck in data processing. To mitigate this, GPU-accelerated data reduction algorithms have been introduced. However, challenges remain in their integration into exascale workflows, including limited portability across different GPU architectures, substantial memory transfer overhead, and reduced scalability on dense multi-GPU systems. To address these challenges, we propose HPDR, a high-performance and portable data reduction framework. HPDR is designed to enable the execution of state-of-the-art reduction algorithms across diverse processor architectures while reducing memory transfer overhead to 2.3 % of the original, resulting in up to 3.5× faster throughput compared to existing solutions. It also achieves up to 96% of the theoretical speedup in multi-GPU settings. In addition, evaluations on accelerating I/O operations at scale up to 1,024 nodes of the Frontier supercomputer demonstrate that HPDR can achieve up to 103 TB/s reduction throughput, providing up to 4× acceleration in parallel I/O performance compared to existing data reduction routines. This work highlights the potential of HPDR to significantly enhance data reduction efficiency in exascale computing environments.

Chen, Jieyang [University of Oregon]↗

Data reduction for low energy nuclear physics experiments using data frames

Low energy nuclear physics experiments are transitioning towards fully digital data acquisition systems. Realizing the gains in flexibility afforded by these systems relies on equally flexible data reduction techniques. In this paper, methods utilizing data frames and in-memory techniques to work with data, including data from self-triggering, digital data acquisition systems, are discussed within the context of a Python package, sauce. It is shown that data frame operations can encompass common analysis needs and allow interactive data analysis. Two event building techniques, dubbed referenced and referenceless event building, are shown to provide a means to transform raw list mode data into correlated multi-detector events. These techniques are demonstrated in the analysis of two example data sets.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Data reduction considerations for the burning velocity of spherical constant volume flames of R32 (CH 2 F 2 ) with air

Here, the present work explores data reduction techniques for the measurement of the laminar burning velocities of R32(CH 2 F 2 )-air mixtures using a constant volume combustion device, in which the pressure-time history is the only measured parameter. To allow clear assessment of the accuracy of the data reduction methods, the pressure-time histories used for analysis are synthetically generated via a detailed numerical simulation employing full kinetics and with and without an optically-thin radiation model. Various data reduction models are employed, including a two-zone model and two multi-zone models, and these are compared with the results from the burning velocity obtained from the output of the numerical simulation. The data reduction schemes are shown to be accurate if the same radiation model is employed in the data reduction as was used in the flame simulation to generate the pressure trace used for post-processing. If the incorrect radiation model is employed, however, the errors can be quite large. The effects of stretch, radiation, and different data post-processing methodologies are explored and the errors quantified. Stretch is shown to be important for the early stages and the selected data range that is used for extrapolation has a significant effect on the extrapolated burning velocity. However, with an appropriate choice of data considered for extrapolation, the prediction of the unstretched burning velocity can be quite accurate.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Performance Improvements on SNS and HFIR Instrument Data Reduction Workflows Using Mantid

Performance of data reduction workflows at the High Flux Isotope Reactor (HFIR) and the Spallation Neutron Source (SNS) at Oak Ridge National Laboratory (ORNL) is mainly determined by the time spent loading raw measurement events stored in large and sparse datasets. This paper describes: (1) our long-term view to leverage SNS and HFIR data management needs with our experience at ORNL’s world-class high performance computing (HPC) facilities, and (2) our short-term efforts to speed up current workflows using Mantid, a data analysis and reduction community framework used across several neutron scattering facilities. We show that minimally invasive short-term improvements in metadata management have a moderate impact in speeding up current production workflows. We propose a more disruptive domain-specific solution: the No Cost Input Output (NCIO) framework, we provide an overview, the risks and challenges in NCIO’s adoption by HFIR and SNS stakeholders.

Godoy, William↗

Overview of IMPACT Data Acquisition System and Data Reduction Process

This report documents the development of the data acquisition system (DAS) and data reduction methodologies for the Irradiated Material Property Accelerated Characterization Test (IMPACT) experiment at the Advanced Test Reactor (ATR). The IMPACT experiment is designed to enable in-pile measurement of thermal conductivity in metallic nuclear fuels, specifically U-10Zr, using an instrumented thermal conductivity probe. The DAS supports both passive temperature monitoring and active thermal interrogation of the probe through controlled AC and DC excitation. Significant modifications to laboratory-scale systems were required to accommodate the higher resistance paths associated with the in-pile application. Custom electronics and relay-controlled measurement sequencing were developed to enable the measurement and sufficient power delivery to the sensing region. A reduced-order, axisymmetric thermal model based on the thermal quadrupoles method is presented to support data interpretation. This model enables efficient evaluation of transient heat transfer behavior and facilitates solution of the inverse problem required to extract thermal properties from measured signals. Multiple boundary condition formulations are discussed to address varying experimental time scales and geometries. Additionally, machine learning techniques are introduced to support data reduction and improve confidence in inverse solutions. Convolutional neural networks are applied to identify the presence of gas gaps and other evolving geometric features that significantly impact thermal response during irradiation. These efforts contribute to the broader integration of digital twin frameworks and real-time modeling capabilities within the Advanced Fuels Campaign.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

On single-crystal total scattering data reduction and correction protocols for analysis in direct space

Data reduction and correction steps and processed data reproducibility in the emerging single-crystal total-scattering-based technique of three-dimensional differential atomic pair distribution function (3D-ΔPDF) analysis are explored. All steps from sample measurement to data processing are outlined using a crystal of CuIr 2 S 4 as an example, studied in a setup equipped with a high-energy X-ray beam and a flat-panel area detector. Computational overhead as pertains to data sampling and the associated data-processing steps is also discussed. Various aspects of the final 3D-ΔPDF reproducibility are explicitly tested by varying the data-processing order and included steps, and by carrying out a crystal-to-crystal data comparison. Situations in which the 3D-ΔPDF is robust are identified, and caution against a few particular cases which can lead to inconsistent 3D-ΔPDFs is noted. Although not all the approaches applied herein will be valid across all systems, and a more in-depth analysis of some of the effects of the data-processing steps may still needed, the methods collected herein represent the start of a more systematic discussion about data processing and corrections in this field.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Data reduction through optimized scalar quantization for more compact neural networks

Raw data generation for several existing and planned large physics experiments now exceeds TB/s rates, generating untenable data sets in very little time. Those data often demonstrate high dimensionality while containing limited information. Meanwhile, Machine Learning algorithms are now becoming an essential part of data processing and data analysis. Those algorithms can be used offline for post processing and post data analysis, or they can be used online for real time processing providing ultra low latency experiment monitoring. Both use cases would benefit from data throughput reduction while preserving relevant information: one by reducing the offline storage requirements by several orders of magnitude and the other by allowing ultra fast online inferencing with low complexity Machine Learning models. Moreover, reducing the data source throughput also reduces material cost, power and data management requirements. In this work we demonstrate optimized nonuniform scalar quantization for data source reduction. This data reduction allows lower dimensional representations while preserving the relevant information of the data, thus enabling high accuracy Tiny Machine Learning classifier models for online fast inferences. We demonstrate this approach with an initial proof of concept targeting the CookieBox, an array of electron spectrometers used for angular streaking, that was developed for LCLS-II as an online beam diagnostic tool. We used the Lloyd-Max algorithm with the CookieBox dataset to design an optimized nonuniform scalar quantizer. Optimized quantization lets us reduce input data volume by 69% with no significant impact on inference accuracy. When we tolerate a 2% loss on inference accuracy, we achieved 81% of input data reduction. Finally, the change from a 7-bit to a 3-bit input data quantization reduces our neural network size by 38%.

97 MATHEMATICS AND COMPUTING↗

Efficient Data Management in Neutron Scattering Data Reduction Workflows at ORNL

Oak Ridge National Laboratory (ORNL) experimental neutron science facilities produce 1.2 TB a day of raw event-based data that is stored using the standard metadata-rich NeXus schema built on top of the HDF5 file format. Performance of several data reduction workflows is largely determined by the amount of time spent on the loading and processing algorithms in Mantid, an open-source data analysis framework used across several neutron sciences facilities around the world. The present work introduces new data management algorithms to address identified input output (I/O) bottlenecks on Mantid. First, we introduce an in-memory binary-tree metadata index that resemble NeXus data access patterns to provide a scalable search and extraction mechanism. Second, data encapsulation in Mantid algorithms is optimally redesigned to reduce the total compute and memory runtime footprint associated with metadata I/O reconstruction tasks. Results from this work show speed ups in wall-clock time on ORNL data reduction workflows, ranging from 11% to 30% depending on the complexity of the targeted instrument-specific data. Nevertheless, we highlight the need for more research to address reduction challenges as experimental data volumes increase.

Godoy, William↗

Machine Learning-Based Extreme Data Reduction for Prompt Supernova Pointing at DUNE

One of the goals of the Deep Underground Neutrino Experiment (DUNE) is to use the massive underground liquid argon time projection chamber (LArTPC) detectors at its far site for multimessenger astronomy (MMA), in the detection of neutrinos from core-collapse supernovae (SNe). Its current baseline trigger strategy detects activity in the detector that is consistent with supernova (SN) neutrinos and saves the raw data for further offline analysis but provides no prompt pointing information crucial for optical follow-ups by other observatories. This approach is based on the assumption that prompt pointing determination using raw data is computationally prohibitive. In this article, we demonstrate a proof-of-concept based on applying extreme data reduction on the buffered SN data in the DUNE data acquisition (DAQ) system’s front-end computers using a machine learning (ML) workflow. This reduces the data by ~5 orders of magnitude, allowing a full track reconstruction to be carried out quickly on a single server. The total time to perform the ML-based data reduction and the full track reconstruction is less than the time to transfer the SN data back to Fermilab or a high-performance computing (HPC) center. This shows that prompt processing of raw SN data is possible and, in fact, trivial once the data have been reduced to reject radiological backgrounds, paving the way to a high-quality SN pointing trigger that is based on fully reconstructed data instead of trigger primitives (TPs).

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Data Reduction for Science: Brochure from the Advanced Scientific Computing Research Workshop

Data reduction for science holds promise for addressing the challenges of moving, storing, and processing massive data sets produced by the scientific community. Pursuing the PRDs outlined here will enable advances in data streaming, fast feedback and/or autonomous control of experiments, and faster time to scientific insight. These advances will result in a significant improvement in the ability to transport, store, process and interpret experimental, observational, and computational data.

97 MATHEMATICS AND COMPUTING↗

Hybrid learning techniques for scientific data reduction with performance guarantees

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

BioXTAS RAW 2 : new developments for a free open-source program for small-angle scattering data reduction and analysis

BioXTAS RAW is a free open-source program for reduction, analysis and modelling of biological small-angle scattering data. Here, the new developments in RAW version 2 are described. These include improved data reduction using pyFAI ; updated automated Guinier fitting and D max finding algorithms; automated series ( e.g. size-exclusion chromatography coupled small-angle X-ray scattering or SEC-SAXS) buffer- and sample-region finding algorithms; linear and integral baseline correction for series; deconvolution of series data using regularized alternating least squares ( REGALS ); creation of electron-density reconstructions using electron density via solution scattering ( DENSS ); a comparison window showing residuals, ratios and statistical comparisons between profiles; and generation of PDF reports with summary plots and tables for all analysis. Furthermore, there is now a RAW API, which can be used without the graphical user interface (GUI), providing full access to all of the functionality found in the GUI. In addition to these new capabilities, RAW has undergone significant technical updates, such as adding Python 3 compatibility, and has entirely new documentation available both online and in the program.

97 MATHEMATICS AND COMPUTING↗

Variational autoencoders for at-source data reduction and anomaly detection in high energy particle detectors

Detectors in next-generation high-energy physics experiments face several daunting requirements, such as high data rates, damaging radiation exposure, and stringent constraints on power, space, and latency. To address these challenges, machine learning in readout electronics can be leveraged for smart detector designs, enabling intelligent inference and data reduction at-source. Variational autoencoders (VAEs) offer a variety of benefits for front-end readout; an on-sensor encoder can perform efficient lossy data compression while simultaneously providing a latent space representation that can be used for anomaly detection. Results are presented from low-latency and resource-efficient VAEs for front-end data processing in a futuristic silicon pixel detector. Encoder-based data compression is found to preserve good performance of off-detector analysis while significantly reducing the off-detector data rate as compared to a similarly sized data filtering approach. Furthermore, the latent space information is found to be a useful discriminator in the context of real-time sensor defect monitoring. Together, these results highlight the multifaceted utility of autoencoder-based front-end readout schemes and motivate their consideration in future detector designs.

47 OTHER INSTRUMENTATION↗

Data reduction considerations for spherical R-32(CH 2 F 2 )-air flame experiments

The burning velocities of mixtures of refrigerant R-32 (CH 2 F 2 ) with air over a range of equivalence ratios are studied via shadowgraph images of spherically expanding flames (SEFs) in a large, optically accessible spherical chamber at constant pressure. Numerical simulations of the 1-D, unsteady, spherical flames incorporating an optically thin radiation model and detailed kinetics accurately predict experimental results for a range of equivalence ratios. For these low burning velocity flames, the effects of stretch and radiation occur simultaneously and make extraction of the unstretched burning velocity from the experimental data difficult. Different data reduction approaches are shown to have large effects on the burning velocities inferred from the experiments and simulations. A new flame radius tracking approach for the experimental images is shown to provide improved agreement of the burned gas velocity variation with stretch predicted by the simulations and helps to compensate for mild flame distortion due to buoyancy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗