Search NASASearch

SEARCH · Search NASA

Results for “Data reduction methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

HPDR: High-Performance Portable Scientific Data Reduction Framework

The rapid growth in scientific data generation is outpacing advancements in computing systems necessary for efficient storage, transfer, and analysis, particularly in the context of exascale computing. With the deployment of first-generation exascale computing systems and next-generation experimental facilities, this gap is widening and necessitates effective data reduction techniques to manage enormous data volumes. Over the past decade, various data reduction methods, including lossless compression, error-controlled lossy compression, and data refactoring, have been developed to accelerate I/O in scientific workflows. Despite significant reductions in data volume, these methods introduce considerable computational overhead, which can become the new bottleneck in data processing. To mitigate this, GPU-accelerated data reduction algorithms have been introduced. However, challenges remain in their integration into exascale workflows, including limited portability across different GPU architectures, substantial memory transfer overhead, and reduced scalability on dense multi-GPU systems. To address these challenges, we propose HPDR, a high-performance and portable data reduction framework. HPDR is designed to enable the execution of state-of-the-art reduction algorithms across diverse processor architectures while reducing memory transfer overhead to 2.3 % of the original, resulting in up to 3.5× faster throughput compared to existing solutions. It also achieves up to 96% of the theoretical speedup in multi-GPU settings. In addition, evaluations on accelerating I/O operations at scale up to 1,024 nodes of the Frontier supercomputer demonstrate that HPDR can achieve up to 103 TB/s reduction throughput, providing up to 4× acceleration in parallel I/O performance compared to existing data reduction routines. This work highlights the potential of HPDR to significantly enhance data reduction efficiency in exascale computing environments.

Chen, Jieyang [University of Oregon]

Reactor Containment Passive Safety Analysis: Steam Condensation in Presence of Non-condensable Gas Scaled Experiment and Modeling

This study presents steam condensation scaled experiments and semi-empirical models in presence of nitrogen (N)—a noncondensable gas (NCG), simulating air in the reactor containment—to support water-cooled small modular reactors (SMRs) passive containment cooling system (PCCS) design and analysis. Previous experimental studies on PCCS are focused on fixed and smaller tube (mostly 2-in.) geometries and specific test condition variations, bringing challenges with geometric scaling and mismatching with SMR prototypic design. To address these challenges, this study presents steam condensation test dataset obtained from three scaled test sections of 1-, 2-, and 4-in.-diameter steam condensers with an annular/jacket cooling of 2-, 3-, and 6 in.-diameter tubes, respectively. Test data were collected for steam ranges from 58 to 63 kg/hr., and NCG flow of 4.4 to 13.3 kg/hr. Annular cooling water flow was varied to obtain required testing conditions of saturated steam inlet and fully condensed outlet. Axial temperature test data of bulk cooling water, steam and condensate were collected by thermocouples for three test sections and various steam-NCG mixing/testing conditions. A standard data reduction method was adopted—utilizing iterative and nodalized mass and heat transfer calculation—to estimate axial local heat fluxes, heat transfer coefficients (HTCs), condensation rates, film thickness, and Nusselt number. Based on the obtained dataset semi-empirical model results—a ratio of experimental and Nusselt’s theoretical HTC are presented. Such results and findings are supportive of developing scaled-up testing facility, to enable model validations and accelerate next generation of reactors development and deployment

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS

Presentation: Reactor Containment Passive Safety Analysis: Steam Condensation in Presence of Non-condensable Gas Scaled Experiment and Modeling

This study presents steam condensation scaled experiments and semi-empirical models in presence of nitrogen--a noncondensable gas (NCG), simulating air in the reactor containment--to support water-cooled small modular reactors (SMRs) passive containment cooling system (PCCS) design and analysis. Previous experimental studies on PCCS are focused on fixed and smaller tube (mostly 2-in.) geometries and specific test condition variations, bringing challenges with geometric scaling and mismatching with SMR prototypic design. To address these challenges, this study presents steam condensation test dataset obtained from three scaled test sections of 1-, 2-, and 4-in.-diameter steam condensers with an annular/jacket cooling of 2-, 3-, and 6 in.-diameter tubes, respectively. Test data were collected for steam ranges from 58 to 63 kg/hr., and NCG flow of 4.4 to 13.3 kg/hr. Annular cooling water flow was varied to obtain required testing conditions of saturated steam inlet and fully condensed outlet. Axial temperature test data of bulk cooling water, steam and condensate were collected by thermocouples for three test sections and various steam-NCG mixing/testing conditions. A standard data reduction method was adopted--utilizing iterative and nodalized mass and heat transfer calculation to estimate axial local heat fluxes, heat transfer coefficients (HTCs), condensation rates, film thickness, and Nusselt number. Based on the obtained dataset semi-empirical model results--a ratio of experimental and Nusselt's theoretical HTC are presented. Such results and findings are supportive of developing scaled-up testing facility, to enable model validations and accelerate next generation of reactors development and deployment.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS

DENNIS: a design and analysis tool for dynamic material x-ray diffraction experiments

We present DENNIS (Diffraction Experiment desigN and aNalysiS): a graphical software tool useful for the design and analysis of dynamic x-ray diffraction experiments, such as those performed on the Z Pulsed Power Facility, Thor Pulsed Power Generator, and Dynamic Compression Sector (DCS) of the Advanced Photon Source. DENNIS provides rapid powder and single-crystal diffraction pattern predictions and powder diffraction pattern image integration in three-dimensional geometries. Additional features include crystallographic information file reading, image processing, and synthetic diffraction pattern image generation. We overview the software's capabilities, detail the prediction and integration methodologies, and provide example implementations on Z and DCS experiments.

47 OTHER INSTRUMENTATION

Embedded FPGA developments in 130 nm and 28 nm CMOS for machine learning in particle detector readout

Embedded field programmable gate array (eFPGA) technology allows the implementation of reconfigurable logic within the design of an application-specific integrated circuit (ASIC). This approach offers the low power and efficiency of an ASIC along with the ease of FPGA configuration, particularly beneficial for the use case of machine learning in the data pipeline of next-generation collider experiments. An open-source framework called "FABulous" was used to design eFPGAs using 130 nm and 28 nm CMOS technology nodes, which were subsequently fabricated and verified through testing. The capability of an eFPGA to act as a front-end readout chip was assessed using simulation of high energy particles passing through a silicon pixel sensor. A machine learning-based classifier, designed for reduction of sensor data at the source, was synthesized and configured onto the eFPGA. A successful proof-of-concept was demonstrated through reproduction of the expected algorithm result on the eFPGA with perfect accuracy. Finally, further development of the eFPGA technology and its application to collider detector readout is discussed.

47 OTHER INSTRUMENTATION

Bringing Different Views Together: A Hybrid Cooperative Perception Framework for Connected Autonomous Vehicles

Cooperative perception will be essential for connected autonomous vehicles to enhance object recognition and optimize path planning by sending data information about the surrounding environment. However, an inherent challenge in existing systems is the high bandwidth cost of transmitting information in real-time, which restricts cooperative perception’s practicality. Here, this work presents a hybrid cooperative perception fusion framework aimed at mitigating this issue by optimizing data transmission according to available bandwidth or through data reduction techniques. Our methods ensure that vehicles can rapidly transmit high-confidence data without overwhelming the network. Experimental results indicate that our methodology substantially diminishes data transmission sizes while maintaining object detection accuracy. For cooperative perception in autonomous vehicle systems, our approach provides a scalable and effective way to get past the bandwidth barrier.

Carrillo, Dominic [Univ. of North Texas, Denton, T

Data reduction for low energy nuclear physics experiments using data frames

Low energy nuclear physics experiments are transitioning towards fully digital data acquisition systems. Realizing the gains in flexibility afforded by these systems relies on equally flexible data reduction techniques. In this paper, methods utilizing data frames and in-memory techniques to work with data, including data from self-triggering, digital data acquisition systems, are discussed within the context of a Python package, sauce. It is shown that data frame operations can encompass common analysis needs and allow interactive data analysis. Two event building techniques, dubbed referenced and referenceless event building, are shown to provide a means to transform raw list mode data into correlated multi-detector events. These techniques are demonstrated in the analysis of two example data sets.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Codebase release 2.0 for sauce

Low energy nuclear physics experiments are transitioning towards fully digital data acquisition systems. Realizing the gains in flexibility afforded by these systems relies on equally flexible data reduction techniques. In this paper, methods utilizing data frames and in-memory techniques to work with data, including data from self-triggering, digital data acquisition systems, are discussed within the context of a Python package, sauce. It is shown that data frame operations can encompass common analysis needs and allow interactive data analysis. Two event building techniques, dubbed referenced and referenceless event building, are shown to provide a means to transform raw list mode data into correlated multi-detector events. These techniques are demonstrated in the analysis of two example data sets.

Marshall, Caleb (ORCID:0000000211942920)

Transfer learning nonlinear plasma dynamic transitions in low dimensional embeddings via deep neural networks

Deep learning algorithms provide a new paradigm to study high-dimensional dynamical behaviors, such as those in fusion plasma systems. Development of novel, data-driven model reduction methods, coupled with detection of abnormal modes with plasma physics, opens a unique opportunity to identify plasma instabilities through automated construction of parsimonious models that can be tuned to balance accuracy and cost. Our fusion transfer learning (FTL) model demonstrates success in rapidly reconstructing nonlinear kink mode structures by learning from a limited amount of nonlinear simulation data. The knowledge transfer process leverages a pre-trained neural encoder–decoder network, initially trained on linear simulations, to effectively capture nonlinear dynamics. The low-dimensional embeddings extract the coherent structures of interest, while preserving the inherent dynamics of the complex system. Experimental results highlight FTL’s capacity to capture transitional behaviors and dynamical features in plasma dynamics—a task often challenging for conventional methods. The model developed in this study is generalizable and can be extended broadly through transfer learning to address various magnetohydrodynamics modes.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

A High-Quality Workflow for Multi-Resolution Scientific Data Reduction and Visualization

Multi-resolution methods such as Adaptive Mesh Refinement (AMR) can enhance storage efficiency for HPC applications generating vast volumes of data. However, their applicability is limited and cannot be universally deployed across all applications. Furthermore, integrating lossy compression with multi-resolution techniques to further boost storage efficiency encounters significant barriers. To this end, we introduce an innovative workflow that facilitates high-quality multi-resolution data compression for both uniform and AMR simulations. Initially, to extend the usability of multi-resolution techniques, our workflow employs a compression-oriented Region of Interest (ROI) extraction method, transforming uniform data into a multi-resolution format. Subsequently, to bridge the gap between multi-resolution techniques and lossy compressors, we optimize three distinct compressors, ensuring their optimal performance on multi-resolution data. These optimizations can improve the compression ratio of SOTA approaches by up to 3.3× under the same data quality loss. Lastly, we incorporate an advanced uncertainty visualization method into our workflow to understand the potential impacts of lossy compression. Experimental evaluation demonstrates that our workflow achieves significant compression quality improvements.

Wang, Daoce

Complexity Reduction Methods for Large-Scale Spatially Explicit Biofuels Network Design

The size and complexity of energy system optimization models have increased significantly in recent years, driven by the availability of high-resolution spatial data. We present complexity reduction and solution methods that enable us to efficiently represent high-resolution spatial data in the network design of large-scale energy systems. We aim to reduce the size and enhance the computational efficiency of network design models without sacrificing solution accuracy. Specifically, we first present how to aggregate highly granular data into larger resolutions without averaging out their specific properties through a composite-curve-based approach and then develop a method to linearly represent these curves. Second, we utilize a general clustering method to determine groups of geographically proximate biomass fields and establish a single transportation arc for all of them, reducing the number of transportation-related variables while maintaining an accurate representation of the system. Finally, we introduce a two-step algorithm that decomposes large-scale network design problems into two smaller, more manageable subproblems. We demonstrate the application of our methods using a case study of switchgrass-to-biofuels network design in the eight states of the U.S. Midwest, using realistic and highly explicit spatial data.

09 BIOMASS FUELS

Developing ML/AI Methods for High-Throughput Characterization of Multiple-Sensor Streams of Tokamak Dynamics for High-Speed Control (Final Report)

This project evaluated and developed new mathematical and algorithmic techniques capable of handling (in real-time) the growing amounts of data generated by modern fusion research. While existing numerical linear algebra (NLA) methods provide the backbone to classical data analysis and algorithms, these methods fundamentally do not port to distributed architectures nor do they allow low-latency data reduction for control. Motivated by the needs for modern fusion reactors, this project explored and implemented new numerical methods to characterize plasma dynamics, respond in real-time to discharge evolution, and to process massive-scale data accurately and rapidly more fully. This project links expertise in multiple-sensor diagnostics of tokamak plasma dynamics from Columbia University’s Plasma Physics Laboratory with expertise in massive-scale data reduction and extreme data control algorithms at Columbia University’s Data Science Institute. This interdisciplinary project (i) applied machine learning methods, (ii) implemented a properly-trained neural-network for very fast processing of high-speed plasma videography, and (ii) developed the applied mathematical methods, based on randomized-NLA (rNLA) routines, for data analysis, reduction, and real-time control. The Columbia University High Beta Tokamak-Extended Pulse (HBT-EP) facility provided data to test new algorithms and partnership with Columbia University's Data Sciences Institute evaluated the broader use of new algorithms for many challenging control applications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Online learning of quadratic manifolds from streaming data for nonlinear dimensionality reduction and nonlinear model reduction

Here, this work introduces an online greedy method for constructing quadratic manifolds from streaming data, designed to enable in situ analysis of numerical simulation data on the Petabyte scale. Unlike traditional batch methods, which require all data to be available upfront and take multiple passes over the data, the proposed online greedy method incrementally updates quadratic manifolds in one pass as data points are received, eliminating the need for expensive disk input/output operations as well as storing and loading data points once they have been processed. A range of numerical examples demonstrate that the online greedy method learns accurate quadratic manifold embeddings while being capable of processing data that far exceed common disk input/output capabilities and volumes as well as main-memory sizes.

97 MATHEMATICS AND COMPUTING

User Guide for Sample Reduction at GP-SANS

This manual is intended as a quick guide for data reduction of GP-SANS data. It includes all necessary steps to do the data reduction based on absolute calibration using the open beam method and how to transfer the reduced data to the personal computer system. If any errors are coming up so that the reduction script is not functioning as intended, please contact the instrument scientist.

97 MATHEMATICS AND COMPUTING

SCULPT (Supervised Clustering and Uncovering Latent Patterns with Training) v1

SCULPT (Supervised Clustering and Uncovering Latent Patterns with Training) is a comprehensive data visualization and analysis application focused on working with COLTRIMS (COLd Target Recoil Ion Momentum Spectroscopy) data, which is used in atomic and molecular physics experiments. The application offers several powerful features: - Data uploading and processing capabilities for COLTRIMS files - Multiple visualization methods using UMAP (Uniform Manifold Approximation and Projection) for dimensionality reduction - Interactive selection of data points across multiple views - Feature engineering through various methods: - Manual feature selection from calculated physics parameters - Deep autoencoder for dimension reduction - Genetic programming for discovering meaningful features - Mutual information-based feature selection - Multiple clustering approaches (DBSCAN, KMeans, Agglomerative) - Quality metrics for evaluating clustering results - Export capabilities for selections and generated features

Daoud, Hazem [Lawrence Berkeley National Laborator

Beam Loss Assessment Through Use of Photomultiplier Tubes

The first machine in the Fermilab Accelerator chain, the Linac, delivers a 400 MeV proton beam. The first portion of the Fermilab Linac, the Drift Tube Linac, lacks the degree of instrumentation necessary for beam tuning. To compensate for this, photomultiplier tube (PMT s) based loss monitors were installed on either side of the first two drift tube tanks, but are not yet operational. One of the main goals in this is to make PMT's operational beam loss monitors for tuning. Noise reduction and peak finding on the PMT data is a requirement for this. A method for noise reduction and peak finding has been developed and implemented to produce a consistent and stable output. Future work includes integration with ACNET to automate input and output of data for analysis of beam loss.

Waggoner, Alexander

Archetype-based Redshift Estimation for the Dark Energy Spectroscopic Instrument Survey

We present a computationally efficient galaxy archetype-based redshift estimation and spectral classification method for the Dark Energy Survey Instrument (DESI) survey. The DESI survey currently relies on a redshift fitter and spectral classifier using a linear combination of principal component analysis–derived templates, which is very efficient in processing large volumes of DESI spectra within a short time frame. However, this method occasionally yields unphysical model fits for galaxies and fails to adequately absorb calibration errors that may still be occasionally visible in the reduced spectra. Our proposed approach improves upon this existing method by refitting the spectra with carefully generated physical galaxy archetypes combined with additional terms designed to absorb data reduction defects and provide more physical models to the DESI spectra. We test our method on an extensive data set derived from the survey validation (SV) and Year 1 (Y1) data of DESI. Our findings indicate that the new method delivers marginally better redshift success for SV tiles while reducing catastrophic redshift failure by 10%–30%. At the same time, results from millions of targets from the main survey show that our model has relatively higher redshift success and purity rates (0.5%–0.8% higher) for galaxy targets while having similar success for QSOs. These improvements also demonstrate that the main DESI redshift pipeline is generally robust. Additionally, it reduces the false-positive redshift estimation by 5%–40% for sky fibers. We also discuss the generic nature of our method and how it can be extended to other large spectroscopic surveys, along with possible future improvements.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND