Search NASA⌕ Search

SEARCH · Search NASA

Results for “data reduction algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Maximizing efficiency of dataset compression for machine learning potentials with information theory

Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.

36 MATERIALS SCIENCE↗

An investigation on machine learning predictive accuracy improvement and uncertainty reduction using VAE-based data augmentation

The confluence of ultrafast computers with large memory, rapid progress in Machine Learning (ML) algorithms, and the availability of large datasets place multiple engineering fields at the threshold of dramatic progress. However, a unique challenge in nuclear engineering is data scarcity because experimentation on nuclear systems is usually more expensive and time-consuming than most other disciplines. One potential way to resolve the data scarcity issue is deep generative learning, which uses certain ML models to learn the underlying distribution of existing data and generate synthetic samples that resemble the real data. In this way, one can significantly expand the dataset to train more accurate predictive ML models. In this study, our objective is to evaluate the effectiveness of data augmentation using variational autoencoder (VAE)-based deep generative models. We investigated whether the data augmentation leads to improved accuracy in the predictions of a deep neural network (DNN) model trained using the augmented data. Additionally, the DNN prediction uncertainties are quantified using Bayesian Neural Networks (BNN) and conformal prediction (CP) to assess the impact on predictive uncertainty reduction. To test the proposed methodology, we used TRACE simulations of steady-state void fraction data based on the NUPEC Boiling Water Reactor Full-size Fine-mesh Bundle Test (BFBT) benchmark. Here, we found that augmenting the training dataset using VAEs has improved the DNN model’s predictive accuracy, improved the prediction confidence intervals, and reduced the prediction uncertainties.

Bayesian neural network↗

AEOLUS: Advances in Experimental Design, Optimal Control, and Learning for Uncertain Complex Systems

Sustained advances in the mathematics of modeling and simulation have resulted in the capability today for routine simulation of a number of large scale complex DOE-relevant systems. As remarkable as this capability for solving the so-called forward problem is, it is typically only the first step-an inner loop within an outer loop that explores the simulation model's parameter space and decision space to characterize uncertainty in the model's predictions, learn unknown model parameters from data, design the most informative experiments, determine optimal control strategies, and create optimal designs. Broadly, what unifies all of these outer loop problems is that they are, in one form or another, optimization problems over parameter/control/design space that are constrained by complex uncertain models. To fully realize the power of scientific simulation as a basis for scientific discovery, technological innovation, and rational decision-making, it is imperative to move beyond simulation to tackle the outer loop of optimization for learning from data, experimental design, and control with complex uncertain models. When the models under consideration are large-scale and complex, and when the optimization variable and uncertain parameter spaces are high (or infinite) dimensional, this constitutes a grand challenge of the highest order, and is intractable with conventional methods. To overcome these challenges, the AEOLUS Center was established to develop a unified mathematical, computational, and statistical framework for (1) Learning predictive models from complex data via Bayesian inference and optimization, and (2) Optimizing experiments, processes, and designs using the resulting uncertain models. These problems are intractable with conventional methods, for several reasons: (1) The simulation problems that govern the inner loops of the optimization problems are expensive to execute (due to severe nonlinearity, heterogeneity, multiphysics/multiscale coupling); (2) The optimization variable and uncertain parameter spaces are high dimensional, often stemming from discretizations of infinite dimensional fields such as initial conditions, sources, or material properties. We argue that the key to overcoming these challenges is to develop new mathematical, computational, and statistical methods that exploit the structure of the Bayesian inference and optimization problems mediated by their underlying complex uncertain models. This structure includes the regularity, sparsity, geometry, low intrinsic dimensionality, and multifidelity nature of the maps from uncertain parameter/optimization variable spaces to the specific objectives targeted: Bayesian inference, optimal experimental design, and optimal control design. Black box methods developed as generic tools are incapable of exploiting this structure. To be successful, we must create, integrate, and cross-fertilize ideas across multiple areas of applied math--including approximation theory, Bayesian inference, data science, experimental design, information theory, machine learning, model reduction, optimal control theory, parallel algorithms, PDE-constrained optimization, randomized algorithms, stochastic optimization, and uncertainty quantification--all while exploiting the structure of the problems at hand. With this goal in mind, we have marshaled a team of leading authorities in these areas. While the methods we develop will be broadly applicable across a wide spectrum of DOE problems in which experiments inform models and the systems those models describe must be optimized under uncertainty, we have chosen a specific area, advanced manufacturing and materials, to drive our work. AMM is characterized by complex models across multiple scales, and is a rich source of challenging problems in inference, experimental design, and optimal control, requiring multifaceted and integrated advances in applied mathematics. As such, AMM serves as an excellent vehicle to motivate and demonstrate the advances in applied mathematics developed by our center.

97 MATHEMATICS AND COMPUTING↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗

Framework of compressive sensing and data compression for 4D-STEM

Four-dimensional Scanning Transmission Electron Microscopy (4D-STEM) is a powerful technique for high-resolution and high-precision materials characterization at multiple length scales, including the characterization of beam-sensitive materials. However, the field of view of 4D-STEM is relatively small, which in absence of live processing is limited by the data size required for storage. Furthermore, the rectilinear scan approach currently employed in 4D-STEM places a resolution- and signal-dependent dose limit for the study of beam sensitive materials. Improving 4D-STEM data and dose efficiency, by keeping the data size manageable while limiting the amount of electron dose, is thus critical for broader applications. Here we introduce a general method for reconstructing 4D-STEM data with subsampling in both real and reciprocal spaces at high fidelity. The approach is first tested on the subsampled datasets created from a full 4D-STEM dataset, and then demonstrated experimentally using random scan in real-space. The same reconstruction algorithm can also be used for compression of 4D-STEM datasets, leading to a large reduction (100 times or more) in data size, while retaining the fine features of 4D-STEM imaging, for crystalline samples.

4D-STEM↗

Scalable Tensor Methods for Nonuniform Hypergraphs

While multilinear algebra appears natural for studying the multiway interactions modeled by hypergraphs, tensor methods for general hypergraphs have been stymied by theoretical and practical barriers. A recently proposed adjacency tensor is applicable to nonuniform hypergraphs, but is prohibitively costly to form and analyze in practice. We develop tensor times same vector (TTSV) algorithms for this tensor which improve complexity from $O(n^r)$ to a low-degree polynomial in $r$, where $n$ is the number of vertices and $r$ is the maximum hyperedge size. Our algorithms are implicit, avoiding formation of the order $r$ adjacency tensor. Here, we demonstrate the flexibility and utility of our approach in practice by developing tensor-based hypergraph centrality and clustering algorithms. We also show these tensor measures offer complementary information to analogous graph-reduction approaches on data, and are also able to detect higher-order structure that many existing matrix-based approaches provably cannot.

97 MATHEMATICS AND COMPUTING↗

Identification of tau leptons using a convolutional neural network with domain adaptation

A tau lepton identification algorithm,DeepTau, based on convolutional neural network techniques, has been developed in the CMS experiment to discriminate reconstructed hadronic decays of tau leptons (τ h ) from quark or gluon jets and electrons and muons that are misreconstructed as τ h candidates. The latest version of this algorithm, v2.5, includes domain adaptation by backpropagation, a technique that reduces discrepancies between collision data and simulation in the region with the highest purity of genuine τh candidates. Additionally, a refined training workflow improves classification performance with respect to the previous version of the algorithm, with a reduction of 30–50% in the probability for quark and gluon jets to be misidentified as τ h candidates for given reconstruction and identification efficiencies. This paper presents the novel improvements introduced in theDeepTau algorithm and evaluates its performance in LHC proton-proton collision data at √(s) = 13 and 13.6 TeV collected in 2018 and 2022 with integrated luminosities of 60 and 35 fb -1 , respectively. Techniques to calibrate the performance of the τ h identification algorithm in simulation with respect to its measured performance in real data are presented, together with a subset of results among those measured for use in CMS physics analyses.

Large detector-systems performance↗

Event Detection and Classification Using Machine Learning Applied to PMU Data for the Western US Power System

Smart grid technology enhances our comprehension and reliability of the power grid, leveraging Phasor Measurement Unit (PMU) data—time-synchronized, high-frequency measurements gathered across the US power grid. This paper employs machine learning techniques to effectively analyze the vast PMU data in Wide Area Monitoring Systems (WAMS) for power grid event detection and classification. Analyzing several months of real-world PMU data, the paper focuses on machine learning for fast, precise event detection and classification, corroborated by utility event logs. Practical challenges like feature extraction, dimensionality reduction, and model selection are addressed. A novel feature yielding improved results is discovered, and a supplementary algorithm for detecting small power grid faults is developed. The final algorithm is validated using a month-long real PMU data set, demonstrating its capability in accurately identifying power grid events in near real-time.

machine learning, event detection, PMU↗

Embedded FPGA developments in 130 nm and 28 nm CMOS for machine learning in particle detector readout

Embedded field programmable gate array (eFPGA) technology allows the implementation of reconfigurable logic within the design of an application-specific integrated circuit (ASIC). This approach offers the low power and efficiency of an ASIC along with the ease of FPGA configuration, particularly beneficial for the use case of machine learning in the data pipeline of next-generation collider experiments. An open-source framework called "FABulous" was used to design eFPGAs using 130 nm and 28 nm CMOS technology nodes, which were subsequently fabricated and verified through testing. The capability of an eFPGA to act as a front-end readout chip was assessed using simulation of high energy particles passing through a silicon pixel sensor. A machine learning-based classifier, designed for reduction of sensor data at the source, was synthesized and configured onto the eFPGA. A successful proof-of-concept was demonstrated through reproduction of the expected algorithm result on the eFPGA with perfect accuracy. Finally, further development of the eFPGA technology and its application to collider detector readout is discussed.

47 OTHER INSTRUMENTATION↗

FedOSAA: Improving Federated Learning with One-Step Anderson Acceleration

Federated learning (FL) is a distributed machine learning approach that enables multiple local clients and a central server to collaboratively train a model while keeping the data on their own devices. First-order methods, particularly those incorporating variance reduction techniques, are the most widely used FL algorithms due to their simple implementation and stable performance. However, these methods tend to be slow and require a large number of communication rounds to reach the global minimizer. We propose FedOSAA, a novel approach that preserves the simplicity of first-order methods while achieving the rapid convergence typically associated with second-order methods. Our approach applies one Anderson acceleration (AA) step following classical local updates based on first-order methods with variance reduction, such as FedSVRG and SCAFFOLD, during local training. This AA step is able to leverage curvature information from the history points and gives a new update that approximates the Newton-GMRES direction, thereby significantly improving the convergence. We establish a local linear convergence rate to the global minimizer of FedOSAA for smooth and strongly convex loss functions. Numerical comparisons show that FedOSAA substantially improves the communication and computation efficiency of the original first-order methods, achieving performance comparable to second-order methods like GIANT.

Feng, Xue [University of California, Davis]↗

In-pixel integration of signal processing and AI/ML based data filtering for particle tracking detectors

We present the first physical realization of in-pixel signal processing with integrated AI-based data filtering for particle tracking detectors. Building on prior work that demonstrated a physics-motivated edge-AI algorithm suitable for ASIC implementation, this work marks a significant milestone toward intelligent silicon trackers. Our prototype readout chip performs real-time data reduction at the sensor level while meeting stringent requirements on power, area, and latency. The chip is taped-out in 28nm TSMC CMOS bulk process, which has been shown to have sufficient radiation hardness for particle experiments. This development represents a key step toward enabling fully on-detector edge AI, with broad implications for data throughput and discovery potential in high-rate, high-radiation environments such as the High-Luminosity LHC.

Parpillon, Benjamin [Fermilab; Illinois U., Chicag↗

Analysis of Rig Parameter Data Using Drilling Process Modeling Constraints, Volume 5: Utah FORGE Well 16B(78)-32

Drill rig parameter measurements are routinely used during deep well construction to monitor and guide drilling conditions for improved performance and reduced costs. While insightful into the drilling process, these measurements are of reduced value without a standard to aid in data evaluation and decision making. In the main body of this work (Volume 1), a method is demonstrated whereby rock reduction model constraints are used to interpret drilling response parameters; the method could be applied in real-time to improve decision-making in the field and to further discern technology performance during post-drilling evaluations. Drilling parameters are evaluated using laboratory-validated rock reduction models for predicting the phenomenological response of drag bits (Detournay and Defourny, 1992) in computational algorithms. The method presented has applicability to development of advanced analytics on future geothermal wells using real-time electronic data recording for improved performance and reduced drilling costs. A drilling cost model is also used to show the tradeoff between rate of penetration and bit life and the influence on interval drilling costs. Details of the bit specifications and performance are cataloged in an independent volume, documented under separate cover, for each of the four wells, and include Volume 2: Utah FORGE 16A(78)-32; Volume 3: Utah FORGE 56-32; Volume 4: Utah FORGE 78B-32 and Volume 5: Utah FORGE 16B(78)-32.

15 GEOTHERMAL ENERGY↗

Analysis of Rig Parameter Data Using Drilling Process Modeling Constraints, Volume 4: Utah FORGE Well 78B-32

Drill rig parameter measurements are routinely used during deep well construction to monitor and guide drilling conditions for improved performance and reduced costs. While insightful into the drilling process, these measurements are of reduced value without a standard to aid in data evaluation and decision making. In the main body of this work (Volume 1), a method is demonstrated whereby rock reduction model constraints are used to interpret drilling response parameters; the method could be applied in real-time to improve decision-making in the field and to further discern technology performance during post-drilling evaluations. Drilling parameters are evaluated using laboratory-validated rock reduction models for predicting the phenomenological response of drag bits (Detournay and Defourny, 1992) in computational algorithms. The method presented has applicability to development of advanced analytics on future geothermal wells using real-time electronic data recording for improved performance and reduced drilling costs. A drilling cost model is also used to show the tradeoff between rate of penetration and bit life and the influence on interval drilling costs. Details of the bit specifications and performance are cataloged in an independent volume, documented under separate cover, for each of the four wells, and include Volume 2: Utah FORGE 16A(78)-32; Volume 3: Utah FORGE 56-32; Volume 4: Utah FORGE 78B-32 and Volume 5: Utah FORGE 16B(78)-32.

15 GEOTHERMAL ENERGY↗

Analysis of Rig Parameter Data Using Drilling Process Modeling Constraints, Volume 3: Utah FORGE Well 56-32

Drill rig parameter measurements are routinely used during deep well construction to monitor and guide drilling conditions for improved performance and reduced costs. While insightful into the drilling process, these measurements are of reduced value without a standard to aid in data evaluation and decision making. In the main body of this work (Volume 1), a method is demonstrated whereby rock reduction model constraints are used to interpret drilling response parameters; the method could be applied in real-time to improve decision-making in the field and to further discern technology performance during post-drilling evaluations. Drilling parameters are evaluated using laboratory-validated rock reduction models for predicting the phenomenological response of drag bits (Detournay and Defourny, 1992) in computational algorithms. The method presented has applicability to development of advanced analytics on future geothermal wells using real-time electronic data recording for improved performance and reduced drilling costs. A drilling cost model is also used to show the tradeoff between rate of penetration and bit life and the influence on interval drilling costs. Details of the bit specifications and performance are cataloged in an independent volume, documented under separate cover, for each of the four wells, and include Volume 2: Utah FORGE 16A(78)-32; Volume 3: Utah FORGE 56-32; Volume 4: Utah FORGE 78B-32 and Volume 5: Utah FORGE 16B(78)-32.

15 GEOTHERMAL ENERGY↗

Analysis of Rig Parameter Data Using Drilling Process Modeling Constraints, Volume 1: Summary of Utah FORGE Wells 16A(78)-32, 56-32, 78B-32 and 16B(78)-32

Drill rig parameter measurements are routinely used during deep well construction to monitor and guide drilling conditions for improved performance and reduced costs. While insightful into the drilling process, these measurements are of reduced value without a standard to aid in data evaluation and decision making. In the main body of this work (Volume 1), a method is demonstrated whereby rock reduction model constraints are used to interpret drilling response parameters; the method could be applied in real-time to improve decision-making in the field and to further discern technology performance during post-drilling evaluations. Drilling parameters are evaluated using laboratory-validated rock reduction models for predicting the phenomenological response of drag bits (Detournay and Defourny, 1992) in computational algorithms. The method presented has applicability to development of advanced analytics on future geothermal wells using real-time electronic data recording for improved performance and reduced drilling costs. A drilling cost model is also used to show the tradeoff between rate of penetration and bit life and the influence on interval drilling costs. Details of the bit specifications and performance are cataloged in an independent volume, documented under separate cover, for each of the four wells, and include Volume 2: Utah FORGE 16A(78)-32; Volume 3: Utah FORGE 56-32; Volume 4: Utah FORGE 78B-32 and Volume 5: Utah FORGE 16B(78)-32.

15 GEOTHERMAL ENERGY↗

Transfer learning nonlinear plasma dynamic transitions in low dimensional embeddings via deep neural networks

Deep learning algorithms provide a new paradigm to study high-dimensional dynamical behaviors, such as those in fusion plasma systems. Development of novel, data-driven model reduction methods, coupled with detection of abnormal modes with plasma physics, opens a unique opportunity to identify plasma instabilities through automated construction of parsimonious models that can be tuned to balance accuracy and cost. Our fusion transfer learning (FTL) model demonstrates success in rapidly reconstructing nonlinear kink mode structures by learning from a limited amount of nonlinear simulation data. The knowledge transfer process leverages a pre-trained neural encoder–decoder network, initially trained on linear simulations, to effectively capture nonlinear dynamics. The low-dimensional embeddings extract the coherent structures of interest, while preserving the inherent dynamics of the complex system. Experimental results highlight FTL’s capacity to capture transitional behaviors and dynamical features in plasma dynamics—a task often challenging for conventional methods. The model developed in this study is generalizable and can be extended broadly through transfer learning to address various magnetohydrodynamics modes.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Machine Learning on Heterogeneous, Edge, and Quantum Hardware for Particle Physics (ML-HEQUPP)

The next generation of particle physics experiments will face a new era of challenges in data acquisition, due to unprecedented data rates and volumes along with extreme environments and operational constraints. Harnessing this data for scientific discovery demands real-time inference and decision-making, intelligent data reduction, and efficient processing architectures beyond current capabilities. Crucial to the success of this experimental paradigm are several emerging technologies, such as artificial intelligence and machine learning (AI/ML) and silicon microelectronics, and the advent of quantum algorithms and processing. Their intersection includes areas of research such as low-power and low-latency devices for edge computing, heterogeneous accelerator systems, reconfigurable hardware, novel codesign and synthesis strategies, readout for cryogenic or high-radiation environments, and analog computing. This white paper presents a community-driven vision to identify and prioritize research and development opportunities in hardware-based ML systems and corresponding physics applications, contributing towards a successful transition to the new data frontier of fundamental science.

Gonski, Julia [SLAC]↗