Search NASASearch

SEARCH · Search NASA

Results for “Data reduction methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Archetype-based Redshift Estimation for the Dark Energy Spectroscopic Instrument Survey

We present a computationally efficient galaxy archetype-based redshift estimation and spectral classification method for the Dark Energy Survey Instrument (DESI) survey. The DESI survey currently relies on a redshift fitter and spectral classifier using a linear combination of principal component analysis–derived templates, which is very efficient in processing large volumes of DESI spectra within a short time frame. However, this method occasionally yields unphysical model fits for galaxies and fails to adequately absorb calibration errors that may still be occasionally visible in the reduced spectra. Our proposed approach improves upon this existing method by refitting the spectra with carefully generated physical galaxy archetypes combined with additional terms designed to absorb data reduction defects and provide more physical models to the DESI spectra. We test our method on an extensive data set derived from the survey validation (SV) and Year 1 (Y1) data of DESI. Our findings indicate that the new method delivers marginally better redshift success for SV tiles while reducing catastrophic redshift failure by 10%–30%. At the same time, results from millions of targets from the main survey show that our model has relatively higher redshift success and purity rates (0.5%–0.8% higher) for galaxy targets while having similar success for QSOs. These improvements also demonstrate that the main DESI redshift pipeline is generally robust. Additionally, it reduces the false-positive redshift estimation by 5%–40% for sky fibers. We also discuss the generic nature of our method and how it can be extended to other large spectroscopic surveys, along with possible future improvements.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Boosting efficiency and reducing graph reliance: Basis adaptation integration in Bayesian multi-fidelity networks

The computational cost of high-fidelity numerical models makes outer-loop analysis, which requires repeated interrogation of the model such as uncertainty quantification, computationally demanding. Multi-fidelity methods, which construct a surrogate model using data from an ensemble of models of varying cost and accuracy, can substantially reduce the cost of outer-loop analysis. However, these methods can be difficult to apply when the model ensemble does not admit a clear hierarchy a priori and the correlations between models are low. Consequently, in this paper, we present a multi-fidelity method that leverages dimension reduction to enhance the correlation between models, thereby reducing the amount of data needed to train a surrogate from an unordered ensemble of models. Our method utilizes basis adaptation to build low-dimensional polynomial chaos expansions of each model and employs Multi-fidelity Networks to encode the relationships among models. We show that the resulting method exhibit two notable advantages over its counterpart: (1) enhanced accuracy (both reduced bias and variance); and (2) reduced dependency on the graph structure encoding relationships among models. We demonstrate the approach on an analytical test problem and a challenging finite element model for a spent nuclear fuel. Our method produces a surrogate model that is significantly more accurate than either a single-fidelity surrogate or a multi-fidelity surrogate constructed without basis adaptation.

42 ENGINEERING

Diesel Fuel Consumption in Prominent U.S. Open-Pit Mines: Site-Level Estimates

This report presents a comprehensive framework for estimating diesel fuel consumption and prices at open-pit mines in the United States. The framework includes transparent methods for calculating site-level diesel energy use when direct reporting is unavailable, and a structured confidence evaluation for each method. The framework is demonstrated to estimate current diesel consumption at 21 open-pit mines in the United States. Initial findings support ongoing efforts to strengthen the competitiveness and security of the U.S. industrial base by supporting data-driven supply chain analysis and decision-making, improved transparency in mining sector energy use, and targeted deployment of energy innovation and cost-reduction strategies. Future updates to the dataset—coupled with expanded data transparency and method validation—will help ensure that the findings remain relevant as the sector continues to evolve.

02 PETROLEUM

Data-Efficient Dimensionality Reduction and Surrogate Modeling of High-Dimensional Stress Fields

Tensor datatypes representing field variables like stress, displacement, velocity, etc., have increasingly become a common occurrence in data-driven modeling and analysis of simulations. Numerous methods [such as convolutional neural networks (CNNs)] exist to address the meta-modeling of field data from simulations. As the complexity of the simulation increases, so does the cost of acquisition, leading to limited data scenarios. Modeling of tensor datatypes under limited data scenarios remains a hindrance for engineering applications. Here, in this article, we introduce a direct image-to-image modeling framework of convolutional autoencoders enhanced by information bottleneck loss function to tackle the tensor data types with limited data. The information bottleneck method penalizes the nuisance information in the latent space while maximizing relevant information making it robust for limited data scenarios. The entire neural network framework is further combined with robust hyperparameter optimization. We perform numerical studies to compare the predictive performance of the proposed method with a dimensionality reduction-based surrogate modeling framework on a representative linear elastic ellipsoidal void problem with uniaxial loading. The data structure focuses on the low-data regime (fewer than 100 data points) and includes the parameterized geometry of the ellipsoidal void as the input and the predicted stress field as the output. The results of the numerical studies show that the information bottleneck approach yields improved overall accuracy and more precise prediction of the extremes of the stress field. Additionally, an in-depth analysis is carried out to elucidate the information compression behavior of the proposed framework.

artificial intelligence

Maximizing efficiency of dataset compression for machine learning potentials with information theory

Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.

36 MATERIALS SCIENCE

In situ midcircuit qubit measurement and reset in a single-species trapped-ion quantum computing system

We implement in situ midcircuit measurement and reset (MCMR) operations on a full-scale trapped-ion quantum computing system by using metastable qubit states in 171 Yb + ions. We compare two methods for isolating data qubits from measured qubits: one shelves the data qubits into the metastable state and the other drives the measured qubit to the metastable state without disturbing the other qubits. We experimentally demonstrate both methods on a crystal of two 171 Yb + ions using both the 𝑆 1/2 ground-state hyperfine clock qubit and the 𝑆 1/2 −𝐷 3/2 optical qubit. These MCMR methods result in errors on the data qubit of about 2% without degrading the measurement fidelity. With straightforward reductions in laser noise, these errors can be suppressed to less than 0.1%. The demonstrated methods allow MCMR to be performed in a single-species ion chain without shuttling or additional qubit-addressing optics, greatly simplifying the system architecture and allowing straightforward integration with existing trapped-ion quantum computers.

coherent control

In-situ mid-circuit qubit measurement and reset in a single-species trapped-ion quantum computing system

We implement in-situ mid-circuit measurement and reset (MCMR) operations on a trapped-ion quantum computing system by using metastable qubit states in $^{171}\textrm{Yb}^+$ ions. We introduce and compare two methods for isolating data qubits from measured qubits: one shelves the data qubit into the metastable state and the other drives the measured qubit to the metastable state without disturbing the other qubits. We experimentally demonstrate both methods on a crystal of two $^{171}\textrm{Yb}^+$ ions using both the $S_{1/2}$ ground state hyperfine clock qubit and the $S_{1/2}$-$D_{3/2}$ optical qubit. These MCMR methods result in errors on the data qubit of about $2\%$ without degrading the measurement fidelity. With straightforward reductions in laser noise, these errors can be suppressed to less than $0.1\%$. The demonstrated method allows MCMR to be performed in a single-species ion chain without shuttling or additional qubit-addressing optics, greatly simplifying the architecture.

Atomic Physics (physics.atom-ph)

Multidimensional scaling informed by F -statistic: Visualizing grouped microbiome data with inference

Multidimensional scaling (MDS) is a widely used dimensionality reduction technique in microbial ecology data analysis that captures the multivariate structure of the data while preserving pairwise distances between samples. While improvements in MDS have enhanced the ability to reveal group-specific data patterns, these MDS-based methods require prior assumptions for inference, limiting their application in general microbiome analysis. Here, in this study, we introduce a new MDS-based ordination method, “F-informed MDS,” which configures the data distribution based on the F-statistic, the ratio of dispersion between groups sharing common and different characteristics. Using semisynthetic datasets, we demonstrate that the proposed method is robust to hyperparameter selection while maintaining statistical significance throughout the ordination process. Various quality metrics for evaluating dimensionality reduction confirm that F-informed MDS is comparable to state-of-the-art methods in preserving both local and global data structures. Its application to a diatom-associated bacterial community suggests the role of this new method in interpreting the community’s response to the host. Our approach offers a well-founded refinement of MDS that aligns with statistical test results, which can be beneficial for broader multidimensional data analyses in microbiology and ecology. This new visualization tool can be incorporated into standard microbiome data analyses.

Biological and medical sciences

Feedback Controllability Components Analysis (FCCA) v1.0

FCCA is a linear dimensionality reduction method that find subspaces of high-dimensional time-series data that are most feedback controllable. The key innovation is to formulate an objective function that quantifies the joint cost of state reconstruction and state regulation that can be evaluated from purely observational data. To do this, it leverages the duality between controllability and observability. We provide analytic results demonstrating the validity of the cost function. We evaluated this method in both synthetic and real neural data (from multiple organisms and brain areas).

Kumar, Ankit

Bayesian Optimized Deep Ensemble for Uncertainty Quantification of Deep Neural Networks: a System Safety Case Study on Sodium Fast Reactor Thermal Stratification Modeling

Deep neural networks (DNNs) are increasingly important to scientific computing and engineering system simulations. Accurate uncertainty quantification (UQ) for DNNs is critical in safety-sensitive engineering domains. Traditional Deep Ensemble (DE) methods, while easy to implement, frequently suffer from poorly calibrated uncertainty estimates and limited predictive accuracy due to reliance on fixed architectures with varied weight initializations. To address these issues, we introduce a workflow that combines Bayesian Optimization (BO) and DE. The workflow is modular, scalable, and integrates parallel BO initialized with Sobol sequences to individually optimize the hyperparameters of each ensemble member. This method enhances ensemble diversity, improves predictive accuracy, and provides reliable uncertainty estimates. We evaluate the proposed BODE approach in a sodium fast reactor thermal stratification modeling case study, where we used a densely connected convolutional neural network to predict turbulent viscosity during the reactor transient with consideration of data noise. We benchmark its performance against several optimization approaches, including baseline deep ensemble, evolutionary algorithm-optimized ensemble, ensemble formed via random search combined with greedy selection, and a BO ensemble using random initialization. Here, our results demonstrate superior performance of the developed BODE approach. In noise-free scenarios, BODE notably reduces incorrect aleatoric uncertainty and significantly enhances predictive accuracy. Under conditions of 5% and 10% Gaussian noise, BODE adaptively quantifies uncertainty proportional to data noise, achieving up to an 80% reduction in root mean square error compared to baseline methods and producing well-calibrated prediction intervals.

Bayesian optimization

FedOSAA: Improving Federated Learning with One-Step Anderson Acceleration

Federated learning (FL) is a distributed machine learning approach that enables multiple local clients and a central server to collaboratively train a model while keeping the data on their own devices. First-order methods, particularly those incorporating variance reduction techniques, are the most widely used FL algorithms due to their simple implementation and stable performance. However, these methods tend to be slow and require a large number of communication rounds to reach the global minimizer. We propose FedOSAA, a novel approach that preserves the simplicity of first-order methods while achieving the rapid convergence typically associated with second-order methods. Our approach applies one Anderson acceleration (AA) step following classical local updates based on first-order methods with variance reduction, such as FedSVRG and SCAFFOLD, during local training. This AA step is able to leverage curvature information from the history points and gives a new update that approximates the Newton-GMRES direction, thereby significantly improving the convergence. We establish a local linear convergence rate to the global minimizer of FedOSAA for smooth and strongly convex loss functions. Numerical comparisons show that FedOSAA substantially improves the communication and computation efficiency of the original first-order methods, achieving performance comparable to second-order methods like GIANT.

Feng, Xue [University of California, Davis]

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis

Analysis of Rig Parameter Data Using Drilling Process Modeling Constraints, Volume 5: Utah FORGE Well 16B(78)-32

Drill rig parameter measurements are routinely used during deep well construction to monitor and guide drilling conditions for improved performance and reduced costs. While insightful into the drilling process, these measurements are of reduced value without a standard to aid in data evaluation and decision making. In the main body of this work (Volume 1), a method is demonstrated whereby rock reduction model constraints are used to interpret drilling response parameters; the method could be applied in real-time to improve decision-making in the field and to further discern technology performance during post-drilling evaluations. Drilling parameters are evaluated using laboratory-validated rock reduction models for predicting the phenomenological response of drag bits (Detournay and Defourny, 1992) in computational algorithms. The method presented has applicability to development of advanced analytics on future geothermal wells using real-time electronic data recording for improved performance and reduced drilling costs. A drilling cost model is also used to show the tradeoff between rate of penetration and bit life and the influence on interval drilling costs. Details of the bit specifications and performance are cataloged in an independent volume, documented under separate cover, for each of the four wells, and include Volume 2: Utah FORGE 16A(78)-32; Volume 3: Utah FORGE 56-32; Volume 4: Utah FORGE 78B-32 and Volume 5: Utah FORGE 16B(78)-32.

15 GEOTHERMAL ENERGY

Analysis of Rig Parameter Data Using Drilling Process Modeling Constraints, Volume 4: Utah FORGE Well 78B-32

Drill rig parameter measurements are routinely used during deep well construction to monitor and guide drilling conditions for improved performance and reduced costs. While insightful into the drilling process, these measurements are of reduced value without a standard to aid in data evaluation and decision making. In the main body of this work (Volume 1), a method is demonstrated whereby rock reduction model constraints are used to interpret drilling response parameters; the method could be applied in real-time to improve decision-making in the field and to further discern technology performance during post-drilling evaluations. Drilling parameters are evaluated using laboratory-validated rock reduction models for predicting the phenomenological response of drag bits (Detournay and Defourny, 1992) in computational algorithms. The method presented has applicability to development of advanced analytics on future geothermal wells using real-time electronic data recording for improved performance and reduced drilling costs. A drilling cost model is also used to show the tradeoff between rate of penetration and bit life and the influence on interval drilling costs. Details of the bit specifications and performance are cataloged in an independent volume, documented under separate cover, for each of the four wells, and include Volume 2: Utah FORGE 16A(78)-32; Volume 3: Utah FORGE 56-32; Volume 4: Utah FORGE 78B-32 and Volume 5: Utah FORGE 16B(78)-32.

15 GEOTHERMAL ENERGY

Analysis of Rig Parameter Data Using Drilling Process Modeling Constraints, Volume 3: Utah FORGE Well 56-32

Drill rig parameter measurements are routinely used during deep well construction to monitor and guide drilling conditions for improved performance and reduced costs. While insightful into the drilling process, these measurements are of reduced value without a standard to aid in data evaluation and decision making. In the main body of this work (Volume 1), a method is demonstrated whereby rock reduction model constraints are used to interpret drilling response parameters; the method could be applied in real-time to improve decision-making in the field and to further discern technology performance during post-drilling evaluations. Drilling parameters are evaluated using laboratory-validated rock reduction models for predicting the phenomenological response of drag bits (Detournay and Defourny, 1992) in computational algorithms. The method presented has applicability to development of advanced analytics on future geothermal wells using real-time electronic data recording for improved performance and reduced drilling costs. A drilling cost model is also used to show the tradeoff between rate of penetration and bit life and the influence on interval drilling costs. Details of the bit specifications and performance are cataloged in an independent volume, documented under separate cover, for each of the four wells, and include Volume 2: Utah FORGE 16A(78)-32; Volume 3: Utah FORGE 56-32; Volume 4: Utah FORGE 78B-32 and Volume 5: Utah FORGE 16B(78)-32.

15 GEOTHERMAL ENERGY

Analysis of Rig Parameter Data Using Drilling Process Modeling Constraints, Volume 1: Summary of Utah FORGE Wells 16A(78)-32, 56-32, 78B-32 and 16B(78)-32

Drill rig parameter measurements are routinely used during deep well construction to monitor and guide drilling conditions for improved performance and reduced costs. While insightful into the drilling process, these measurements are of reduced value without a standard to aid in data evaluation and decision making. In the main body of this work (Volume 1), a method is demonstrated whereby rock reduction model constraints are used to interpret drilling response parameters; the method could be applied in real-time to improve decision-making in the field and to further discern technology performance during post-drilling evaluations. Drilling parameters are evaluated using laboratory-validated rock reduction models for predicting the phenomenological response of drag bits (Detournay and Defourny, 1992) in computational algorithms. The method presented has applicability to development of advanced analytics on future geothermal wells using real-time electronic data recording for improved performance and reduced drilling costs. A drilling cost model is also used to show the tradeoff between rate of penetration and bit life and the influence on interval drilling costs. Details of the bit specifications and performance are cataloged in an independent volume, documented under separate cover, for each of the four wells, and include Volume 2: Utah FORGE 16A(78)-32; Volume 3: Utah FORGE 56-32; Volume 4: Utah FORGE 78B-32 and Volume 5: Utah FORGE 16B(78)-32.

15 GEOTHERMAL ENERGY

Efficient near-field ptychography reconstruction using the Hessian operator

X-ray ptychography is a powerful and robust coherent imaging method providing access to the complex object and probe (illumination). Ptychography reconstruction is typically performed using first-order methods due to their computational efficiency. Higher-order methods, while potentially more accurate, are often prohibitively expensive in terms of computation. In this study, we present a mathematical framework for reconstruction using second-order information derived from an efficient computation of the bilinear Hessian and Hessian operator. The formulation is provided for Gaussian-based models, enabling the simultaneous reconstruction of the object, probe, and object positions. Synthetic data tests, along with experimental near-field ptychography data processing, demonstrate a ten-fold reduction in computation time compared to first-order methods. The derived formulas for computing the Hessians, along with the strategies for incorporating them into optimization schemes, are well-structured and easily adaptable to various ptychography problem formulations.

Carlsson, Marcus [Lund Univ. (Sweden)] (ORCID:0000

Packages of Distributed Energy Technologies Demonstrating Demand Flexibility at Community Scale

The combination of increased electric load growth across all sectors, deferred electrical infrastructure investment, and other factors resulting in variable electric power supply, has created technical challenges to maintaining a resilient and reliable grid. Many federal, regional, and local efforts are in play to modernize the electric grid, including advancing building technologies and distributed energy resources (DERs) that are utilizing smarter controls to become responsive to both occupant and grid needs. This report reviews ten pilot projects demonstrating how groups of buildings combined with behind-the-meter (BTM) DERs such as electric vehicle (EV) charging, battery storage, flexible HVAC and domestic hot water systems, and photovoltaic systems can reliably and cost effectively provide grid services. Each of the ten pilot projects aim to deliver both energy efficiency and demand flexibility (DF) while supporting load growth. The ten demonstration teams are piloting flexible DER packages across diverse communities of residential and commercial buildings to address a variety of regional grid needs. The outcomes of these pilot projects will be used to inform future scaling through utility program development. This paper characterizes the ten teams, showcasing the decision-making process used by each group to develop their packages (Section 2), the grid services they plan to deliver (Section 3), the types of DER packages selected for deployment within building sectors (Section 4) and trends between building sector, DER types, and grid services In order to achieve community scale benefits, the pilot projects must utilize aggregated control mechanisms for coordinating buildings and DERs together. Several types of coordinated control architectures have evolved amongst the teams, influenced by use type, existing market conditions, and integration type. Three coordinated controls architectures have been characterized, highlighting their use cases, benefits, challenges, and tradeoffs in their design. These insights can aid utilities, control vendors, and developers in scaling community-level energy systems (Paul, 2024). Ultimately, the technology packages selected by the ten teams will be coordinated to provide power system services, also known as grid services. Insights from these demonstrations will be useful for grid operators, regulators, aggregators and other stakeholders as they look to deploy demand flexible resources as grid services in the future. The grid services that each team is targeting for demonstration are described in Section 3 and Section 4. Methods for evaluating the grid services have been described in the paper Metrics for Evaluating Grid Service Provision from Communities of Grid-interactive and Efficient Buildings and other DER (MacDonald, 2023). To identify technology packages for demonstration, Section 2 shows that project teams used a range of analysis approaches, including building energy modeling, AMI data analysis, cost-benefit frameworks, and utility pilot data. Some teams emphasized technical modeling to quantify grid impacts and demand reduction potential, while others prioritized economic evaluations, stakeholder input, or exploratory pilots to inform deployment decisions. This diversity reflects the need to tailor selection methods to project goals, available data, and organizational context. Section 5 discusses trends between the DER technologies deployed and the grid service provisions from each team. Residential buildings (multifamily and single family) lean towards technologies that enhance energy efficiency (e.g. weatherization upgrades, smart thermostats) and onsite power generation integration (e.g. solar PV). Commercial building demonstrations prioritize technologies that ensure operational reliability (e.g. battery storage) and centralized energy management systems and optimization solutions. Teams that are deploying controllable storage-based technologies are more likely to provide grid services that require a near real-time response. Teams incorporating load shifting technologies like smart thermostats with HEMs are likely to include energy markets participation and customer bill management offerings. Campus demonstrations are adopting diverse sets of DERs to emphasize renewable generation, paired with centralized control. This section also describes technologies that were considered during project planning but ultimately excluded from final deployment. These demonstrations reveal that effective DER package design should be tailored to building type, customer segment, and construction vintage. Multifamily buildings benefit from centralized HVAC upgrades and supervisory controls, while single-family homes are well-suited for individualized technologies like solar, storage, and smart home energy monitors. Commercial and campus settings prioritize EMIS integration and load optimization. New construction enables cost-effective integration of DER-ready infrastructure, whereas retrofits require deployments aligned with owner and tenant value streams. For utility program planners, early coordination with developers and building owners, paired with segmented and modular program offerings, can improve adoption, scalability, and grid impact.

24 POWER TRANSMISSION AND DISTRIBUTION