Search NASA⌕ Search

SEARCH · Search NASA

Results for “data reduction algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Machine Learning to Select Experiments Driven by Fundamental Science and Applications for Targeted Nuclear Data Improvement

This work describes a blueprint for a process that accelerates progress in science by quantitatively answering the following question: What is the optimal combination of fundamental-science and application-driven experiments to maximally reduce pertinent data uncertainties? Answering this question entails solving a high-dimensional and complex optimization problem that is best solved with advanced statistic techniques often classified as machine learning. We apply this process within the framework of nuclear data with the aim to select an experiment combination that will reduce uncertainties in 239 Pu nuclear data for neutron energies between 1 and 600 keV. In this field, fundamental-physics driven data, called differential, look at one nuclear physics observable at a time. They are contrasted to application-driven, integral, data where one or few resulting values inform a broad set of nuclear data across several nuclides and energies. The candidates for integral experiments are criticality measurements that were refined by a genetic algorithm to be maximally sensitive to 239 Pu fission cross sections in the desired energy range. Twenty-three candidate differential experiments were investigated and span multiple nuclear physics observables (e.g., total, capture cross sections) for isotopes appearing in the integral experiments. The optimal combination among these candidate experiments was investigated via generalized least squares fitting, augmented with Gaussian processes to ameliorate statistical irregularities in data, and the D-optimality criterion. The latter evaluates for each pair of candidates the joint reduction in uncertainties of all 12200 nuclear data appearing in the integral experiments compared to the knowledge we have from 168 past experiments, theory, and nuclear data. We chose as differential measurements those that investigate 63 Cu and 239 Pu total cross sections, based on D-optimality rank and feasibility constraints. Two integral (criticality) experiments were selected: An experiment with Al 2 ⁢O 3 and graphite interleaved with Pu and a thick Cu reflector explores 1–30 keV, while we target the 30–600 keV range with an experiment that swaps boron in place of graphite with a different geometry.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Privacy-Preserving Artificial Intelligence on Edge Devices: A Homomorphic Encryption Approach

Recent advancements in privacy-preserving artificial intelligence (AI) have paved the way for enhanced privacy in computational processes. A standing challenge, however, is the robust privacy preservation in AI algorithms, especially when integrated into edge devices and Internet-of-Thing (IoT) infrastructures. Most prevailing solutions have adopted traditional encryption methods which, though secure, often introduce significant overhead and potential dips in accuracy. In this study, we put forth an innovative approach, utilizing the CKKS encryption scheme, aiming to harmoniously balance computational efficiency with stringent data privacy. By harnessing the capabilities of Full Homomorphic Encryption (FHE) under the CKKS scheme, we ensure the preservation of privacy, successfully curbing the inherent noise traditionally linked with accuracy reductions in similar encryption-oriented solutions. Through comprehensive experiments, our approach showcased its potential as a strong contender for privacy preservation, demonstrating commendable performance across all tests, affirming that FHE is indeed viable for devices with constrained computational power and energy resources.

Khan, Muhammad Jahanzeb↗

Generative learning for slow manifolds and bifurcation diagrams

In dynamical systems characterized by separation of time scales, the approximation of so called “slow manifolds”, on which the long term dynamics lie, is a useful step for model reduction. Initializing on such slow manifolds is a useful step in modeling, since it circumvents fast transients, and is crucial in multiscale algorithms (like the equation-free approach) alternating between fine scale (fast) and coarser scale (slow) simulations. In a similar spirit, when one studies the infinite time dynamics of systems depending on parameters, the system attractors (e.g., its steady states) lie on bifurcation diagrams (curves for one-parameter continuation, and more generally, on manifolds in state parameter space. Sampling these manifolds gives us representative attractors (here, steady states of ODEs or PDEs) at different parameter values. Algorithms for the systematic construction of these manifolds (slow manifolds, bifurcation diagrams) are required parts of the “traditional” numerical nonlinear dynamics toolkit. In more recent years, as the field of Machine Learning develops, conditional score-based generative models (cSGMs) have been demonstrated to exhibit remarkable capabilities in generating plausible data from target distributions that are conditioned on some given label. It is tempting to exploit such generative models to produce samples of data distributions (points on a slow manifold, steady states on a bifurcation surface) conditioned on (consistent with) some quantity of interest (QoI, observable). In this work, we present a framework for using cSGMs to quickly (a) initialize on a low-dimensional (reduced-order) slow manifold of a multi-time-scale system consistent with desired value(s) of a QoI (a “label”) on the manifold, and (b) approximate steady states in a bifurcation diagram consistent with a (new, out-of-sample) parameter value. This conditional sampling can help uncover the geometry of the reduced slow-manifold and/or approximately “fill in” missing segments of steady states in a bifurcation diagram. Finally, the quantity of interest, which determines how the sampling is conditioned, is either known a priori or identified using manifold learning-based dimensionality reduction techniques applied to the training data.

Dynamical systems↗

Simulation driven adaptive sampling for neutron-diffraction based strain mapping of additively manufactured parts

Neutron diffraction based strain mapping is a useful technique for measuring residual strains in additively manufactured (AM) metal parts. The measurement is traditionally done by scanning the sample in a point-wise raster pattern to extract the strain at each position. Since the overall scan can span several hours, adaptive sampling approaches using Bayesian optimization based on Gaussian process (BO-GP) regression have been introduced—demonstrating that even with a fraction of the typically made measurements the dominant strain patterns in the sample can be reconstructed. However, the parameters of the BO-GP algorithm have to be carefully chosen for best performance, and the movement time between arbitrary points can offset the time savings from a reduced number of measurement locations. In this paper, we propose algorithms to refine the BO-GP based methods by using simulations of strain patterns in AM parts based on the materials and the process used to print them. We demonstrate that the simulated strain patterns can be used to help choose better parameters for the BO-GP based framework—leading to low reconstruction error for the final strain pattern. Furthermore, we show that the strain mapping experiment can be initialized with a sampling pattern learnt from the simulation data and ordered to reduce movement time, dramatically enabling reduction in the overall time required to run the baseline BO-GP method.

Gaussian process regression↗

On optimizing the sensor spacing for pressure measurements on wind turbine airfoils

This research article presents a robust approach to optimizing the layout of pressure sensors around an airfoil. A genetic algorithm and a sequential quadratic programming algorithm are employed to derive a sensor layout best suited to represent the expected pressure distribution and, thus, the lift force. The fact that both optimization routines converge to almost identical sensor layouts suggests that an optimum exists and is reached. By comparing against a cosine-spaced sensor layout, it is demonstrated that the underlying pressure distribution can be captured more accurately with the presented layout optimization approach. Conversely, a 39 %–55 % reduction in the number of sensors compared to cosine spacing is achievable without loss in lift prediction accuracy. Given these benefits, an optimized sensor layout improves the data quality, reduces unnecessary equipment and saves cost in experimental setups. While the optimization routine is demonstrated based on the generic example of the IEA 15 MW reference wind turbine, it is suitable for a wide range of applications requiring pressure measurements around airfoils.

17 WIND ENERGY↗

Status of the Measurement of Proton Scattering on Carbon Nuclei in EMPHATIC for Neutrino Flux Uncertainty Reduction

In long-baseline neutrino oscillation experiments, Monte Carlo (MC) simulations based on hadron interactions and decays are used to predict the neutrino flux. The 10%-level systematic uncertainty of the predicted neutrino fluxes from these simulations is dominated by uncertainties in hadron interaction cross sections due to limited hadron scattering data. EMPHATIC aims to reduce the neutrino flux uncertainty by providing additional data. Using a table-top-sized spectrometer located at the Fermilab Test Beam Facility (FTBF), its physics program includes precise measurements of hadron scattering and production cross sections at various beam momenta and target species that are relevant for GeV-scale neutrino production. Using simulation, we have developed a simple single-track reconstruction algorithm that has a momentum resolution of 3-4\%. We will demonstrate the progress in developing one of EMPHATIC’s first track reconstruction algorithms – an important step in making a new single-track forward scattering measurement (p + C $\rightarrow$ p + C at several beam momenta) using Phase 1 data collected between 2022 and 2023.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

PV Performance Modeling and Stakeholder Engagement (Final Technical Report)

This core capability project’s objective is to increase the value of photovoltaic (PV) performance models by improving their functionality, demonstrating, and quantifying their validity, and offering a wide range of stakeholder engagement opportunities. In FY22-24, we developed new and improved modeling algorithms and functions to represent PV performance more accurately in a variety of environments and conditions. The “Model parameter toolkit” was developed and includes functions to translate between different module temperature models, incidence angle modifier models, and single-diode models. A new modeling capability named “PV Atlas” was also developed leveraging Sandia’s High Performance Computing resources. This capability allows us to investigate several questions and provide climate-specific best practices and geographic data files; all these are hosted on an interactive website on Sandia’s GitHub and can be used for training, system optimization, or to provide best practices for uncertainty reduction. For model validation, we published high-quality PV performance, and weather data; these data are well documented, filtered, and processed for quality and include examples on how to run PV simulations. We also developed well documented, standardized methods for validating PV models and ran independent model validation and 2 blind modeling intercomparisons engaging with 49 organizations from 17 countries. We co-led and contributed to a growing, well documented and maintained suite of open-source functions for PV modeling (i.e., the pvlib-python) and we outreached to the PV modeling stakeholders via the PVPMC workshops and web resources. In addition, this project supported US representation and leadership for the International Energy Agency (IEA) PVPS Task 13; specifically, members of our team led and supported 3 subtasks on: 1) Best practices for the optimization of bifacial photovoltaic tracking, 2) Extreme weather events and their multiple impact on PV power plants: Risks, failure mechanisms and mitigation strategies, and 3) Best practice guidelines for the use of economic and technical Key Performance Indicators (KPIs). This project resulted in the publications of 14 peer reviewed journal papers, 37 conference presentations, 6 SAND reports, 5 public datasets and 6 new webpages on the PVPMC website. It supported the release of 13 pvlib-python versions where 28 enhancements were from this PV Performance Modeling project. We co-organized 5 PVPMC workshops in FY22-24 with the participation of 214 unique institutions and around 700 participants. The PVPMC website was redesigned, and its reliability was improved; it receives over 50,000 visitors/year from 202 unique countries.

14 SOLAR ENERGY↗

FAIR Data and Interpretable AI Framework for Architectured Metamaterials (Final Report)

This research program established a transformative framework for the discovery and design of mechanical metamaterials, which are architected structures engineered to control physical phenomena like sound and vibration in ways natural materials cannot. To overcome the traditional reliance on trial-and-error, the project developed an interpretable Artificial Intelligence (AI) framework that moves beyond "black box" models to reveal the specific geometric patterns—such as "unit-cell templates"—that govern a material’s performance. A major breakthrough was the development of a hierarchical design method, which allows a single material to block vibrations across multiple frequency ranges simultaneously by layering patterns at different scales without them interfering with one another. This was further expanded to include irregular, graph-based designs that use spanning tree algorithms to ensure structural connectivity while allowing for customized, direction-dependent properties like stiffness and acoustic impedance. Beyond design, the project addressed the practicalities of real-world production by developing uncertainty quantification techniques that account for manufacturing defects and material variability, reducing the need for expensive physical testing by orders of magnitude. To speed up the discovery process, the team implemented Gaussian Process Regression and other surrogate models that provide accurate performance predictions at a fraction of the traditional computational cost. The AI-generated designs were successfully validated through fabrication of physical samples and wave propagation experiments, confirming their ability to accurately guide or reflect waves as predicted. By contributing these tools and high-quality FAIR benchmark datasets to the wider scientific community, this work provides a scalable foundation for advancing technologies in aerospace vibration control, medical imaging, and noise reduction.

36 MATERIALS SCIENCE↗

Optimal Control of Differentially Private EV Charging: A Scalable Learning Approach Under Uncertainty

Internet of Things (IoT)-enabled electric vehicles (IoEVs) enable intelligent charging coordination that accounts for grid congestion. However, increased data exchange raises privacy concerns, as charging patterns can reveal sensitive driver behavior to grid operators. Here, we propose a differentially private (DP) EV charging framework that enables coordinated control while protecting driver data with theoretical privacy guarantees. Nevertheless, integrating DP inevitably introduces uncertainty into the control strategy for EVs, which can lead to infeasible solutions. To tackle this challenge, we develop a feasible and scalable control algorithm based on constrained reinforcement learning (CRL) and convex hulls. While our framework is designed to handle the uncertainty introduced by DP, it is general and also applicable to other sources of uncertainty in EV charging, such as the stochastic nature of driver behavior and renewable variability. This ensures feasible and privacy-preserving coordination of EV charging at scale. Our method constructs convex hulls within the action space to guarantee feasibility under stochastic constraints and incorporates constraint reduction techniques to improve scalability. Case studies based on IEEE benchmark systems demonstrate that the proposed approach effectively balances feasibility under uncertainty, scalability, and privacy in large-scale EV charging control.

Engineering - Power transmission and distribution↗

Frozen Freedom: Unleashing Grocery Store Demand Flexibility: Preprint

Grocery stores consumed approximately 3% of total electricity used by commercial buildings in the U.S. in 2018 (EIA 2018), representing a unique end-use load profile characterized by the critical use of refrigerated display cases. Exploring demand response (DR) scenarios in grocery stores presents an opportunity to enhance the efficiency and sustainability of surrounding communities. In addition, recent studies demonstrate that implementing control algorithms considering demand flexibility strategies can lead to load and peak reductions in standalone refrigerated display cases. Because small business grocery stores operate on thin margins, the energy bill cost savings DR might provide could make a positive difference toward continued operations. Still, uncertainty remains about the extent of demand flexibility potential controls could provide when coupling refrigeration with whole building operation. To enhance economic viability and grid stability, it is essential to quantify the load flexibility capability of grocery stores. Advanced controls can optimize energy consumption by responding to load shedding, shifting, and DR events, as well as daily Time-of-Use (TOU) rates without compromising food safety. Using both quantitative data and interviews with community-based organizations, we developed a full-size store model and two small store models with controlled refrigerated cases, HVAC, and lighting systems based on actual grocery store properties. Through simulations, we have assessed load flexibility strategies with varied DR events. The results highlight potential for energy and peak reduction with advanced or basic controls. However, interviews and data indicate that more support is needed to make DR strategies consistently accessible to small grocery stores.

demand flexibility↗

Randomized Preconditioned Solvers for Strong Constraint 4D-Var Data Assimilation

The Strong Constraint 4D Variational (SC-4DVAR) data assimilation method is widely used in climate and weather applications. SC-4DVAR involves solving a minimization problem to compute the maximum a posteriori estimate, which we tackle using the Gauss-Newton method. The computation of the descent direction is expensive since it involves the solution of a large-scale and potentially ill-conditioned linear system, solved using the preconditioned conjugate gradient (PCG) method. Here, to address this cost, we efficiently construct scalable preconditioners using three different randomization techniques, which all rely on a certain low-rank structure involving the Gauss-Newton Hessian. The proposed techniques come with theoretical guarantees on the condition number, and at the same time, are amenable to parallelization. We also develop an adaptive approach to estimate the sketch size and choose between the reuse or recomputation of the preconditioner. We demonstrate the performance and effectiveness of our methodology on two representative model problems—the Burgers and barotropic vorticity equation—showing a drastic reduction in both the number of PCG iterations and the number of Gauss-Newton Hessian products after including the preconditioner construction cost.

Gauss-Newton↗

A Digital Twin Framework Utilizing Machine Learning for Robust Predictive Maintenance: Enhancing Tire Health Monitoring

We introduce a novel digital twin (DT) framework for the predictive maintenance of long-term physical systems. Using monitoring tire health as an application, we show how the DT framework can be used to enhance automotive safety and efficiency, and how the technical challenges can be overcome using a three-step approach. First, to manage the data complexity over a long operation span, we employ data reduction techniques to concisely represent physical tires using historical performance and usage data. Relying on these data, for fast real-time prediction, we train a transformer-based model offline on our concise dataset to predict future tire health over time, represented as remaining casing potential (RCP). Based on our architecture, our model quantifies both epistemic and aleatoric uncertainties, providing reliable confidence intervals around predicted RCP. Second, to incorporate real-time data, we update the predictive model in the DT framework, ensuring its accuracy throughout its lifespan with the aid of hybrid modeling and the use of the discrepancy function. Third, to assist decision-making in predictive maintenance, we implement a tire state decision algorithm, which strategically determines the optimal timing for tire replacement based on RCP forecasted by our transformer model. This approach ensures that our DT accurately predicts system health, continually refines its digital representation, and supports predictive maintenance decisions. Furthermore, our framework effectively embodies a physical system, leveraging big data and machine learning (ML) for predictive maintenance, model updates, and decision-making.

advanced computing infrastructure↗

FedEFsz: Fair Cross-Silo Federated Learning System With Error-Bounded Lossy Compression

Cross-Silo federated learning systems have been identified as an efficient approach to scaling DNN training across geographically-distributed data silos to preserve the privacy of the training data. Communication efficiency and fairness are two major issues that need to be both satisfied when federated learning systems are deployed in practice. Simultaneously guaranteeing both of them, however, is exceptionally difficult because simply combining communication reduction and fairness optimization approaches often causes non-converged training or drastic accuracy degradation. Here, to bridge this gap, we propose FedEFsz. On the one hand, it integrates the state-of-the-art error-bounded lossy compressor SZ3 into cross-silo federated learning systems to significantly reduce communication traffic during the training. On the other hand, it achieves a high fairness (i.e., rather consistent model accuracy and performance across different clients) through a carefully designed heuristic algorithm that can tune the error-bound of SZ3 for different clients during the training. Extensive experimental results based on a GPU cluster with 65 GPU cards show that FedEFsz improves the fairness across different benchmarks by up to 60.88% and meanwhile reduces the communication traffic by up to 315×.

Cross-Silo Federated Learning Systems↗

Spectroscopic Online Monitoring: Using a Multi-Track Visible Spectrometer to Facilitate a Mass Balance Study in a Simulated TALSPEAK Process

Nuclear energy is a promising low-carbon energy candidate to meet the increased demand for green energy, where the integration of fuel recycling can have significant benefits for material usage and waste reduction. Utilizing in situ monitoring tools can provide ample opportunities to better control and safeguard nuclear material recycle processes while also offering knowledge and insight into real-time solution properties. The simultaneous measurement of analytical targets in multiple process locations can enable real-time mass balance and material accountancy calculations. This is demonstrated here with a mass balance study of Nd 3+ on countercurrent aqueous/organic metal extraction within a single centrifugal contactor. The Nd 3+ concentration was simultaneously monitored at the inlets and outlets of both aqueous and organic phases using a visible absorbance detector that allowed for the simultaneous measurement of up to six locations. The Nd 3+ concentration was calculated by using chemical data science algorithms, where model training sets were collected on a single track of the detector. The discussion includes addressing the challenges of using a model collected on a single track and applying it as a model across the other tracks on the detector. Each track of the detector corresponds to one measurement location on the contactor. The difference in the integrated moles of Nd 3+ between the inlet and outlet at the end of the experiment was near zero, indicating that the mass balance of this experiment was maintained. Overall, the online spectroscopic monitoring was able to follow changing solution conditions and accurately measure the concentration of Nd 3+ in different locations within the contactor system.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The detection of marine microseismic activity with the CUORE tonne-scale cryogenic experiment

Vibrations from experimental setups and the environment are a persistent source of noise for low-temperature calorimeters searching for rare events, including neutrinoless double beta ( 0νββ ) decay or dark matter interactions. Such noise can significantly limit experimental sensitivity to the physics case under investigation. Here, we report the detection of marine microseismic vibrations using mK-scale calorimeters. This study employs a multi-device analysis correlating data from CUORE, the leading experiment in the search for 0νββ decay with mK-scale calorimeters, and the Copernicus Earth Observation program, revealing the seasonal impact of Mediterranean Sea activity on CUORE’s energy thresholds, resolution, and sensitivity over four years. The detection of marine microseisms underscores the need to address faint environmental noise in ultra-sensitive experiments. Understanding how such noise couples to the detector and developing mitigation strategies is essential for next-generation experiments. We demonstrate one such strategy: a noise decorrelation algorithm implemented in CUORE using auxiliary sensors, which reduces vibrational noise and improves detector performance. Enhancing sensitivity to 0νββ decay and to rare events with low-energy signatures requires identifying unresolved noise sources, advancing noise reduction methods, and improving vibration suppression systems, all of which inform the design of next-generation rare event experiments.

experimental nuclear physics↗

Hybrid data-driven cement-stabilized soil design: An integration of machine learning, multi-objective optimization, and life cycle assessment

Soil stabilization is crucial in geotechnical engineering, yet conventional methods are often time-consuming, resource-intensive, and environmentally unsustainable. Despite growing interest in Machine Learning (ML) and optimization tools for mix design, few studies integrate these methods with decision-making techniques and environmental assessment to support practical implementation. This study proposes a hybrid data-driven framework for predicting strength, optimizing mix compositions, and evaluating environmental impacts via life cycle assessment of cement-stabilized soft soils. Six ML models were evaluated, and the top-performing eXtreme Gradient Boosting (XGB) model was further improved using the Grey Wolf Optimizer (GWO). The optimized XGB-GWO model, integrated with a polynomial cost function, served as the objective function in a multi-objective optimization problem solved via the Non-Dominated Sorting Genetic Algorithm II (NSGA-II), with final mix selection guided by the entropy-weighted TOPSIS method. Validation through a case study produced mix designs offering superior strength-cost trade-offs, with the optimal mix achieving 2243.2 kPa unconfined compressive strength and a 16.07 % reduction in carbon emissions compared to the highest-cost design. In conclusion, this study offers a sustainable, scalable approach to soil stabilization and supports informed decision-making in construction.

Life cycle assessment↗

Complexity Reduction Methods for Large-Scale Spatially Explicit Biofuels Network Design

The size and complexity of energy system optimization models have increased significantly in recent years, driven by the availability of high-resolution spatial data. We present complexity reduction and solution methods that enable us to efficiently represent high-resolution spatial data in the network design of large-scale energy systems. We aim to reduce the size and enhance the computational efficiency of network design models without sacrificing solution accuracy. Specifically, we first present how to aggregate highly granular data into larger resolutions without averaging out their specific properties through a composite-curve-based approach and then develop a method to linearly represent these curves. Second, we utilize a general clustering method to determine groups of geographically proximate biomass fields and establish a single transportation arc for all of them, reducing the number of transportation-related variables while maintaining an accurate representation of the system. Finally, we introduce a two-step algorithm that decomposes large-scale network design problems into two smaller, more manageable subproblems. We demonstrate the application of our methods using a case study of switchgrass-to-biofuels network design in the eight states of the U.S. Midwest, using realistic and highly explicit spatial data.

09 BIOMASS FUELS↗