Search NASA⌕ Search

SEARCH · Search NASA

Results for “Derived quantities preservation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Online and Scalable Data Compression Pipeline with Guarantees on Quantities of Interest

Data compression is becoming critical for data-intensive scientific applications. Scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Prior work has shown that a pipeline can be built to guarantee error on the primary data (PD) within user-defined bounds and achieve near-floating point QoI errors. In this paper, we present novel computational approaches for accelerating the pipeline and demonstrate results that enable concurrent execution of compression in parallel with the simulation nodes. This allows compression, including the writing of the required compression data, for the previous time step to be completed while the simulation proceeds with the current time step. Overall, the approach presented in this paper results in a 6–8 times improvement in computational overhead compared to previous work. These results were obtained using data generated by a large-scale fusion code called XGC, which produces hundreds of terabytes of data in a single day.

Banerjee, Tania↗

Scalable Hybrid Learning Techniques for Scientific Data Compression

Data compression is becoming critical for storing scientific data because many scientific applications need to store large amounts of data and post process this data for scientific discovery. Unlike image and video compression algorithms that limit errors to primary data (PD), scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Here, this article presents a physics-informed compression technique implemented as an end-to-end, scalable, GPU-based pipeline for data compression that addresses this requirement. Our hybrid compression technique combines machine learning techniques and standard compression methods. Specifically, we combine an autoencoder, an error-bounded lossy compressor to provide guarantees on raw data error, and a constraint satisfaction post-processing step to preserve the QoIs within a minimal error (generally less than floating point error). The effectiveness of the data compression pipeline is demonstrated by compressing nuclear fusion simulation data generated by a large-scale fusion code, XGC, which produces hundreds of terabytes of data in a single day. Our approach works within the ADIOS framework and results in compression by a factor of more than 150 while requiring only a few percent of the computational resources necessary for generating the data, making the overall approach highly effective for practical scenarios.

ITER↗

Fast Algorithms for Scientific Data Compression

Many scientific simulations and experiments generate terabytes to petabytes of data daily, necessitating data compression techniques. Unlike video and image compression, scientists require methods that accurately preserve primary data (PD) and derived quantities of interest (QoIs). In our previous work, we demonstrated the effectiveness of hybrid compression techniques that combine machine learning with traditional approaches. This paper presents innovative computational techniques aimed at expediting the compression pipeline. Our experiments, conducted on two distinct platforms with a large-scale XGC-based fusion simulation, demonstrate that the overhead incurred by these new approaches is less than one percent of the computational resources needed for the simulation.

Banerjee, Tania↗

Deep learning symmetries and their Lie groups, algebras, and subalgebras from first principles

Abstract We design a deep-learning algorithm for the discovery and identification of the continuous group of symmetries present in a labeled dataset. We use fully connected neural networks to model the symmetry transformations and the corresponding generators. The constructed loss functions ensure that the applied transformations are symmetries and the corresponding set of generators forms a closed (sub)algebra. Our procedure is validated with several examples illustrating different types of conserved quantities preserved by symmetry. In the process of deriving the full set of symmetries, we analyze the complete subgroup structure of the rotation groups SO (2), SO (3), and SO (4), and of the Lorentz group S O ( 1 , 3 ) . Other examples include squeeze mapping, piecewise discontinuous labels, and SO (10), demonstrating that our method is completely general, with many possible applications in physics and data science. Our study also opens the door for using a machine learning approach in the mathematical study of Lie groups and their properties.

97 MATHEMATICS AND COMPUTING↗

Integrability Breaking from Backscattering

Herein we analyze the onset of diffusive hydrodynamics in the one-dimensional hard-rod gas subject to stochastic backscattering. While this perturbation breaks integrability and leads to a crossover from ballistic to diffusive transport, it preserves infinitely many conserved quantities corresponding to even moments of the velocity distribution of the gas. In the limit of small noise, we derive the exact expressions for the diffusion and structure factor matrices, and show that they generically have off diagonal components. We find that the particle density structure factor is non-Gaussian and singular near the origin, with a return probability showing logarithmic deviations from diffusion.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Hybrid learning techniques for scientific data reduction with performance guarantees

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Including the vacuum energy in stellarator coil design

Being three-dimensional, stellarators have the advantage that plasma currents are not essential for creating rotational-transform; however, the external current-carrying coils in stellarators can have strong geometrical shaping, which can complicate the construction. Reducing the inter-coil electromagnetic forces acting on strongly shaped 3D coils and the stress on the support structure while preserving the favorable properties of the magnetic field is a design challenge. In this work, we recognize that the inter-coil forces are the gradient of the vacuum magnetic energy. We introduce an objective functional built on the usual quadratic flux on a prescribed target surface together with a weighed penalty on the vacuum energy. The Euler–Lagrange equation for stationary states is derived, and numerical illustrations are computed using a modern stellarator optimization framework. A study of the effect of the energy functional on the inter-coil forces is conducted and the energy is shown to be a promising quantity in producing coils with low forces.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Unified Workflow for Sensitivity-Based Kinetic Analysis in Microkinetic Models

Degrees of rate control (DRC), apparent activation energies, and apparent reaction orders are established local sensitivity diagnostics for interpreting microkinetic models, but applying them routinely to large mechanisms often requires substantial reaction-specific bookkeeping, perturbation design, and postprocessing. Here, in this study, we present a unified derivative-based workflow that evaluates these quantities from a single compiled reaction-network model and target-rate definition. For any user-provided microkinetic model, the workflow compiles the mechanism into stoichiometrically consistent mass-action rate equations, solves the surface dynamics, and uses automatic differentiation to compute sensitivities with respect to rate constants, temperature, and gas partial pressures. By combining their calculations in the same framework, the workflow clearly demonstrates the relationships between different DRCs and the apparent activation energy. Using existing examples of propylene partial oxidation and methane oxidation on Pd(100), we verify expected transient redistribution of rate control, distinguish net Campbell DRCs from one-sided directional sensitivities, and show how apparent activation energy can be reconstructed either from one-sided DRCs or from state-based DRCs while critical mechanistic insights are obtained consistently. In the methane oxidation case, a pathway-subset test further illustrates how a simplified mechanism preserves key kinetic signatures of a full model, showing the potential of our user-friendly tool for model construction beyond kinetic analysis.

36 MATERIALS SCIENCE↗

A Class of Sparse Johnson–Lindenstrauss Transforms and Analysis of their Extreme Singular Values

The Johnson–Lindenstrauss (JL) lemma is a powerful tool for dimensionality reduction in modern algorithm design. The lemma states that any set of high-dimensional points in a Euclidean space can be projected into lower dimensions while approximately preserving pairwise Euclidean distances. Random matrices satisfying this lemma are called JL transforms (JLTs). Inspired by existing $s$-hashing JLTs with exactly $s$ nonzero elements on each column, the present work introduces an ensemble of sparse matrices encompassing so-called $s$-hashing-like matrices whose expected number of nonzero elements on each column is $s$. The independence of the sub-Gaussian entries of these matrices and the knowledge of their exact distribution play an important role in their analyses. Using properties of independent sub-Gaussian random variables, these matrices are demonstrated to be JLTs, and their smallest nontrivial singular values and largest singular values are estimated nonasymptotically using a technique from geometric functional analysis. As the dimensions of the matrix grow to infinity, these singular values are proved to converge almost surely to fixed quantities (by using the universal Bai–Yin law) and in distribution to the Gaussian orthogonal ensemble Tracy–Widom law after proper rescalings. Understanding the behaviors of extreme singular values is important in general because they are often used to define a measure of stability of matrix algorithms. For example, JLTs were recently used in derivative-free optimization algorithmic frameworks to select random subspaces in which are constructed random models or poll directions to achieve scalability, and hence estimating their smallest singular value in particular helps determine the dimension of these subspaces.

97 MATHEMATICS AND COMPUTING↗

High precision tests of QCD without scale or scheme ambiguities: The 40th anniversary of the Brodsky–Lepage–Mackenzie method

A key issue in making precise predictions in QCD is the uncertainty in setting the renormalization scale μ r and thus determining the correct values of the QCD running coupling α s (μ r ) at each order in the perturbative expansion of a QCD observable. It has often been conventional to simply set the renormalization scale to the typical scale of the process Q and vary it in the range μ r $\in$ [Q/2, 2Q] in order to estimate the theoretical error. This is the practice of Conventional Scale Setting (CSS). The resulting CSS prediction will however depend on the theorist’s choice of renormalization scheme and the resulting pQCD series will diverge factorially. It will also disagree with renormalization scale setting used in QED and electroweak theory thus precluding grand unification. A solution to the renormalization scale-setting problem is offered by the Principle of Maximum Conformality (PMC), which provides a systematic way to eliminate the renormalization scale-and-scheme dependence in perturbative calculations. The PMC method has rigorous theoretical foundations, it satisfies Renormalization Group Invariance (RGI) and preserves all self-consistency conditions derived from the renormalization group. The PMC cancels the renormalon growth, reduces to the Gell-Mann–Low scheme in the N c → 0 Abelian limit and leads to scale- and scheme-invariant results. The PMC has now been successfully applied to many high-energy processes. In this article we summarize recent developments and results in solving the renormalization scale and scheme ambiguities in perturbative QCD. In particular, we present a recently developed method the PMC ∞ and its applications, comparing the results with CSS. The method preserves the property of renormalizable SU(N)/U(1) gauge theories defined as Intrinsic Conformality (iCF). This property underlies the scale invariance of physical observables and leads to a remarkably efficient method to solve the conventional renormalization scale ambiguity at every order in pQCD. This new method reflects the underlying conformal properties displayed by pQCD at NNLO, eliminates the scheme dependence of pQCD predictions and is consistent with the general properties of the PMC. A new method to identify conformal and β-terms, which can be applied either to numerical or to theoretical calculations is also shown. We present results for the thrust and C-parameter distributions in e + e - annihilation showing errors and comparison with the CSS. We also show results for a recent innovative comparison between the CSS and the PMC ∞ applied to the thrust distribution investigating both the QCD conformal window and the QED N c → 0 limit. In order to determine the thrust distribution along the entire renormalization group flow from the highest energies to zero energy, we consider the number of flavors near the upper boundary of the conformal window. In this flavor-number regime the theory develops a perturbative infrared interacting fixed point. These results show that PMC ∞ leads to higher precision and introduces new interesting features in the PMC. In fact, this method preserves with continuity the position of the peak, showing perfect agreement with the experimental data already at NNLO. We also show a detailed comparison of the PMC ∞ with the other PMC approaches: the multi-scale-setting approach (PMCm) and the single-scale-setting approach (PMCs) by comparing their predictions for three important fully integrated quantities R e+e- , R$_Τ$ and Γ (H→$b\bar{b}$) up to the four-loop accuracy.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗