Search NASASearch

SEARCH · Search NASA

Results for “MATHEMATICAL STATISTICS”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Information and Statistics in Nuclear Experiment and Theory (ISNET)

As with all empirical sciences, nuclear physics operates in the virtuous cycle of the scientific method: observations inspire theoretical models; models lead to new predictions; predictions are tested in experiments; experiments lead to new observations; and so on. Evaluating what we are inferring, and how certain we are of it, is key to this process. These requirements, and a general interest in applying novel statistical, mathematical, and computational techniques, led to the formation of a dedicated research community entitled “Information and Statistics in Nuclear Experiment and Theory (ISNET)” (https://isnet-series.github.io/), which now includes more than 300 members. While the community’s interests lean toward nuclear theory, the unifying theme for this group is the inference of knowledge from data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Basic Research Needs for Inverse Methods for Complex Systems under Uncertainty

Inverse problems, which aim to infer unknown properties of a system using experimental and observational data, are central to addressing many of the U.S. Department of Energy’s (DOE) most critical scientific and engineering challenges. Accurate, computationally efficient, and data-efficient solutions to inverse problems are essential for advancing DOE mission-critical science drivers, including analyzing data from large-scale experimental facilities, optimizing fusion reactor performance, accelerating materials discovery, enhancing geophysical imaging, improving wildfire predictions, and enabling autonomous systems and digital twins. However, these problems are becoming increasingly complex, often involving nonlinear, highdimensional, and interconnected systems and models that span multiple physics and scales, while relying on data with varying quantity, quality, and information content. Compounding these challenges is the uncertainty inherent in DOE-relevant systems, where errors in inputs, noise in data, incompleteness of data, and discrepancies between models and reality constrain the accuracy and precision of solutions. At the same time, the convergence of recent scientific computing trends—scientific machine learning, artificial intelligence, and computing advances such as exascale computing—is creating unprecedented opportunities for tackling these challenges. The cross-cutting nature of inverse problems, combined with their growing complexity and rapidly evolving data and algorithmic demands, strongly motivates the formulation of a prioritized research agenda to maximize their capabilities and impact. In response to this need, DOE’s Advanced Scientific Computing Research (ASCR) program in the Office of Science convened the Workshop on Basic Research Needs for Inverse Problems for Complex Systems Under Uncertainty in June 2025. This workshop brought together experts across disciplines to identify grand challenges and major opportunities in the field. Through collaborative discussions, the workshop defined transformative research directions aimed at addressing the mathematical, statistical, and computational challenges posed by inverse problems under uncertainty. As a result of these efforts, four priority research directions (PRDs) were identified to guide future research and development in this area. These PRDs, summarized below, represent a roadmap for advancing the foundational science and mathematics of inverse problems, enabling robust, scalable, and uncertainty-aware solutions that are critical for DOE applications.

97 MATHEMATICS AND COMPUTING

Basic Research Needs for Inverse Methods for Complex Systems under Uncertainty [Brochure]

The four priority research directions outlined in this brochure represent a cohesive vision for advancing the science of inverse problems for complex systems under uncertainty. Together, they address the critical challenges of: discovering, exploiting, and preserving physical and problem structure; overcoming model limitations; integrating disparate, multimodal, and/or dynamic data; and tailoring the solution of inverse problems to downstream tasks. While each PRD focuses on a distinct aspect of inverse-problem research, their interconnected nature highlights the importance of a holistic approach that leverages progress across all areas to achieve transformative solutions. This agenda calls for research across mathematics, statistics, and computer science disciplines, which are guided and complemented by rapid advances in artificial intelligence, high-performance computing, and experimental facilities, to unlock new capabilities, maximize scientific impact, and meet the growing demands of inverse problems that arise across applications that are critical to DOE's mission.

97 MATHEMATICS AND COMPUTING

Database and deep-learning scalability of anharmonic phonon properties by automated brute-force first-principles calculations

Understanding the anharmonic phonon properties of crystal compounds—such as phonon lifetimes and thermal conductivities—is essential for investigating and optimizing their thermal transport behaviors. These properties also impact optical, electronic, and magnetic characteristics through interactions between phonons and other quasiparticles and fields. In this study, we develop an automated first-principles workflow to calculate anharmonic phonon properties and build a comprehensive database encompassing more than 6500 inorganic compounds. Utilizing this dataset, we train a graph neural network model to predict thermal conductivity values and spectra from structural parameters, demonstrating a scaling law in which prediction accuracy improves with increasing training data size. High-throughput screening with the model enables the identification of materials exhibiting extreme thermal conductivities—both high and low. The resulting database offers valuable insights into the anharmonic behavior of phonons, thereby accelerating the design and development of advanced functional materials.

Ohnishi, Masato [University of Tokyo (Japan); Inst

sOPTICS: a modified density-based algorithm for identifying galaxy groups/clusters and brightest cluster galaxies

A direct approach to studying the galaxy–halo connection is to analyse groups and clusters of galaxies that trace the underlying dark matter haloes, emphasizing the importance of identifying galaxy clusters and their associated brightest cluster galaxies (BCGs). In this work, we test and propose a robust density-based clustering algorithm that outperforms the traditional Friends-of-Friends (FoF) algorithm in the currently available galaxy group/cluster catalogues. Our new approach is a modified version of the Ordering Points To Identify the Clustering Structure (OPTICS) algorithm, which accounts for line-of-sight positional uncertainties due to redshift space distortions by incorporating a scaling factor, and is thereby referred to as sOPTICS. When tested on both a galaxy group catalogue based on semi-analytic galaxy formation simulations and observational data, our algorithm demonstrated robustness to outliers and relative insensitivity to hyperparameter choices. In total, we compared the results of eight clustering algorithms. The proposed density-based clustering method, sOPTICS, outperforms FoF in accurately identifying giant galaxy clusters and their associated BCGs in various environments with higher purity and recovery rate, also successfully recovering 115 BCGs out of 118 reliable BCGs from a large galaxy sample. Furthermore, when applied to an independent observational catalogue without extensive re-tuning, sOPTICS maintains high recovery efficiency, confirming its flexibility and effectiveness for large-scale astronomical surveys.

79 ASTRONOMY AND ASTROPHYSICS

The hierarchical growth of bright central galaxies and intracluster light as traced by the magnitude gap

Using a sample of 2800 galaxy clusters identified in the Dark Energy Survey across the redshift range 0.20 < z < 0.60, we characterize the hierarchical assembly of bright central galaxies (BCGs) and the surrounding intracluster light (ICL). To quantify hierarchical formation we use the stellar mass–halo mass (SMHM) relation, comparing the halo mass, estimated via the mass–richness relation, to the stellar mass within the BCG + ICL system. Moreover, we incorporate the magnitude gap (M14), the difference in brightness between the BCG (measured within 30 kpc) and fourth brightest cluster member galaxy within 0.5 $R_{200,c}$, as a third parameter in this linear relation. The inclusion of M14, which traces BCG hierarchical growth, increases the slope and decreases the intrinsic scatter, highlighting that it is a latent variable within the BCG + ICL SMHM relation. Moreover, the correlation with M14 decreases at large radii. However, the stellar light within the BCG + ICL transition region (30 –80 kpc) most strongly correlates with halo mass and has a statistically significant correlation with M14. Since the transition region and M14 are independent measurements, the transition region may grow due to the BCG’s hierarchical formation. Additionally, as M14 and ICL result from hierarchical growth, we use a stacked sample and find that clusters with large M14 values are characterized by larger ICL and BCG + ICL fractions, which illustrates that the merger processes that build the BCG stellar mass also grow the ICL. Furthermore, this may suggest that M14 combined with the ICL fraction can identify dynamically relaxed clusters.

79 ASTRONOMY AND ASTROPHYSICS

Stacked reverberation mapping of high-redshift quasars in DESI. I. Feasibility analysis

The broad-line region of quasars has long been probed by reverberation mapping techniques that measure time lags between continuum and broad emission-line variations. Stacked reverberation mapping has been proposed as a less observationally expensive alternative to traditional methods. This ensemble approach also reduces biases from small-number statistics. The Dark Energy Spectroscopic Instrument (DESI) is conducting the most extensive spectroscopic survey of quasars to date. We create mock light curves emulating expected DESI quasar observations at redshifts $1.48\lt z\lt 5.2$ and luminosities $44.68 \le \log \lambda L_{1350 \mathring{\rm A}{}} / \mathrm{erg\, s^{-1}} \le 45.99$ to test stacked reverberation mapping feasibility using sparse spectroscopic data paired with well-sampled photometric data. The pipeline, using the lag estimation code JAVELIN (Just Another Vehicle for Estimating Lags In Nuclei), successfully recovers the simulated C IV lags within 1σ of the true values using spectroscopic light curves composed of only a few spectral epochs (2–10) with irregular cadences. We investigate how observational factors, including C IV flux error magnitude, number of stacked quasars, and spectral epoch count, affect performance. This work motivates a pathway for future stacked reverberation mapping projects with large-scale spectroscopic surveys of quasars having $\ge 2$ spectroscopic observations. Our results suggest an economical alternative for constraining and extending the radius–luminosity relation to higher redshifts and luminosities. Subsequently, this relation can be employed more reliably in single-epoch black hole mass measurements and quasar cosmology in these distant regimes.

quasars: general, quasars: supermassive black hole

Generalization error guaranteed auto-encoder-based nonlinear model reduction for operator learning

Many physical processes in science and engineering are naturally represented by operators between infinite-dimensional function spaces. The problem of operator learning, in this context, seeks to extract these physical processes from empirical data, which is challenging due to the infinite or high dimensionality of data. An integral component in addressing this challenge is model reduction, which reduces both the data dimensionality and problem size. In this paper, we utilize low-dimensional nonlinear structures in model reduction by investigating Auto-Encoder-based Neural Network (AENet). AENet first learns the latent variables of the input data and then learns the transformation from these latent variables to corresponding output data. Our numerical experiments validate the ability of AENet to accurately learn the solution operator of nonlinear partial differential equations. Furthermore, we establish a mathematical and statistical estimation theory that analyzes the generalization error of AENet. Finally, our theoretical framework shows that the sample complexity of training AENet is intricately tied to the intrinsic dimension of the modeled process, while also demonstrating the robustness of AENet to noise.

Auto-encoder

Lattice expansion due to hydrogen absorption into β-rhombohedral boron

β-Rhombohedral boron (β-boron) represents one of the most popular and important allotropic forms of elemental boron. The unit cell of β-boron crystal consists of B 106.6 with a complicated and relatively open structure. We have previously reported the abrupt lattice expansion and shrinkage of β-boron crystal by thermal treatment above 700 K and photoirradiation at room temperature. Our recent studies, combined with X-ray diffraction and atom probe tomography experiments, suggest that the lattice expansion and shrinkage are related to the absorption and release of hydrogen into the structure. In conclusion, the results lead us to a new application of β-boron as a photo-switchable hydrogen storage material.

Atom probe tomography

The cluster decomposition of the configurational energy of multicomponent alloys

Abstract The cluster expansion method (CEM) is a widely used lattice-based technique in the study of multicomponent alloys. Despite its prevalent use, a clear understanding of expansion terms is lacking. We present a modern mathematical formalism of the CEM and introduce thecluster decomposition—a unique and basis-independent decomposition for functions of the atomic configuration in a crystal. We identify the cluster decomposition as an invariant ANOVA decomposition; and demonstrate how functional analysis of variance and sensitivity analysis can be used to interpret interactions among species. Furthermore, we show how the mathematical structure of the cluster decomposition enables numerical evaluation that scales with the number of clusters and is independent of the number of species. Overall, our work enables rigorous interpretations of interactions among species, provides opportunities to explore parameter estimation beyond linear regression, introduces a numerical efficient implementation, and enables analysis of cluster expansions based on established mathematical and statistical principles.

Chemistry

Phonon Olympics: Phonon property and lattice thermal conductivity benchmarking from open-source packages

Three widely used open-source packages for determining phonon properties and lattice thermal conductivities (ALAMODE, phono3py, and ShengBTE) are benchmarked by teams of expert users and the package developers. The phonons for Ge, RbBr, monolayer MoSe 2 , and AlN are modeled at zero temperature, and they scatter through three-phonon and phonon-isotope processes, with thermal conductivities obtained from the linearized Peierls–Boltzmann transport equation with input from density functional theory calculations. Over a wide range of temperatures, the thermal conductivities calculated by the teams fall within at most ±15% of their mean values for each of the four materials. The phonon frequencies, obtained from the harmonic force constants, do not show large differences between the calculations, indicating that the modal heat capacities and group velocities are not responsible for the thermal conductivity variations. It is the lifetimes associated with three-phonon scattering, obtained from the cubic force constants, that drive the variations. The many decisions required to calculate the cubic force constants (e.g., supercell size, atomic displacement, neighbor cutoff, and application of symmetries) make identification of the precise origin of the thermal conductivity variations challenging. The calculated thermal conductivities do not generally show agreement with experimental measurements, which is attributed to the limitations of the density functional theory calculations. Guidance for the development of best practices is provided, which will help to standardize protocols needed for building thermal conductivity databases. The results provide a baseline for future benchmarking of other packages and more advanced calculations.

McGaughey, Alan J. H. [Carnegie Mellon Univ., Pitt

Recent advances in plasma control and physics research in the Large Helical Device

The Large Helical Device (LHD), the largest superconducting helical system in the world, is equipped with advanced heating and diagnostic tools, facilitating plasma control and physics research. Data assimilation was employed for electron temperature control using a real-time Thomson scattering system and real time prediction code. A virtual LHD environment enabled visualization of escaping high-energy tritium ions and demonstrated that these ions impact the rear side of the divertor plate. Pioneering results crucial to plasma control have also been achieved. Real-time wall conditioning using Lithium granule dropping improved bulk ion energy and particle transport while simultaneously enhancing the heavy impurity transport. Progress has also been made in the investigation of turbulence-driven transport. At the confinement bifurcation, ion-scale turbulence decreased, while electron-scale turbulence increased. A change in the anisotropy of turbulent eddies was also observed at the confinement bifurcation. Coexistence of local and non-local turbulence was identified in electron-scale turbulence. Non-local turbulence exhibited the rapid spatial propagation of perturbations throughout the plasma, while local turbulence followed the temperature gradient. A transition between drift-wave turbulence and magnetohydrodynamics (MHD) turbulence was observed with the turbulence minimized at the transition condition. Machine learning analysis was employed to evaluate the temperate and density conditions of this turbulence transition. Then, real-time control of fueling and heating was applied to maintain the turbulence transition condition, improving the energy confinement enhancement factor by 20%. In addition, evidence was obtained for collisionless ion heating by energetic-ion-driven geodesic acoustic modes and MHD bursts. These achievements represent unique contributions to the development of fusion reactors.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

OpenUniverse2024: a shared, simulated view of the sky for the next generation of cosmological surveys

The OpenUniverse2024 simulation suite is a cross-collaboration effort to produce matched simulated imaging for multiple surveys as they would observe a common simulated sky. Both the simulated data and associated tools used to produce it are intended to uniquely enable a wide range of studies to maximize the science potential of the next generation of cosmological surveys. We have produced simulated imaging for approximately 70 deg 2 of the Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) Wide-Fast-Deep survey and the Nancy Grace Roman Space Telescope High-Latitude Wide-Area Survey, as well as overlapping versions of the ELAIS-S1 Deep-Drilling Field for LSST and the High-Latitude Time-Domain Survey for Roman. OpenUniverse2024 includes (i) an early version of the updated extragalactic model called Diffsky, which substantially improves the realism of optical and infrared photometry of objects, compared to previous versions of these models; (ii) updated transient models that extend through the wavelength range probed by Roman and Rubin; and (iii) improved survey, telescope, and instrument realism based on up-to-date survey plans and known properties of the instruments. It is built on a new and updated suite of simulation tools that improves the ease of consistently simulating multiple observatories viewing the same sky. The approximately 400 TB of synthetic survey imaging and simulated universe catalogs are publicly available, and we preview some scientific uses of the simulations.

large-scale structure of Universe

AEOLUS: Advances in Experimental Design, Optimal Control, and Learning for Uncertain Complex Systems

Sustained advances in the mathematics of modeling and simulation have resulted in the capability today for routine simulation of a number of large scale complex DOE-relevant systems. As remarkable as this capability for solving the so-called forward problem is, it is typically only the first step-an inner loop within an outer loop that explores the simulation model's parameter space and decision space to characterize uncertainty in the model's predictions, learn unknown model parameters from data, design the most informative experiments, determine optimal control strategies, and create optimal designs. Broadly, what unifies all of these outer loop problems is that they are, in one form or another, optimization problems over parameter/control/design space that are constrained by complex uncertain models. To fully realize the power of scientific simulation as a basis for scientific discovery, technological innovation, and rational decision-making, it is imperative to move beyond simulation to tackle the outer loop of optimization for learning from data, experimental design, and control with complex uncertain models. When the models under consideration are large-scale and complex, and when the optimization variable and uncertain parameter spaces are high (or infinite) dimensional, this constitutes a grand challenge of the highest order, and is intractable with conventional methods. To overcome these challenges, the AEOLUS Center was established to develop a unified mathematical, computational, and statistical framework for (1) Learning predictive models from complex data via Bayesian inference and optimization, and (2) Optimizing experiments, processes, and designs using the resulting uncertain models. These problems are intractable with conventional methods, for several reasons: (1) The simulation problems that govern the inner loops of the optimization problems are expensive to execute (due to severe nonlinearity, heterogeneity, multiphysics/multiscale coupling); (2) The optimization variable and uncertain parameter spaces are high dimensional, often stemming from discretizations of infinite dimensional fields such as initial conditions, sources, or material properties. We argue that the key to overcoming these challenges is to develop new mathematical, computational, and statistical methods that exploit the structure of the Bayesian inference and optimization problems mediated by their underlying complex uncertain models. This structure includes the regularity, sparsity, geometry, low intrinsic dimensionality, and multifidelity nature of the maps from uncertain parameter/optimization variable spaces to the specific objectives targeted: Bayesian inference, optimal experimental design, and optimal control design. Black box methods developed as generic tools are incapable of exploiting this structure. To be successful, we must create, integrate, and cross-fertilize ideas across multiple areas of applied math--including approximation theory, Bayesian inference, data science, experimental design, information theory, machine learning, model reduction, optimal control theory, parallel algorithms, PDE-constrained optimization, randomized algorithms, stochastic optimization, and uncertainty quantification--all while exploiting the structure of the problems at hand. With this goal in mind, we have marshaled a team of leading authorities in these areas. While the methods we develop will be broadly applicable across a wide spectrum of DOE problems in which experiments inform models and the systems those models describe must be optimized under uncertainty, we have chosen a specific area, advanced manufacturing and materials, to drive our work. AMM is characterized by complex models across multiple scales, and is a rich source of challenging problems in inference, experimental design, and optimal control, requiring multifaceted and integrated advances in applied mathematics. As such, AMM serves as an excellent vehicle to motivate and demonstrate the advances in applied mathematics developed by our center.

97 MATHEMATICS AND COMPUTING

Modeling Protein–Protein and Protein–Ligand Interactions by the ClusPro Team in CASP16

ABSTRACT In the CASP16 experiment, our team employed hybrid computational strategies to predict both protein–protein and protein–ligand complex structures. For protein–protein docking, we combined physics‐based sampling—using ClusPro FFT docking and molecular dynamics—with AlphaFold (AF)‐based sampling, followed by AF‐based refinement. Our method produced numerous high‐accuracy complex models, including cases where AF alone failed, underscoring the critical role of physics‐based sampling alongside deep learning‐based refinement. For protein–ligand docking, we integrated the ClusPro LigTBM template‐based approach with a machine learning‐based confidence model for rescoring. The method preserves conserved interaction fragments derived from homologous complexes, followed by local resampling using physics‐based sampling and a diffusion model. Our template‐based strategy achieved a mean lDDT‐PLI of 0.69 across 233 targets, which was highly competitive. These results demonstrate that combining physics‐based modeling with AI‐driven refinement can significantly enhance the accuracy of both protein–protein and protein–ligand structure predictions.

Ashizawa, Ryota [Department of Applied Mathematics

Multiple Changepoint Detection for Non‐Gaussian Time Series

ABSTRACT This article combines methods from existing techniques to identify multiple changepoints in non‐Gaussian autocorrelated time series. A transformation is used to convert a Gaussian series into a non‐Gaussian series, enabling penalized likelihood methods to handle non‐Gaussian scenarios. When the marginal distribution of the data is continuous, the methods essentially reduce to the change of variables formula for probability densities. When the marginal distribution is count‐oriented, Hermite expansions and particle filtering techniques are used to quantify the scenario. Simulations demonstrating the efficacy of the methods are given and two data sets are analyzed: 1) the proportion of home runs hit by Major League Baseball batters from 1920 to 2023 and 2) a six‐dimensional series of tropical cyclone counts from the Earth's basins of generation from 1980 to 2023. In the first series, beta marginal distributions are used to describe the proportions; in the second, Poisson marginal distributions seem appropriate.

Lund, Robert [Department of Statistics University

Image Distinguishability Analysis Testing Through Principal Components and Its Application to Hot Spot Scale Invariance

Hot spots are spatial regions of intense energy localization that govern initiation of secondary high explosives. Studies that characterize or compare simulated hot spots are frequently either qualitatively descriptive or resort to quantitative distribution functions that neglect stochastic variations and spatial correlations—effects that are also neglected in common comparison tests like the Kolmogorov–Smirnov test. To this end, we develop an image distinguishability analysis (IDA) test based on principal component (PC) analysis that makes pixel-by-pixel comparisons between small, for example, O(<10), image data sets. The IDA test makes comparisons through a generalized distance metric in the PC space and a test statistic that is derived to calculate mathematical equation-values. Here, we derive a statistical distribution and criticality criterion to determine whether images are distinguishable from established baselines. We apply the IDA test on images generated from molecular dynamics simulations of hot spots from pore collapse in TATB to assess scale invariance in the complex patterns of hot spots that form in a representative high explosive crystal. The IDA test shows that TATB hot spot spatial temperature fields and their derived temperature histograms exhibit scale-invariant features over specific intervals of shock orientation, strength, and initial pore diameter. However, the IDA test also shows that qualitatively different conclusions regarding invariance can be reached depending on whether the hot spot is treated as a spatially correlated field as opposed to a distribution function that lacks spatial information.

organic