Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical techniques”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Improvement and generalization of ABCD method with Bayesian inference

To find New Physics or to refine our knowledge of the Standard Model at the LHC is an enterprise that involves many factors, such as the capabilities and the performance of the accelerator and detectors, the use and exploitation of the available information, the design of search strategies and observables, as well as the proposal of new models. We focus on the use of the information and pour our effort in re-thinking the usual data-driven ABCD method to improve it and to generalize it using Bayesian Machine Learning techniques and tools. We propose that a dataset consisting of a signal and many backgrounds is well described through a mixture model. Signal, backgrounds and their relative fractions in the sample can be well extracted by exploiting the prior knowledge and the dependence between the different observables at the event-by-event level with Bayesian tools. We show how, in contrast to the ABCD method, one can take advantage of understanding some properties of the different backgrounds and of having more than two independent observables to measure in each event. In addition, instead of regions defined through hard cuts, the Bayesian framework uses the information of continuous distribution to obtain soft-assignments of the events which are statistically more robust. To compare both methods we use a toy problem inspired by pp\to hh\to b\bar b b \bar b p p → h h → b b ‾ b b ‾ , selecting a reduced and simplified number of processes and analysing the flavor of the four jets and the invariant mass of the jet-pairs, modeled with simplified distributions. Taking advantage of all this information, and starting from a combination of biased and agnostic priors, leads us to a very good posterior once we use the Bayesian framework to exploit the data and the mutual information of the observables at the event-by-event level. We show how, in this simplified model, the Bayesian framework outperforms the ABCD method sensitivity in obtaining the signal fraction in scenarios with 1% and 0.5% true signal fractions in the dataset. We also show that the method is robust against the absence of signal. We discuss potential prospects for taking this Bayesian data-driven paradigm into more realistic scenarios.

Alvarez, Ezequiel↗

Imaging and structure analysis of ferroelectric domains, domain walls, and vortices by scanning electron diffraction

Direct electron detectors in scanning transmission electron microscopy give unprecedented possibilities for structure analysis at the nanoscale. In electronic and quantum materials, this new capability gives access to, for example, emergent chiral structures and symmetry-breaking distortions that underpin functional properties. Quantifying nanoscale structural features with statistical significance, however, is complicated by the subtleties of dynamic diffraction and coexisting contrast mechanisms, which often results in a low signal-to-noise ratio and the superposition of multiple signals that are challenging to deconvolute. Here we apply scanning electron diffraction to explore local polar distortions in the uniaxial ferroelectric Er(Mn,Ti)O 3 . Using a custom-designed convolutional autoencoder with bespoke regularization, we demonstrate that subtle variations in the scattering signatures of ferroelectric domains, domain walls, and vortex textures can readily be disentangled with statistical significance and separated from extrinsic contributions due to, e.g., variations in specimen thickness or bending. The work demonstrates a pathway to quantitatively measure symmetry-breaking distortions across large areas, mapping structural changes at interfaces and topological structures with nanoscale spatial resolution.

36 MATERIALS SCIENCE↗

Universal energy-speed-accuracy trade-offs in driven nonequilibrium systems

The connection between measure theoretic optimal transport and dissipative nonequilibrium dynamics provides a language for quantifying nonequilibrium control costs, leading to a collection of thermodynamic speed limits, which rely on the assumption that the target probability distribution is perfectly realized. This is almost never the case in experiments or numerical simulations, so here we address the situation in which the external controller is imperfect. We obtain a lower bound for the dissipated work in generic nonequilibrium control problems that (1) is asymptotically tight and (2) matches the thermodynamic speed limit in the case of optimal driving. Along with analytically solvable examples, we refine this imperfect driving notion to systems in which the controlled degrees of freedom are slow relative to the nonequilibrium relaxation rate, and identify independent energy contributions from fast and slow degrees of freedom. Furthermore, we develop a strategy for optimizing minimally dissipative protocols based on optimal transport flow matching, a generative machine learning technique. Furthermore, this latter approach ensures the scalability of both the theoretical and computational framework we put forth. Crucially, we demonstrate that we can compute the terms in our bound numerically using efficient algorithms from the computational optimal transport literature and that the protocols we learn saturate the bound.

59 BASIC BIOLOGICAL SCIENCES↗

Agentic AI vs ML-Based Autotuning: A Comparative Study for Loop Reordering Optimization

High Performance Computing (HPC) applications rely heavily on code optimizations to achieve good performance on modern CPU and GPU architectures. Traditional Machine Learning auto-tuning approaches have demonstrated success in exploring high-dimensional spaces, but they often require expensive compile-run evaluations and lack adaptability for large HPC applications. The recent advances in Large Language Models (LLMs) and Agentic AI systems raise intriguing questions about the potential of these approaches to address specific optimization methodologies. This work aims to answer an essential question for the HPC community: “How Agentic AI Systems Compare to Traditional ML Autotuning Techniques?” To address this question, we present a comparative analysis between a traditional ML-based optimization approach and an Agentic AI system, evaluating their respective capabilities and limitations for loop-level optimization. In addition, we introduced a new Agentic AI system named LoopGen-AI using three different Large Language Models: GPT-4.1, Claude 4.0, and Gemini 2.5. A key finding is that LoopGen-AI achieves competitive per-formance with only a few program runs, the reasoning logs from the agents revealed that their decisions rely heavily on the combination of semantic understanding of the target kernel with dynamic feedback from the environment, highlighting a promising new dimension in performance tuning. In contrast, ML-based autotuners focus on statistical exploration, and require orders of magnitude more runs to reach peak performance. Additionally, our analysis shows that prompt engineering, particularly using Persona + Context Manager patterns, significantly impacts the effectiveness of Agentic AI. Our results indicate that while Agentic AI systems are not yet a complete replacement for ML-based autotuners, it can effectively complement traditional methods.

Rosas, Miguel Romero↗

Critical impact of experimentally-driven strut level anisotropic material models in advanced stress analysis of additively manufactured lattice structures

The rapid acceleration in materials discovery may overshadow the importance of thoroughly understanding the mechanical performance of newly developed materials in demanding environments. The recent interest in combining parametric studies with machine learning techniques to explore how changes in specific processing parameters or model inputs affect the overall behavior of a material system can only be truly beneficial if the governing constitutive relations describing material behavior are accurately established. In this study, we demonstrate the critical impact of accurately representing strut-level anisotropic material behavior in advanced stress analysis of additively manufactured lattice structures (AMLS). We introduce a systematic experimental and modeling approach for developing strut-level anisotropic elastoplastic material models that account for the influence of microstructural features such as porosity, texture, and surface roughness on the development of local anisotropic mechanical properties, which vary with strut orientation relative to the build direction (BD). As a result the presented material model captures and relates the statistics of spatially varying struts’ microstructural features to the local stress distribution. Our findings suggest that incorporating strut-level anisotropic material behavior into unit cell analysis significantly influences the load distribution and evolution of local stresses within the structure. Therefore, accounting for this anisotropy is critical for developing an understanding of unit cell behavior and performance, including subsequent topology/component design optimization based on this analysis.

Sahoo, Subhadip [University of Arizona]↗

Real-Time event reconstruction for Nuclear Physics Experiments using Artificial Intelligence

Charged track reconstruction is a critical task in nuclear physics experiments, enabling the identification and analysis of particles produced in high-energy collisions. Machine learning (ML) has emerged as a powerful tool for this purpose, addressing the challenges posed by complex detector geometries, high event multiplicities, and noisy data. Traditional methods rely on pattern recognition algorithms like the Kalman filter, but ML techniques, such as neural networks, graph neural networks (GNNs), and recurrent neural networks (RNNs), offer improved accuracy and scalability. By learning from simulated and real detector data, ML models can identify and classify tracks, predict trajectories, and handle ambiguities caused by overlapping or missing hits. Moreover, ML-based approaches can process data in near-real-time, enhancing the efficiency of experiments at large-scale facilities like the Large Hadron Collider (LHC) and Jefferson Lab (JLAB). As detector technologies and computational resources evolve, ML-driven charged track reconstruction continues to push the boundaries of precision and discovery in nuclear physics. In these proceedings, we highlight advancements in charged track identification leveraging Artificial Intelligence within the CLAS12 detector, achieving a notable enhancement in experimental statistics compared to traditional methods. Additionally, we showcase real-time event reconstruction capabilities, including the inference of charged particle properties, such as momentum, direction, and species identification, at speeds matching data acquisition rates. These innovations enable the extraction of physics observables directly from the experiment in real-time.

Gavalian, Gagik (ORCID:0000000267385457)↗

Multiple Changepoint Detection for Non‐Gaussian Time Series

ABSTRACT This article combines methods from existing techniques to identify multiple changepoints in non‐Gaussian autocorrelated time series. A transformation is used to convert a Gaussian series into a non‐Gaussian series, enabling penalized likelihood methods to handle non‐Gaussian scenarios. When the marginal distribution of the data is continuous, the methods essentially reduce to the change of variables formula for probability densities. When the marginal distribution is count‐oriented, Hermite expansions and particle filtering techniques are used to quantify the scenario. Simulations demonstrating the efficacy of the methods are given and two data sets are analyzed: 1) the proportion of home runs hit by Major League Baseball batters from 1920 to 2023 and 2) a six‐dimensional series of tropical cyclone counts from the Earth's basins of generation from 1980 to 2023. In the first series, beta marginal distributions are used to describe the proportions; in the second, Poisson marginal distributions seem appropriate.

Lund, Robert [Department of Statistics University ↗

Benchmarking large language models for materials synthesis: The case of atomic layer deposition

In this work, we introduce an open-ended question benchmark, ALDbench, to evaluate the performance of large language models (LLMs) in materials synthesis, and, in particular, in the field of atomic layer deposition, a thin film growth technique used in energy applications and microelectronics. Our benchmark comprises questions with a level of difficulty ranging from the graduate level to domain expert current with the state of the art in the field. Human experts reviewed the questions along the criteria of difficulty and specificity, and the model responses along four different criteria: overall quality, specificity, relevance, and accuracy. We ran this benchmark on an instance of OpenAI’s GPT-4o. The responses from the model received a composite quality score of 3.7 on a 1–5 scale, consistent with a passing grade. However, 36% of the questions received at least one below average score. An in-depth analysis of the responses identified at least five instances of suspected hallucination. Finally, we observed statistically significant correlations between the difficulty of the question and the quality of the response, the difficulty of the question and the relevance of the response, the specificity of the question, and the accuracy of the response as graded by the human experts. Furthermore, this emphasizes the need to evaluate LLMs across multiple criteria beyond difficulty or accuracy.

Artificial intelligence↗

Tripling the Census of Dwarf AGN Candidates Using DESI Early Data

Using early data from the Dark Energy Spectroscopic Instrument (DESI) survey, we search for active galactic nuclei (AGN) signatures in 410,757 line-emitting galaxies. By employing the BPT emission-line ratio diagnostic diagram, we identify AGNs in 75,928/296,261 (≈25.6%) high-mass ($\mathrm{log}({M}_{\star }/{M}_{\odot })\,\gt$ 9.5) and 2444/114,496 (≈2.1%) dwarf ($\mathrm{log}({M}_{\star }/{M}_{\odot })\,\leqslant$ 9.5) galaxies. Of these AGN candidates, 4181 sources exhibit a broad Hα component, allowing us to estimate their BH masses via virial techniques. This study more than triples the census of dwarf AGNs and doubles the number of intermediate-mass black hole (M BH ≤ 10 6 M ⊙ ) candidates, spanning a broad discovery space in stellar mass (7 $\lt \mathrm{log}({M}_{\star }/{M}_{\odot })\,\lt$ 12) and redshift (0.001 < z < 0.45). The observed AGN fraction in dwarf galaxies (≈2.1%) is nearly four times higher than prior estimates, primarily due to DESI’s smaller fiber size, which enables the detection of lower-luminosity dwarf AGN candidates. We also extend the M BH –M⋆ scaling relation down to ${\rm{log}}({M}_{\star }/{M}_{\odot })\,\approx$ 8.5 and $\mathrm{log}({M}_{\mathrm{BH}}/{M}_{\odot })\,\approx$ 4.4, with our results aligning well with previous low-redshift studies. The large statistical sample of dwarf AGN candidates from current and future DESI releases will be invaluable for enhancing our understanding of galaxy evolution at the low-mass end of the galaxy mass function.

79 ASTRONOMY AND ASTROPHYSICS↗

Search for dijet resonances with data scouting in proton-proton collisions at $\sqrt{s}=13$ TeV

A search is presented for narrow resonances, with a mass between 0.6 and 1.8 TeV, decaying to pairs of jets, in proton-proton collisions at $\sqrt{s}=13$ TeV. The search is performed using dijets that are reconstructed, selected, and recorded in a compact form by the high-level trigger in a technique referred to as “data scouting”, from data collected in 2016–2018 corresponding to an integrated luminosity of 117 fb −1 . The dijet mass spectra are well described by a smooth parameterization, and no significant evidence for the production of new particles is observed. Model-independent upper limits are presented on the product of the cross section, branching fraction, and acceptance for the individual cases of narrow quark-quark, quark-gluon, and gluon-gluon resonances, and are compared to the predictions from a variety of models of narrow dijet resonance production. The upper limit on the coupling of a dark matter mediator to quarks is presented as a function of the mediator mass. The sensitivity of this search goes beyond what is expected from statistical scaling with the integrated luminosity alone, as a consequence of the use of fewer parameters in the background function within a more robust statistical procedure.

beyond Standard Model↗

Precision Computations in Strongly Coupled Conformal Field Theories (Final Technical Report)

Conformal Field Theories (CFTs) are quantum field theories that are invariant under the conformal symmetry group (which includes translations and rotations, but also local rescalings of spacetime). They are building blocks of general quantum field theories, and appear in many areas of physics, including statistical physics, condensed matter physics, particle physics, and quantum gravity. Because of their extra symmetries, the mathematical structure of CFTs is tightly constrained, and this leads to the idea of the ``conformal bootstrap," which is to use these mathematical structures to constrain, and in some cases determine, CFT observables. A new numerical implementation of the conformal bootstrap idea appeared in 2008 with the work of Rattazzi, Rychkov, Tonni, and Vichi. Their observation was that certain bootstrap constraints (conformal symmetry and unitarity) could be combined to yield a convex optimization problem that constraints CFT data. By solving this convex optimization problem on a computer, one could obtain bounds on observables like critical exponents and operator product expansion (OPE) coefficients. Over the course of this award, the PI has improved numerical bootstrap techniques by optimizing known algorithms and finding new ones for performing the required convex optimization computations. The PI has applied these techniques to compute high-precision observables in several important strongly-coupled systems. The PI has also explored both analytical and numerical bootstrap methods for constraining the space of low energy effective field theories of quantum gravity, and developed new analytical techniques for CFT and QFT more broadly.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will produce unprecedented volumes of heterogeneous astronomical data (images, catalogs, and alerts) that challenge traditional analysis pipelines. The LSST Dark Energy Science Collaboration (DESC) aims to derive robust constraints on dark energy and dark matter from these data, requiring methods that are statistically powerful, scalable, and operationally reliable. Artificial intelligence and machine learning (AI/ML) are already embedded across DESC science workflows, from photometric redshifts and transient classification to weak lensing inference and cosmological simulations. Yet their utility for precision cosmology hinges on trustworthy uncertainty quantification, robustness to covariate shift and model misspecification, and reproducible integration within scientific pipelines. This white paper surveys the current landscape of AI/ML across DESC's primary cosmological probes and cross-cutting analyses, revealing that the same core methodologies and fundamental challenges recur across disparate science cases. Since progress on these cross-cutting challenges would benefit multiple probes simultaneously, we identify key methodological research priorities, including Bayesian inference at scale, physics-informed methods, validation frameworks, and active learning for discovery. With an eye on emerging techniques, we also explore the potential of the latest foundation model methodologies and LLM-driven agentic AI systems to reshape DESC workflows, provided their deployment is coupled with rigorous evaluation and governance. Finally, we discuss critical software, computing, data infrastructure, and human capital requirements for the successful deployment of these new methodologies, and consider associated risks and opportunities for broader coordination with external actors.

Aubourg, Eric [APC, Paris] (ORCID:000000025592023X↗

Galaxy bispectrum in the spherical Fourier-Bessel basis

The bispectrum, the three-point correlation in Fourier space, is a crucial statistic for studying many effects targeted by the next-generation galaxy surveys, such as primordial non-Gaussianity (PNG) and general relativistic (GR) effects on large scales. In this work we develop a formalism for the bispectrum in the spherical Fourier-Bessel (SFB) basis—a natural basis for computing correlation functions on the curved sky, as it diagonalizes the Laplacian operator in spherical coordinates. Working in the SFB basis allows for line-of-sight effects such as redshift space distortions and GR to be accounted for exactly, i.e., without having to resort to perturbative expansions to go beyond the plane-parallel approximation. Only analytic results for the SFB bispectrum exist in the literature given the intensive computations needed. We numerically calculate the SFB bispectrum for the first time, enabled by a few techniques: We implement a template decomposition of the redshift-space kernel Z 2 into Legendre polynomials, and separately treat the PNG and velocity-divergence terms. We derive an identity to integrate a product of three spherical harmonics connected by a Dirac delta function as a simple sum and use it to investigate the limit of a homogeneous and isotropic Universe. Furthermore, we present a formalism for convolving the signal with separable window functions and use a toy spherically symmetric window to demonstrate the computation and give insights into the properties of the observed bispectrum signal. While our implementation remains computationally challenging, it is a step toward a feasible full extraction of information on large scales via a SFB bispectrum analysis.

79 ASTRONOMY AND ASTROPHYSICS↗

Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis

HPC datacenters offer a backbone to the modern digital society. Increasingly, they run Machine Learning (ML) jobs next to generic, compute-intensive workloads, supporting science, business, and other decision-making processes. However, understanding how ML jobs impact the operation of HPC datacenters, relative to generic jobs, remains desirable but understudied. In this work, we leverage long-term operational data, collected from a national-scale production HPC datacenter, and statistically compare how ML and generic jobs can impact the performance, failures, resource utilization, and energy consumption of HPC datacenters. Our study provides key insights, e.g., ML-related power usage causes GPU nodes to run into temperature limitations, median/mean runtime and failure rates are higher for ML jobs than for generic jobs, both ML and generic jobs exhibit highly variable arrival processes and resource demands, significant amounts of energy are spent on unsuccessfully terminating jobs, and concurrent jobs tend to terminate in the same state. We open-source our cleaned-up data traces on Zenodo (https://doi. org/10.5281/zenodo.13685426), and provide our analysis toolkit as software hosted on GitHub (https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization). This study offers multiple benefits for data center administrators, who can improve operational efficiency, and for researchers, who can further improve system designs, scheduling techniques, etc.

crossanalysis↗

Identifying neutron sources using recoil and time-of-flight spectroscopy

Identification of neutron sources is central to nuclear physics and its applications, from planetary science to nuclear security, yet direct source discrimination from measured neutron spectra remains fundamentally elusive. Here, we introduce a Bayesian protocol that directly infers source ensembles from measured neutron spectra by combining full-spectrum template matching with probabilistic evidence evaluation. Applying this protocol to recoil and time-of-flight spectroscopy, we recover single- and two-source configurations with strong statistical significance (beyond 4⁢𝜎) at event counts as low as ∼10 3 . These results demonstrate that neutron spectral signatures can be leveraged for robust source identification, opening a new observational window for both fundamental research and operationally driven applications.

neutron physics↗

A Machine Learning Framework for Modeling Ensemble Properties of Atomically Disordered Materials

Atomic disorder can strongly influence material properties such as charge transport, optical response, and catalytic activity. However, efficiently modeling these disorder effects remains challenging for first-principles methods due to the cost of sampling large configurational spaces and computing complex physical quantities. Recent advances of machine learning techniques, particularly graph neural networks (GNNs), has enabled the efficient and accurate predictions of complex material properties, offering promising tools for studying disordered systems. In this work, we present a general machine-learning-assisted computational framework that integrates equivariant GNNs with Monte Carlo simulations to compute the thermodynamic and ensemble-averaged functional properties of disordered materials. Using the surface-termination-disordered MXene monolayer Ti 3 C 2 T 2–x as a representative system, we find that electrical conductivity exhibits an emergent peak near the order–disorder phase transition temperature due to the interplay between electron scattering and doping. In contrast, optical conductivity remains largely insensitive to local atomic disorder and reflects the global surface chemical composition. These results highlight the role of atomic disorder in affecting material properties and demonstrate the potential of our approach for statistically modeling disorder effects in a wide range of materials such as high-entropy alloys and spin liquids.

MXene↗

The cluster decomposition of the configurational energy of multicomponent alloys

Abstract The cluster expansion method (CEM) is a widely used lattice-based technique in the study of multicomponent alloys. Despite its prevalent use, a clear understanding of expansion terms is lacking. We present a modern mathematical formalism of the CEM and introduce thecluster decomposition—a unique and basis-independent decomposition for functions of the atomic configuration in a crystal. We identify the cluster decomposition as an invariant ANOVA decomposition; and demonstrate how functional analysis of variance and sensitivity analysis can be used to interpret interactions among species. Furthermore, we show how the mathematical structure of the cluster decomposition enables numerical evaluation that scales with the number of clusters and is independent of the number of species. Overall, our work enables rigorous interpretations of interactions among species, provides opportunities to explore parameter estimation beyond linear regression, introduces a numerical efficient implementation, and enables analysis of cluster expansions based on established mathematical and statistical principles.

Chemistry↗

Post-irradiation Examination of Eurofer-97 Steel Irradiated to 20 dpa at 200–400°C in HFIR under the EUROfusion (ORNL-KIT) Collaboration Program

The Oak Ridge National Laboratory/Karlsruhe Institute of Technology (ORNL/KIT) collaboration focuses on research involving irradiation experiments and post-irradiation examinations (PIE) of isotopically modified Eurofer-97 steels for fusion reactor applications. This collaboration capitalizes on ORNL’s expertise in radiation effects on materials and its capabilities in irradiation and PIE. The research aims to qualify Fe-9Cr-based Eurofer-97 steel under simulated fusion conditions, which include high doses and high transmutation rates. A unique isotopic modification technique is utilized in the research, wherein high-transmutation isotopes such as Fe-54 and Ni-58 are alloyed into the Eurofer steel to align with the expected helium production rates in fusion reactor conditions. This report presents the results of post-irradiation mechanical testing activities for the EUROFER-97/2 specimens after irradiation to ~20 dpa at various irradiation temperatures. The irradiation doses and temperatures for the ES capsules ranged from 18.4–20.4 dpa and 202–256°C, respectively. The mechanical property datasets, obtained through baseline testing and PIE of the tensile and fracture specimens, include microhardness data from 93 irradiated and non-irradiated tensile specimens and 28 irradiated bend bar specimens, uniaxial tensile property data from 45 irradiated and non-irradiated tensile specimens, and fracture toughness data from the bend bar specimens. Further data analyses provide statistical information on the microhardness values and tensile properties, as well as the reference ductile-brittle transition temperature (T 0 ) data.

36 MATERIALS SCIENCE↗