Search NASASearch

SEARCH · Search NASA

Results for “Statistical methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

The Information Length Concept Applied to Plasma Turbulence

A methodology to study statistical properties of anomalous transport in fusion plasma is investigated. Three time traces generated by the full-f gyrokinetic code GKNET are analyzed for this purpose. The time traces consist of heat flux as a function of the radial position, which is studied in a novel manner using statistical methods. The simulation data exhibit transport processes with both medium and long correlation length along the radius. A typical example of a phenomenon with long correlation length is avalanches. In order to investigate the evolution of the turbulent state, two basic configurations are studied, one flux-driven and one gradient-driven with decaying turbulence. The information length concept in tandem with Boltzmann–Gibbs and Tsallis entropy is used in the investigation. It is found that the dynamical states in both flux-driven and gradient-driven cases are surprisingly similar, but the Tsallis entropy reveals differences between them. This indicates that the types of probability distribution function are nevertheless quite different since the higher moments are significantly different.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

An R shiny graphical user interface for highprecision mass spectrometric data analysis

• There is currently a lack of software that meets the needs for the analysis of raw data produced by modern isotope ratio mass spectrometers for both R&D and routine use at SRNL and other US national labs • Needs to accommodate multiple isotope systems, instruments, and manufacturers • Include modern statistical methods and handling/visualization of uncertainty • Flexible software with transparent (no “black box”) and reproducible methods • This project is inspired by existing discipline-specific data analysis software (e.g., Tripoli1 , ET_Redux2, IsoplotR3) used in the geochemical community • Our goal is to build an open source data analysis software package that focuses on flexibility, transparency, and reproducibility

LABONE, ELIZABETH

Contrasting Time-Frequency Representations for Unknown Waveform Detection

Identifying unseen electromagnetic waveforms is critical for many applications, like interference management, electronic warfare and spectrum management. Traditionally this is done using statistical methods for anomaly detection, which has evolved to deep learning models for identifying the unseen data, formally termed as open set recognition. Some prior methods use a generative model to emulate open set data, which face challenges in generating synthetic samples for open set while simultaneously selecting an optimal discriminator for accurate classification. To alleviate this issue, we propose a discriminative model that effectively combines time and frequency domain features of communication signals for accurate predictions. We further introduce a cosine similarity loss that makes the domain specific features unique to enhance the prediction rate. Additionally, our model avoids generic feature vectors by extracting class-specific features during training, resulting in improved class representation. The experiment results show that this combined feature approach with cosine loss outperforms single-domain models and improves accuracy by 10% over models without cosine loss.

99 - GENERAL AND MISCELLANEOUS

Intern Poster

Large Language Models (LLMs) have skyrocketed in popularity after the release of ChatGPT in late 2022. Although LLMs are powerful tools, they can be subject to hallucinations, which is when an LLM (or any AI model) produces misleading/ nonsensical information. The objective is to determine if statistical methods can be used to detect hallucinations as an LLM generates its answer token by token (essentially word by word).

97 - MATHEMATICS AND COMPUTING

Classification of events from α -induced reactions in the MUSIC detector via statistical and ML methods

The Multi-Sampling Ionization Chamber (MUSIC) detector is typically used to measure nuclear reaction cross sections relevant for nuclear astrophysics, fusion studies, and other applications. From the MUSIC data produced in one experiment scientists carefully extract an order of 10 3 events of interest from about 10 9 total events, where each event can be represented by an 18-dimensional vector. However, the standard data classification process is based on expert driven, manually intensive data analysis techniques that require several months to identify patterns and classify the relevant events from the collected data. Here, to address this issue, we present a method for the classification of events originating from specific α-induced reactions by combining statistical and machine learning methods that require significantly less input from the domain scientist, relative to the standard technique. Here, we applied the new method to two experimental data sets and compared our results with those obtained using traditional methods. With few exceptions, the number of events classified by our method agrees within ±20% with the results obtained using traditional methods. With the present method, which is the first of its kind for the MUSIC data, we have established the foundation for the automated extraction of physical events of interest from experiments using the MUSIC detector.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Bayesian model mixing with multireference energy density functional

Reliably predicting nuclear properties across the entire chart of isotopes is important for applications ranging from nuclear astrophysics to superheavy science to nuclear technology. To this day, however, all the theoretical models that can scale at the level of the chart of isotopes remain semiphenomenological. Because they are fitted locally, their predictive power can vary significantly; different versions of the same theory provide different predictions. Bayesian model mixing takes advantage of such imperfect models to build a local mixture of a set of models to make improved predictions. Earlier attempts to use Bayesian model mixing for mass table calculations relied on models treated at single-reference energy density functional level, which fail to capture some of the correlations caused by configuration mixing or the restoration of broken symmetries. In this study we have applied Bayesian model mixing techniques within a multireference energy density functional (MR-EDF) framework. We considered predictions of two-particle separation energies from particle number projection or angular momentum projection with four different energy density functionals—a total of eight different MR-EDF models. We used a hierarchical Bayesian stacking framework with a Dirichlet prior distribution over weights together with an inverse log-ratio transform to enable positive correlations between different models. We found that Bayesian model mixing provides significantly improved predictions compared to the participating models. Published by the American Physical Society 2025

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Bayesian event categorization matrix approach for explosion monitoring

Current efforts to correctly categorize natural events from suspected explosion sources with data that is collected by ground- or space-based sensors presents historical challenges that remain unaddressed by the Event Categorization Matrix (ECM) model. Smaller historical events (lower yield explosions) may have data available from fewer measurement techniques than are available today, and therefore, a historical event record can lack a complete set of discriminants. The covariance structures can also differ between such observations of event (source-type) categories. Both obstacles are problematic for the classic ECM model. Our work addresses this gap and presents a Bayesian update to the previous ECM model, termed the Bayesian Event Categorization Matrix model, which can be trained on partial observations and does not rely on a pooled covariance structure. We further augment the ECM model with Bayesian Decision Theory so that false negative or false positive rates of an event categorization can be reduced in an intuitive manner. To demonstrate improved categorization rates for the Bayesian Event Categorization Matrix model, we compare an array of Bayesian and classic models with multiple performance metrics using Monte Carlo experiments. We use both synthetic and real data. Our Bayesian models show consistent gains in overall accuracy and lower false negative rates relative to the classic ECM model. Here, we propose future avenues to improve Bayesian Event Categorization Matrix models’ decision making and predictive capability.

58 GEOSCIENCES

Greedy emulators for nuclear two-body scattering

Applications of reduced basis method emulators are increasing in low-energy nuclear physics because they enable fast and accurate sampling of high-fidelity calculations, enabling robust uncertainty quantification. Here, in this paper, we develop, implement, and test two model-driven emulators based on the (Petrov-)Galerkin projection using the prototypical test case of two-body scattering with the Minnesota potential and a more realistic local chiral potential. The high-fidelity scattering equations are solved with the matrix Numerov method, a reformulation of the popular Numerov recurrence relation for solving special second-order differential equations as a linear system of coupled equations. A novel error estimator based on reduced-space residuals is applied to an active learning approach (a greedy algorithm) to choosing training samples (“snapshots”) for the emulator and contrasted with a proper orthogonal decomposition (POD) approach. Both approaches allow for computationally efficient offline-online decompositions, but the greedy approach requires many fewer snapshot calculations. These developments set the groundwork for emulating scattering observables based on chiral nucleon-nucleon and three-nucleon interactions and optical models, where computational speed-ups are necessary for Bayesian uncertainty quantification. Our emulators and error estimators are widely applicable to linear systems.

Bayesian methods

Active learning emulators for nuclear two-body scattering in momentum space

In this work we extend the active learning emulators for two-body scattering in coordinate space with error estimation, recently developed by Maldonado et al. [Phys. Rev. C 112, 024002], to coupled-channel scattering in momentum space. Our full-order model (FOM) solver is based on the Lippmann-Schwinger integral equation for the scattering t-matrix as opposed to the radial Schrödinger equation. We use (Petrov-)Galerkin projections and high-fidelity calculations at a few snapshots across the parameter space of the interaction to construct efficient reduced-order models (ROMs), trained by a greedy algorithm for locally optimal snapshot selection. Both the FOM solver and the corresponding ROMs are implemented efficiently in Python using Google's JAX library. We present results for emulating scattering phase shifts in coupled and uncoupled channels and cross sections, and assess the accuracy of the developed ROMs and their computational speedup factors. We also develop emulator error estimation for both the t-matrix and the total cross section. The software framework for reproducing and extending our results is publicly available. Together with our recent advances in developing active-learning emulators for three-body scattering, these emulator frameworks set the stage for full Bayesian calibrations of chiral nuclear interactions and optical models against scattering data with quantified emulator errors.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

An implementation of neural simulation-based inference for parameter estimation in ATLAS

Neural simulation-based inference (NSBI) is a powerful class of machine-learning-based methods for statistical inference that naturally handles high-dimensional parameter estimation without the need to bin data into low-dimensional summary histograms. Such methods are promising for a range of measurements, including at the Large Hadron Collider, where no single observable may be optimal to scan over the entire theoretical phase space under consideration, or where binning data into histograms could result in a loss of sensitivity. This work develops a NSBI framework for statistical inference, using neural networks to estimate probability density ratios, which enables the application to a full-scale analysis. It incorporates a large number of systematic uncertainties, quantifies the uncertainty due to the finite number of events in training samples, develops a method to construct confidence intervals, and demonstrates a series of intermediate diagnostic checks that can be performed to validate the robustness of the method. As an example, the power and feasibility of the method are assessed on simulated data for a simplified version of an off-shell Higgs boson couplings measurement in the four-lepton final states. This approach represents an extension to the standard statistical methodology used by the experiments at the Large Hadron Collider, and can benefit many physics analyses.

frequentist statistics

Out-of-Distribution Detection and Radiological Data Monitoring Using Statistical Process Control

Abstract Machine learning (ML) models often fail with data that deviates from their training distribution. This is a significant concern for ML-enabled devices as data drift may lead to unexpected performance. This work introduces a new framework for out of distribution (OOD) detection and data drift monitoring that combines ML and geometric methods with statistical process control (SPC). We investigated different design choices, including methods for extracting feature representations and drift quantification for OOD detection in individual images and as an approach for input data monitoring. We evaluated the framework for both identifying OOD images and demonstrating the ability to detect shifts in data streams over time. We demonstrated a proof-of-concept via the following tasks: 1) differentiating axial vs. non-axial CT images, 2) differentiating CXR vs. other radiographic imaging modalities, and 3) differentiating adult CXR vs. pediatric CXR. For the identification of individual OOD images, our framework achieved high sensitivity in detecting OOD inputs: 0.980 in CT, 0.984 in CXR, and 0.854 in pediatric CXR. Our framework is also adept at monitoring data streams and identifying the time a drift occurred. In our simulations tracking drift over time, it effectively detected a shift from CXR to non-CXR instantly, a transition from axial to non-axial CT within few days, and a drift from adult to pediatric CXRs within a day—all while maintaining a low false positive rate. Through additional experiments, we demonstrate the framework is modality-agnostic and independent from the underlying model structure, making it highly customizable for specific applications and broadly applicable across different imaging modalities and deployed ML models.

Zamzmi, Ghada

Joint inference of multiplicative and additive systematics in galaxy density fluctuations and clustering measurements

Galaxy clustering measurements are a key probe of the matter density field in the Universe. With the era of precision cosmology upon us, surveys rely on precise measurements of the clustering signal for meaningful cosmological analysis. However, the presence of systematic contaminants can bias the observed galaxy number density, and thereby bias the galaxy two-point statistics. As the statistical uncertainties get smaller, correcting for these systematic contaminants becomes increasingly important for unbiased cosmological analysis. We present and validate a new method for understanding and mitigating both additive and multiplicative systematics in galaxy clustering measurements (two-point function) by joint inference of contaminants in the galaxy overdensity field (one-point function) using a maximum-likelihood estimator (MLE). We test this methodology with Kilo-Degree Survey-like mock galaxy catalogues and synthetic systematic template maps. We estimate the cosmological impact of such mitigation by quantifying uncertainties and possible biases in the inferred relationship between the observed and the true galaxy clustering signal. Our method robustly corrects the clustering signal to the sub-percent level and reduces numerous additive and multiplicative systematics from 1.5σ to less than 0.1σ for the scenarios we tested. In addition, we provide an empirical approach to identifying the functional form (additive, multiplicative, or other) by which specific systematics contaminate the galaxy number density. Even though this approach is tested and geared towards systematics contaminating the galaxy number density, the methods can be extended to systematics mitigation for other two-point correlation measurements.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Effect of threshold parameters on infrared segmentation methods for porosity detection in electron beam powder bed fusion

In-situ process monitoring has seen significant interest in additive manufacturing to address qualification and certification goals. This is especially prevalent in metal powder bed fusion processes such as electron beam powder bed fusion (PBF-EB), with layer-wise infrared imaging being commonly used to detect defects. Here, this work compares two different segmentation methods (static thresholding and statistical thresholding) used for detecting porosity from in-situ infrared imaging data for PBF-EB. Samples were manufactured at a variety of focus offset values to induce porosity. Then, the segmented infrared images were compared to ex-situ X-ray computed tomography scans, which served as a ground-truth reference for objective evaluation. Through this analysis framework, the influential parameters, static threshold and N-value (number of standard deviations above the mean pixel value), respectively, for both image segmentation methods were analyzed and compared for their effects on porosity detection. With optimal parameter settings, the two methods had similar porosity detection performance, but the statistical method performed better under a larger variety of parameter settings.

Infrared imaging

The development and application of the stirred‐reactor coupon analysis (SRCA) test method

A new technique, termed the stirred‐reactor coupon analysis (SRCA) method, has been developed to measure the rate of glass dissolution in forward‐rate conditions. Monolithic glass coupons are partially masked with an inert material before placement in a large volume of well‐mixed solution with known chemistry and temperature for a predetermined duration. After the test, the mask is removed, and the difference in step height between the protected area and the exposed corroded portions of the sample coupon is measured to determine the extent of glass dissolution. The step height is converted to a rate measurement using the test duration and glass density. Test parameters such as sample surface preparation and test duration were evaluated to determine their effects on the measured rates. Additionally, results from an interlaboratory study (ILS) consisting of 12 laboratories from 11 different institutions are presented, where each laboratory performed 12 independent tests. When removing experimental outlier data, the 95% reproducibility limits for the SRCA method has no statistical difference with previously published standardized test methods used to determine the forward rate of glass dissolution. Overall, this paper describes steps necessary to perform the test method and provides the statistical calculations to evaluate test accuracy.

chemical durability

Evaluating downscaled products with expected hydroclimatic co-variances

Abstract. There has been widespread adoption of downscaled products amongst practitioners and stakeholders to ascertain risk from climate hazards at the local scale (e.g., ∼ 5 km resolution). Such products must nevertheless be consistent with physical laws to be credible and of value to users. Here we evaluate statistically and dynamically downscaled products by examining local co-evolution of downscaled temperature and precipitation during convective and frontal precipitation events (two mechanisms testable with just temperature and precipitation). We find that two widely used statistical downscaling techniques (Localized Constructed Analogs version 2, LOCA2, and Seasonal Trends and Analysis of Residuals Empirical Statistical Downscaling Model, STAR-ESDM) generally preserve expected co-variances during convective precipitation events over the historical and future projected intervals as compared to European Centre for Medium-Range Weather Forecasts Reanalysis v5 (ERA5) and two observation-based data products (Livneh and nClimGrid-Daily). However, both techniques dampen future intensification of frontal precipitation that is otherwise robustly captured in global climate models (i.e., prior to downscaling) and with process-based dynamical downscaling across five different regional climate models. In the case of LOCA2, this leads to appreciable underestimation of future frontal precipitation event intensity. This study is one of the first to quantify a likely ramification of the stationarity assumption underlying statistical downscaling methods and identify a phenomenon where projections of future change diverge depending on data production method employed. Finally, our work proposes expected co-variances during convective and frontal precipitation as useful evaluation diagnostics that can be universally applied to a wide range of statistically downscaled products.

54 ENVIRONMENTAL SCIENCES

Nonparametric Multiparticle Set Methods for Interpreting Environmental Samples

Collection and analysis of environmental samples is commonly used by a range of stakeholders in nuclear safeguards and security contexts. While the ubiquity of samples and their transport in the environment allow regular collection, developing and demonstrating methods for analyzing these samples is difficult. In this work, an environmental sample consists of a set of one or more individual particles. Recent advances in reactor simulation have allowed us to generate data that are more representative of real-world environmental samples, enabling statistically defensible method development and testing. The most notable of these advances is a drastic increase in the number of material depletion regions, which allows our simulations to capture the variation in isotopic composition seen at length scales consistent with environmental samples. Traditional approaches for handling multiparticle samples treat each particle in the sample individually, estimating the quantity of interest (e.g., core-average burnup) resulting from measurement and analysis of signatures (e.g., nuclide assays) from each individual particle. Individual estimates are then averaged to generate a single estimate of the quantity of interest over the entire sample. In this presentation, we introduce two novel approaches for interpreting environmental samples that comprise of multiple particles: (1) the Quantile-Quantile Comparator, which uses a multivariate generalization of quantile-quantile plots for comparing unknown statistical distributions, and (2) the Set Transformer, an attention-based neural network module designed to model interactions among elements (particles) in the input set (sample). Statistically representative sampling cannot be guaranteed as samples are passively collected and are beholden to what particles are available in the environment. These new analysis methods for set-input problems are expected to be more robust than traditional approaches to issues of sampling bias where particles are not uniformly distributed throughout regions of interest, as well as generally outperform traditional approaches by jointly considering all elements in the set. We will present results comparing the performance of traditional single particle approaches and the novel Quantile-Quantile Comparator and Set Transformer for interpretation of simulated environmental samples.

Phathanapirom, Birdy

Development of a 95-Year Solar Dataset for Resource Adequacy Studies

Long-term high-resolution solar data provides enhanced understanding of variability of solar generation and enhances our ability to develop strategies for a resilient and reliable electric grid under high deployment of solar energy. Therefore, it is important to develop long-term synthetic datasets that can provide multiple occurrences of various severe weather scenarios that are expected to test the limits of resource adequacy under scenarios contain various energy generation sources. Examples of such scenarios could be long periods of high temperatures when demand for electricity is high or periods where high winds could lead to a shut-down of transmission lines for long periods of time to ensure fire safety. NREL has developed the first version of such a dataset covering a 95-year period covering 2006-2100 at a 4km hourly resolution. This dataset contains all variables necessary to calculate solar generation. During development of this dataset, we focused on creating unbiased, high-resolution solar irradiance through statistical downscaling methods, using Regional Climate Model (RCM) simulations from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX) as input. The National Solar Radiation Database (NSRDB) containing over 25 years of observations was used to calibrate the statistical downscaling models. This presentation will outline the primary steps in developing this dataset, including (1) regridding RCM data to a common grid at 20-km resolution, (2) correcting RCM biases with NSRDB, (3) applying temporal and spatial downscaling methods to generate high-resolution (4-km, hourly) solar and ancillary data. Additionally, we will present an evaluation of the downscaled data against the NSRDB across various zones in the CONUS. Lastly, we will present a user guide for accessing the datasets.

14 SOLAR ENERGY

From chiral effective field theory to perturbative QCD: A Bayesian model mixing approach to symmetric nuclear matter

Constraining the equation of state (EOS) of strongly interacting, dense matter is the focus of intense experimental, observational, and theoretical effort. Chiral effective field theory (𝜒⁢EFT ) can describe the EOS between the typical densities of nuclei and those in the outer cores of neutron stars, while perturbative QCD (pQCD) can be applied to properties of deconfined quark matter, both with quantified theoretical uncertainties. However, describing the full range of densities in between with a single EOS that has well-quantified uncertainties is a challenging problem. Bayesian multimodel inference from 𝜒⁢EFT and pQCD can help bridge the gap between the two theories. In this work, we introduce a correlated Bayesian model mixing framework that uses a Gaussian process (GP) to assimilate different information into a single QCD EOS for symmetric nuclear matter. The present implementation uses a stationary GP to infer this mixed EOS solely from the EOSs of 𝜒⁢EFT and pQCD while accounting for the truncation errors of each theory. The GP is trained on the pressure as a function of number density in the low- and high-density regions where 𝜒⁢EFT and pQCD are, respectively, valid. We impose priors on the GP kernel hyperparameters to suppress unphysical correlations between these regimes. This, together with the assumption of stationarity, results in smooth 𝜒⁢EFT-to-pQCD curves for both the pressure and the speed of sound. We show that using uncorrelated mixing requires uncontrolled extrapolation of at least one of 𝜒⁢EFT or pQCD into regions where the perturbative series breaks down and leads to an acausal EOS. Here, we also discuss extensions of this framework to nonstationary and less differentiable GP kernels, its future application to neutron-star matter, and the incorporation of additional constraints from nuclear theory, experiment, and multimessenger astronomy.

Bayesian methods