Search NASASearch

SEARCH · Search NASA

Results for “statistical methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Development of a 95-Year Solar Dataset for Resource Adequacy Studies

Long-term high-resolution solar data provides enhanced understanding of variability of solar generation and enhances our ability to develop strategies for a resilient and reliable electric grid under high deployment of solar energy. Therefore, it is important to develop long-term synthetic datasets that can provide multiple occurrences of various severe weather scenarios that are expected to test the limits of resource adequacy under scenarios contain various energy generation sources. Examples of such scenarios could be long periods of high temperatures when demand for electricity is high or periods where high winds could lead to a shut-down of transmission lines for long periods of time to ensure fire safety. NREL has developed the first version of such a dataset covering a 95-year period covering 2006-2100 at a 4km hourly resolution. This dataset contains all variables necessary to calculate solar generation. During development of this dataset, we focused on creating unbiased, high-resolution solar irradiance through statistical downscaling methods, using Regional Climate Model (RCM) simulations from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX) as input. The National Solar Radiation Database (NSRDB) containing over 25 years of observations was used to calibrate the statistical downscaling models. This presentation will outline the primary steps in developing this dataset, including (1) regridding RCM data to a common grid at 20-km resolution, (2) correcting RCM biases with NSRDB, (3) applying temporal and spatial downscaling methods to generate high-resolution (4-km, hourly) solar and ancillary data. Additionally, we will present an evaluation of the downscaled data against the NSRDB across various zones in the CONUS. Lastly, we will present a user guide for accessing the datasets.

14 SOLAR ENERGY

From chiral effective field theory to perturbative QCD: A Bayesian model mixing approach to symmetric nuclear matter

Constraining the equation of state (EOS) of strongly interacting, dense matter is the focus of intense experimental, observational, and theoretical effort. Chiral effective field theory (𝜒⁢EFT ) can describe the EOS between the typical densities of nuclei and those in the outer cores of neutron stars, while perturbative QCD (pQCD) can be applied to properties of deconfined quark matter, both with quantified theoretical uncertainties. However, describing the full range of densities in between with a single EOS that has well-quantified uncertainties is a challenging problem. Bayesian multimodel inference from 𝜒⁢EFT and pQCD can help bridge the gap between the two theories. In this work, we introduce a correlated Bayesian model mixing framework that uses a Gaussian process (GP) to assimilate different information into a single QCD EOS for symmetric nuclear matter. The present implementation uses a stationary GP to infer this mixed EOS solely from the EOSs of 𝜒⁢EFT and pQCD while accounting for the truncation errors of each theory. The GP is trained on the pressure as a function of number density in the low- and high-density regions where 𝜒⁢EFT and pQCD are, respectively, valid. We impose priors on the GP kernel hyperparameters to suppress unphysical correlations between these regimes. This, together with the assumption of stationarity, results in smooth 𝜒⁢EFT-to-pQCD curves for both the pressure and the speed of sound. We show that using uncorrelated mixing requires uncontrolled extrapolation of at least one of 𝜒⁢EFT or pQCD into regions where the perturbative series breaks down and leads to an acausal EOS. Here, we also discuss extensions of this framework to nonstationary and less differentiable GP kernels, its future application to neutron-star matter, and the incorporation of additional constraints from nuclear theory, experiment, and multimessenger astronomy.

Bayesian methods

Anomaly Detection and Approximate Similarity Searches of Transients in Real-time Data Streams

Abstract We present Lightcurve Anomaly Identification and Similarity Search ( LAISS ), an automated pipeline to detect anomalous astrophysical transients in real-time data streams. We deploy our anomaly detection model on the nightly Zwicky Transient Facility (ZTF) Alert Stream via the ANTARES broker, identifying a manageable ∼1–5 candidates per night for expert vetting and coordinating follow-up observations. Our method leverages statistical light-curve and contextual host galaxy features within a random forest classifier, tagging transients of rare classes ( spectroscopic anomalies), of uncommon host galaxy environments ( contextual anomalies), and of peculiar or interaction-powered phenomena ( behavioral anomalies). Moreover, we demonstrate the power of a low-latency (∼ms) approximate similarity search method to find transient analogs with similar light-curve evolution and host galaxy environments. We use analogs for data-driven discovery, characterization, (re)classification, and imputation in retrospective and real-time searches. To date, we have identified ∼50 previously known and previously missed rare transients from real-time and retrospective searches, including but not limited to superluminous supernovae (SLSNe), tidal disruption events, SNe IIn, SNe IIb, SNe I-CSM, SNe Ia-91bg-like, SNe Ib, SNe Ic, SNe Ic-BL, and M31 novae. Lastly, we report the discovery of 325 total transients, all observed between 2018 and 2021 and absent from public catalogs (∼1% of all ZTF Astronomical Transient reports to the Transient Name Server through 2021). These methods enable a systematic approach to finding the “needle in the haystack” in large-volume data streams. Because of its integration with the ANTARES broker, LAISS is built to detect exciting transients in Rubin data.

79 ASTRONOMY AND ASTROPHYSICS

A Framework Using Applied Process Analysis Methods to Assess Water Security in the Vu Gia–Thu Bon River Basin, Vietnam

The Vu Gia–Thu Bon (VG–TB) river basin is facing numerous challenges to water security, particularly in light of the increasing impacts of climate change. These challenges, including salinity intrusion, shifts in rainfall patterns, and reduced water supply in downstream areas, are of great concern. This study comprehensively assessed the current state of water security in the basin using robust statistical analysis methods such as the Process Analysis Method (PAM), SMART principle, and Analytic Hierarchy Process (AHP). This resulted in the development of a comprehensive assessment framework for water security in the VG–TB river basin. This framework identified five key dimensions, with basin development activities (0.32), the ability to meet water needs (0.24), and natural disaster resilience (0.19) being the most crucial and water resource potential being the least crucial (0.11) according to the AHP methodology. The latter also highlighted 15 indicators, four of which are particularly influential, including waste resources (0.54), flood (0.53), water storage capacity (0.45), and basin governance (0.42). Furthermore, 28 variables with high weight factors were identified. This framework aligns with the UN-Water water security definition and addresses the global water sustainability criteria outlined in Sustainable Development Goal 6 (SDG6). It enables the computation of a comprehensive Water Security Index (WSI) for specific regions, providing a strong foundation for decision-making and policy formulation. It aims to enhance water security in the context of climate change and support sustainable basin development, thereby guiding future research and policy decisions in water resource management.

54 ENVIRONMENTAL SCIENCES

PNNL-Predictive-Phenomics/ProteoMeter

ProteoMeter is a Python package that assists in the statistical analysis of global proteomics, protein post-translation modification (PTM), and limited proteolysis (LiP) data. It contains batch correction, normalization, and statistical testing methods, as well as functions that "roll up" peptide-level data to the single-site level. It has a robust user configuration system, allowing it to flexibly integrate different types of experiment designs. For basic usage, a simple configuration file provides the essential functionality. Advanced users have access to the entire statistical pipeline for fine-tuning analyses. Processed data is easily exported to many common spreadsheet and data-frame formats.

Rozum, Jordan [Pacific Northwest National Lab]

Comparative Analyses of Bioequivalence Assessment Methods for In Vitro Permeation Test Data

ABSTRACT For topical, dermatological drug products, an in vitro option to determine bioequivalence (BE) between test and reference products is recommended. In particular, in vitro permeation test (IVPT) data analysis uses a reference‐scaled approach for two primary endpoints, cumulative penetration amount (AMT) and maximum flux ( J max ), which takes the within donor variability into consideration. In 2022, the Food and Drug Administration (FDA) published a draft IVPT guidance that includes statistical analysis methods for both balanced and unbalanced cases of IVPT study data. This work presents a comprehensive evaluation of various methodologies used to estimate critical parameters essential in assessing BE. Specifically, we investigate the performance of the FDA draft IVPT guidance approach alongside alternative empirical and model‐based methods utilizing mixed‐effects models. Our analyses include both simulated scenarios and real‐world studies. In simulated scenarios, empirical formulas consistently demonstrate robustness in approximating the true model, particularly in effectively addressing treatment–donor interactions. Conversely, the effectiveness of model‐based approaches heavily relies on precise model selection, which significantly influences their results. The research emphasizes the importance of accurate model selection in model‐based BE assessment methodologies. It sheds light on the advantages of empirical formulas, highlighting their reliability compared to model‐based approaches and offers valuable implications for BE assessments. Our findings underscore the significance of robust methodologies and provide essential insights to advance their understanding and application in the assessment of BE, employed in IVPT data analysis.

Leon, Sami

Advanced measurement techniques in quantum Monte Carlo: The permutation matrix representation approach

In a typical finite temperature quantum Monte Carlo (QMC) simulation, estimators for simple static observables such as specific heat and magnetization are known. With a great deal of system-specific manual labor, one can sometimes also derive more complicated non-local or even dynamic observable estimators. In contrast, we show that arbitrary static observables can be estimated within the permutation matrix representation (PMR) flavor for any Hamiltonian. We then generalize these results to general imaginary-time correlation functions and non-trivial integrated susceptibilities thereof. Finally, we demonstrate the practical versatility of our method by estimating various non-local, random observables for the transverse-field Ising model on a square lattice and a toy random model.

Permutation matrix representation

Statistical data analysis of x-ray spectroscopy data enabled by neural network accelerated Bayesian inference

Bayesian inference applied to x-ray spectroscopy data analysis enables uncertainty quantification necessary to rigorously test theoretical models. However, when comparing to data, detailed atomic physics and radiation transfer calculations of x-ray emission from non-uniform plasma conditions are typically too slow to be performed in line with statistical sampling methods, such as Markov Chain Monte Carlo sampling. Furthermore, differences in transition energies and x-ray opacities often make direct comparisons between simulated and measured spectra unreliable. Here, we present a spectral decomposition method that allows for corrections to line positions and bound–bound opacities to best fit experimental data, with the goal of providing quantitative feedback to improve the underlying theoretical models and guide future experiments. In this work, we use a neural network (NN) surrogate model to replace spectral calculations of isobaric hot-spots created in Kr-doped implosions at the National Ignition Facility. The NN was trained on calculations of x-ray spectra using an isobaric hot-spot model post-processed with Cretin, a multi-species atomic kinetics and radiation code. The speedup provided by the NN model to generate x-ray emission spectra enables statistical analysis of parameterized models with sufficient detail to accurately represent the physical system and extract the plasma parameters of interest.

47 OTHER INSTRUMENTATION

A Method for Producing Hierarchical and Statistically Calibrated Predictions of Nuclear Material Properties from Existing Models

Computer vision-based analysis of micrographs of nuclear materials is an emerging technique for property prediction, synthetic route identification, and other material analysis tasks. These analysis tasks play a pivotal role in many material characterization applications such as signature development for treaty verification, process optimization, etc. The backbone in many of the recent computer vision-based techniques is a deep learning model, which takes a fixed-size set of pixels and provides a class prediction for that set of pixels. For example, previous work developed a deep convolutional neural network (CNN) to predict the synthetic route from a 256 px x 256 px patch taken from a larger image of uranium ore concentrates. In this work, we present several methods for first calibrating these models in a manner that they can provide accurate probabilities of their predictions’ veracity, and several methods of combining these probabilities. Overall, the combination of these two steps into a pipeline allows for full-image and even full-sample (where a sample has many images) predictions with associated confidence values. Finally, we show that one can also use the patch predictions and confidence to produce a visualization to map predicted constituents through the image. Results and examples for predicting and mapping uranium ore concentrates’ synthetic process from imagery will be presented.

artificial intelligence

Normalizing flows for domain adaptation when identifying Λ hyperon events

Here this study focuses on the application of a normalizing flow as a method of domain adaptation when classifying physics data. Normalizing flows offer a way to transform data points between two different distributions. The present study investigates a novel method of transforming latent representations of physics data to a normal distribution and then to a physics distribution again. The final distribution models a simulated distribution. After being transformed, the data can be classified by a neural network trained on labeled simulation data. The present study succeeds in training two normalizing flows that can transform between data (or simulation) and a Gaussian distribution.

47 OTHER INSTRUMENTATION

Evaluation of a high-throughput method for processing sponge-stick samples to detect viable, non-spore-forming biothreat agents

After a bioterrorism incident, surface sampling is often used to determine the extent of contamination and exposure, guiding decontamination efforts and decisions for re-occupancy of affected sites. The sponge-stick (SS) is a preferred and commonly used device for sample collection to detect both spore-forming and non-spore-forming biothreat agents from non-porous surfaces. Here, in this study, a recently developed high-throughput method (HTM) for processing SS samples to detect viable Bacillus anthracis spores was adapted for detection of non-spore-forming biothreat agents, Yersinia pestis and Francisella tularensis. The scalable HTM was used to process up to 20 SS samples simultaneously, compared to the current stomacher-based method which processes one SS at a time. Comparisons of the HTM and the stomacher-based method were statistically indistinguishable for most experiments (P > 0.05) with HTM recoveries of 37–60 % for Y. pestis inoculated at 102–103 cells/SS and held 48 h at 4 °C to mimic sample transport/storage. The HTM was integrated with Rapid Viability-Polymerase Chain Reaction (RV-PCR) analysis to detect viable Y. pestis in the presence of particulate contamination (Arizona Test Dust, ATD). This approach detected Y. pestis inoculated at 20 cells/SS and ATD did not impact detection (P > 0.05). F. tularensis showed significantly lower recoveries between no-hold time and 48-h hold time (4 °C, P < 0.05) using the HTM, which further testing showed could be due to toxicity of the neutralizing buffer used for SS pre-wetting. With modifications, this method could enhance throughput capacity while maintaining similar recovery efficiencies to current methods for other non-spore-forming bacterial pathogens.

Biological and medical sciences

Probabilistic inference of the structure and orbit of Milky Way satellites with semi-analytic modelling

Semi-analytic modelling furnishes an efficient avenue for characterizing dark matter haloes associated with satellites of Milky Way-like systems, as it easily accounts for uncertainties arising from halo-to-halo variance, the orbital disruption of satellites, baryonic feedback, and the stellar-to-halo mass (SMHM) relation. We use the SatGen semi-analytic satellite generator, which incorporates both empirical models of the galaxy–halo connection as well as analytic prescriptions for the orbital evolution of these satellites after accretion onto a host to create large samples of Milky Way-like systems and their satellites. By selecting satellites in the sample that match observed properties of a particular dwarf galaxy, we can infer arbitrary properties of the satellite galaxy within the cold dark matter paradigm. For the Milky Way’s classical dwarfs, we provide inferred values (with associated uncertainties) for the maximum circular velocity v max and the radius r max at which it occurs, varying over two choices of baryonic feedback model and two prescriptions for the SMHM relation. While simple empirical scaling relations can recover the median inferred value for v max and r max , this approach provides realistic correlated uncertainties and aids interpretability. We also demonstrate how the internal properties of a satellite’s dark matter profile correlate with its orbit, and we show that it is difficult to reproduce observations of the Fornax dwarf without strong baryonic feedback. Furthermore, the technique developed in this work is flexible in its application of observational data and can leverage arbitrary information about the satellite galaxies to make inferences about their dark matter haloes and population statistics.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

Examining the Robustness of Weakened Orographic Influence on Precipitation in Downscaled Climate Projections Over the Western US

Assessing local climate change impacts often requires downscaling coarse global climate model (GCM) output to finer resolution. Two main approaches exist: dynamical downscaling using high-resolution regional climate models, and statistical downscaling based on historical relationships between large-scale and local variables. In a recent analysis of five dynamically downscaled simulations over the western United States, Koszuta et al. (2024, https://doi.org/10.1029/2023gl107298) found that warming weakens orographic influence on winter precipitation, damping increases on windward slopes and amplifying them in rain-shadowed regions. Here we show that this effect is robust across seasons and multiple dynamically downscaled ensembles, and is more pronounced at higher model resolutions. However, it is absent in projections from a widely used statistical model (LOCA2), even when trained on high-resolution future simulations (LOCA2-Hybrid). This highlights a key limitation of many statistical downscaling methods: their preservation of parent GCM trends, which usually fail to capture emergent changes in orographic precipitation patterns.

54 ENVIRONMENTAL SCIENCES

Neural posterior unfolding

Differential cross section measurements are the currency of scientific exchange in particle and nuclear physics. A key challenge for these analyses is the correction for detector distortions, known as deconvolution or unfolding. Binned unfolding of cross section measurements traditionally rely on the regularized inversion of the response matrix that represents the detector response, mapping pre-detector (`particle level') observables to post-detector (`detector level') observables. In this paper we introduce Neural Posterior Unfolding, a modern, Bayesian approach that leverages normalizing flows for unfolding. By using normalizing flows for neural posterior estimation, NPU offers several key advantages including implicit regularization through the neural network architecture, fast amortized inference that eliminates the need for repeated retraining, and direct access to the full uncertainty in the unfolded result. In addition to introducing NPU, we implement a classical Bayesian unfolding method called Fully Bayesian Unfolding (FBU) in modern Python so it can also be studied. These tools are validated on simple Gaussian examples and then tested on simulated jet substructure examples from the Large Hadron Collider (LHC). We find that the Bayesian methods are effective and worth additional development to be analysis ready for cross section measurements at the LHC and beyond.

Analysis and statistical methods

Tools for unbinned unfolding

Machine learning has enabled differential cross section measurements that are not discretized. Going beyond the traditional histogram-based paradigm, these unbinned unfolding methods are rapidly being integrated into experimental workflows. Here, in order to enable widespread adaptation and standardization, we develop methods, benchmarks, and software for unbinned unfolding. For methodology, we demonstrate the utility of boosted decision trees for unfolding with a relatively small number of high-level features. This complements state-of-the-art deep learning models capable of unfolding the full phase space. To benchmark unbinned unfolding methods, we develop an extension of existing dataset to include acceptance effects, a necessary challenge for real measurements. Additionally, we directly compare binned and unbinned methods using discretized inputs for the latter in order to control for the binning itself. Lastly, we have assembled two software packages for the OmniFold unbinned unfolding method that should serve as the starting point for any future analyses using this technique. One package is based on the widely-used RooUnfold framework and the other is a standalone package available through the Python Package Index (PyPI).

47 OTHER INSTRUMENTATION

AI for nuclear physics: the EXCLAIM project

An overview of the recent activity of the newly funded EXCLusives with AI and Machine learning (EXCLAIM) collaboration is presented. The main goal of the collaboration is to develop a framework to implement AI and machine learning techniques in problems emerging from the phenomenology of high energy exclusive scattering processes from nucleons and nuclei, maximizing the information that can be extracted from various sets of experimental data, while implementing theoretical constraints from lattice QCD. A specific perspective embraced by EXCLAIM is to use the methods of theoretical physics to understand the working of ML, beyond its standardized applications to physics analyses which most often rely on industrially provided tools, in an automated way.

Analysis and statistical methods

Optimizing spin dressing sensitivity for the nEDMSF experiment

nEDMSF aims to measure the neutron electric dipole moment (d n ) with unprecedented precision. In this paper we explore the experiment's sensitivity when operating with an implementation of the critical dressing method in which the angle between the neutron and Helium-3 spins (ϕ 3n ) is subjected to a square modulation by an amount ϕ d (the “dressing angle”). Several parameters can be tuned to optimize sensitivity. We find roughly 10% improvement over a previous estimate, resulting primarily from the addition of a waiting period between the π/2 pulse that initiates d n -driven ϕ 3n growth and the start of ϕ3n modulation. We find negligible further improvement by allowing ϕ d to vary continuously over the course of a run, and no degradation resulting from the addition of an in situ background measurement into each ϕ3n modulation sequence. A complete simulation confirms a 300 live-day sensitivity ofσ = 1.45×10 -28 e ·cm. At this level of sensitivity, σ ϕ3n0 = 1 mrad precision on the initial n/ 3 He angle difference is not negligible.

47 OTHER INSTRUMENTATION

Characterization of a SiPM-based monolithic neutron scatter camera using dark counts

The Single Volume Scatter Camera (SVSC) Collaboration aims to develop portable neutron imaging systems for a variety of applications in nuclear non-proliferation. Conventional double-scatter neutron imagers are composed of several separate detector volumes organized in at least two planes. A neutron must scatter in two of these detector volumes for its initial trajectory to be reconstructed. As such, these systems typically have a large footprint and poor geometric efficiency. We report on the design and characterization of a prototype monolithic neutron scatter camera that is intended to significantly improve upon the geometrical shortcomings of conventional neutron cameras. The detector consists of a 50 mm×56 mm× 60 mm monolithic block of EJ-204 plastic scintillator instrumented on two faces with arrays of 64 Hamamatsu S13360-6075PE silicon photomultipliers (SiPMs). The electronic crosstalk is limited to < 5% between adjacent channels and < 0.1% between all other channel pairs. SiPMs introduce a significantly elevated dark count rate over PMTs, as well as correlated noise from after-pulsing and optical crosstalk. In this article, we characterize the dark count rate and optical crosstalk and present a modified event reconstruction likelihood function that accounts for them. We find that the average dark count rate per SiPM is 4.3 MHz with a standard deviation of 1.5 MHz among devices. The analysis method we employ to measure internal optical crosstalk also naturally yields the mean and width of the single-electron pulse height. Here, we calculate separate contributions to the width of the single-electron pulse-height from electronic noise and avalanche fluctuations. We demonstrate a timing resolution for a single-photon pulse to be (128 ± 4) ps. Finally, coincidence analysis is employed to measure external (pixel-to-pixel) optical crosstalk. We present a map of the average external crosstalk probability between 2×4 groups of SiPMs, as well as the in-situ timing characteristics extracted from the coincidence analysis. Further work is needed to characterize the performance of the camera at reconstructing single- and double-site interactions, as well as image reconstruction.

47 OTHER INSTRUMENTATION