Search NASA⌕ Search

SEARCH · Search NASA

Results for “Bayesian Statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

A flexible class of priors for orthonormal matrices with basis function-specific structure

Statistical modeling of high-dimensional matrix-valued data motivates the use of a low-rank representation that simultaneously summarizes key characteristics of the data and enables dimension reduction. Low-rank representations commonly factor the original data into the product of orthonormal basis functions and weights, where each basis function represents an independent feature of the data. However, the basis functions in these factorizations are typically computed using algorithmic methods that cannot quantify uncertainty or account for basis function correlation structure a priori. While there exist Bayesian methods that allow for a common correlation structure across basis functions, empirical examples motivate the need for basis function-specific dependence structure. We propose a prior distribution for orthonormal matrices that can explicitly model basis function-specific structure. The prior is used within a general probabilistic model for singular value decomposition to conduct posterior inference on the basis functions while accounting for measurement error and fixed effects. We discuss how the prior specification can be used for various scenarios and demonstrate favorable model properties through synthetic data examples. Finally, we apply our method to two-meter air temperature data from the Pacific Northwest, enhancing our understanding of the Earth system’s internal variability.

97 MATHEMATICS AND COMPUTING↗

Analysis Methods and Results for Weak Gamma-Ray Bursts in the BATSE Data

We report initial results on the statistical properties of the dimmest gamma-ray bursts (GRBs) observed with the Burst and Transient Source Experiment (BATSE), using new ground-based methods to obtain a sample of GRBs from 502 days of BATSE data. Using the most sensitive ground-based detection of GRBs, the sample extends to GRBs much fainter than those detected by the on-board trigger, but because of the temporal resolution of the data, the sample is limited to GRBs of duration of at least 2(approx.)s. For each detected event, Bayesian probabilities are calculated for the event to belong to each of seven classes of differing physical origins. The sample of GRB candidates is defined by the requirement that the Bayesian probability for belonging to the GRB class is higher than 0.5. The intensity distribution of the GRB sample is corrected using a Monte Carlo simulation of the post-flight detection efficiency. The dimmest BATSE bursts of the sample continue the hardness-intensity trend seen in brighter GREs and are consistent with isotropy.

Mitrofanov, I. G.↗

Ripening of Rh Nanoparticle Catalysts in Reverse Water–Gas Shift via a Data-Driven Model Combining Physics, Theory, and Experiment

Degradation via sintering is an ongoing challenge that impedes the broad commercial success of supported metallic nanoparticle catalysts. To mitigate degradation via informed catalyst design and process operations, here we aim to disambiguate the underlying mechanisms of sintering by combining theory and experiment in a quantitative framework. While mechanistic sintering models exist, they only model a single sintering pathway, even though multiple sintering mechanisms can occur simultaneously or dominate at different stages of the process. Data-driven machine learning models have emerged as a means to represent complex processes through data regression. However, machine learning models have very large data needs and lack mechanistic insights due to their black-box encoding. To develop an interpretive model of catalyst degradation via sintering, we constructed a hybrid model combining mechanistic “physics-based” models and data-driven methods to obtain both reliable predictions and mechanistic insights regarding experimentally observed sintering phenomena. Focusing on nanoparticle sintering in the Rh–TiO 2 catalyst for the reverse water–gas shift (RWGS) reaction, the hybrid model couples a mechanistic term for Ostwald ripening with energy values calculated via density functional theory (DFT) with a parametric, data-driven discrepancy function term for unmodeled mechanisms. The hybrid model is trained using Bayesian inference with data collected from small-angle X-ray scattering (SAXS) in situ experiments wherein average nanoparticle diameter versus time was measured at three relevant operating temperatures. The calibrated hybrid model results show that an Ostwald ripening-only model parameterized with fixed DFT energies does not fully capture the time and temperature dependence of the SAXS-observed sintering kinetics, and that an additional functional contribution, or DFT energy calibration, is required to reconcile simulation and experiment. Analysis of the hybrid-model error confirms that the hybrid model outperforms both the purely mechanistic and purely data-driven alternatives in terms of expected predictive accuracy for time-evolving average particle sizes. Furthermore, the results support the hypothesis that the Ostwald ripening mechanism is less important for explaining the sintering phenomena as operating temperature increases under an assumed fixed DFT parameterization. This could be explained in one of two ways: either latent, unmodeled sintering mechanisms dominate at higher temperatures, or the DFT uncertainty increases with temperature. The proposed modeling approach directly links theory to experiments and simulations via a statistical hybrid modeling framework and can be extended to other catalytic systems to improve predictive models and mechanistic understanding.

Bayesian hybrid modeling↗

A statistical technique for determining rainfall over land employing Nimbus 6 ESMR measurements

Statistical analysis of the Nimbus 6 ESMR measurements for remote monitoring of active rainfall data over land is presented. Horizontally and vertically polarized brightness temperature pairs from ESMR 6 were sampled for areas of rainfall over land as determined from the rain recording stations and the WSR 57 radar, and wet and dry ground over the southeastern U.S. These three categories of brightness temperatures were significantly different so that the possibilities of the mean vectors of any two populations coinciding were less than 1 in 100, so that classification algorithms were then developed. The Fisher linear classifier, the Bayesian quadratic classifier, and a non-parametric linear classifier were examined, and the Bayesian algorithm performed best. It was concluded that a rainfall area delineated by the Bayesian classifier coincided well with the synoptic-scale rainfall area mapped by ground recording rain data and radar echoes.

Rodgers, E.↗

Identifying microbial drivers in biological phenotypes with a Bayesian network regression model

Abstract In Bayesian Network Regression models, networks are considered the predictors of continuous responses. These models have been successfully used in brain research to identify regions in the brain that are associated with specific human traits, yet their potential to elucidate microbial drivers in biological phenotypes for microbiome research remains unknown. In particular, microbial networks are challenging due to their high dimension and high sparsity compared to brain networks. Furthermore, unlike in brain connectome research, in microbiome research, it is usually expected that the presence of microbes has an effect on the response (main effects), not just the interactions. Here, we develop the first thorough investigation of whether Bayesian Network Regression models are suitable for microbial datasets on a variety of synthetic and real data under diverse biological scenarios. We test whether the Bayesian Network Regression model that accounts only for interaction effects (edges in the network) is able to identify key drivers (microbes) in phenotypic variability. We show that this model is indeed able to identify influential nodes and edges in the microbial networks that drive changes in the phenotype for most biological settings, but we also identify scenarios where this method performs poorly which allows us to provide practical advice for domain scientists aiming to apply these tools to their datasets. BNR models provide a framework for microbiome researchers to identify connections between microbes and measured phenotypes. We allow the use of this statistical model by providing an easy‐to‐use implementation which is publicly available Julia package at https://github.com/solislemuslab/BayesianNetworkRegression.jl .

59 BASIC BIOLOGICAL SCIENCES↗

The Analysis of the Contribution of Human Factors to the In-Flight Loss of Control Accidents

In-flight loss of control (LOC) is currently the leading cause of fatal accidents based on various commercial aircraft accident statistics. As the Next Generation Air Transportation System (NextGen) emerges, new contributing factors leading to LOC are anticipated. The NASA Aviation Safety Program (AvSP), along with other aviation agencies and communities are actively developing safety products to mitigate the LOC risk. This paper discusses the approach used to construct a generic integrated LOC accident framework (LOCAF) model based on a detailed review of LOC accidents over the past two decades. The LOCAF model is comprised of causal factors from the domain of human factors, aircraft system component failures, and atmospheric environment. The multiple interdependent causal factors are expressed in an Object-Oriented Bayesian belief network. In addition to predicting the likelihood of LOC accident occurrence, the system-level integrated LOCAF model is able to evaluate the impact of new safety technology products developed in AvSP. This provides valuable information to decision makers in strategizing NASA's aviation safety technology portfolio. The focus of this paper is on the analysis of human causal factors in the model, including the contributions from flight crew and maintenance workers. The Human Factors Analysis and Classification System (HFACS) taxonomy was used to develop human related causal factors. The preliminary results from the baseline LOCAF model are also presented.

Ancel, Ersin↗

More Data Needed for Failure Rate Estimation, Validation, and Uncertainty Reduction

The currently planned schedule for advanced Environmental Control and Life Support System (ECLSS) development and test activities to support human exploration missions is unlikely to generate sufficient data to enable statistically-supportable, precise Orbital Replacement Unit (ORU) failure rate estimates to meet existing crew safety expectations. Accurate and precise failure rate estimates are critical for missions beyond Low Earth Orbit (LEO) because current risk mitigation approaches –namely regular resupply and rapid abort capabilities –will not be available. Safe operations will depend on mission planners’ ability to accurately forecast spares demand and efficiently provide the necessary resources. However, even after more than a decade of International Space Station (ISS) ECLSS operations, a significant amount of uncertainty remains in failure rate estimates. Uncertain or inaccurate failure rates result in increased risk and spares mass. A Bayesian estimation approach, such as the one currently implemented by the ISS Program, can reduce uncertainty by incorporating engineering judgement into failure rate estimates. However, experience on the ISS and with other complex systems shows that these prior failure rate estimates are often inaccurate. In addition, prior estimates are typically point values; some level of uncertainty must be added to convert these into probability distributions for Bayesian updating, and there are several potential methods for doing so. Due to the low rate of data collection, any inaccuracy in theseprior estimates currently hasa strong influence on the end result. This paper examines the challenges associated with failure rate estimation, validation, and uncertainty reduction in the context of ECLSS development for beyond-LEO missions. A variety of techniques for generating and updating Bayesian priors are discussed and evaluated using both real-world and simulated data. Potential solutions for improving failure rate estimation, including testing additional units, are analyzed and discussed, and a set of recommendations are provided for next-generation system development activities.

Reliability↗

More Data Needed for Failure Rate Estimation, Validation, and Uncertainty Reduction

The currently planned schedule for advanced Environmental Control and Life Support System (ECLSS) development and test activities to support human exploration missions is unlikely to generate sufficient data to enable statistically-supportable, precise Orbital Replacement Unit (ORU) failure rate estimates to meet existing crew safety expectations. Accurate and precise failure rate estimates are critical for missions beyond Low Earth Orbit (LEO) because current risk mitigation approaches –namely regular resupply and rapid abort capabilities –will not be available. Safe operations will depend on mission planners’ ability to accurately forecast spares demand and efficiently provide the necessary resources. However, even after more than a decade of International Space Station (ISS) ECLSS operations, a significant amount of uncertainty remains in failure rate estimates. Uncertain or inaccurate failure rates result in increased risk and spares mass. A Bayesian estimation approach, such as the one currently implemented by the ISS Program, can reduce uncertainty by incorporating engineering judgement into failure rate estimates. However, experience on the ISS and with other complex systems shows that these prior failure rate estimates are often inaccurate. In addition, prior estimates are typically point values; some level of uncertainty must be added to convert these into probability distributions for Bayesian updating, and there are several potential methods for doing so. Due to the low rate of data collection, any inaccuracy in theseprior estimates currently hasa strong influence on the end result. This paper examines the challenges associated with failure rate estimation, validation, and uncertainty reduction in the context of ECLSS development for beyond-LEO missions. A variety of techniques for generating and updating Bayesian priors are discussed and evaluated using both real-world and simulated data. Potential solutions for improving failure rate estimation, including testing additional units, are analyzed and discussed, and a set of recommendations are provided for next-generation system development activities.

Reliability↗

The Error Distribution of BATSE GRB Location

We develop empirical probability models for BATSE GRB location errors by a Bayesian analysis of the separations between BATSE GRB locations and locations obtained with the InterPlanetary Network (IPN). Models are compared and their parameters estimated using 394 GRBs with single IPN annuli and 20 GRBs with intersecting IPN annuli. Most of the analysis is for the 4B (rev) BATSE catalog; earlier catalogs are also analyzed. The simplest model that provides a good representation of the error distribution has 78% of the locations in a 'core' term with a systematic error of 1.85 degrees and the remainder in an extended tail with a systematic error of 5.36 degrees, implying a 68% confidence region for bursts with negligible statistical errors of 2.3 degrees. There is some evidence for a more complicated model in which the error distribution depends on the BATSE datatype that was used to obtain the location. Bright bursts are typically located using the CONT datatype, and according to the more complicated model, the 68% confidence region for CONT-located bursts with negligible statistical errors is 2.0 degrees.

Briggs, Michael S.↗

The Error Distribution of BATSE Gamma-Ray Burst Locations

Empirical probability models for BATSE gamma-ray burst (GRB) location errors are developed via a Bayesian analysis of the separations between BATSE GRB locations and locations obtained with the Interplanetary Network (IPN). Models are compared and their parameters estimated using 392 GRBs with single IPN annuli and 19 GRBs with intersecting IPN annuli. Most of the analysis is for the 4Br BATSE catalog; earlier catalogs are also analyzed. The simplest model that provides a good representation of the error distribution has 78% of the probability in a "core" term with a systematic error of 1.85 deg and the remainder in an extended tail with a systematic error of 5.1 deg, which implies a 68% confidence radius for bursts with negligible statistical uncertainties of 2.2 deg. There is evidence for a more complicated model in which the error distribution depends on the BATSE data type that was used to obtain the location. Bright bursts are typically located using the CONT data type, and according to the more complicated model, the 68% confidence radius for CONT-located bursts with negligible statistical uncertainties is 2.0 deg.

Briggs, Michael S.↗

More Data Needed for Failure Rate Estimation, Validation, and Uncertainty Reduction

Current Environmental Control and Life Support System (ECLSS) development and test activities are not generating data fast enough to provide statistically-supportable precise Orbital Replacement Unit (ORU) failure rate estimates for future missions. Accurate and precise failure rate estimates are critical for missions beyond Low Earth Orbit (LEO) because current risk mitigation approaches – namely regular resupply and rapid abort capabilities – will not be available. Safe operations will depend on mission planners’ ability to accurately forecast spares demand and efficiently provide the necessary resources. However, even after more than a decade of operations on board the International Space Station (ISS), a significant amount of uncertainty remains in failure rate estimates. Uncertain or inaccurate failure rates result in increased risk and spares mass for future missions. A Bayesian failure rate estimation approach, such as the one currently implemented by the ISS Program, can help reduce uncertainty by incorporating engineering judgement into failure rate estimates. However, experience on the ISS and with other complex systems shows that these prior failure rate estimates are often inaccurate. In addition, prior failure rate estimates are typically point values; some level of uncertainty must be added to convert these into probability distributions for Bayesian updating, and there are several potential methods for doing so. Due to the low rate of data collection, these subjective (and often inaccurate) prior estimates currently have a strong influence on the end result. This paper examines the challenges associated with failure rate estimation, validation, and uncertainty reduction in the context of ECLSS development for beyond-LEO missions. A variety of techniques for generating and updating Bayesian priors are discussed and evaluated using both real-world and simulated data. Potential solutions for improving failure rate estimation, including testing additional units, are analyzed and discussed, and a set of recommendations are provided for next-generation system development activities.

Supportability↗

Derivation of Failure Rates and Probability of Failures for the International Space Station Probabilistic Risk Assessment Study

National Aeronautics and Space Administration s (NASA) International Space Station (ISS) Program uses Probabilistic Risk Assessment (PRA) as part of its Continuous Risk Management Process. It is used as a decision and management support tool to not only quantify risk for specific conditions, but more importantly comparing different operational and management options to determine the lowest risk option and provide rationale for management decisions. This paper presents the derivation of the probability distributions used to quantify the failure rates and the probability of failures of the basic events employed in the PRA model of the ISS. The paper will show how a Bayesian approach was used with different sources of data including the actual ISS on orbit failures to enhance the confidence in results of the PRA. As time progresses and more meaningful data is gathered from on orbit failures, an increasingly accurate failure rate probability distribution for the basic events of the ISS PRA model can be obtained. The ISS PRA has been developed by mapping the ISS critical systems such as propulsion, thermal control, or power generation into event sequences diagrams and fault trees. The lowest level of indenture of the fault trees was the orbital replacement units (ORU). The ORU level was chosen consistently with the level of statistically meaningful data that could be obtained from the aerospace industry and from the experts in the field. For example, data was gathered for the solenoid valves present in the propulsion system of the ISS. However valves themselves are composed of parts and the individual failure of these parts was not accounted for in the PRA model. In other words the failure of a spring within a valve was considered a failure of the valve itself.

Vitali, Roberto↗

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will produce unprecedented volumes of heterogeneous astronomical data (images, catalogs, and alerts) that challenge traditional analysis pipelines. The LSST Dark Energy Science Collaboration (DESC) aims to derive robust constraints on dark energy and dark matter from these data, requiring methods that are statistically powerful, scalable, and operationally reliable. Artificial intelligence and machine learning (AI/ML) are already embedded across DESC science workflows, from photometric redshifts and transient classification to weak lensing inference and cosmological simulations. Yet their utility for precision cosmology hinges on trustworthy uncertainty quantification, robustness to covariate shift and model misspecification, and reproducible integration within scientific pipelines. This white paper surveys the current landscape of AI/ML across DESC's primary cosmological probes and cross-cutting analyses, revealing that the same core methodologies and fundamental challenges recur across disparate science cases. Since progress on these cross-cutting challenges would benefit multiple probes simultaneously, we identify key methodological research priorities, including Bayesian inference at scale, physics-informed methods, validation frameworks, and active learning for discovery. With an eye on emerging techniques, we also explore the potential of the latest foundation model methodologies and LLM-driven agentic AI systems to reshape DESC workflows, provided their deployment is coupled with rigorous evaluation and governance. Finally, we discuss critical software, computing, data infrastructure, and human capital requirements for the successful deployment of these new methodologies, and consider associated risks and opportunities for broader coordination with external actors.

Aubourg, Eric [APC, Paris] (ORCID:000000025592023X↗

Multimodality in the Search for New Physics in Pulsar Timing Data and the Case of Kination-amplified Gravitational-wave Background from Inflation

We investigate the kination-amplified inflationary gravitational-wave background (GWB) interpretation of the signal recently reported by various pulsar timing array (PTA) experiments. Kination is a post-inflationary phase in the expansion history dominated by the kinetic energy of some scalar field, characterized by a stiff equation of state w = 1. Within the inflationary GWB model, we identify two modes that can fit the current data sets (NANOGrav and EPTA) with equal likelihood: the kination-amplification (KA) mode and the ordinary, no-kination-amplification (no-KA) mode. The multimodality of the likelihood motivates a Bayesian analysis with nested sampling. We analyze the free spectra of current PTA data and mock free spectra constructed with higher signal-to-noise ratios using nested sampling. The analysis of the mock spectrum designed to be consistent with the best fit to the NANOGrav 15 yr (NG15) data successfully reveals the expected bimodal posterior for the first time while excluding the reheating mode that appears in the fit to the current NG15 data, making a case for our correct and comprehensive treatment of potential multimodal posteriors arising from future PTA data sets. The resultant Bayes factor is $\mathcal{B}$ $\equiv$ Z no–KA /Z KA = 2.9 ± 1.9, indicating comparable statistical significance between the two modes. Given the theoretical model-building challenges of producing highly blue-tilted primordial tensor spectra, the KA mode has the advantage of requiring less blue primordial spectra, compared with the no-KA mode. The synergy between future cosmic microwave background polarization, pulsar timing, and laser interferometer measurements of gravitational waves will help resolve the ambiguity implied by the multimodal posterior in PTA-only searches.

Cosmology↗

Assimilation of Microwave Observations in the Rainbands of Tropical Cyclones

We propose a novel Bayesian Monte Carlo Integration (BMCI) technique to retrieve the profiles of temperature, water vapor, and cloud liquid/ice water content from microwave cloudy measurements in the presence of tropical cyclones (TC). These retrievals then can either be directly used by meteorologists to analyze the structure of TCs or be assimilated into numerical models to provide accurate initial conditions for the NWP (Numerical Weather Prediction) models. The BMCI technique is applied to the data from the Advanced Technology Microwave Sounder (ATMS) onboard Suomi National Polar-orbiting Partnership (NPP) and Global Precipitation Measurement (GPM) Microwave Imager (GMI). The retrieved profiles are then assimilated into Hurricane WRF (Weather Research and Forecasting) using the GSI (Gridpoint Statistical Interpolation) data assimilation system.

Moradi, Isaac↗

Comparison of Likelihood Methods for Generalized Linear Mixed Models with Application to Quiet Supersonic Flights 2018 Data

Repeated measurement will be a feature of the survey data collected during the Quesst missionX-59 community response tests (CRT). Since each participant will report his or her categorical level of annoyance in response to multiple events, the responses from any single individual may be correlated with one another. Several models within the class of generalized linear mixed models (GLMM) are pertinent to the analysis of correlated categorical outcomes; the random intercept logistic regression model is one example. Both Bayesian and frequentist methods for fitting these models are available, with frequentist methods relying on some form of approximation (of either an integral or the integrand) that appears in the marginal likelihood function. Given several anticipated similarities of the X-59 CRT data to data collected during a past risk reduction, Quiet Supersonic Flights 2018 (QSF18), this short note is intended to create awareness. It documents an instance in which a reported population average dose-response relationship derived from QSF18 single event data was distorted by the integral approximation applied in likelihood-based methods. We review some of the available literature on the topic, compare the outputs of several different computational approaches implemented in available statistical software, and present simple corrective actions that may be useful during the Quesst mission.

dose-response model↗

Assay-based background projection for the Majorana Demonstrator using Monte Carlo uncertainty propagation

The background index (BI) is an important quantity to project and calculate the half-life sensitivity of neutrinoless double-𝛽 decay (0⁢𝜈⁢𝛽⁢𝛽) experiments. An analysis framework is presented to calculate the BI using the specific activities, masses, and simulated efficiencies of an experiments components as distributions. This Bayesian framework includes a unified approach to combine specific activities from assay. Monte Carlo uncertainty propagation is used to build a BI distribution from the specific activity, mass, and efficiency distributions. This method is applied to the M AJORANA D EMONSTRATOR , which deployed arrays of high-purity Ge detectors enriched in 76 Ge to search for 0⁢𝜈⁢𝛽⁢𝛽. The original assay-based projection is requantified in the new framework, using the as-built geometry of the Demonstrator and additional assay information. While 47% higher than the original projection, the resulting BI of [8.95±0.36]×10 −4 cts/(keVkgyr) from the 232 Th and 238 U decay chains does not account for the higher-than-expected BI observed by the D EMONSTRATOR . Finally, this method enables us to demonstrate the statistical incompatibility between the D EMONSTRATOR 's observed background and the assay results.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Topics in inference and decision-making with partial knowledge

Two essential elements needed in the process of inference and decision-making are prior probabilities and likelihood functions. When both of these components are known accurately and precisely, the Bayesian approach provides a consistent and coherent solution to the problems of inference and decision-making. In many situations, however, either one or both of the above components may not be known, or at least may not be known precisely. This problem of partial knowledge about prior probabilities and likelihood functions is addressed. There are at least two ways to cope with this lack of precise knowledge: robust methods, and interval-valued methods. First, ways of modeling imprecision and indeterminacies in prior probabilities and likelihood functions are examined; then how imprecision in the above components carries over to the posterior probabilities is examined. Finally, the problem of decision making with imprecise posterior probabilities and the consequences of such actions are addressed. Application areas where the above problems may occur are in statistical pattern recognition problems, for example, the problem of classification of high-dimensional multispectral remote sensing image data.

Safavian, S. Rasoul↗