Search NASA⌕ Search

SEARCH · Search NASA

Results for “likelihood”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Common Cause Case Study: An Estimated Probability of Four Solid Rocket Booster Hold-Down Post Stud Hang-ups

Until Solid Rocket Motor ignition, the Space Shuttle is mated to the Mobil Launch Platform in part via eight (8) Solid Rocket Booster (SRB) hold-down bolts. The bolts are fractured using redundant pyrotechnics, and are designed to drop through a hold-down post on the Mobile Launch Platform before the Space Shuttle begins movement. The Space Shuttle program has experienced numerous failures where a bolt has hung up. That is, it did not clear the hold-down post before liftoff and was caught by the SRBs. This places an additional structural load on the vehicle that was not included in the original certification requirements. The Space Shuttle is currently being certified to withstand the loads induced by up to three (3) of eight (8) SRB hold-down experiencing a "hang-up". The results of loads analyses performed for (4) stud hang-ups indicate that the internal vehicle loads exceed current structural certification limits at several locations. To determine the risk to the vehicle from four (4) stud hang-ups, the likelihood of the scenario occurring must first be evaluated. Prior to the analysis discussed in this paper, the likelihood of occurrence had been estimated assuming that the stud hang-ups were completely independent events. That is, it was assumed that no common causes or factors existed between the individual stud hang-up events. A review of the data associated with the hang-up events, showed that a common factor (timing skew) was present. This paper summarizes a revised likelihood evaluation performed for the four (4) stud hang-ups case considering that there are common factors associated with the stud hang-ups. The results show that explicitly (i.e. not using standard common cause methodologies such as beta factor or Multiple Greek Letter modeling) taking into account the common factor of timing skew results in an increase in the estimated likelihood of four (4) stud hang-ups of an order of magnitude over the independent failure case.

Cross, Robert↗

Hardware Implementation of Serially Concatenated PPM Decoder

A prototype decoder for a serially concatenated pulse position modulation (SCPPM) code has been implemented in a field-programmable gate array (FPGA). At the time of this reporting, this is the first known hardware SCPPM decoder. The SCPPM coding scheme, conceived for free-space optical communications with both deep-space and terrestrial applications in mind, is an improvement of several dB over the conventional Reed-Solomon PPM scheme. The design of the FPGA SCPPM decoder is based on a turbo decoding algorithm that requires relatively low computational complexity while delivering error-rate performance within approximately 1 dB of channel capacity. The SCPPM encoder consists of an outer convolutional encoder, an interleaver, an accumulator, and an inner modulation encoder (more precisely, a mapping of bits to PPM symbols). Each code is describable by a trellis (a finite directed graph). The SCPPM decoder consists of an inner soft-in-soft-out (SISO) module, a de-interleaver, an outer SISO module, and an interleaver connected in a loop (see figure). Each SISO module applies the Bahl-Cocke-Jelinek-Raviv (BCJR) algorithm to compute a-posteriori bit log-likelihood ratios (LLRs) from apriori LLRs by traversing the code trellis in forward and backward directions. The SISO modules iteratively refine the LLRs by passing the estimates between one another much like the working of a turbine engine. Extrinsic information (the difference between the a-posteriori and a-priori LLRs) is exchanged rather than the a-posteriori LLRs to minimize undesired feedback. All computations are performed in the logarithmic domain, wherein multiplications are translated into additions, thereby reducing complexity and sensitivity to fixed-point implementation roundoff errors. To lower the required memory for storing channel likelihood data and the amounts of data transfer between the decoder and the receiver, one can discard the majority of channel likelihoods, using only the remainder in operation of the decoder. This is accomplished in the receiver by transmitting only a subset consisting of the likelihoods that correspond to time slots containing the largest numbers of observed photons during each PPM symbol period. The assumed number of observed photons in the remaining time slots is set to the mean of a noise slot. In low background noise, the selection of a small subset in this manner results in only negligible loss. Other features of the decoder design to reduce complexity and increase speed include (1) quantization of metrics in an efficient procedure chosen to incur no more than a small performance loss and (2) the use of the max-star function that allows sum of exponentials to be computed by simple operations that involve only an addition, a subtraction, and a table lookup. Another prominent feature of the design is a provision for access to interleaver and de-interleaver memory in a single clock cycle, eliminating the multiple clock-cycle latency characteristic of prior interleaver and de-interleaver designs.

Moision, Bruce↗

Statistical inference of static analysis rules

Various apparatus and methods are disclosed for identifying errors in program code. Respective numbers of observances of at least one correctness rule by different code instances that relate to the at least one correctness rule are counted in the program code. Each code instance has an associated counted number of observances of the correctness rule by the code instance. Also counted are respective numbers of violations of the correctness rule by different code instances that relate to the correctness rule. Each code instance has an associated counted number of violations of the correctness rule by the code instance. A respective likelihood of the validity is determined for each code instance as a function of the counted number of observances and counted number of violations. The likelihood of validity indicates a relative likelihood that a related code instance is required to observe the correctness rule. The violations may be output in order of the likelihood of validity of a violated correctness rule.

Engler, Dawson Richards↗

Addressing Human System Risks to Future Space Exploration

NASA is contemplating future human exploration missions to destinations beyond low Earth orbit, including the Moon, deep-space asteroids, and Mars. While we have learned much about protecting crew health and performance during orbital space flight over the past half-century, the challenges of these future missions far exceed those within our current experience base. To ensure success in these missions, we have developed a Human System Risk Board (HSRB) to identify, quantify, and develop mitigation plans for the extraordinary risks associated with each potential mission scenario. The HSRB comprises research, technology, and operations experts in medicine, physiology, psychology, human factors, radiation, toxicology, microbiology, pharmacology, and food sciences. Methods: Owing to the wide range of potential mission characteristics, we first identified the hazards to human health and performance common to all exploration missions: altered gravity, isolation/confinement, increased radiation, distance from Earth, and hostile/closed environment. Each hazard leads to a set of risks to crew health and/or performance. For example the radiation hazard leads to risks of acute radiation syndrome, central nervous system dysfunction, soft tissue degeneration, and carcinogenesis. Some of these risks (e.g., acute radiation syndrome) could affect crew health or performance during the mission, while others (e.g., carcinogenesis) would more likely affect the crewmember well after the mission ends. We next defined a set of design reference missions (DRM) that would span the range of exploration missions currently under consideration. In addition to standard (6-month) and long-duration (1-year) missions in low Earth orbit (LEO), these DRM include deep space sortie missions of 1 month duration, lunar orbital and landing missions of 1 year duration, deep space journey and asteroid landing missions of 1 year duration, and Mars orbital and landing missions of 3 years duration. We then assessed the likelihood and consequences of each risk against each DRM, using three levels of likelihood (Low: less than or equal to 0.1%; Medium: 0.1%–1.0%; High: greater than or equal to 1.0%) and four levels of consequence ranging from Very Low (temporary or insignificant) to High (death, loss of mission, or significant reduction to length or quality of life). Quantitative evidence from clinical, operational, and research sources were used whenever available. Qualitative evidence was used when quantitative evidence was unavailable. Expert opinion was used whenever insufficient evidence was available. Results: A set of 30 risks emerged that will require further mitigation efforts before being accepted by the Agency. The likelihood by consequence risk assessment process provided a means of prioritizing among the risks identified. For each of the high priority risks, a plan was developed to perform research, technology, or standards development thought necessary to provide suitable reduction of likelihood or consequence to allow agency acceptance. Conclusion: The HSRB process has successfully identified a complete set of risks to human space travelers on planned exploration missions based on the best evidence available today. Risk mitigation plans have been established for the highest priority risks. Each risk will be reassessed annually to track the progress of our risk mitigation efforts.

Paloski, W. H.↗

A New Metric for Indian Monsoon Rainfall Extremes

Extreme monsoon rainfall in India has disastrous consequences, including significant socio- economic impacts. However, little is known about the overall trends and climate factors associated with extreme rainfall because rainfall greatly varies across India and because few appropriate methods are available to measure extreme rainfall in the context of such heterogeneity. To provide a comprehensive assessment of extreme monsoon rainfall, we developed a metric using record rainfall data to measure the changes in the likelihood of extreme high and extreme low rainfall over time; this metric is independent of the characteristics of the underlying rainfall distributions. Hence, the metric is ideally suited to aggregate extreme rainfall information across heterogeneous regions covering India. We found that from 1930 to 2013, the likelihood of extreme high and extreme low rainfall increases 2-fold and 4-fold, respectively. These overall trend increases are driven by anomalous increases, particularly in the early 2000s; the likelihood of extreme high and extreme low rainfall increases 5-fold and 18-fold in 2005 and 2002, respectively. These findings imply a broadening of the underlying monsoon rainfall distribution over the past century. We also show that the time patterns of the likelihood of extreme rainfall in recent decades are correlated with the El Nino Southern Oscillation, Indian Ocean Dipole, and surface air temperature in the Northern Hemisphere.

Rain↗

Toward The Development of Hailstorm Climatologies Derived From Reanalyses and Infared/Passive Microwave Satellite Imagers

Geostationary satellite imagers, such as those of the Geostationary Operational Environmental Satellite (GOES) and Meteosat series, provide both historical and near-real-time observations of cloud top patterns that are commonly associated with severe convection. Environmental conditions favorable for severe weather are thought to be represented well by reanalyses. Predicting exactly where convection and costly storm hazards like hail will occur using models or satellite imagery alone, however, is extremely challenging. The multivariate combination of satellite-observed cloud patterns with reanalysis environmental parameters, linked to United States Next Generation Weather Radar- (NEXRAD-) estimated Maximum Expected Size of Hail (MESH) using a deep neural network (DNN), enables estimation of potentially severe hail likelihood for any observed storm cell. These estimates are specifically designed to make hail likelihood distinctions based on satellite-indicated points of deep convection within environments favorable for storm development. We seek an approach that can be used to estimate climatological hailstorm frequency and risk throughout the historical satellite data record. This presentation demonstrates that statistical distributions of convective parameters from satellite and reanalysis show separation between non-severe/severe hailstorm classes for predictors including overshooting cloud top temperature and area characteristics, convective available potential energy, vertical wind shear, 500 hPa temperature, mid-level lapse rate, precipitable water, and convective inhibition. These complex, multivariate predictor relationships are exploited within a DNN to produce a hail likelihood metric with a critical success index of 0.504 and Heidke skill score of 0.403, which is exceptional among recent analogous hail studies. Furthermore, applications of the DNN to select case studies demonstrate good qualitative agreement between hail likelihood and MESH. These hail classifications are aggregated across an 11-year GOES-12/13 image database to derive a hail frequency and severity climatology, which denotes the Central Plains, the Midwest, and northwestern Mexico as being the most hail-prone regions within the domain studied. Opportunities for training and applying DNN-based hailstorm predictions to recently developed GOES-8/10/12/13/16 and Meteosat Second Generation convective storm detection and characterization climatologies over South America and South Africa, respectively, will also be presented.

Kristopher Bedka↗

Multiyear Dry Periods in Southern Africa

Characteristics and physical features related to low precipitation across many years in Southern Africa that lead to societal disruptions are diagnosed using observed analyses and an ensemble of historical coupled climate model simulations during 1921 to 2014. Four regions are evaluated, as identified through a hierarchical clustering algorithm applied to the Standardized Precipitation Index (SPI) during the October–April precipitation season. Although dryness spanning many October–April occurs periodically in each region, they seldom occur simultaneously, consistent with largely insignificant SPI cross-correlations between them. However, characteristics relevant to low precipitation across many years are generalizable between the four regions, including the serial persistence of October–April precipitation, the likelihood of consecutive dry October–April, and the likelihood of dry October–April in temporal extents of up to 10 consecutive such 7-month seasons. Systematic precipitation persistence is not a feature in any of the four Southern Africa regions, as serial correlations of October–April SPI are not statistically significant at any time lags. It follows that there is an exponential-folding decay in the likelihood of consecutive October–April for various SPI thresholds and that there is a large spread in the likelihood of low October–April SPI across many years. In terms of physical features, low October–April SPI in each Southern Africa region is closely related to local atmospheric circulations; however, they are not as closely related to sea surface temperatures (SSTs). These results suggest that dryness spanning many years is determined primarily by persistent local circulations related to atmospheric variability and to a lesser extent variability related to SST anomalies, including the El Niño–Southern Oscillation.

subtropical Indian Ocean dipole↗

Extravehicular Activity on the Lunar Surface: Mapping Mitigation Risk Consequence for Crew Needing Assistance or Rescue

Extravehicular activity (EVA) on the lunar surface presents unique risks to crew with possibility for injury. Without appropriate assistance or rescue capability, inability to nominally ambulate and return to a lander, especially during early Artemis missions, could have catastrophic consequences. Mapping likelihood and consequence safety risk associated with identified injury scenarios establishes a baseline from which to assess potential mitigation solutions to ensure crew health and safety. Causes leading to the need for incapacitated crew rescue (ICR) during EVA on the lunar surface were previously identified and classified using an ICR/Acute Injury scenario spectrum. Severe scenarios are those when the affected astronaut requires either partial or full continuous assistance from the rescuer. Evaluation of these continual reliance conditions included calculating event probabilities (likelihoods) associated with an early Artemis mission and mapping them to established Exploration System Directorate (ESD) probability thresholds; safety consequences were analyzed and correlated to defined ESD personnel safety categories. These resulting likelihood and consequence values served as a baseline for assessing risk reduction of three mitigation capabilities: crew assistance (rescuer crew) only, walking assist devices, and a wheeled transport device. Of the twenty-five continual reliance conditions, ten were evaluated as “catastrophic” (Level 5, loss of life) during EVA on the lunar surface with probabilities ranging from moderate to very low during an early Artemis mission. Crew assistance only and walking assist devices showed similar potential for risk reduction, with four of the ten causes decreasing to Level 4. A wheeled transport device further increased risk reduction with six of the ten conditions decreasing to Level 4. Given the catastrophic consequence of several identified conditions, assessments should be performed to determine the feasibility of mitigation capabilities. It is currently unknown whether a rescuer astronaut could effectively provide continuous assistance to enable both crew to return safely to the lander from the standpoint of both suit geometry and human performance. Although resulting in an increase in resources, providing a wheeled transport provides the highest risk reduction potential, and walking assist devices may have prevention as well as mitigation benefits.

lunar surface↗

Persistent Sampling: Enhancing the Efficiency of Sequential Monte Carlo

Sequential Monte Carlo (SMC) samplers are powerful tools for Bayesian inference but suffer from high computational costs due to their reliance on large particle ensembles for accurate estimates. We introduce persistent sampling (PS), an extension of SMC that systematically retains and reuses particles from all prior iterations to construct a growing, weighted ensemble. By leveraging multiple importance sampling and resampling from a mixture of historical distributions, PS mitigates the need for excessively large particle counts, directly addressing key limitations of SMC such as particle impoverishment and mode collapse. Crucially, PS achieves this without additional likelihood evaluations-weights for persistent particles are computed using cached likelihood values. This framework not only yields more accurate posterior approximations but also produces marginal likelihood estimates with significantly lower variance, enhancing reliability in model comparison. Furthermore, the persistent ensemble enables efficient adaptation of transition kernels by leveraging a larger, decorrelated particle pool. Experiments on high-dimensional Gaussian mixtures, hierarchical models, and non-convex targets demonstrate that PS consistently outperforms standard SMC and related variants, including recycled and waste-free SMC, achieving substantial reductions in mean squared error for posterior expectations and evidence estimates, all at reduced computational cost. PS thus establishes itself as a robust, scalable, and efficient alternative for complex Bayesian inference tasks.

Karamanis, Minas↗

Distinguishing Single and Linked Ruptures in the Laboratory and Nature

Earthquakes can grow either monotonically from a single, stressed patch or through linking multiple stressed regions. The distinction has implications for magnitude predictability with single ruptures requiring knowledge of the local stress state, while linked ruptures require knowing the global stress and energy distribution. Here, we use a laboratory fault that allows direct observation of slip to determine which fault conditions promote linked ruptures, and how to observationally evaluate their likelihood. Higher concentrations of normal stress due to increased normal force, applied stress asperities, or larger heterogeneity between the samples, all lead to significant increases in linked rupture likelihood. The mean radiated energy enhancement factor (REEF) of large events is an excellent proxy for linked event likelihood, and a combination of REEF and duration can identify the rupture style of some but not all events. The results imply that globally observed variations of REEF can be interpreted as variations in stress concentration.

earthquakes↗

Feasibility of peak temperature targets in light of institutional constraints

Despite faster-than-expected progress in clean energy technology deployment, global annual CO 2 emissions have increased from 2020 to 2023. The feasibility of limiting warming to 1.5 °C is therefore questioned. Here we present a model intercomparison study that accounts for emissions trends until 2023 and compares cost-effective scenarios to alternative scenarios with institutional, geophysical and technological feasibility constraints and enablers informed by previous literature. Our results show that the most ambitious mitigation trajectories with updated climate information still manage to limit peak warming to below 1.6 °C (‘low overshoot’) with around 50% likelihood. However, feasibility constraints, especially in the institutional dimension, decrease this maximum likelihood considerably to 5–45%. Accelerated energy demand transformation can reduce costs for staying below 2 °C but have only a limited impact on further increasing the likelihood of limiting warming to 1.6 °C. Our study helps to establish a new benchmark of mitigation scenarios that goes beyond the dominant cost-effective scenario design.

54 ENVIRONMENTAL SCIENCES↗

Characterization of uncertainties in electron-argon collision cross sections

Abstract The predictive capability of a plasma discharge model depends on accurate representations of electron-impact collision cross sections, which determine the corresponding reaction rates and electron transport properties. The values of cross sections can be known only approximately either through experiments or simulations and are thus subject to uncertainties. Quantifying the uncertainties in plasma simulations allows us to assess the reliability of simulations and to provide a basis for interpreting discrepancies between simulations and experiments. For such uncertainty quantification of plasma simulations, it is essential to quantify the uncertainties of the underlying cross sections. Although much effort has been committed to calibrate the cross section values, their uncertainties are not well investigated. We characterize uncertainties in electron-argon atom collision cross sections using a Bayesian framework. Six collision processes—elastic momentum transfer, ionization, and four excitations—are characterized with semi-empirical models, which effectively capture the features important to the macroscopic properties of the plasma. A probability model for the uncertain parameters of these semi-empirical models is developed. Specifically, a Gaussian-process likelihood model is proposed to capture discrepancies among data sets, as well as the model-form inadequacies of the semi-empirical models. Two other likelihood models are compared with the proposed Gaussian-process model, to illustrate the importance of the choice of the likelihood model. The cross section models are calibrated using the electron-beam experiments and ab-inito quantum simulations. The resulting calibrated uncertainties capture well the scattering among the data sets. The calibrated cross section models are further validated against swarm-parameter experiments and zero-dimensional Boltzmann equation simulations of widely used cross section datasets.

Chung, Seung Whan (ORCID:0000000302501549)↗

Supercharging simulation-based inference for Bayesian optimal experimental design

Abstract Bayesian optimal experimental design (BOED) seeks to maximize the expected information gain (EIG) of experiments. This requires a likelihood estimate, which in many settings is intractable. Simulation-based inference (SBI) provides powerful tools for this regime. However, existing work explicitly connecting SBI and BOED is restricted to a single contrastive EIG bound. We show that the EIG admits multiple formulations which can directly leverage modern SBI density estimators, encompassing neural posterior, likelihood, and ratio estimation. Building on this perspective, we define a novel EIG estimator using neural likelihood estimation. Further, we identify optimization as a key bottleneck of gradient based EIG maximization and show that a simple multi-start parallel gradient ascent procedure can substantially improve reliability and performance. With these innovations, our SBI-based BOED methods are able to match or outperform by up to 22% existing state-of-the-art approaches across standard BOED benchmarks.

97 MATHEMATICS AND COMPUTING↗

Cosmological implications of DESI DR2 BAO measurements in light of the latest ACT DR6 CMB data

We report cosmological results from the Dark Energy Spectroscopic Instrument (DESI) measurements of baryon acoustic oscillations (BAO) when combined with recent data from the Atacama Cosmology Telescope (ACT). By jointly analyzing ACT and Planck data and applying conservative cuts to overlapping multipole ranges, we assess how different 𝑃⁢𝑙⁢𝑎⁢𝑛⁢𝑐⁢𝑘 + ACT dataset combinations affect consistency with DESI. While ACT alone exhibits a tension with DESI exceeding 3⁢𝜎 within the Λ⁢ CDM model, this discrepancy is reduced when ACT is analyzed in combination with Planck. For our baseline DESI DR2 BAO + 𝑃⁢𝑙⁢𝑎⁢𝑛⁢𝑐⁢𝑘 P⁢R⁢4 + ACT likelihood combination, the preference for evolving dark energy over a cosmological constant is about 3⁢𝜎, increasing to over 4⁢𝜎 with the inclusion of type Ia supernova data. While the dark energy results remain quite consistent across various combinations of Planck and ACT likelihoods with those obtained by the DESI collaboration, the constraints on neutrino mass are more sensitive, ranging from ∑𝑚 𝜈 < 0.061 eV in our baseline analysis, to ∑𝑚 𝜈 < 0.077 eV (95% confidence level) in the CMB likelihood combination chosen by ACT when imposing the physical prior ∑𝑚 𝜈 > 0 eV.

79 ASTRONOMY AND ASTROPHYSICS↗

Unified nonparametric equation-of-state inference from the neutron-star crust to perturbative-QCD densities

Perturbative quantum chromodynamics (pQCD), while valid only at densities exceeding those found in the cores of neutron stars, could provide constraints on the dense-matter equation of state (EOS). Here, in this work, we examine the impact of pQCD information on the inference of the EOS using a nonparametric framework based on Gaussian processes (GPs). We examine the application of pQCD constraints through a ``pQCD likelihood,'' and verify the findings of previous works; namely, a softening of the EOS at the central densities of the most massive neutron stars and a reduction in the maximum neutron-star mass. Although the pQCD likelihood can be easily integrated into existing EOS inference frameworks, this approach requires an arbitrary selection of the density at which the constraints are applied. The EOS behavior is also treated differently on either side of the chosen density. To mitigate these issues, we extend the EOS model to higher densities, thereby constructing a ``unified'' description of the EOS from the neutron-star crust to densities relevant for pQCD. In this approach the pQCD constraints effectively become part of the prior. Since the EOS is unconstrained by any calculation or data between the densities applicable to neutron stars and pQCD, we argue for maximum modeling flexibility in that regime. We compare the unified EOS with the traditional pQCD likelihood, and although we confirm the EOS softening, we do not see a reduction in the maximum neutron-star mass or any impact on macroscopic observables. Though residual model dependence cannot be ruled out, we find that pQCD suggests the speed of sound in the densest neutron-star cores has already started decreasing toward the asymptotic limit; we find that the speed of sound squared at the center of the most massive neutron star has an upper bound of $\sim 0.5$ at the 90% level.

equations of state of nuclear matter↗

Neural network based emulation of galaxy power spectrum covariances: A reanalysis of BOSS DR12 data

We train neural networks to quickly generate redshift-space galaxy power spectrum covariances from a given parameter set (cosmology and galaxy bias). This covariance emulator utilizes a combination of traditional fully connected network layers and transformer architecture to accurately predict covariance matrices for the high redshift, north galactic cap sample of the BOSS DR12 galaxy catalog. We run simulated likelihood analyses with emulated and brute-force computed covariances, and we quantify the network’s performance via two different metrics: (1) difference in Χ 2 and (2) likelihood contours for simulated BOSS DR 12 analyses. We find that the emulator returns excellent results over a large parameter range. We then use our emulator to perform a reanalysis of the BOSS HighZ NGC galaxy power spectrum, and find that varying covariance with cosmology along with the model vector produces Ω m = $0.27⁢6$$^{+0.013}_{–0.015}$, H 0 = 70.2 ± 1.9 km/s/Mpc, and σ 8 = $0.67⁢4$$^{+0.058}_{–0.077}$. These constraints represent an average 0.46⁢σ shift in best-fit values and a 5% increase in constraining power compared to fixing the covariance matrix (Ω m = 0.293 ± 0.017, H 0 = 70.3 ± 2.0 km/s/Mpc, σ 8 = $0.70⁢2$$^{+0.063}_{–0.075}$). As a result, this work demonstrates that emulators for more complex cosmological quantities than second-order statistics can be trained over a wide parameter range at sufficiently high accuracy to be implemented in realistic likelihood analyses.

79 ASTRONOMY AND ASTROPHYSICS↗

Calibration verification for stochastic agent-based disease spread models

Accurate disease spread modeling is crucial for identifying the severity of outbreaks and planning effective mitigation efforts. To be reliable when applied to new outbreaks, model calibration techniques must be robust. However, current methods frequently forgo calibration verification (a stand-alone process evaluating the calibration procedure) and instead use overall model validation (a process comparing calibrated model results to data) to check calibration processes, which may conceal errors in calibration. In this work, we develop a stochastic agent-based disease spread model to act as a testing environment as we test two calibration methods using simulation-based calibration, which is a synthetic data calibration verification method. The first calibration method is a Bayesian inference approach using an empirically-constructed likelihood and Markov chain Monte Carlo (MCMC) sampling, while the second method is a likelihood-free approach using approximate Bayesian computation (ABC). Simulation-based calibration suggests that there are challenges with the empirical likelihood calculation used in the first calibration method in this context. These issues are alleviated in the ABC approach. Despite these challenges, we note that the first calibration method performs well in a synthetic data model validation test similar to those common in disease spread modeling literature. We conclude that stand-alone calibration verification using synthetic data may benefit epidemiological researchers in identifying model calibration challenges that may be difficult to identify with other commonly used model validation techniques.

60 APPLIED LIFE SCIENCES↗

ASCR Workshop Position Paper: Challenges and Opportunities in High Energy Physics

High energy particle physics and cosmology concern themselves with estimating fundamental parameters of nature, such as the masses and interactions of fundamental particles like the Higgs boson and the rate of expansion of the universe. In doing so, they analyze exabyte-scale datasets, some of the largest in all of science, and face many challenges in subsequent data analysis. These challenges are shared between the two disciplines, but we focus on particle physics to highlight one specific domain. In particle physics, the standard method for estimating parameters involves performing Monte Carlo (MC) integration as a function of both parameters of interest and nuisance parameters using an expensive simulator, counting the number of observed collision events (i.i.d. samples) from an experiment in the corresponding integration domains, and forming a Poisson likelihood function. This likelihood function is then used in a Frequentist manner to construct a maximum likelihood point estimate (MLE) and confidence set for the parameters. To sufficiently populate the high-dimensional integration domains, simulators consume billions of CPU-hours annually and produce hundreds of petabytes of intermediate output data. Several techniques have been developed to: optimize definitions of the integration domains so as to be maximally sensitive to a particular subset of parameters, efficiently estimate the integrals, and build robust surrogate models by interpolating between integral evaluations at different parameter points. One can view this whole endeavor as classical Simulation-Based Inference (SBI).

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗