Search NASA⌕ Search

SEARCH · Search NASA

Results for “Likelihood”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29

Towards a Methodology and Tooling for Model-Based Probabilistic Risk Assessment (PRA)

A Probabilistic Risk Assessment (PRA) aims to identify and assess potential risks to system technical performance requirements for the purpose of furnishing risk insights into project decisions. PRAs have traditionally been conducted manually using software with an isolated data model. As system complexity rises it becomes difficult to ensure consistency between a PRA, the evolving system design, and other engineering analyses; the techniques for conducting PRAs must evolve to meet this challenge. This work presents progress towards a methodology and tooling for conducting a PRA by leveraging data in the system model, embedded for other purposes and analyses, to conduct a PRA. An approach for identifying the appropriate probabilistic equation for each risk scenario from a standard library is presented, which is a significant step towards the quantification of the likelihood of a risk scenario occurrence. The final calculation of the likelihood of occurrence is left as an item of future work. We also present the development of preliminary tooling to carry out the methodology on a well-formed system model. The information needed to conduct the PRA is embedded in a consistent manner in a system model, so the model-based PRA can be regularly executed as the system model changes. The ability to modify the PRA in concert with lifecycle evolution affords a project the opportunity to track the extent to which system modification impacts compliance with requirements. These aspects of this model-based PRA methodology make it capable of managing risk in increasingly complex technical systems.

Schreiner, Samuel S.↗

Method for Tracking and Communicating Aggregate Risk Through the Use of Model-Based Systems Engineering (MBSE) Tools

Large, complex projects can identify a significant number and variety of risks, throughout the project life cycle. These risks are analyzed, mitigated, closed or accepted as independent uncertainties. Once closed or accepted, it is easy for projects to lose awareness of their impact. In reality, each of these risks contributes some amount to the overall risk posture of the project. The ability to track and effectively communicate this aggregate risk has represented a challenge to project management. There have been previous attempts to create a schema to communicate the aggregate effect of risks, without notable success. Most of these attempts have centered on some additive metric derived from the scoring of likelihood and consequence values. This, in and of itself, is a logical approach, but all too often the scores were then aggregated to a level where all context was lost. One weakness has been a lack of attempt to create linkages or logical groups of the risks upon which useful aggregation could then occur. The overall move to model-based (systems) engineering (MBSE) has opened up a vast frontier of opportunities to better integrate all project data. MBSE provides an underlying layer that links data items to each other. Objectives link to requirements, which then link to functions, functions to physical architecture items, and so on, as far down as projects want to model. While it started with a focus on modeling requirements based on things like use cases, efforts are now underway to integrate safety and mission assurance (S&MA) information and analyses, such as risks. This effort, called Model Based Mission Assurance (MBMA), is yielding models that are more useful and are a more accurate representations of the systems. MBSE models, with this ability to link related items, provide a new means of tracking and communicating aggregate risks. In the proposed method, risks are added into the models as distinct items, having attributes that communicate a scoring derived from the likelihood and consequence values as charted on the standard NASA 5x5 risk matrix. Like earlier efforts, each box in the 5x5 has an associated scoring, which may include both a current score and potential post-mitigation/control score. The risk items are then linked to elements of the model, such as system objectives/goals, requirements, functions, or physical architecture items, with "Risk to" relationships. These risks will then be communicated by use of reports generated from the model, detailing all risks and/or hazards linked to model elements. These reports can include aggregate impacts, including a current scoring and potential future state scoring based on the planned mitigations and/or controls. These reports will show all risks, open, accepted, and closed, linked to project objectives or requirements. When run as part of an upcoming risk acceptance discussion, these reports will serve to remind the team of all previous risks that relate to the effected portion of the system. When included as part of periodic program or project reviews, risk reviews, and safety reviews, this method can improve the overall understanding of the system's true risk posture. This proposed method takes full advantage of the advances that modern modeling techniques provide, with a minimal investment of additional time. Utilizing the model environment also enables a near constant access to current state of aggregate risks.

model based mission assurance↗

An Operational Algorithm for Evaluating Satellite Collision Consequence

Risk is properly considered as the combination of likelihood and consequence; but conjunction assessment has usually limited itself to the consideration of only collision likelihood. When considered from an orbital regime protection perspective, the focus shifts to the question of the amount of debris that a collision might produce (the “consequence”). The present paper presents an operational algorithm for determining the expected amount of debris production should a conjunction result in a collision, and an assessment of the algorithm’s fidelity against a database of characterized objects.

Lechtenberg, Travis↗

An Operational Algorithm for Evaluating Satellite Collision Consequence

Risk is properly considered as the combination of likelihood and consequence; but conjunction assessment has usually limited itself to the consideration of only collision likelihood. When considered from an orbital regime protection perspective, the focus shifts to the question of the amount of debris that a collision might produce (the “consequence”). The present paper presents an operational algorithm for determining the expected amount of debris production should a conjunction result in a collision, and an assessment of the algorithm’s fidelity against a database of characterized objects.

Lechtenberg, Travis↗

Assessment of Precipitation Anomalies in California Using TRMM and MERRA Data

After more than a decade of moderate seasonal deviations from the expected climate, it is easy to forget that California is actually prone to instabilities in precipitation patterns that occur on various scales. Using modern satellite and reanalysis data we reassess certain aspects of the precipitation climate in California from the past three decades. California has a well-pronounced rain season that peaks in December-February. However, the 95% confidence interval around the climatological precipitation during these months imply that deviations on the order of 60% of the expected amounts are very likely during the most important period of the rain season. While these positive and negative anomalies alternate almost every year and tend to cancel each other, severe multi-year declines of precipitation in California seem to appear on decadal scales. The 1986-1994 decline of precipitation was similar to the current one that started in 2011, and is apparent in the reanalysis data. In terms of accumulated deficits of precipitation, that episode was no less severe than the current one. While El Niño (the warm phase of the El Niño Southern Oscillation, ENSO) is frequently cited as the natural forcing expected to bring a relief, our assessment is that ENSO has been driving at best only 6% of precipitation variability in California in the past three decades. It means El Niño needs to be stronger and longer, in order to have a higher likelihood of a positive impact, and the current one does not match these criteria. Using fractional risk analysis of precipitation populations during normal and dry periods, we show that the likelihood of losing the most intensive precipitation events drastically increases during the multi-year drying events. Since storms delivering up to 50% of precipitation in California are driven by atmospheric rivers making landfall, thus the importance of their suppression and blockage by persistent ridges of atmospheric pressure in the northeast Pacific.

Savtchenko, Andrey K.↗

Numerical and Experimental Study of Local Resin Pressure for the Manufacturing of Composite Structures and Their Effect on Porosity

Porosity can have significant impact on the mechanical performance of composite structures. The primary sources of voids during the cure of composites include air entrapped during lay-up, bag or tool leaks, and the off gassing of volatiles. Capturing the physics of void evolution during composite processing is a challenge due to the number of sources and changing phenomena that give rise to voids as the resin cures. Local changes in resin pressure due to geometric features or cure shrinkage have been experimentally linked to a higher likelihood of void formation. In this work, a preliminary model was developed to predict local resin pressure which can be used to identify regions more susceptible to porosity. Experiments were conducted to validate the model accuracy. Ultimately, this model can be used as a tool to minimize the likelihood of porosity by optimizing the material selection, part design, layup, and cure process.

Bedayat, Houman↗

Adaptive Stress Testing of Trajectory Predictions in Flight Management Systems

To find failure events and their likelihoods in flight-critical systems, we investigate the use of an advanced black-box stress testing approach called adaptive stress testing. We analyze a trajectory predictor from a developmental commercial flight management system which takes as input a collection of lateral waypoints and en-route environmental conditions. Our aim is to search for failure events relating to inconsistencies in the predicted lateral trajectories. The intention of this work is to find likely failures and report them back to the developers so they can address and potentially resolve shortcomings of the system before deployment. To improve search performance, this work extends the adaptive stress testing formulation to be applied more generally to sequential decision-making problems with episodic reward by collecting the state transitions during the search and evaluating at the end of the simulated rollout. We use a modified Monte Carlo tree search algorithm with progressive widening as our adversarial reinforcement learner. The performance is compared to direct Monte Carlo simulations and to the cross-entropy method as an alternative importance sampling baseline. The goal is to find potential problems otherwise not found by traditional requirements-based testing. Results indicate that our adaptive stress testing approach finds more failures and finds failures with higher likelihood relative to the baseline approaches.

adaptive stress testing↗

3D Representation of UAV-obstacle Collision Risk Under Off-nominal Conditions

Safe operations of autonomous unmanned aerial vehicles (UAVs) in low-altitude airspace with beyond visual line-of-sight (BVLOS) flights demand robust risk monitoring of airspace as well as of people and property on ground. One of the safety critical factors for UAV flights is the risk of collision with static and dynamic obstacles in proximity to its flight path. This paper presents a detailed formulation of risk of obstacle collision incorporating the effects of off-nominal conditions introduced by component failures, degraded controllability and environmental disturbances such as wind gusts. The risk is represented in terms of a matrix with rows corresponding to the likelihood of occurrence of collision and columns representing severity of collision to the vehicle and surrounding structures. Risk likelihood is generated using a Bayesian Belief Network (BBN) that compiles knowledge from related Failure Modes and Effects Analysis (FMEAs) and Subject Matter Experts (SMEs) to determine the probability of collision based on on-board sensor measurements indicative of vehicle health and controllability. Risk severity is computed utilizing a point-mass 3D kinematic model of the vehicle in presence of wind. The proposed risk factor is demonstrated on real flight data from experimental flights of an octocopter at NASA Langley Research Center in presence of simulated obstacles and wind conditions. Effect of varying wind conditions, level of controllability and obstacle measurement noise on the risk factor is demonstrated. The proposed approach enables risk-informed decision making for timely mitigation of current and future unsafe events in autonomous systems.

risk analysis↗

3D Representation of UAV-obstacle Collision Risk under off-nominal conditions

Safe operations of autonomous unmanned aerial vehicles (UAVs) in low-altitude airspace with beyond visual line-of-sight (BVLOS) flights demand robust risk monitoring of airspace as well as of people and property on ground. One of the safety critical factors for UAV flights is the risk of collision with static and dynamic obstacles in proximity to its flight path. This paper presents a detailed formulation of risk likelihood of obstacle collision incorporating the effects of off-nominal conditions introduced by component failures, degraded controllability and environmental disturbances such as wind gusts. The deviation in the planned trajectory caused due to wind is computed utilizing a point-mass 3D kinematic simulation model of the vehicle. Likelihood of risk for the flight plan is then analyzed based on generating the probability of collision for each point in the trajectory. The proposed risk factor is demonstrated on real flight data from experimental flights of an octocopter at NASA Langley Research Center in presence of simulated obstacles and wind conditions. Effect of varying wind conditions, distance from obstacles, level of controllability and obstacle measurement noise on the risk factor is demonstrated. The proposed approach enables risk-informed decision making for timely mitigation of current and future unsafe events in autonomous systems.

Portia Banerjee↗

A Bayesian Analysis of SDSS J0914+0853, a Low-mass Dual AGN Candidate

We present the first results from Bayesian AnalYsis of Multiple AGN in X-rays (BAYMAX), a tool that uses a Bayesian framework to quantitatively evaluate whether a given Chandra observation is more likely a single or dual point source. Although the most robust method of determining the presence of dual active galactic nuclei (AGNs) is to use X-ray observations, only sources that are widely separated relative to the instrumentʼs point-spread function are easy to identify. It becomes increasingly difficult to distinguish dual AGNs from single AGNs when the separation is on the order of Chandraʼs angular resolution (<1″). Using likelihood models for single and dual point sources, BAYMAX quantitatively evaluates the likelihood of an AGN for a given source. Specifically, we present results from BAYMAX analyzing the lowest-mass dual AGN candidate to date, SDSS J0914+0853, where archival Chandra data shows a possible secondary AGN ∼ 0"3 from the primary. Analyzing a new 50 ks Chandra observation, results from BAYMAX shows that SDSS J0914+0853 is most likely a single AGN with a Bayes factor of 13.5 in favor of a single point source model. Further, posterior distributions from the dual point source model are consistent with emission from a single AGN. We find a very low probability of SDSS J0914+0853 being a dual AGN system with a flux ratio f>0.3 and separation r>0"3. Overall, BAYMAX will be an important tool for correctly classifying candidate dual AGNs in the literature, as well as studying the dual AGN population where past spatial resolution limits have prevented systematic analyses.

Active galaxies↗

Peter Pan Disks: Long-lived Accretion Disks Around Young M Stars

WISEA J080822.18–644357.3, an M star in the Carina association, exhibits extreme infrared excess and accretion activity at an age greater than the expected accretion disk lifetime. We consider J0808 as the prototypical example of a class of M star accretion disks at ages ≳20 Myr, which we call "Peter Pan" disks, because they apparently refuse to grow up. We present four new Peter Pan disk candidates identified via the Disk Detective citizen science project, coupled with Gaia astrometry. We find that WISEA J044634.16–262756.1 and WISEA J094900.65–713803.1 both exhibit significant infrared excess after accounting for nearby stars within the Two Micron All Sky Survey (2MASS) beams. The J0446 system has >95% likelihood of Columba membership. The J0949 system shows >95% likelihood of Carina membership. We present new Gemini Multi-Object Spectrograph optical spectra of all four objects, showing possible accretion signatures on all four stars. We present ground-based and TESS light curves of J0808 and 2MASS J0501–4337, including a large flare and aperiodic dipping activity on J0808, and strong periodicity on J0501. We find Paβ and Brγ emission indicating ongoing accretion in near-IR spectroscopy of J0808. Using observed characteristics of these systems, we discuss mechanisms that lead to accretion disks at ages ≳20 Myr, and find that these objects most plausibly represent long-lived CO-poor primordial disks, or "hybrid" disks, exhibiting both debris and primordial-disk features. The question remains: why have gas-rich disks persisted so long around these particular stars?

Steven M. Silverberg↗

Probabilistic Blast Damage Modeling Uncertainties and Sensitivities

Blast overpressure is the predominant source of ground damage posed by potentially hazardous asteroid strikes. Estimates of the extent, severity, and likelihoods of potential blast damage regions will be one of the key metrics needed to mount civil defense or disaster response plans in the face of an impending impact. However, there are many inherent sources of uncertainty in evaluating the damage, both in characterizing the properties of the incoming object and in the approaches used to model the entry/impact and resulting damage, which make it difficult to produce a single ‘accurate’ or ‘best guess’ prediction of ground damage. The current 2021 PDC hypothetical impact scenario poses a particular challenge due to its short warning time. The need for rapid disaster response to prepare for an immanent impact, combined with lack of observational opportunities to refine basic knowledge about the object’s basic size and properties, make understanding the range and relative likelihood of consequences particularly critical. The potential damage caused by these blasts can be evaluated using a range of modeling and simulation approaches and levels of fidelity. Fast-running engineering-level models can be used to run large numbers of probabilistically sampled cases covering wide variations of uncertain properties or parameters. High-fidelity simulations, on the other hand, can capture more detailed/accurate blast physics, but can only be performed for a small selection of specific cases, requiring many assumptions to be made about the initial object and its unpredictable entry/breakup characteristics. In order to provide a more complete picture of the potential threat for effective disaster response, both types of analysis need to be employed together. In this approach, high-fidelity simulations are used to refine and anchor engineering models, and the probabilistic engineering models are used to evaluate broad parameters spaces and guide selection of the most pertinent simulation cases for a given scenario. This presentation expands upon the probabilistic asteroid impact risk assessments being performed as part of the 2021 PDC hypothetical impact exercise, focusing on key aspects of blast damage modeling uncertainties and sensitivities. We review the current modeling and simulation approaches employed in the current assessment, compare the relative levels of uncertainty stemming from each main element of the problem (i.e., knowledge of the asteroid properties, modeling of the atmospheric entry/breakup and airburst, and estimates of the ground damage from the resulting blasts waves), and highlight any notable trends and sensitivities for the current scenario case.

SMD↗

Reliability and Safety Assessment of Urban Air Mobility Concept Vehicles

The primary objective of this research effort is to identify failure modes and hazards associated with several configurations of multicopter concept vehicles supplied by NASA. Functional hazard analyses (FHA) and failure modes and effects criticality analyses (FMECA) are performed for each of the eight vehicle configurations under review. Conceptual design of notional powertrain configurations (turboshaft, electric, hybrid electric), notional thrust control systems (rpm control and collective control), and navigation control systems for the concept vehicles were to support the reliability and safety analysis and to assess whether a mission can be completed safely. Two kinds of analyses are performed: static safety analysis which enable the quantification of the likelihood of individual events, and dynamic safety analyses which allows the investigation of multiple time-dependent failures. Their objective is to quantify the likelihood of catastrophic failures.

Urban Air Mobility↗

A NICER View of the Massive Pulsar PSR J0740+6620 Informed by Radio Timing and XMM-Newton Spectroscopy

We report on Bayesian estimation of the radius, mass, and hot surface regions of the massive millisecond pulsar PSR J0740+6620, conditional on pulse-profile modeling of Neutron Star Interior Composition Explorer X-ray Timing Instrument event data. We condition on informative pulsar mass, distance, and orbital inclination priors derived from the joint North American Nanohertz Observatory for Gravitational Waves and Canadian Hydrogen Intensity Mapping Experiment/Pulsar wideband radio timing measurements of Fonseca et al. We use XMM-Newton European Photon Imaging Camera spectroscopic event data to inform our X-ray likelihood function. The prior support of the pulsar radius is truncated at 16 km to ensure coverage of current dense matter models. We assume conservative priors on instrument calibration uncertainty. We constrain the equatorial radius and mass of PSR J0740+6620 to be-+12.390.981.30km and-+2.0720.0660.067Me respectively, each reported as the posterior credible interval bounded by the 16% and 84% quantiles, conditional on surface hot regions that are non-overlapping spherical caps of fully ionized hydrogen atmosphere with uniform effective temperature; a posteriori, the temperature is=-+TlogK5.99100.060.05([])for each hot region. All software for the X-ray modeling framework is open-source and all data, model, and sample information is publicly available, including analysis notebooks and model modules in the Python language. Our marginal likelihood function of mass and equatorial radius is proportional to the marginal joint posterior density of those parameters(within the prior support)and can thus be computed from the posterior samples.

Millisecond pulsars↗

The Cooling Loop A Anomaly of 2013: A Case Study in Human-Systems Resilience

Throughout the history of human spaceflight, NASA has employed an operational paradigm of 24/7 dependence on experts in Mission Control Center (MCC). In addition to nominal flight control and mission operations, these 85+ experts per shift manage anomaly detection, diagnosis, and response, and support the crew in real-time in performing maintenance and repair, procedure execution, and other complex mission operations. Future long-duration exploration missions (LDEMs) beyond low-Earth orbit (LEO) will not operate successfully using this same Human-Systems Integration Architecture (HSIA) where crew rely on ground controllers, have ready access to resupply, and have a fallback plan of evacuation. As distance from Earth increases and the communication delay grows, crews will need to respond independently and adequately to time-critical vehicle malfunctions. It will not always be sufficient or even possible to ‘safe the system’ and then wait upon ground intervention. A new and radically different HSIA is needed to accommodate the paradigm shift of deep-space travel. Historical International Space Station (ISS) data show that for a 30-day mission, the likelihood of a high-consequence vehicle anomaly of uncertain origin that requires rapid response is greater than 10%. The likelihood of such an event is 50% by the fourth month of the mission, and it grows exponentially with time. Our team has conducted in-depth investigations into these events and their corresponding anomaly resolution activities. Using MCC and Mission Evaluation Room (MER) anomaly resolution artifacts (including meeting summaries, caution and warning data, and ISS daily summaries), we created timelines detailing ground actions and in-orbit events for two significant anomalies. We then mapped these timelines onto Mars transit conditions, introducing a ground-crew communications time delay and shifting immediate response, time-critical task execution, and vehicle commanding to the crew. In detailing successful anomaly resolution in transit to Mars, the timelines highlight where effective resolution requires drastically evolved onboard capabilities. Though this research has yielded a rich data set based on ground response in past missions, there is still insufficient knowledge to assess the potential impact of inflight anomalies on a small autonomous crew on future LDEMs beyond LEO. To begin building an evidence base that will inform future HSIA standards and requirements, we are developing an approach to systematically capture crew anomaly response and procedure execution during early Artemis missions. Being the first human spaceflight beyond LEO since Apollo, early Artemis missions provide a rare and unique opportunity to serve as a testbed for Mars missions. Our work aims to capitalize on planned data collection to derive crew operational responses to anomalous events in real-time. Our team is also researching the level of simulation fidelity required for empirically validating proposed HSIA standards and evaluating HSIA implementations for LDEMs beyond LEO. This work will produce a trade space study of HSIA simulation objectives and fidelity requirements. Ultimately, these research efforts will assist in developing the standards and technologies needed to build a next-generation HSIA for LDEMs beyond LEO.

human-systems integration architecture↗

Augmenting Landsat time series with Harmonized Landsat Sentinel-2 data products: Assessment of spectral correspondence

An increase in the temporal revisit of satellite data is often sought to increase the likelihood of obtaining cloud- and shadow-free observations as well as to improve mapping of rapidly- or seasonally-changing features. Currently, as a tandem, Landsat-7 Enhanced Thematic Mapper Plus (ETM+) and −8 Operational Land Imager (OLI) provide an acquisition opportunity on an 8-day revisit interval. Sentinel-2A and -2B MultiSpectral Instrument (MSI), with a wider swath, have a 5-day revisit interval at the equator. Due to robust pre- and post-launch cross-calibration, it has been possible for NASA to produce the Harmonized Landsat Sentinel-2 (HLS) data product from Landsat-8 OLI and Sentinel-2 MSI: L30 and S30, respectively. Knowledge of the agreement of HLS outputs (especially S30) with historic Landsat surface reflectance products will inform the ability to integrate historic time-series information with new and more frequent measures as delivered by HLS. In this research, we control for acquisition date and data source to cross-compare the HLS data (L30, S30) with established Landsat-8 OLI surface-reflectance measures as delivered by the USGS (hereafter BAP, Best Available Pixel). S30 and L30 were found to have high agreement (R = 0.87–0.96) for spectral channels and an r = 0.99 for Normalized Burn Ratio (NBR) with low relative root-mean-square difference values (1.7%–3.3%). Agreement between L30 and BAP was lower, with R values ranging from 0.85 to 0.92 for spectral channels and R = 0.94 for NBR. S30 and BAP had the lowest agreement, with R values ranging from 0.71 to 0.85 for spectral channels and r = 0.90 for NBR. Comparisons indicated a stronger agreement at latitudes above 55° N. Some dependency between spectral agreement and land cover was found, with stronger correspondence for non-vegetated cover types. The level of agreement between S30 and BAP reported herein would enable integration of HLS outputs with historic Landsat data. The resulting increased temporal frequency of data allows for improvements to current cloud screening practices and increases data density and the likelihood of temporal proximity to target date for pixel compositing approaches. Furthermore, additional within-year observations will enable change products with a higher temporal fidelity and allow for the incorporation of phenological trends into land cover classification algorithms.

Michael A. Wulder↗

NASA Physics of Failure (PoF) for Reliability

An item’s reliability or longevity is dependent not only on its design but also on how it is used, manufactured, tested, and the stresses it has or will experience. Stresses include operational and environmental exposures to thermal, voltage, current, age/exposure, mechanical, and radiation mechanisms. Therefore, in reliability analysis, it is important to consider the contributions of all of these factors when predicting the failure rates of components. Historically, there has been a reliance on handbook data (e.g., MIL-HDBK-217), but experience has shown that these values and distributions are not representative of actual performance (1,2). Therefore, to make more credible reliability and risk assessments for its missions, NASA must transition to estimating likelihoods of failure based on an item’s reliability/longevity factors (or the physical susceptibilities and strengths impacting the design’s performance) has or will experience, whenever possible. To facilitate this transition a “Handbook on Methodology for Physics of Failure Based Reliability Assessments” has been developed by NASA to assist in applying physics experiences or experiment physics for empirical analysis and conceptualized physics exposures or theoretical physics for deterministic analysis, to develop and aggregate realistic likelihoods of failure leading to more credible forecasts of item performance and longevity. In addition, since it is NASA’s intention that this document continues to evolve based on community lessons learned and the introduction of new assessment methodologies, NASA is encouraging and appreciates the contributions of current and future authors to maintain and enhance this handbook and its supporting case studies.

Physics of Failure↗

Adding GPU Support to the Markov Chain Monte Carlo Code Catmip

In geophysics, we are confronted with many under-determined inverse problems. For example, all of our observations of earthquakes are made at the Earth’s surface. So, when we try to infer how slip during an earthquake evolves in space and time, we find that there are many potential slip histories that are consistent with our limited observations and our understanding of earthquake physics. One way to approach these problems is with Bayesian analysis which allows us to infer the ensemble of all potential slip models that satisfy the observations and our prior knowledge of earthquake physics. In Bayesian analysis, our prior knowledge is known as the prior probability density function or prior PDF, the fit to the data is known as the data likelihood, and the target PDF that satisfies both the prior PDF and data likelihood is known as the posterior PDF. However, simulating the posterior PDF typically requires using Markov Chain Monte Carlo (MCMC) to draw tens of billions of random realizations of earthquake slip models, which may not be computationally feasible. To make this and similar geophysical inversions computationally tractable, we developed the Cascading Adaptive Transitional Metropolis In Parallel (CATMIP) algorithm. CATMIP is an efficient parallel Markov Chain Monte Carlo (MCMC) sampler that is used for model fitting and uncertainty quantification in geophysics. Example use cases are earthquake rupture modeling, determining mineral composition on Mars, reconstructing the history of ocean salinity, and historical earthquake relocation. CATMIP employs many parallel instances of the Metropolis algorithm for sampling in a transitioning framework. Transitioning is a process in which a set of random samples at equilibrium with a known probability density function (PDF) are used as seeds for the Markov chains to sample successive target PDFs that incrementally move the distribution from the starting seeds to the final desired PDF that describes the relative plausibility of potential values for the model parameters. The algorithm is implemented as a Master-Worker model employing MPI for communication. The worker processes are loosely coupled with global parameters periodically optimized by the master process. This provides a very high amount of parallelism with little communication between updates. During the presentation we will discuss the history of the algorithm and elaborate the earthquake rupture modeling use case for the CATMIP package. Our first step toward GPU optimization was to optimize the code for the CPU. CPU profiling revealed that most of the compute time is spent in calls to level 2 BLAS routines and calls to GSL random number generators. We revised the algorithm to employ level 3 BLAS routines instead. In our presentation we will describe how this was accomplished. Adding GPU support to CATMIP consisted mostly of replacing the calls to GSL with calls to GPU vendor-provided library routines. A small number of loops were directly implemented in CUDA. In the presentation will provide implementation details. Finally, we will discuss methods for profiling and opportunities for further optimizing GPU execution. By creating a code with the flexibility to run on either a CPU or GPU architecture, CATMIP can be used on systems ranging from large CPU-based HPC environments to single servers with GPU acceleration and everything in between.

HECC↗