Search NASA⌕ Search

SEARCH · Search NASA

Results for “MATHEMATICAL STATISTICS”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Bayesian Framework for Bioburden Density Estimation in Planetary Protection

To comply with the international planetary protection policy set forth by the Committee on Space Research and NASA Agency level requirements, spacecraft destined to biologically sensitive planetary bodies have to minimize terrestrial biological contamination. Analysis, testing and inspection are the standard forward verification activities that are used to demonstrate compliance with the biological contamination requirements. For testing of spacecraft surface areas, a swab or wipe sample is collected from surfaces prior to last access and subsequently processed in the lab using NASA Approved Planetary Protection Methods for Culture Based Assays. Raw data resulting from this assay is then statistically treated employing a mathematical paradigm stemming from the 1970’s Viking Lander Project to generate the bioburden density and total microbial bioburden present. This standard approach arbitrarily accounts for error and provides an upper conservative bound as it reports the maximum number of spores estimated to be present on flight hardware surfaces. A bioburden density estimate factors in the following variables: the observed bioburden count, representative volume processed, sampling efficiencies. Notably, to account for error in the approach, a 0 observed count is arbitrarily changed to a count of 1 for each hardware grouping. The data generated by spacecraft bioburden verification campaigns in the past have resulted in <80% of wipes and <90% of swabs containing a bioburden count of 0. As such, having a robust and well documented statistical approach for dealing with the probability of low incident rates is necessary to be able to estimate spacecraft bioburden. Being able to statistically describe the bioburden distribution and associated confidence level is a gamechanger for the development of bioburden allocations during mission design and will allow for tighter management of risk throughout spacecraft build. Thus, Empirical Bayes statistical approach was evaluated to estimate the microbial bioburden on spacecraft to mitigate the aforementioned mathematical concerns and provide a probabilistic bioburden distribution of the flight hardware surface. For application of this approach to performing bioburden calculations, a range of non-informative prior assumptions on hardware surfaces are explored for Bayesian analyses while informative priors using posterior distributions from prior assays are utilized for Empirical Bayes analyses. Several non-informative priors are currently under investigation to assess fitness including use of these priors to serve as a foundation to build off of NASA specification values or a basis of risk to account for unknowns during the integration and testing process. Informative priors under consideration are generated using sampled bioburden values from hardware originating within like processing environments (e.g. vendor cleaning process or similar assembly process), temporal spacecraft status events as a prediction for hardware cleanliness of future samples, and heritage system bioburden actuals to predict allocation for subsequent missions. Informative priors and probabilistic bioburden distributions are then validated using data sets from the Mars Exploration Rover, Mars Science Laboratory, and InSight missions. Using Empirical Bayes approach to generate a probabilistic bioburden distribution as demonstrated through mission use cases provides a valid approach for use in the end-to-end requirements verification process.

97 - MATHEMATICS AND COMPUTING↗

Data-driven particle dynamics: Structure-preserving coarse-graining for emergent behavior in non-equilibrium systems

Multiscale systems are ubiquitous in science and technology, but are notoriously challenging to simulate as short spatiotemporal scales must be appropriately linked to emergent bulk physics. When expensive high-dimensional dynamical systems are coarse-grained into low-dimensional models, the entropic loss of information leads to emergent physics which are dissipative, history-dependent, and stochastic. To machine learn coarse-grained dynamics from time-series observations of particle trajectories, we propose a framework using the metriplectic bracket formalism that preserves these properties by construction; most notably, the framework guarantees discrete notions of the first and second laws of thermodynamics, conservation of momentum, and a discrete fluctuation-dissipation balance crucial for capturing non-equilibrium statistics. We introduce the mathematical framework abstractly before specializing to a particle discretization. As labels are generally unavailable for entropic state variables, we introduce a novel self-supervised learning strategy to identify emergent structural variables. We validate the method on benchmark systems and demonstrate its utility on two challenging examples: (1) coarse-graining star polymers at challenging levels of coarse-graining while preserving non-equilibrium statistics, and (2) learning models from high-speed video of colloidal suspensions that capture coupling between local rearrangement events and emergent stochastic dynamics. We provide open-source implementations in both PyTorch and LAMMPS, enabling large-scale inference and extensibility to diverse particle-based systems.

Computational Engineering, Finance, and Science (c↗

Changes in the Regional Water Cycle and Their Impact on Societies

ABSTRACT Changes in “blue water”, which is the total supply of fresh water available for human extraction over land, are quite closely related to changes in runoff or equivalently precipitation minus evaporation, . This article examines how climate change‐driven recent past and future changes in the regional water cycle relate to blue water availability and changes in human blue water demand. Although at the largest scales theoretical and numerical model predictions are in broad agreement with observations, at continental scales and below models predict large ranges of possible future and runoff especially at the scale of individual river catchments and for shorter timescale subseasonal floods and droughts. Nevertheless, it is expected that the occurrence and severity of floods will increase and that of droughts may increase, possibly compounded by human‐driven non‐climatic changes such as changes in land use, dam water impoundment, irrigation and extraction of groundwater. Contemporary assessments predict that increases in 21st century human water extraction in many highly‐populated regions are unlikely to be sustainable given projections of future . To reduce uncertainty in future predictions, there is an urgent need to improve modeling of atmospheric, land surface and human processes and how these components are coupled. This should be supported by maintaining the observing network and expanding it to improve measurements of land surface, oceanic and atmospheric variables. This includes the development of satellite observations stable over multiple decades and suitable for building reanalysis datasets appropriate for model evaluation.

54 ENVIRONMENTAL SCIENCES↗

Higher‐Order Gravity Waves and Traveling Ionospheric Disturbances From the Polar Vortex Jet on 11–15 January 2016: Modeling With HIAMCM‐SAMI3 and Comparison With Observations in the Thermosphere and Ionosphere

Abstract In Vadas et al. (2024, https://doi.org/10.1029/2024ja032521 ), we modeled the atmospheric gravity waves (GWs) during 11–14 January 2016 using the HIAMCM, and found that the polar vortex jet generates medium to large‐scale, higher‐order GWs in the thermosphere. In this paper, we model the traveling ionospheric disturbances (TIDs) generated by these GWs using the HIAMCM‐SAMI3 and compare with ionospheric observations from ground‐based Global Navigation Satellite System (GNSS) receivers, Incoherent Scatter Radars (ISR) and the Super Dual Auroral Radar Network (SuperDARN). We find that medium to large‐scale TIDs are generated worldwide by the higher‐order GWs from this event. Many of the TIDs over Europe and Asia have concentric ring/arc‐like structure, and most of those over North/South America have planar wave structure and occur during the daytime. Those over North/South America propagate southward and are generated by higher‐order GWs from Europe/Asia which propagate over the Arctic. These latter TIDs can be misidentified as arising from geomagnetic forcing. We find that the higher‐order GWs that propagate to Africa and Brazil from Europe may aid in the formation of equatorial plasma bubbles (EPBs) there. We find that the simulated GWs, TIDs and EPBs agree with EISCAT, PFISR, GNSS, and SuperDARN measurements. We find that the higher‐order GWs are concentrated at N at 200 km, in agreement with GOCE and CHAMP data. Thus the polar vortex jet is important for generating TIDs in the northern winter ionosphere via multi‐step vertical coupling through GWs.

Vadas, Sharon L. [Northwest Research Associates Bo↗

Robust error calibration for serial crystallography

Serial crystallography is an important technique with unique abilities to resolve enzymatic transition states, minimize radiation damage to sensitive metalloenzymes and perform de novo structure determination from micrometre-sized crystals. This technique requires the merging of data from thousands of crystals, making manual identification of errant crystals unfeasible. cctbx.xfel.merge uses filtering to remove problematic data. However, this process is imperfect, and data reduction must be robust to outliers. We add robustness to cctbx.xfel.merge at the step of uncertainty determination for reflection intensities. This step is a critical point for robustness because it is the first step where the data sets are considered as a whole, as opposed to individual lattices. Robustness is conferred by reformulating the error-calibration procedure to have fewer and less stringent statistical assumptions and incorporating the ability to down-weight low-quality lattices. We then apply this method to five macromolecular XFEL data sets and observe the improvements to each. The appropriateness of the intensity uncertainties is demonstrated through internal consistency. This is performed through theoretical CC 1/2 and I /σ relationships and by weighted second moments, which use Wilson's prior to connect intensity uncertainties with their expected distribution. This work presents new mathematical tools to analyze intensity statistics and demonstrates their effectiveness through the often underappreciated process of uncertainty analysis.

Mittan-Moreau, David W.↗

Mapping Incidence and Prevalence Peak Data for SIR Modeling Applications

Infectious disease modeling and forecasting have played a key role in helping assess and respond to epidemics and pandemics. Recent work has leveraged data on disease peak infection and peak hospital incidence to fit compartmental models for the purpose of forecasting and describing the dynamics of a disease outbreak. Incorporating these data can greatly stabilize a compartmental model fit on early observations, where slight perturbations in the data may lead to model fits that forecast wildly unrealistic peak infection. We introduce a new method for incorporating historic data on the value and time of peak incidence of hospitalization into the fit for a Susceptible-Infectious-Recovered (SIR) model by formulating the relationship between an SIR model’s starting parameters and peak incidence as a system of two equations that can be solved computationally. We demonstrate how to calculate SIR parameter estimates – which describe disease dynamics such as transmission and recovery rates – using this method, and determine that there is a noticeable loss in accuracy whenever prevalence data is misspecified as incidence data. To exhibit the modeling potential, we update the Dirichlet-Beta State Space modeling framework to use hospital incidence data, as this framework was previously formulated to incorporate only data on total infections. This approach is assessed for practicality in terms of accuracy and speed of computation via simulation.

97 MATHEMATICS AND COMPUTING↗

Data-driven prediction of scaling and ignition of inertial confinement fusion experiments

Recent advances in inertial confinement fusion (ICF) at the National Ignition Facility (NIF), including ignition and energy gain, are enabled by a close coupling between experiments and high-fidelity simulations. Neither simulations nor experiments can fully constrain the behavior of ICF implosions on their own, meaning pre- and postshot simulation studies must incorporate experimental data to be reliable. Linking past data with simulations to make predictions for upcoming designs and quantifying the uncertainty in those predictions has been an ongoing challenge in ICF research. We have developed a data-driven approach to prediction and uncertainty quantification that combines large ensembles of simulations with Bayesian inference and deep learning. The approach builds a predictive model for the statistical distribution of key performance parameters, which is jointly informed by past experiments and physics simulations. The prediction distribution captures the impact of experimental uncertainty, expert priors, design changes, and shot-to-shot variations. We have used this new capability to predict a 10× increase in ignition probability between Hybrid-E shots driven with 2.05 MJ compared to 1.9 MJ, and validated our predictions against subsequent experiments. We describe our new Bayesian postshot and prediction capabilities, discuss their application to NIF ignition and validate the results, and finally investigate the impact of data sparsity on our prediction results.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Stochastic Error Cancellation in Analog Quantum Simulation

Analog quantum simulation is a promising path towards solving classically intractable problems in many-body physics on near-term quantum devices. However, the presence of noise limits the size of the system and the length of time that can be simulated. In our work, we consider an error model in which the actual Hamiltonian of the simulator differs from the target Hamiltonian we want to simulate by small local perturbations, which are assumed to be random and unbiased. We analyze the error accumulated in observables in this setting and show that, due to stochastic error cancellation, with high probability the error scales as the square root of the number of qubits instead of linearly. We explore the concentration phenomenon of this error as well as its implications for local observables in the thermodynamic limit. Moreover, we show that stochastic error cancellation also manifests in the fidelity between the target state at the end of time-evolution and the actual state we obtain in the presence of noise. This indicates that, to reach a certain fidelity, more noise can be tolerated than implied by the worst-case bound if the noise comes from many statistically independent sources.

Analog quantum simulation↗

Bayesian Adaptive Polynomial Chaos Expansions

Polynomial chaos expansions (PCEs) are widely used for uncertainty quantification (UQ) tasks, particularly in the applied mathematics community. However, PCE has received comparatively less attention in the statistics literature, and fully Bayesian formulations remain rare—especially with implementations in R. Motivated by the success of adaptive Bayesian machine learning models such as BART, BASS and BPPR, we develop a new fully Bayesian adaptive PCE method with an efficient and accessible R implementation: khaos. Our approach includes a novel proposal distribution that enables data-driven interaction selection and supports a modified g-prior tailored to PCE structure. Through simulation studies and real-world UQ applications, we demonstrate that the Bayesian adaptive PCE provides competitive performance for surrogate modeling, global sensitivity analysis and ordinal regression tasks.

97 MATHEMATICS AND COMPUTING↗

Quantum mechanical closure of partial differential equations with symmetries

We develop a statistical framework for the dynamical closure of spatiotemporal dynamics governed by partial differential equations. Employing the mathematical framework of quantum mechanics to embed the original classical dynamics into a quantum mechanical representation, we use the space of quantum density operators to model the unresolved degrees of freedom of the original dynamics in a statistical sense, and the framework of quantum measurement to predict their contributions to the resolved dynamics. The embedded dynamics is discretized by a positivity preserving process, leading to a compressed representation that is invariant under the dynamical symmetries of the resolved dynamics. We present a data based formulation of the closure scheme and apply it to a closure problem for the shallow water equations. The numerical results demonstrate that our closure model can accurately predict the main features of the true dynamics, including for out of sample initial conditions.

Delay embedding↗

Simultaneous inference of equation of state parameters and unknown data errors with uncertainty quantification via hierarchical Bayesian posterior maximization

Equations of state (EOSs) are a key component in running hydrodynamic simulations as they relate the thermodynamic states for the material. The Davis reactants EOS is commonly used for modeling high explosives (HEs), and the EOS model parameters are calibrated using material specific data. The calibrations are often performed with uncertainty quantification via Bayesian inference to account for uncertainty in the data and generate ensembles of likely parameters. However, there are relatively few HE data sets to use for calibration and many are historical and lack error information. In this work, we simultaneously calibrate the Davis reactants EOS model parameters and unknown data error terms for the high explosive PBX 9501. To quantify the uncertainty in the models and the data, we use a Bayesian framework for the calibration and compute the hierarchical Bayesian posterior distribution with both a posteriori maximization approach and Markov Chain Monte Carlo. In general, we find that, given our assumptions, the two approaches result in similar calibrated parameters, posterior covariance matrices, and insights about the parameters but that the posterior maximization requires far less computational resources.

97 MATHEMATICS AND COMPUTING↗

Modern chemical graph theory

Abstract Graph theory has a long history in chemistry. Yet as the breadth and variety of chemical data is rapidly changing, so too do graph encoding methods and analyses that yield qualitative and quantitative insights. Using illustrative cases within a basic mathematical framework, we showcase modern chemical graph theory's utility in Chemists' analysis and model development toolkit. The encoding of both experimental and simulation data is discussed at various levels of granularity of information. This is followed by a discussion of the two major classes of graph theoretical analyses: identifying connectivity patterns and partitioning methods. Measures, metrics, descriptors, and topological indices are then introduced with an emphasis upon enhancing interpretability and incorporation into physical models. Challenging data cases are described that include strategies for studying time dependence. Throughout, we incorporate recent advancements in computer science and applied mathematics that are propelling chemical graph theory into new domains of chemical study. This article is categorized under: Molecular and Statistical Mechanics > Molecular Dynamics and Monte‐Carlo Methods Structure and Mechanism > Computational Materials Science Structure and Mechanism > Molecular Structures

Leite, Leonardo S. G.↗

Data-driven Mori–Zwanzig modeling of Lagrangian particle dynamics in turbulent flows

The dynamics of Lagrangian particles in turbulence play a crucial role in mixing, transport, and dispersion in complex flows. Their trajectories exhibit highly nontrivial statistical behavior, motivating the development of surrogate models that can reproduce these trajectories without incurring the high computational cost of direct numerical simulations of the full Eulerian field. This task is particularly challenging because reduced-order models typically lack access to the full set of interactions with the underlying turbulent field. Novel data-driven machine learning techniques can be powerful in capturing and reproducing complex statistics of the reduced-order/surrogate dynamics. In this work, we show how one can learn a surrogate dynamical system that is able to evolve a turbulent Lagrangian trajectory in a way that is point-wise accurate for short-time predictions (with respect to Kolmogorov time) and stable and statistically accurate at long times. This approach is based on the Mori–Zwanzig formalism, which prescribes a mathematical decomposition of the full dynamical system into resolved dynamics that depend on the current state and the past history of a reduced set of observables, and the unresolved orthogonal dynamics due to unresolved degrees of freedom of the initial state. We show how by training this reduced order model on a point-wise error metric on short time-prediction, we are able to correctly learn the dynamics of Lagrangian turbulence, such that also the long-time statistical behavior is stably recovered at test time. This opens up a range of applications, for example, for the control of active Lagrangian agents in turbulence.

97 MATHEMATICS AND COMPUTING↗

On the statistical theory of self-gravitating collisionless dark matter flow: High order kinematic and dynamic relations

Dark matter, if it exists, accounts for five times as much as ordinary baryonic matter. To better understand the self-gravitating collisionless dark matter flow on different scales, a statistical theory involving kinematic and dynamic relations must be developed for different types of flow, e.g., incompressible, constant divergence, and irrotational flow. This is mathematically challenging because of the intrinsic complexity of dark matter flow and the lack of a self-closed description of flow velocity. Here, this paper extends our previous work on second-order statistics Xu to kinematic relations of any order for any type of flow. Dynamic relations were also developed to relate statistical measures of different orders. The results were validated by N-body simulations. On large scales, we found that (i) third-order velocity correlations can be related to density correlation or pairwise velocity; (ii) the pth-order velocity correlations follow ∝ a (p+2)/2 for odd p and ∝ a p/2 for even p, where a is the scale factor; (iii) the overdensity δ is proportional to density correlation on the same scale, $\langle$δ$\rangle$∝$\langle$δδ'$\rangle$; (iv) velocity dispersion on a given scale r is proportional to the overdensity on the same scale. On small scales, (i) a self-closed velocity evolution is developed by decomposing the velocity into motion in haloes and motion of haloes; (ii) the evolution of vorticity and enstrophy are derived from the evolution of velocity; (iii) dynamic relations are derived to relate second- and third-order correlations; (iv) while the first moment of pairwise velocity follows $\langle$Δu L $\rangle$=-Har (H is the Hubble parameter), the third moment follows $\langle$(Δu L ) 3 $\rangle$ ∝ ε u ar that can be directly compared with simulations and observations, where ε u ≈ 10 -7 m 2 /s 3 is the constant rate for energy cascade; (v) the pth order velocity correlations follow ∝ a (3p-5)/4 for odd p and ∝ a 3p/4 for even p. Finally, the combined kinematic and dynamic relations lead to exponential and one-fourth power-law velocity correlations on large and small scales, respectively.

79 ASTRONOMY AND ASTROPHYSICS↗

Long-Term Statistical Process Monitoring of an Ultrafiltration Water Treatment Process

As water treatment technology has improved, the amount of available process data has substantially increased, making real-time, data-driven fault detection a reality. One shortcoming of the fault detection literature is that methods are usually evaluated by comparing their performance on hand-picked, short-term case studies, which yields no insight into long-term performance. In this work, we first evaluate multiple statistical and machine learning approaches for detrending process data. Then, we evaluate the performance of a PCA-based fault detection approach, applied to the detrended data, to monitor influent water quality, filtrate quality, and membrane fouling of an ultrafiltration membrane system for indirect potable reuse. Based on two short case studies, the adaptive lasso detrending method is selected, and the performance of the multivariate approach is evaluated over more than a year. The method is tested for different sets of three critical tuning parameters, and we find that for long-term, autonomous monitoring to be successful, these parameters should be carefully evaluated. However, in comparison with industry standards of simpler, univariate monitoring or daily pressure decay tests, multivariate monitoring produces substantial benefits in long-term testing.

ammonia↗

Elliptically-Contoured Tensor-variate Distributions with Application to Image Learning

Statistical analysis of tensor-valued data has largely used the tensor-variate normal (TVN) distribution that may be inadequate for data arising from distributions with heavier or lighter tails. We study a general family of elliptically contoured (EC) TV distributions and derive its characterizations, moments, marginal, and conditional distributions. We describe procedures for maximum likelihood estimation from data that are (1) uncorrelated draws from an EC distribution, (2) from a scale mixture of the TVN distribution, and (3) from an underlying but unknown EC distribution, for which we extend Tyler’s robust estimator. A detailed simulation study highlights the benefits of choosing an EC distribution over the TVN for heavier-tailed data. We develop TV classification rules using discriminant analysis and EC errors and show that they better predict cats and dogs from images in the Animal Faces-HQ dataset than the TVN-based rules. A novel tensor-on-tensor regression and TV analysis of variance (TANOVA) framework under EC errors is also demonstrated to better characterize gender, age, and ethnic origin than the usual TVN-based TANOVA in the celebrated labeled faces of the wild dataset.

97 MATHEMATICS AND COMPUTING↗

Using data-science approaches to unravel insights for enhanced transport of lithium ions in single-ion conducting polymer electrolyte

Solid polymer electrolytes have yet to achieve the an ionic conductivity > 1 mS/cm at room temperature for realistic applications. This target implies the need to reduce the effective energy barriers of ion transport in polymer electrolytes to around 20 kJ/mol. In this work, we combine information extracted from existing experimental results with theoretical calculations to provide insights into ion transport in single-ion conductors (SICs) with a focus on lithium ion SICs. Through the analysis of temperature-dependent ionic conductivity data obtained from the literature, we evaluate different methods of extracting energy barriers for lithium transport. The traditional Arrhenius fit to the temperature-dependent ionic conductivity data indicates that the Meyer-Neldel rule holds for SICs. However, the values of the fitting parameters remain unphysical. Our modified approach based on recent work (Macromolecules, 56, 15, 6051(2023)), which incorporates a fixed pre-exponential factor, reveals that the energy barriers exhibit temperature dependence over a wide range of temperatures. Using this approach, we identify a series of anions leading to the energy barriers less than 30 kJ/mol, which include trifluoromethane sulfonimide (TFSI), fluoromethane sulfonimide (FSI), and boron-based organic anions. In our efforts to design the next generation of anions, which can exhibit the energy barriers less than 20 kJ/mol, we focused on boron-containing SICs, and performed density functional theory (DFT) based calculations to connect the chemical structures via the binding energy of cation (lithium)-anion pairs with the experimentally derived effective energy barriers for ion transport. Not only have we identified a correlation between the binding energy and the energy barriers, but we also propose a strategy to design new boron-based anions by using the correlation. This combined approach involving experiments and theoretical calculations is capable of facilitating the identification of promising new anions, which can exhibit ionic conductivity $> 1$ mS/cm near room temperature, thereby expediting the development of novel superionic single-ion conducting polymer electrolytes. The published datasets include all the temperature-dependent ionic conductivity collected from the literature with literature DOIs, DFT calculated binding energies, and python scripts to analyze data, construct statistical models, and generate plots.

36 MATERIALS SCIENCE↗