Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical Computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Cosmic Reionization on Computers: Statistics, Physical Properties, and Environments of Lyman Limit Systems at z ~ 6

Lyman limit systems (LLSs) are dense hydrogen clouds with high enough H i column densities to absorb Lyman continuum photons emitted from distant quasars. Their high column densities imply an origin in dense environments; however, the statistics and distribution of LLSs at high redshifts still remain uncertain. In this paper, we use self-consistent radiative transfer cosmological simulations from the Cosmic Reionization on Computers (CROC) project to study the physical properties of LLSs at the tail end of cosmic reionization at z ~ 6. We generate 3000 synthetic quasar sight lines to obtain a large number of LLS samples in the simulations. In addition, with the high physical fidelity and resolution of CROC, we are able to quantify the association between these LLS samples and nearby galaxies. Our results show that the fraction of LLSs spatially associated with nearby galaxies increases with H i column density. Moreover, we find that LLSs that are not near any galaxy typically reside in filamentary structures connecting neighboring galaxies in the intergalactic medium (IGM). This quantification of the distribution and association of LLSs to large-scale structure informs our understanding of the IGM–galaxy connection during the "Epoch of Reionization," and provides a theoretical basis for interpreting future observations.

79 ASTRONOMY AND ASTROPHYSICS↗

Cosmic Reionization on Computers: Statistical Properties of the Distributions of Mean Opacities

Quasar absorption lines provide a unique window to the relationship between galaxies and the intergalactic medium during the Epoch of Reionization. In particular, high redshift quasars enable measurements of the neutral hydrogen content of the universe. However, the limited sample size of observed quasar spectra, particularly at the highest redshifts, hampers our ability to fully characterize the intergalactic medium during this epoch from observations alone. In this work, we characterize the distributions of mean opacities of the intergalactic medium in simulations from the Cosmic Reionization on Computers (CROC) project. We find that the distribution of mean opacities along sightlines follows a non-trivial distribution that cannot be easily approximated by a known distribution. When comparing the cumulative distribution function of mean opacities measurements in subsamples of sample sizes similar to observational measurements from the literature, we find consistency between CROC and observations at redshifts $z\lesssim 5.7$. However, at higher redshifts ($z\gtrsim5.7$), the cumulative distribution function of mean opacities from CROC is notably narrower than those from observed quasar sightlines implying that observations probe a systematically more opaque intergalactic medium at higher redshifts than the intergalactic medium in CROC boxes at these same redshifts. This is consistent with previous analyses that indicate that the universe is reionized too early in CROC simulations.

79 ASTRONOMY AND ASTROPHYSICS↗

The Data Mine model for accessible partnerships in data science

Abstract The Data Mine at Purdue University is a pioneering experiential learning community for undergraduate and graduate students of any background to learn data science. The first data‐intensive experience embedded in a large learning community, The Data Mine had nearly 1300 students in academic year (AY) 2022–2023 and nearly 1700 students for AY 2023–2024. The Data Mine embodies data‐infused education, research, and collaboration. Students learn Python, R, SQL, and shell‐scripting, while working on weekly projects within a high‐performance computing (HPC) cluster. In the Corporate Partners cohort, students work on teams of 5–15 students, led by a paid student team leader. Each cohort follows an Agile approach, working on data‐intensive projects provided by industry partners and mentored by company employees. Students develop professional and data skills throughout the academic year, from August through April. Many students return in subsequent years to the program, increasing their tenure with a Corporate Partner. Student teams are inherently interdisciplinary; students from 133 different majors are involved in the program, ranging from new incoming students through PhD level students. These interdisciplinary teams of students bring new perspectives to challenging problems in which data science is a key part of the solution. The interdisciplinary teams foster an environment of synthesis with ideas and solutions. Students come together with different life experiences, different levels of technical skill, but also varying ways they navigate paths to solutions because of the variety of majors represented, resulting in a more creative and robust solution than a traditional data science program. This article is categorized under: Applications of Computational Statistics > Education in Computational Statistics

Betz, Margaret A.↗

Localization of infrasonic sources via Bayesian back projection

SUMMARY A Bayesian framework is investigated for event-specific localization of infrasonic sources using back projection ray tracing. Direction-of-arrival information from array-based detection analysis is used to initialize a back projection ray path originating from the detecting array location and quantifying propagation characteristics from hypothetical source locations. The Fisher statistic, computed from the array’s beam coherence, is mapped into uncertainty in the launch angles of the ray path. Auxiliary parameters previously introduced for solving the Transport equation to compute geometric spreading along ray paths are used to map uncertainty in the ray launch angles into spatial and temporal uncertainties in the ray path. An atmospheric ensemble approach is applied to account for atmospheric uncertainty, and the relation between uncertainties in the atmospheric state and confidence in estimated localization are evaluated using several ensembles with specified variances. The method is evaluated using a synthetic event in the western United States constructed via forward propagation simulations as well as a single-station, multi-arrival detection from a surface explosion in the western United States. Localization results using this event-specific approach are more accurate and exhibit improved precision than existing Bayesian localization methods that leverage generalized, pre-computed propagation statistics.

58 GEOSCIENCES↗

Exploring the fragmentation efficiency of proteins analyzed by MALDI-TOF-TOF tandem mass spectrometry using computational and statistical analyses

Matrix-assisted laser desorption/ionization time-of-flight-time-of-flight (MALDI-TOF-TOF) tandem mass spectrometry (MS/MS) is a rapid technique for identifying intact proteins from unfractionated mixtures by top-down proteomic analysis. MS/MS allows isolation of specific intact protein ions prior to fragmentation, allowing fragment ion attribution to a specific precursor ion. However, the fragmentation efficiency of mature, intact protein ions by MS/MS post-source decay (PSD) varies widely, and the biochemical and structural factors of the protein that contribute to it are poorly understood. With the advent of protein structure prediction algorithms such as Alphafold2, we have wider access to protein structures for which no crystal structure exists. In this work, we use a statistical approach to explore the properties of bacterial proteins that can affect their gas phase dissociation via PSD. We extract various protein properties from Alphafold2 predictions and analyze their effect on fragmentation efficiency. Our results show that the fragmentation efficiency from cleavage of the polypeptide backbone on the C-terminal side of glutamic acid (E) and asparagine (N) residues were nearly equal. In addition, we found that the rearrangement and cleavage on the C-terminal side of aspartic acid (D) residues that result from the aspartic acid effect (AAE) were higher than for E- and N-residues. From residue interaction network analysis, we identified several local centrality measures and discussed their implications regarding the AAE. We also confirmed the selective cleavage of the backbone at D-proline bonds in proteins and further extend it to N-proline bonds. Finally, we note an enhancement of the AAE mechanism when the residue on the C-terminal side of D-, E- and N-residues is glycine. To the best of our knowledge, this is the first report of this phenomenon. Our study demonstrates the value of using statistical analyses of protein sequences and their predicted structures to better understand the fragmentation of the intact protein ions in the gas phase.

59 BASIC BIOLOGICAL SCIENCES↗

Uncertainty estimation of bifurcated solutions in the Rayleigh–Bénard problem for advanced nuclear reactors applications

Multiphysics models of nuclear reactors frequently comprise nonlinear systems of equations. The nonlinear nature of these models could lead to solution bifurcations, where a small change in a certain parameter, e.g., the thermophysical properties of the coolant, can lead to a sudden change in the system’s behavior. At the point in parameter space where this happens, called a critical point, the Jacobian matrix of the model’s nonlinear operator becomes singular potentially permitting multiple solutions to coexist. In this paper, we perform uncertainty estimation (UE) in a parameter range that includes bifurcated solutions within the context of Rayleigh–Bénard problem. We perform this analysis assuming uncertain temperature difference, and tilt angle for the iterative solution algorithm with a unit Prandtl number (Pr = 1). Also, we perform this analysis under uncertain thermophysical properties for both FLiBe molten salt and liquid sodium as working fluid. We deploy two approaches to compute statistical moments for the resulting distributions of selected flow-field variables. The first approach is the blind computation of the mean and the standard deviation without any consideration of solution bifurcation, while the second approach utilizes k-means clustering to cluster each branch’s solutions together and compute separate statistical moments for each branch. The statistical distributions are obtained by perturbing the selected parameters about nominal values that correspond to a solution on one of the valid branches, and that solution is used as initial guess for the iterative solution algorithm. We found that perturbation of any parameter when its nominal value is close to its critical point always leads to branch jumping, i.e., the iterations converge to a solution on a branch different from the branch of the initial guess. This produces a statistical ensemble comprised of fundamentally different solutions leading to wrong mean values and uncertainty estimates, whereas clustering provides an efficient way to deal with this type of computation. This work is important for developing Gen IV nuclear systems because many of these systems rely on natural convection for cooling especially in accident conditions.

97 - MATHEMATICS AND COMPUTING↗

Risk Ratio and Risk Difference Estimation in Case-cohort Studies

Background: In case-cohort studies with binary outcomes, ordinary logistic regression analyses have been widely used because of their computational simplicity. However, the resultant odds ratio estimates cannot be interpreted as relative risk measures unless the event rate is low. The risk ratio and risk difference are more favorable outcome measures that are directly interpreted as effect measures without the rare disease assumption. Methods: We provide pseudo-Poisson and pseudo-normal linear regression methods for estimating risk ratios and risk differences in analyses of case-cohort studies. These multivariate regression models are fitted by weighting the inverses of sampling probabilities. Also, the precisions of the risk ratio and risk difference estimators can be improved using auxiliary variable information, specifically by adapting the calibrated or estimated weights, which are readily measured on all samples from the whole cohort. Finally, we provide computational code in R (R Foundation for Statistical Computing, Vienna, Austria) that can easily perform these methods. Results: Through numerical analyses of artificially simulated data and the National Wilms Tumor Study data, accurate risk ratio and risk difference estimates were obtained using the pseudo-Poisson and pseudo-normal linear regression methods. Also, using the auxiliary variable information from the whole cohort, precisions of these estimators were markedly improved. Conclusion: The ordinary logistic regression analyses may provide uninterpretable effect measure estimates, and the risk ratio and risk difference estimation methods are effective alternative approaches for case-cohort studies. These methods are especially recommended under situations in which the event rate is not low.

60 APPLIED LIFE SCIENCES↗

Infrared Fingerprint and Unimolecular Decay Dynamics of the Hydroperoxyalkyl Intermediate (•QOOH) in Cyclopentane Oxidation

A transient carbon-centered hydroperoxyalkyl intermediate (•QOOH) in the oxidation of cyclopentane is identified by IR action spectroscopy with time-resolved unimolecular decay to hydroxyl (OH) radical products that are detected by UV laser-induced fluorescence. Two nearly degenerate •QOOH isomers, β- and γ-QOOH, are generated by H atom abstraction of the cyclopentyl hydroperoxide precursor. Fundamental and first overtone OH stretch transitions and combination bands of •QOOH are observed and compared with anharmonic frequencies computed by second-order vibrational perturbation theory. An OH stretch transition is also observed for a conformer arising from torsion about a low-energy CCOO barrier. Definitive identification of the β-QOOH isomer relies on its significantly lower transition state (TS) barrier to OH products, which results in rapid unimolecular decay and near unity branching to OH products. A benchmarking approach is utilized to compute high-accuracy stationary point energies, most importantly TS barriers, for cyclopentane oxidation (C 5 H 9 O 2 ), building on higher level reference calculations for ethane oxidation (C 2 H 5 O 2 ). The experimental OH product appearance rates are compared with computed statistical microcanonical rates using RRKM theory, including heavy-atom tunneling, thereby validating the computed TS barrier. The results are extended to thermal unimolecular decay rate constants at temperatures and pressures relevant to cyclopentane combustion via master-equation modeling. Furthermore, the various torsional and ring puckering states of the wells and transition states are explicitly considered in these calculations.

Chemical calculations↗

Inexact iterative numerical linear algebra for neural network-based spectral estimation and rare-event prediction

Understanding dynamics in complex systems is challenging because there are many degrees of freedom, and those that are most important for describing events of interest are often not obvious. The leading eigenfunctions of the transition operator are useful for visualization, and they can provide an efficient basis for computing statistics, such as the likelihood and average time of events (predictions). Here, we develop inexact iterative linear algebra methods for computing these eigenfunctions (spectral estimation) and making predictions from a dataset of short trajectories sampled at finite intervals. We demonstrate the methods on a low-dimensional model that facilitates visualization and a high-dimensional model of a biomolecular system. Implications for the prediction problem in reinforcement learning are discussed.

Chemistry↗

Quantum Simulations of Radiation Damage in a Molecular Polyethylene Analog

Abstract An atomic‐level understanding of radiation‐induced damage in simple polymers like polyethylene is essential for determining how these chemical changes can alter the physical and mechanical properties of important technological materials such as plastics. Ensembles of quantum simulations of radiation damage in a polyethylene analog are performed using the Density Functional Tight Binding method to help bind its radiolysis and subsequent degradation as a function of radiation dose. Chemical degradation products are categorized with a graph theory approach, and occurrence rates of unsaturated carbon bond formation, crosslinking, cycle formation, chain scission reactions, and out‐gassing products are computed. Statistical correlations between product pairs show significant correlations between chain scission reactions, unsaturated carbon bond formation, and out‐gassing products, though these correlations decrease with increasing atom recoil energy. The results present relatively simple chemical descriptors as possible indications of network rearrangements in the middle range of excitation energies. Ultimately, the work provides a computational framework for determining the coupling between nonequilibrium chemistry in polymers and potential changes to macro‐scale properties that can aid in the interpretation of future radiation damage experiments on plastic materials.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Mathematical nuances of Gaussian process-driven autonomous experimentation

Abstract The fields of machine learning (ML) and artificial intelligence (AI) have transformed almost every aspect of science and engineering. The excitement for AI/ML methods is in large part due to their perceived novelty, as compared to traditional methods of statistics, computation, and applied mathematics. But clearly, all methods in ML have their foundations in mathematical theories, such as function approximation, uncertainty quantification, and function optimization. Autonomous experimentation is no exception; it is often formulated as a chain of off-the-shelf tools, organized in a closed loop, without emphasis on the intricacies of each algorithm involved. The uncomfortable truth is that the success of any ML endeavor, and this includes autonomous experimentation, strongly depends on the sophistication of the underlying mathematical methods and software that have to allow for enough flexibility to consider functions that are in agreement with particular physical theories. We have observed that standard off-the-shelf tools, used by many in the applied ML community, often hide the underlying complexities and therefore perform poorly. In this paper, we want to give a perspective on the intricate connections between mathematics and ML, with a focus on Gaussian process-driven autonomous experimentation. Although the Gaussian process is a powerful mathematical concept, it has to be implemented and customized correctly for optimal performance. We present several simple toy problems to explore these nuances and highlight the importance of mathematical and statistical rigor in autonomous experimentation and ML. One key takeaway is that ML is not, as many had hoped, a set of agnostic plug-and-play solvers for everyday scientific problems, but instead needs expertise and mastery to be applied successfully. Graphical abstract

97 MATHEMATICS AND COMPUTING↗

Uncertainty Quantification Enabled by Automatic Differentiation for Hydrodynamic Simulation of Shock‐to‐Detonation Transition in High Explosives

Quantifying the effects of uncertainty in a reactive burn model on the run-to-detonation time in high explosives (HEs) provides a robust methodology for assessing the probability of an HE failing the IHE qualification standard. Moreover, uncertainty quantification helps evaluate whether the model calibration accurately represents data outside the calibration set. This study uses a specialized hydrodynamic simulation code for modeling detonation to determine the run-to-detonation time of the HE PBX 9502 for various impact velocities. To quickly approximate uncertainties in the model, a surrogate was constructed using a Taylor series expansion centered at the mean of the input parameters. To obtain the sensitivities required for constructing the Taylor series, HYP-percomplex Automatic Differentiation (HYPAD) was implemented. HYPAD is a methodology for infusing existing codes with automatic differentiation capabilities by augmenting variables with one or more imaginary units to compute step-size independent partial derivatives. These derivatives are accurate to machine precision with respect to the implemented numerical algorithm, meaning their accuracy reflects that of the underlying method (e.g., integration or discretization schemes). Using reduced order modeling techniques, the mean and standard deviation of the run-to-detonation time of a shock within PBX 9502 were computed for a number of initial impact velocities. A weighted least squares regression was then performed to obtain a best fit curve and prediction interval for the computed statistics. Historical data points from explosively driven wedge tests were utilized to validate the prediction interval, ensuring its reliability in predicting future outcomes. With this prediction interval and a known safety constraint curve, the most probable point of failure and the probability of failure for the HE PBX 9502 were determined.

97 MATHEMATICS AND COMPUTING↗

Fundamental limit of jet tagging

Identifying the origin of high-energy hadronic jets (jet tagging) has been a critical benchmark problem for machine learning in particle physics. Jets are ubiquitous at colliders and are complex objects that serve as prototypical examples of collections of particles to be categorized. Over the last decade, machine learning-based classifiers have replaced classical observables as the state of the art in jet tagging. Increasingly complex machine learning models are leading to increasingly more effective tagger performance. Our goal is to address the question of convergence—are we getting close to the fundamental limit on jet tagging or is there still potential for computational, statistical, and physical insights for further improvements? We address this question using state-of-the-art generative models to create a realistic, synthetic dataset with a known jet tagging optimum. Various state-of-the-art taggers are deployed on this dataset, showing that there is a significant gap between their performance and the optimum. Our dataset and software are made public to provide a benchmark task for future developments in jet tagging and other areas of particle physics.

Artificial intelligence↗

Wormholes with ends of the world

We study classical wormhole solutions in 3D gravity with end-of-the-world (EOW) branes, conical defects, kinks, and punctures. These solutions compute statistical averages of an ensemble of boundary conformal field theories (BCFTs) related to universal asymptotics of OPE data extracted from the 2D conformal bootstrap. Conical defects connect BCFT bulk operators; branes join BCFT boundary intervals with identical boundary conditions; kinks (1D defects along branes) link BCFT boundary operators; and punctures (0D defects) are endpoints where conical defects terminate on branes. We provide evidence for a correspondence between the gravity theory and the ensemble. In particular, the agreement of the g-function dependence results from an underlying topological aspect of the on-shell EOW brane action, from which a BCFT analog of the Schlenker-Witten theorem also follows.

AdS-CFT Correspondence↗

Isomer-resolved unimolecular dynamics of the hydroperoxyalkyl intermediate (•QOOH) in cyclohexane oxidation

The oxidation of cycloalkanes is important in the combustion of transportation fuels and in atmospheric secondary organic aerosol formation. A transient carbon-centered radical intermediate (•QOOH) in the oxidation of cyclohexane is identified through its infrared fingerprint and time- and energy-resolved unimolecular dissociation dynamics to hydroxyl (OH) radical and bicyclic ether products. Although the cyclohexyl ring structure leads to three nearly degenerate •QOOH isomers (β-, γ-, and δ-QOOH), their transition state (TS) barriers to OH products are predicted to differ considerably. Selective characterization of the β-QOOH isomer is achieved at excitation energies associated with the lowest TS barrier, resulting in rapid unimolecular decay to OH products that are detected. A benchmarking approach is employed for the calculation of high-accuracy stationary point energies, in particular TS barriers, for cyclohexane oxidation (C 6 H 11 O 2 ), building on higher-level reference calculations for the smaller ethane oxidation (C 2 H 5 O 2 ) system. The isomer-specific characterization of β-QOOH is validated by comparison of experimental OH product appearance rates with computed statistical microcanonical rates, including significant heavy-atom tunneling, at energies in the vicinity of the TS barrier. Master-equation modeling is utilized to extend the results to thermal unimolecular decay rate constants at temperatures and pressures relevant to cyclohexane combustion.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Effects of renormalon scheme and perturbative scale choices on determinations of the strong coupling from e + e − event shapes

We study the role of renormalon cancellation schemes and perturbative scale choices in extractions of the strong coupling constant α s ( m Z ) and the leading nonperturbative shift parameter Ω 1 from resummed predictions of the e + e − event shape thrust. We calculate the thrust distribution to N L 3 L ′ resummed accuracy in soft-collinear effective theory (SCET) matched to the fixed-order O ( α s 2 ) prediction, and perform a new high-statistics computation of the O ( α s 3 ) matching in , although we do not include the latter in our final α s fits due to some observed systematics that require further investigation. We are primarily interested in testing the phenomenological impact sourced from varying amongst three renormalon cancellation schemes and two sets of perturbative scale profile choices. We then perform a global fit to available data spanning center-of-mass energies between 35–207 GeV in each scenario. Relevant subsets of our results are consistent with prior SCET-based extractions of α s ( m Z ) , but we are also led to a number of novel observations. Notably, we find that the combined effect of altering the renormalon cancellation scheme and profile parameters can lead to few-percent-level impacts on the extracted values in the α s − Ω 1 plane, indicating a potentially important systematic theory uncertainty that should be accounted for. We also observe that fits performed over windows dominated by dijet events are typically of a higher quality than those that extend into the far tails of the distributions, possibly motivating future fits focused more heavily in this region. Finally, we discuss how different estimates of the three-loop soft matching coefficient c S ˜ 3 can also lead to measurable changes in the fitted { α s , Ω 1 } values. Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

MultiLoad-GAN: A GAN-Based Synthetic Load Group Generation Method Considering Spatial-Temporal Correlations

This paper presents a deep-learning framework, Multi-load Generative Adversarial Network (MultiLoad-GAN), for generating a group of synthetic load profiles (SLPs) simultaneously. The main contribution of MultiLoad-GAN is the capture of spatial-temporal correlations among a group of loads that are served by the same distribution transformer. This enables the generation of a large amount of correlated SLPs required for microgrid and distribution system studies. Here, the novelty and uniqueness of the MultiLoad-GAN framework are three-fold. First, to the best of our knowledge, this is the first method for generating a group of load profiles bearing realistic spatial- temporal correlations simultaneously. Second, two complementary realisticness metrics for evaluating generated load profiles are developed: computing statistics based on domain knowledge and comparing high-level features via a deep-learning classifier. Third, to tackle data scarcity, a novel iterative data augmentation mechanism is developed to generate training samples for enhancing the training of both the classifier and the MultiLoad-GAN model. Simulation results show that MultiLoad- GAN can generate more realistic load profiles than existing approaches, especially in group level characteristics. With little finetuning, MultiLoad-GAN can be readily extended to generate a group of load or PV profiles for a feeder or a service area.

24 POWER TRANSMISSION AND DISTRIBUTION↗

pnnl/frequency_sensitivity

This research code base includes functions to compute statistics of image distributions in (spatial DFT) frequency space, train deep learning image classifiers on these distributions with variable depth and weight decay, and finally measure the sensitivity of the trained models to perturbations along (spatial DFT) frequency components.

Central, PNNL Developer↗