Search NASA⌕ Search

SEARCH · Search NASA

Results for “high dimensional statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Data-driven high-dimensional statistical inference with generative models

Crucial to many measurements at the LHC is the use of correlated multi-dimensional information to distinguish rare processes from large backgrounds, which is complicated by the poor modeling of many of the crucial backgrounds in Monte Carlo simulations. In this work, we introduce HI-SIGMA, a method to perform unbinned high-dimensional statistical inference with data-driven background distributions. In contradistinction to many applications of Simulation Based Inference in High Energy Physics, HI-SIGMA relies on generative ML models, rather than classifiers, to learn the signal and background distributions in the high-dimensional space. These ML models allow for interpretable inference while also incorporating model errors and other sources of systematic uncertainties. We showcase this methodology on a simplified version of a di-Higgs measurement in the bbγγ final state, where the di-photon resonance allows for background interpolation from sidebands into the signal region. We demonstrate that HI-SIGMA provides improved sensitivity as compared to standard classifier-based methods, and that systematic uncertainties can be straightforwardly incorporated by extending methods which have been used for histogram based analyses.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Hierarchical Bayesian Inverse Problems: A High-Dimensional Statistics Viewpoint

This paper analyzes hierarchical Bayesian inverse problems using techniques from highdimensional statistics. Furthermore, our analysis leverages a property of hierarchical Bayesian regularizers that we call approximate decomposability to obtain non-asymptotic bounds on the reconstruction error attained by maximum a posteriori estimators. The new theory explains how hierarchical Bayesian models that exploit sparsity, group sparsity, and sparse representations of the unknown parameter can achieve accurate reconstructions in high-dimensional settings.

MAP estimation↗

Ambient-temperature liquid jet targets for high-repetition-rate HED discovery science

High-power lasers can generate energetic particle beams and astrophysically relevant pressure and temperature states in the high-energy-density (HED) regime. Recently-commissioned high-repetition-rate (HRR) laser drivers are capable of producing these conditions at rates exceeding 1 Hz. However, experimental output from these systems is often limited by the difficulty of designing targets that match these repetition rates. To overcome this challenge, we have developed tungsten microfluidic nozzles, which produce a continuously replenishing jet that operates at flow speeds of approximately 10 m/s and can sustain shot frequencies up to 1 kHz. The ambient-temperature planar liquid jets produced by these nozzles can have thicknesses ranging from hundreds of nanometers to tens of micrometers. In this work, we illustrate the operational principle of the microfluidic nozzle and describe its implementation in a vacuum environment. Further, we provide evidence of successful laser-driven ion acceleration using this target and discuss the prospect of optimizing the ion acceleration performance through an in situ jet thickness scan. Future applications for the jet throughout HED science include shock compression and studies of strongly heated nonequilibrium plasmas. When fielded in concert with HRR-compatible laser, diagnostic, and active feedback technology, this target will facilitate advanced automated studies in HRR HED science, including machine learning-based optimization and high-dimensional statistical analysis.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

The galactic globular cluster system

We explore correlations between various properties of Galactic globular clusters, using a database on 143 objects. Our goal is identify correlations and trends which can be used to test and constrain theoretical models of cluster formation and evolution. We use a set of 13 cluster parameters, 9 of which are independently measured. Several arguments suggest that the number of clusters still missing in the obscured regions of the Galaxy is of the order of 10, and thus the selection effects are probably not severe for our sample. Known clusters follow a power-law density distribution with a slope approximately -3.5 to -4, and an apparent core with a core radius approximately 1 kpc. Clusters show a large dynamical range in many of their properties, more so for the core parameters (which are presumably more affected by dynamical evolution) than for the half-light parameters. There are no good correlations with luminosity, although more luminous clusters tend to be more concentrated. When data are binned in luminosity, several trends emerge: more luminous clusters tend to have smaller and denser cores. We interpret this as a differential survival effect, with more massive clusters surviving longer and reaching more evolved dynamical states. Cluster core parameters and concentrations also correlate with the position in the Galaxy, with clusters closer to the Galactic center or plane being more concentrated and having smaller and denser cores. These trends are more pronounced for the fainter (less massive) clusters. This is in agreement with a picture where tidal shocks form disk or bulge passages accelerate dynamical evolution of clusters. Cluster metallicities do not correlate with any other parameter, including luminosity and velocity dispersion; the only detectable trend is with the position in the Galaxy, probably reflecting Zinn's disk-halo dichotomy. This suggests that globular clusters were not self-enriched systems. Velocity dispersions show excellent correlations with luminosity and surface brightness. Their origin is not well understood, but they may well reflect initial conditions of cluster formation, and perhaps even be used to probe the initial density perturbation spectrum on a approximately 10(exp 6) solar mass scale. Core radii and concentrations play a role of a 'second parameter' in these correlations. While a global manifold of cluster properties has a high statistical dimensionality (D greater than 4), a subset of structural, photometric, and dynamical parameters forms a statistically three-dimensional family, as expected from objects following King models; we propose to call this set of quantities the King Manifold. Some of the observed correlations may be usable as distance indicator relations for globular clusters.

Djorgovski, S.↗

Analyzing High-Dimensional Multispectral Data

In this paper, through a series of specific examples, we illustrate some characteristics encountered in analyzing high- dimensional multispectral data. The increased importance of the second-order statistics in analyzing high-dimensional data is illustrated, as is the shortcoming of classifiers such as the minimum distance classifier which rely on first-order variations alone. We also illustrate how inaccurate estimation or first- and second-order statistics, e.g., from use of training sets which are too small, affects the performance of a classifier. Recognizing the importance of second-order statistics on the one hand, but the increased difficulty in perceiving and comprehending information present in statistics derived from high-dimensional data on the other, we propose a method to aid visualization of high-dimensional statistics using a color coding scheme.

Lee, Chulhee↗

Analyzing high dimensional data

Problems encountered in analyzing high dimensional data are discussed and possible solutions are proposed. The increased importance of second-order statistics in analyzing high dimensional data and the shortcoming of the minimum distance classifier in high dimensional data are recognized. By investigating characteristics of high dimensional data, it is shown that second-order statistics must be taken into account in high dimensional data. There is a need to represent second order statistics effectively. As the data dimensionality increases, it becomes more difficult to perceive and compare information present in statistics derived from data. In order to overcome this problem, a method to visualize statistics using color code is proposed. By representing statistics using a color code, the first and the second statistics can be more readily compared.

Lee, Chulhee↗

Feature extraction and classification algorithms for high dimensional data

Feature extraction and classification algorithms for high dimensional data are investigated. Developments with regard to sensors for Earth observation are moving in the direction of providing much higher dimensional multispectral imagery than is now possible. In analyzing such high dimensional data, processing time becomes an important factor. With large increases in dimensionality and the number of classes, processing time will increase significantly. To address this problem, a multistage classification scheme is proposed which reduces the processing time substantially by eliminating unlikely classes from further consideration at each stage. Several truncation criteria are developed and the relationship between thresholds and the error caused by the truncation is investigated. Next an approach to feature extraction for classification is proposed based directly on the decision boundaries. It is shown that all the features needed for classification can be extracted from decision boundaries. A characteristic of the proposed method arises by noting that only a portion of the decision boundary is effective in discriminating between classes, and the concept of the effective decision boundary is introduced. The proposed feature extraction algorithm has several desirable properties: it predicts the minimum number of features necessary to achieve the same classification accuracy as in the original space for a given pattern recognition problem; and it finds the necessary feature vectors. The proposed algorithm does not deteriorate under the circumstances of equal means or equal covariances as some previous algorithms do. In addition, the decision boundary feature extraction algorithm can be used both for parametric and non-parametric classifiers. Finally, some problems encountered in analyzing high dimensional data are studied and possible solutions are proposed. First, the increased importance of the second order statistics in analyzing high dimensional data is recognized. By investigating the characteristics of high dimensional data, the reason why the second order statistics must be taken into account in high dimensional data is suggested. Recognizing the importance of the second order statistics, there is a need to represent the second order statistics. A method to visualize statistics using a color code is proposed. By representing statistics using color coding, one can easily extract and compare the first and the second statistics.

Lee, Chulhee↗

Convex relaxation for Fokker–Planck equation

We propose an approach to directly estimate the moments or marginals for a high-dimensional equilibrium distribution in statistical mechanics by solving the high-dimensional Fokker–Planck equation in terms of low-order cluster moments or marginals. With this approach, we bypass the exponential complexity of estimating the full high-dimensional distribution and directly solve the simplified partial differential equations for low-order moments/marginals. Moreover, the proposed moment/marginal relaxation is fully convex and can be solved via off-the-shelf solvers. We further propose a time-dependent version of the convex programs to study non-equilibrium dynamics. In a specific setting, we show the proposed method can recover a mean-field-type equilibrium density. Numerical results are provided to demonstrate the performance of the proposed algorithm for high-dimensional systems.

Chen, Yian↗

Kronecker-structured covariance models for multiway data

Many applications produce multiway data of exceedingly high dimension. Modeling such multi-way data is important in multichannel signal and video processing where sensors produce multi-indexed data, e.g. over spatial, frequency, and temporal dimensions. We will address the challenges of covariance representation of multiway data and review some of the progress in statistical modeling of multiway covariance over the past two decades, focusing on tensor-valued covariance models and their inference. We will illustrate through a space weather application: predicting the evolution of solar active regions over time.

97 MATHEMATICS AND COMPUTING↗

Supervised Classification Techniques for Hyperspectral Data

The recent development of more sophisticated remote sensing systems enables the measurement of radiation in many mm-e spectral intervals than previous possible. An example of this technology is the AVIRIS system, which collects image data in 220 bands. The increased dimensionality of such hyperspectral data provides a challenge to the current techniques for analyzing such data. Human experience in three dimensional space tends to mislead one's intuition of geometrical and statistical properties in high dimensional space, properties which must guide our choices in the data analysis process. In this paper high dimensional space properties are mentioned with their implication for high dimensional data analysis in order to illuminate the next steps that need to be taken for the next generation of hyperspectral data classifiers.

Jimenez, Luis O.↗

A Localized Ensemble Kalman Smoother

Numerous geophysical inverse problems prove difficult because the available measurements are indirectly related to the underlying unknown dynamic state and the physics governing the system may involve imperfect models or unobserved parameters. Data assimilation addresses these difficulties by combining the measurements and physical knowledge. The main challenge in such problems usually involves their high dimensionality and the standard statistical methods prove computationally intractable. This paper develops and addresses the theoretical convergence of a new high-dimensional Monte-Carlo approach called the localized ensemble Kalman smoother.

recursive estimation↗

System and Safety Analysis with SysAI A Statistical Learning Framework

This is a tutorial on how to use the SYSAI (System Analysis using Statistical AI), a flexible statistical learning framework for the V&V and analysis of complex and high-dimensional Aerospace systems with DNN and AI components. SYSAI provides functionality for a variety of analyses and V&V tasks, including statistical data analysis, high dimensional safety-envelope and time-series analysis, property checking, as well as intelligent test-case generation. The tutorial will demonstrate SYSAI with our industrial partner’s Autonomous Centerline Tracking system, which uses a DNN to enable autonomous aircraft taxiing as an example. Video & Tutorial

Statistical V&V for Complex safety-critical system↗

A flexible class of priors for orthonormal matrices with basis function-specific structure

Statistical modeling of high-dimensional matrix-valued data motivates the use of a low-rank representation that simultaneously summarizes key characteristics of the data and enables dimension reduction. Low-rank representations commonly factor the original data into the product of orthonormal basis functions and weights, where each basis function represents an independent feature of the data. However, the basis functions in these factorizations are typically computed using algorithmic methods that cannot quantify uncertainty or account for basis function correlation structure a priori. While there exist Bayesian methods that allow for a common correlation structure across basis functions, empirical examples motivate the need for basis function-specific dependence structure. We propose a prior distribution for orthonormal matrices that can explicitly model basis function-specific structure. The prior is used within a general probabilistic model for singular value decomposition to conduct posterior inference on the basis functions while accounting for measurement error and fixed effects. We discuss how the prior specification can be used for various scenarios and demonstrate favorable model properties through synthetic data examples. Finally, we apply our method to two-meter air temperature data from the Pacific Northwest, enhancing our understanding of the Earth system’s internal variability.

97 MATHEMATICS AND COMPUTING↗

Statistical mechanics of light elements at high pressure. V Three-dimensional Thomas-Fermi-Dirac theory

A numerical technique for solving the Thomas-Fermi-Dirac (TED) equation in three dimensions, for an array of ions obeying periodic boundary conditions, is presented. The technique is then used to calculate deviations from ideal mixing for an alloy of hydrogen and helium at zero temperature and high presures. Results are compared with alternative models which apply perturbation theory to calculation of the electron distribution, based upon the assumption of weak response of the electron gas to the ions. The TFD theory, which permits strong electron response, always predicts smaller deviations from ideal mixing than would be predicted by perturbation theory. The results indicate that predicted phase separation curves for hydrogen-helium alloys under conditions prevailing in the metallic zones of Jupiter and Saturn are very model dependent.

Macfarlane, J. J.↗

An implementation of neural simulation-based inference for parameter estimation in ATLAS

Neural simulation-based inference (NSBI) is a powerful class of machine-learning-based methods for statistical inference that naturally handles high-dimensional parameter estimation without the need to bin data into low-dimensional summary histograms. Such methods are promising for a range of measurements, including at the Large Hadron Collider, where no single observable may be optimal to scan over the entire theoretical phase space under consideration, or where binning data into histograms could result in a loss of sensitivity. This work develops a NSBI framework for statistical inference, using neural networks to estimate probability density ratios, which enables the application to a full-scale analysis. It incorporates a large number of systematic uncertainties, quantifies the uncertainty due to the finite number of events in training samples, develops a method to construct confidence intervals, and demonstrates a series of intermediate diagnostic checks that can be performed to validate the robustness of the method. As an example, the power and feasibility of the method are assessed on simulated data for a simplified version of an off-shell Higgs boson couplings measurement in the four-lepton final states. This approach represents an extension to the standard statistical methodology used by the experiments at the Large Hadron Collider, and can benefit many physics analyses.

frequentist statistics↗

Enhancing Interpretability in Generative Modeling: Statistically Disentangled Latent Spaces Guided by Generative Factors in Scientific Datasets

This study addresses the challenge of statistically extracting generative factors from complex, high-dimensional datasets in unsupervised or semi-supervised settings. We investigate encoder-decoder-based generative models for nonlinear dimensionality reduction, focusing on disentangling low-dimensional latent variables corresponding to independent physical factors. Introducing Aux-VAE, a novel architecture within the classical Variational Autoencoder framework, we achieve disentanglement with minimal modifications to the standard VAE loss function by leveraging prior statistical knowledge through auxiliary variables. These variables guide the shaping of the latent space by aligning latent factors with learned auxiliary variables. We validate the efficacy of Aux-VAE through comparative assessments on multiple datasets, including astronomical simulations.

97 MATHEMATICS AND COMPUTING↗

Discovering Active Subspaces for High-Dimensional Computer Models

Dimension reduction techniques have long been an important topic in statistics, and active subspaces (AS) have received much attention this past decade in the computer experiments literature. The most common approach towards estimating the AS is to use Monte Carlo with numerical gradient evaluation. While sensible in some settings, this approach has obvious drawbacks. Recent research has demonstrated that active subspace calculations can be obtained in closed form, conditional on a Gaussian process (GP) surrogate, which can be limiting in high-dimensional settings for computational reasons. In this paper, we produce the relevant calculations for a more general case when the model of interest is a linear combination of tensor products. These general equations can be applied to the GP, recovering previous results as a special case, or applied to the models constructed by other regression techniques including multivariate adaptive regression splines (MARS). Furthermore, using a MARS surrogate has many advantages including improved scaling, better estimation of active subspaces in high dimensions and the ability to handle a large number of prior distributions in closed form. In one real-world example, we obtain the active subspace of a radiation-transport code with 240 inputs and 9,372 model runs in under half an hour.

97 MATHEMATICS AND COMPUTING↗

Projection pursuit adaptation on polynomial chaos expansions

Here, the present work addresses the issue of accurate stochastic approximations in high-dimensional parametric space using tools from uncertainty quantification (UQ). The basis adaptation method and its accelerated algorithm in polynomial chaos expansions (PCE) were recently proposed to construct low-dimensional approximations adapted to specific quantities of interest (QoI). The present paper addresses one difficulty with these adaptations, namely their reliance on quadrature point sampling, which limits the reusability of potentially expensive samples. Projection pursuit (PP) is a statistical tool to find the “interesting” projections in high-dimensional data and thus bypass the curse-of-dimensionality. In the present work, we combine the fundamental ideas of basis adaptation and projection pursuit regression (PPR) to propose a novel method to simultaneously learn the optimal low-dimensional spaces and PCE representation from given data. While this projection pursuit adaptation (PPA) can be entirely data-driven, the constructed approximation exhibits mean-square convergence to the solution of an underlying governing equation and thus captures the supports and probability distributions associated with the physics constraints. The proposed approach is demonstrated on a borehole problem and a structural dynamics problem, demonstrating the versatility of the method and its ability to discover low-dimensional manifolds with high accuracy with limited data. In addition, the method can learn surrogate models for different quantities of interest while reusing the same data set.

97 MATHEMATICS AND COMPUTING↗