Search NASA⌕ Search

Engineering topics

Dunton, Alec M.

Publications and source records attributed to Dunton, Alec M..

An Analysis of the Johnson-Lindenstrauss Lemma with the Bivariate Gamma Distribution

Probabilistic proofs of the Johnson-Lindenstrauss lemma imply that random projection can reduce the dimension of a data set and approximately preserve pairwise distances. If a distance being approximately preserved is called a success, and the complement of this event is called a failure, then such a random projection likely results in no failures. Assuming a Gaussian random projection, the lemma is proved by showing that the no-failure probability is positive using a combination of Bonferroni's inequality and Markov's inequality. This paper modifies this proof in two ways to obtain a greater lower bound on the no-failure probability. First, Bonferroni's inequality is applied to pairs of failures instead of individual failures. Second, since a pair of projection errors has a bivariate gamma distribution, this probability of a pair of successes is bounded using an inequality from [Jensen, 1969]. If n is the number of points to be embedded and μ is the probability of success, then this leads to an increase in the lower bound on the no-failure probability of $\frac{1}{2}$ ($\genfrac{}{}{0pt}{}{n}{2}$) (1- μ ) 2 is ($\genfrac{}{}{0pt}{}{n}{2}$) is even and $\frac{1}{2}$ (($\genfrac{}{}{0pt}{}{n}{2}$)-1) (1- μ ) 2 if ($\genfrac{}{}{0pt}{}{n}{2}$) is odd. For example, if n =10 5 points are to be embedded in k =10 4 dimensions with a tolerance of ϵ=0.1, then the improvement in the lower bound is on the order of 10 -14 . We also show that further improvement is possible if the inequality in [Jensen, 1969] extends to three successes, though we do not have a proof of this result.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Pass-efficient methods for compression of high-dimensional turbulent flow data

The future of high-performance computing, specifically on future Exascale computers, will presumably see memory capacity and bandwidth fail to keep pace with data generated, for instance, from massively parallel partial differential equation (PDE) systems. Current strategies proposed to address this bottleneck entail the omission of large fractions of data, as well as the incorporation of in situ compression algorithms to avoid overuse of memory. To ensure that post-processing operations are successful, this must be done in a way that a sufficiently accurate representation of the solution is stored. Moreover, in situations where the input/output system becomes a bottleneck in analysis, visualization, etc., or the execution of the PDE solver is expensive, the number of passes made over the data must be minimized. In the interest of addressing this problem, this work focuses on the utility of pass-efficient, parallelizable, low-rank, matrix decomposition methods in compressing high-dimensional simulation data from turbulent flows. Additionally, a particular emphasis is placed on using coarse representation of the data – compatible with the PDE discretization grid – to accelerate the construction of the low-rank factorization. This includes the presentation of a novel single-pass matrix decomposition algorithm for computing the so-called interpolative decomposition. The methods are described extensively and numerical experiments on two turbulent channel flow data are performed. In the first (unladen) channel flow case, compression factors exceeding 400 are achieved while maintaining accuracy with respect to first- and second-order flow statistics. In the particle-laden case, compression factors of 100 are achieved and the compressed data is used to recover particle velocities. These results show that these compression methods can enable efficient computation of various quantities of interest in both the carrier and disperse phases.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗