Search NASA⌕ Search

Engineering topics

Ainsworth, Mark

Publications and source records attributed to Ainsworth, Mark.

Maintaining Trust in Reduction: Preserving the Accuracy of Quantities of Interest for Lossy Compression

As the growth of data sizes continues to outpace computational resources, there is a pressing need for data reduction techniques that can significantly reduce the amount of data and quantify the error incurred in compression. Compressing scientific data presents many challenges for reduction techniques since it is often on non-uniform or unstructured meshes, is from a high-dimensional space, and has many Quantities of Interests (QoIs) that need to be preserved. To illustrate these challenges, we focus on data from a large scale fusion code, XGC. XGC uses a Particle-In-Cell (PIC) technique which generates hundreds of PetaBytes (PBs) of data a day, from thousands of timesteps. XGC uses an unstructured mesh, and needs to compute many QoIs from the raw data, f.One critical aspect of the reduction is that we need to ensure that QoIs derived from the data (density, temperature, flux surface averaged momentums, etc.) maintain a relative high accuracy. We show that by compressing XGC data on the high-dimensional, nonuniform grid on which the data is defined, and adaptively quantizing the decomposed coefficients based on the characteristics of the QoIs, the compression ratios at various error tolerances obtained using a multilevel compressor (MGARD) increases more than ten times. We then present how to mathematically guarantee that the accuracy of the QoIs computed from the reduced f is preserved during the compression. We show that the error in the XGC density can be kept under a user-specified tolerance over 1000 timesteps of simulation using the mathematical QoI error control theory of MGARD, whereas traditional error control on the data to be reduced does not guarantee the accuracy of the QoIs.

Gong, Qian↗

Online data analysis and reduction: An important co-design motif for extreme-scale computers

A growing disparity between supercomputer computation speeds and I/O rates means that it is rapidly becoming infeasible to analyze supercomputer application output only after that output has been written to a file system. Instead, data-generating applications must run concurrently with data reduction and/or analysis operations, with which they exchange information via high-speed methods such as interprocess communications. The resulting parallel computing motif, online data analysis and reduction (ODAR), has important implications for both application and HPC systems design. Here we introduce the ODAR motif and its co-design concerns, describe a co-design process for identifying and addressing those concerns, present tools that assist in the co-design process, and present case studies to illustrate the use of the process and tools in practical settings.

Data Analysis↗

Plateau Phenomenon in Gradient Descent Training of RELU Networks: Explanation, Quantification, and Avoidance

The ability of neural networks to provide ‘best in class’ approximation across a wide range of applications is well-documented. Nevertheless, the powerful expressivity of neural networks comes to naught if one is unable to effectively train (choose) the parameters defining the network. In general, neural networks are trained by gradient descent type optimization methods,a stochastic variant thereof. In practice, such methods result in the loss function decreases rapidly at the beginning of training but then, after a relatively small number of steps, significantly slow down. The loss may even appear to stagnate over the period of a large number of epochs, only to then suddenly start to decrease fast again for no apparent reason. This so-called plateau phenomenon manifests itself in many learning tasks. The present work aims to identify and quantify the root causes of plateau phenomenon.analysis is carried out in the setting of univariate ReLU networks. No assumptions are made on the number of neurons relative to the number of training data, and our results hold for both the lazy and adaptive regimes. Here, the main findings are: plateaux correspond to periods during which activation patterns remain constant, where activation pattern refers to the number of data points that activate a given neuron; quantification of convergence of the gradient flow dynamics; and, characterization stationary points in terms solutions of local least squares regression lines over subsets of the training data. Based on these conclusions, we propose a new iterative training method, the Active Neuron Least Squares (ANLS), characterised by the explicit adjustment of the activation pattern at each step, which is designed to enable a quick exit from a plateau. Illustrative numerical examples are included throughout.

97 MATHEMATICS AND COMPUTING↗

Galerkin Neural Networks: A Framework for Approximating Variational Equations with Error Control

Herein, we present a new approach to using neural networks to approximate the solutions of variational equations, based on the adaptive construction of a sequence of finite-dimensional sub-spaces whose basis functions are realizations of a sequence of neural networks. Here, the finite-dimensional subspaces are then used to define a standard Galerkin approximation of the variational equation. This approach enjoys a number of advantages, including: the sequential nature of the algorithm offers a systematic approach to enhancing the accuracy of a given approximation; the sequential enhancements provide a useful indicator for the error that can be used as a criterion for terminating the sequential updates; the basic approach is largely oblivious to the nature of the partial differential equation under consideration; and, some basic theoretical results are presented regarding the convergence (or otherwise) of the method which are used to formulate basic guidelines for applying the method.

97 MATHEMATICS AND COMPUTING↗

Multistage Magnetic Separator of Cells and Proteins

The multistage electromagnetic separator for purifying cells and magnetic particles (MAGSEP) is a laboratory apparatus for separating and/or purifying particles (especially biological cells) on the basis of their magnetic susceptibility and magnetophoretic mobility. Whereas a typical prior apparatus based on similar principles offers only a single stage of separation, the MAGSEP, as its full name indicates, offers multiple stages of separation; this makes it possible to refine a sample population of particles to a higher level of purity or to categorize multiple portions of the sample on the basis of magnetic susceptibility and/or magnetophoretic mobility. The MAGSEP includes a processing unit and an electronic unit coupled to a personal computer. The processing unit includes upper and lower plates, a plate-rotation system, an electromagnet, an electromagnet-translation system, and a capture-magnet assembly. The plates are bolted together through a roller bearing that allows the plates to rotate with respect to each other. An interface between the plates acts as a seal for separating fluids. A lower cuvette can be aligned with as many as 15 upper cuvette stations for fraction collection during processing. A two-phase stepping motor drives the rotation system, causing the upper plate to rotate for the collection of each fraction of the sample material. The electromagnet generates a magnetic field across the lower cuvette, while the translation system translates the electromagnet upward along the lower cuvette. The current supplied to the electromagnet, and thus the magnetic flux density at the pole face of the electromagnet, can be set at a programmed value between 0 and 1,400 gauss (0.14 T). The rate of translation can be programmed between 5 and 2,000 m/s so as to align all sample particles in the same position in the cuvette. The capture magnet can be a permanent magnet. It is mounted on an arm connected to a stepping motor. The stepping motor rotates the arm to position the capture magnet above the upper cuvette into which a fraction of the sample is collected. The electronic unit includes a power switch, power-supply circuitry that accepts 110-Vac input power, an RS-232 interface, and status lights. The personal computer runs the MAGSEP software and controls the operation of the MAGSEP through the RS-232 interface. The status of the power, the translating electromagnet, the capture magnet, and the rotation of the upper plate are indicated in a graphical user interface on the computer screen.

Barton, Ken↗