Search NASA⌕ Search

SEARCH · Search NASA

Results for “mathematics of arrays”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Threaded Multi-Core GEMM with MoA and Cache-Blocking: Preprint

A threaded multi-core implementation of the high performance dense linear algebra matrix-matrix multiply GEMM kernel is described. This kernel is widely implemented by vendors in the basic linear algebra subroutine BLAS library. The mathematics of arrays (MoA) paradigm due to Mullin (1988) results in contiguous memory accesses by employing outer-product forms. Our performance studies demonstrate that the MoA implementation of double precision DGEMM combined with optimal cache-blocking strategies results in at least a 25% performance gain on the Intel Xeon Skylake processor over the vendor supplied Intel MKL basic linear algebra libraries. Results are presented for the NREL Eagle supercomputer. The multi-core DGEMM achieves over 100 GigaFlops/sec with eight openMP threads.

cache-blocking↗

Improving the Performance of DGEMM with MoA and Cache-Blocking: Preprint

The goal of this paper is to demonstrate performance enhancements of the high performance dense linear algebra matrix-matrix multiply DGEMM kernel, widely implemented by vendors in the basic linear algebra subroutine BLAS library. The mathematics of arrays (MoA) paradigm due to Mullin (1988) results in contiguous memory accesses in combination with Church-Rosser complete language constructs optimized for target processor architectures [3]. Our performance studies demonstrate that the MoA implementation of DGEMM combined with optimal cache-blocking strategies results in at least a 25% performance gain on both Intel Xeon Skylake and IBM Power-9 processors over the vendor supplied Intel MKL and IBM ESSL basic linear algebra libraries. Results are presented for the NREL Eagle and ORNL Summit supercomputers.

cache-blocking↗

Regression of exchangeable relational arrays

Relational arrays represent measures of association between pairs of actors, often in varied contexts or over time. Trade flows between countries, financial transactions between individuals, contact frequencies between school children in classrooms and dynamic protein-protein interactions are all examples of relational arrays. Elements of a relational array are often modelled as a linear function of observable covariates. Uncertainty estimates for regression coefficient estimators, and ideally the coefficient estimators themselves, must account for dependence between elements of the array, e.g., relations involving the same actor. Existing estimators of standard errors that recognize such relational dependence rely on estimating extremely complex, heterogeneous structure across actors. This paper develops a new class of parsimonious coefficient and standard error estimators for regressions of relational arrays. Here we leverage an exchangeability assumption to derive standard error estimators that pool information across actors, and are substantially more accurate than existing estimators in a variety of settings. This exchangeability assumption is pervasive in network and array models in the statistics literature, but not previously considered when adjusting for dependence in a regression setting with relational data. We demonstrate improvements in inference theoretically, via a simulation study, and by analysis of a dataset involving international trade.

97 MATHEMATICS AND COMPUTING↗

Floating Wind Array Ontology and Modeling Framework

While there are many tools for designing and modeling a single floating turbine, array level design and modeling has much more to consider. Designing floating wind arrays requires a coupled approach considering many variables, from bathymetry to installation and maintenance to failure and risk analysis. With all of these considerations, an array-level modeling tool is needed to quickly evaluate array designs. The Floating Array Model (FAModel) tool developed at the National Renewable Energy Laboratory was created to fill this gap in low-fidelity array modeling. FAModel is a python framework created to streamline holistic low-fidelity floating wind modeling for array-level analysis. FAModel integrates site data and models with a variety of open-source modeling tools developed by NREL, including FLORIS, RAFT, MoorPy, and anchor capacity models. The integration of these tools allows users to quickly and holistically design an array by considering forces, area analysis, visualization, annual energy production, failure modeling, and component costs.

17 WIND ENERGY↗

FPGA Computing

The articles in this special section focus on cutting-edge research on topics that relate to field programmable gate arrays (FPGA) in computing.

97 MATHEMATICS AND COMPUTING↗

A minimum assumption approach to MEG sensor array design

Objective. Our objective is to formulate the problem of the magnetoencephalographic (MEG) sensor array design as a well-posed engineering problem of accurately measuring the neuronal magnetic fields. This is in contrast to the traditional approach that formulates the sensor array design problem in terms of neurobiological interpretability the sensor array measurements. Approach. We use the vector spherical harmonics (VSH) formalism to define a figure-of-merit for an MEG sensor array. We start with an observation that, under certain reasonable assumptions, any array of m perfectly noiseless sensors will attain exactly the same performance, regardless of the sensors' locations and orientations (with the exception of a negligible set of singularly bad sensor configurations). We proceed to the conclusion that under the aforementioned assumptions, the only difference between different array configurations is the effect of (sensor) noise on their performance. We then propose a figure-of-merit that quantifies, with a single number, how much the sensor array in question amplifies the sensor noise. Main results. We derive a formula for intuitively meaningful, yet mathematically rigorous figure-of-merit that summarizes how desirable a particular sensor array design is. We demonstrate that this figure-of-merit is well-behaved enough to be used as a cost function for a general-purpose nonlinear optimization methods such as simulated annealing. We also show that sensor array configurations obtained by such optimizations exhibit properties that are typically expected of 'high-quality' MEG sensor arrays, e.g. high channel information capacity. Significance. Our work paves the way toward designing better MEG sensor arrays by isolating the engineering problem of measuring the neuromagnetic fields out of the bigger problem of studying brain function through neuromagnetic measurements.

60 APPLIED LIFE SCIENCES↗

Abelian combinatorial gauge symmetry

Combinatorial gauge symmetry is a principle that allows us to construct lattice gauge theories with two key and distinguishing properties: a) only one- and two-body interactions are needed; and b) the symmetry is exact rather than emergent in an effective or perturbative limit. The ground state exhibits topological order for a range of parameters. This paper is a generalization of the construction to any finite Abelian group. In addition to the general mathematical construction, we present a physical implementation in superconducting wire arrays, which offers a route to the experimental realization of lattice gauge theories with static Hamiltonians.

Yu, Hongji↗

A participant-derived xenograft model of HIV enables long-term evaluation of autologous immunotherapies

HIV-specific CD8+ T cells partially control viral replication and delay disease progression, but they rarely provide lasting protection, largely due to immune escape. Here, we show that engrafting mice with memory CD4+ T cells from HIV+ donors uniquely allows for the in vivo evaluation of autologous T cell responses while avoiding graft-versus-host disease and the need for human fetal tissues that limit other models. Treating HIV-infected mice with clinically relevant HIV-specific T cell products resulted in substantial reductions in viremia. In vivo activity was significantly enhanced when T cells were engineered with surface-conjugated nanogels carrying an IL-15 superagonist, but it was ultimately limited by the pervasive selection of a diverse array of escape mutations, recapitulating patterns seen in humans. By applying mathematical modeling, we show that the kinetics of the CD8+ T cell response have a profound impact on the emergence and persistence of escape mutations. This “participant-derived xenograft” model of HIV provides a powerful tool for studying HIV-specific immunological responses and facilitating the development of effective cell-based therapies.

60 APPLIED LIFE SCIENCES↗

Lax-Oleinik-Type Formulas and Efficient Algorithms for Certain High-Dimensional Optimal Control Problems

Two of the main challenges in optimal control are solving problems with state-dependent running costs and developing efficient numerical solvers that are computationally tractable in high dimension. In this paper, we provide analytical solutions to certain optimal control problems whose running cost depends on the state variable and with constraints on the control. We also provide Lax-Oleinik-type representation formulas for the corresponding Hamilton-Jacobi partial differential equations with state-dependent Hamiltonians. Additionally, we present an efficient, grid-free numerical solver based on our representation formulas, which is shown to scale linearly with the state dimension, and thus, to overcome the curse of dimensionality. Using existing optimization methods and the min-plus technique, we extend our numerical solvers to address more general classes of convex and nonconvex initial costs. We demonstrate the capabilities of our numerical solvers using implementations on a central processing unit (CPU) and a field-programmable gate array (FPGA). In several cases, our FPGA implementation obtains over a 10 times speedup compared to the CPU, which demonstrates the promising performance boosts FPGAs can achieve. Furthermore, our numerical results show that our solvers have the potential to serve as a building block for solving broader classes of high-dimensional optimal control problems in real-time.

97 MATHEMATICS AND COMPUTING↗

Quasar Identification Using Multivariate Probability Density Estimated from Nonparametric Conditional Probabilities

Nonparametric estimation for a probability density function that describes multivariate data has typically been addressed by kernel density estimation (KDE). A novel density estimator recently developed by Farmer and Jacobs offers an alternative high-throughput automated approach to univariate nonparametric density estimation based on maximum entropy and order statistics, improving accuracy over univariate KDE. This article presents an extension of the single variable case to multiple variables. The univariate estimator is used to recursively calculate a product array of one-dimensional conditional probabilities. In combination with interpolation methods, a complete joint probability density estimate is generated for multiple variables. Good accuracy and speed performance in synthetic data are demonstrated by a numerical study using known distributions over a range of sample sizes from 100 to 10 6 for two to six variables. Performance in terms of speed and accuracy is compared to KDE. The multivariate density estimate developed here tends to perform better as the number of samples and/or variables increases. As an example application, measurements are analyzed over five filters of photometric data from the Sloan Digital Sky Survey Data Release 17. The multivariate estimation is used to form the basis for a binary classifier that distinguishes quasars from galaxies and stars with up to 94% accuracy.

79 ASTRONOMY AND ASTROPHYSICS↗

Comparison of NMC estimates with trans-stilbene, EJ-309, and He-3 detection systems

Neutron multiplicity counting (NMC) is a technique for the assay of fissile material. In this work, three detection systems are utilized for active interrogation assay of shells of the Rocky Flats shells (highly enriched uranium, 93% 235U) stacked from 13.25-54.92 kg assemblies. The singles ($R_1$) and doubles ($R_2$) rates are calculated with each system to estimate two sample parameters: $M_L$- the leakage multiplication and F - the sample fission rate. The estimated mass, m, is found by dividing F by a constant activity per unit mass. Since we are interrogating HEU, the α ratio of (α; n) to fission neutrons is taken to be zero. The system of equations to calculate these quantities was originally derived. The Neutron Multiplicity 3 He Array Detector (NOMAD) consists of 15 3 He tubes inside polyethylene and is the traditional, capture-based detection system for NMC. The polyethylene moderates incoming neutrons for thermal capture in individual tubes. The low gamma background, discrete capture signals, and high efficiency of the NOMAD are beneficial for NMC. However, the time to slow down neutrons to thermal energies leads to a system die away time on the scale of microseconds. Comparatively, the Rossi-alpha Measurements – Rapid Organic (n, γ) Discrimination Detector (RAMRODD) and the Organic Scintillator Array (OSCAR) are scatter-based detection systems. RAMRODD consists of 8, 5.08 by 5.08 cm EJ-309 liquid scintillators in 4 pairs, evenly-spaced with 90 degrees separation about the center of each assembly. OSCAR consists of a single array of 12, 5.08 by 5.08 cm trans-stilbene crystals aligned with the center of the assembly. Scatter-based systems detect fast neutrons without any moderation, leading to system die away times on the scale of tens of nanoseconds. This allows much shorter timing gate widths compared to thermal systems, thus increasing counting statistics of true correlated fission events. However, scatter-based systems are susceptible to neutron cross-talk when an incident neutron scatters off one detector and interacts in an adjacent detector, causing two seemingly correlated detection signals. The equations developed adjust for cross-talk to conduct NMC with scatter-based systems.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Searching for fat tails in CRISPR-Cas systems: Data analysis and mathematical modeling

Understanding CRISPR-Cas systems—the adaptive defence mechanism that about half of bacterial species and most of archaea use to neutralise viral attacks—is important for explaining the biodiversity observed in the microbial world as well as for editing animal and plant genomes effectively. The CRISPR-Cas system learns from previous viral infections and integrates small pieces from phage genomes called spacers into the microbial genome. The resulting library of spacers collected in CRISPR arrays is then compared with the DNA of potential invaders. One of the most intriguing and least well understood questions about CRISPR-Cas systems is the distribution of spacers across the microbial population. Here, using empirical data, we show that the global distribution of spacer numbers in CRISPR arrays across multiple biomes worldwide typically exhibits scale-invariant power law behaviour, and the standard deviation is greater than the sample mean. We develop a mathematical model of spacer loss and acquisition dynamics which fits observed data from almost four thousand metagenomes well. In analogy to the classical ‘rich-get-richer’ mechanism of power law emergence, the rate of spacer acquisition is proportional to the CRISPR array size, which allows a small proportion of CRISPRs within the population to possess a significant number of spacers. Our study provides an alternative explanation for the rarity of all-resistant super microbes in nature and why proliferation of phages can be highly successful despite the effectiveness of CRISPR-Cas systems.

59 BASIC BIOLOGICAL SCIENCES↗

Manipulation of Geographic Information in Global Seismology

Geographic data, such as seismic event locations, station locations, etc., are generally given in geographic latitude Φ ’, longitude θ , and depth below sea level, ζ , using the WGS84 ellipsoid as a reference. In software systems that use this type of geographic data, it is necessary to manipulate the data mathematically in order to perform such tasks as finding the angular distance or azimuth from one point to another, to find an array of points along a great circle, to rotate a point about a pole of rotation, to move a point some angular distance in a specified direction, to find the intersections of two great circles or to find the intersections of a great circle and a small circle. In this paper, equations are presented that convert geographic locations first to geocentric coordinates and then to Earth-centered Cartesian coordinates where many mathematical manipulations can be performed conveniently and efficiently.

58 GEOSCIENCES↗

Bridging Python to Silicon: The SODA Toolchain

Systems performing scientific computing, data analysis, and machine learning tasks have a growing demand for application-specific accelerators that can provide high computational performance while meeting strict size and power requirements. However, the algorithms and applications that need to be accelerated are evolving at a rate that is incompatible with manual design processes based on hardware description languages. Agile hardware design tools based on compiler techniques can help by quickly producing an application-specific integrated circuit (ASIC) accelerator starting from a high-level algorithmic description. Here, we present the software-defined accelerator (SODA) synthesizer, a modular and open-source hardware compiler that provides automated end-to-end synthesis from high-level software frameworks to ASIC implementation, relying on multilevel representations to progressively lower and optimize the input code. Our approach does not require the application developer to write any register-transfer level code, and it is able to reach up to 364 giga floating point operations per second (GFLOPS)/W efficiency (32-bit precision) on typical convolutional neural network operators.

97 MATHEMATICS AND COMPUTING↗

Mathematical modeling of novel porous transport layer architectures for proton exchange membrane electrolysis cells

Thin foil based porous transport layers (PTLs) that contain highly structured pore arrays have shown promise as anode PTLs in proton exchange membrane electrolysis cells. These novel PTLs, fabricated with advanced manufacturing techniques, produce thin, tunable, multifunctional layers with reduced flow and interfacial resistances and high thermal and electric conductivities. To further optimize their design, it is important to understand their fundamental impact on the transport of protons, electrons, and liquid/vapor mixtures in the electrode. In this work, we develop a two-dimensional multiphysics model to simulate the coupled electrochemistry and multiphase transport in an electrolysis cell operated with the novel PTL architecture. The results show that larger pores improve access of water to the anode catalyst layer, which is beneficial for both the oxygen evolution reaction and membrane hydration. Larger pore sizes also improve oxygen gas transport from the catalyst layer, because generated oxygen gas is forced to travel in-plane through the anode catalyst layer until it reaches a pore opening that is connected to a channel. The discussed results confirm that the proposed thin foil based PTLs are fundamentally different from conventional PTLs, such as felts or layered meshes. The model developed in this work also provides generalizable insight into fundamental PEMEC phenomena, such as the competition between liquid and gas phase transport, membrane hydration and water management, and nonuniform electrochemical reactions, which are processes relevant to all PEMEC designs.

25 ENERGY STORAGE↗