Search NASA⌕ Search

SEARCH · Search NASA

Results for “distributed parallel computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44

What Multilevel Parallel Programs do when you are not Watching: A Performance Analysis Case Study Comparing MPI/OpenMP, MLP, and Nested OpenMP

With the current trend in parallel computer architectures towards clusters of shared memory symmetric multi-processors, parallel programming techniques have evolved that support parallelism beyond a single level. When comparing the performance of applications based on different programming paradigms, it is important to differentiate between the influence of the programming model itself and other factors, such as implementation specific behavior of the operating system (OS) or architectural issues. Rewriting-a large scientific application in order to employ a new programming paradigms is usually a time consuming and error prone task. Before embarking on such an endeavor it is important to determine that there is really a gain that would not be possible with the current implementation. A detailed performance analysis is crucial to clarify these issues. The multilevel programming paradigms considered in this study are hybrid MPI/OpenMP, MLP, and nested OpenMP. The hybrid MPI/OpenMP approach is based on using MPI [7] for the coarse grained parallelization and OpenMP [9] for fine grained loop level parallelism. The MPI programming paradigm assumes a private address space for each process. Data is transferred by explicitly exchanging messages via calls to the MPI library. This model was originally designed for distributed memory architectures but is also suitable for shared memory systems. The second paradigm under consideration is MLP which was developed by Taft. The approach is similar to MPi/OpenMP, using a mix of coarse grain process level parallelization and loop level OpenMP parallelization. As it is the case with MPI, a private address space is assumed for each process. The MLP approach was developed for ccNUMA architectures and explicitly takes advantage of the availability of shared memory. A shared memory arena which is accessible by all processes is required. Communication is done by reading from and writing to the shared memory.

Jost, Gabriele↗

A Generalized Fluid System Simulation Program to Model Flow Distribution in Fluid Networks

This paper describes a general purpose computer program for analyzing steady state and transient flow in a complex network. The program is capable of modeling phase changes, compressibility, mixture thermodynamics and external body forces such as gravity and centrifugal. The program's preprocessor allows the user to interactively develop a fluid network simulation consisting of nodes and branches. Mass, energy and specie conservation equations are solved at the nodes; the momentum conservation equations are solved in the branches. The program contains subroutines for computing "real fluid" thermodynamic and thermophysical properties for 33 fluids. The fluids are: helium, methane, neon, nitrogen, carbon monoxide, oxygen, argon, carbon dioxide, fluorine, hydrogen, parahydrogen, water, kerosene (RP-1), isobutane, butane, deuterium, ethane, ethylene, hydrogen sulfide, krypton, propane, xenon, R-11, R-12, R-22, R-32, R-123, R-124, R-125, R-134A, R-152A, nitrogen trifluoride and ammonia. The program also provides the options of using any incompressible fluid with constant density and viscosity or ideal gas. Seventeen different resistance/source options are provided for modeling momentum sources or sinks in the branches. These options include: pipe flow, flow through a restriction, non-circular duct, pipe flow with entrance and/or exit losses, thin sharp orifice, thick orifice, square edge reduction, square edge expansion, rotating annular duct, rotating radial duct, labyrinth seal, parallel plates, common fittings and valves, pump characteristics, pump power, valve with a given loss coefficient, and a Joule-Thompson device. The system of equations describing the fluid network is solved by a hybrid numerical method that is a combination of the Newton-Raphson and successive substitution methods. This paper also illustrates the application and verification of the code by comparison with Hardy Cross method for steady state flow and analytical solution for unsteady flow.

Majumdar, Alok↗

Parallel Runtime Interface for Fortran (PRIF) Specification (Rev. 0.6)

This document specifies an interface to support the multi-image parallelism features of Fortran, named the Parallel Runtime Interface for Fortran (PRIF). PRIF is a solution in which a runtime library is primarily responsible for implementing coarray allocation, deallocation and accesses, image synchronization, atomic operations, events, teams and collective subroutines. The Fortran compiler is responsible for transforming the invocation of Fortran-level multi-image parallelism features into procedure calls to the necessary PRIF subroutines. The interface is designed for portability across shared- and distributed-memory machines, different operating systems, and multiple architectures. Implementations of this interface are intended as an augmentation for the compiler's own runtime library. With an implementation-agnostic interface, alternative parallel runtime libraries may be developed that support the same interface. One benefit of this approach is the ability to vary the communication substrate. A central aim of this document is to define a parallel runtime interface in standard Fortran syntax, which enables us to leverage Fortran to succinctly express various properties of the procedure interfaces, including argument attributes.

97 MATHEMATICS AND COMPUTING↗

Containerless Ripple Turbulence

One of the longest standing unsolved problems in physics relates to the behavior of fluids that are driven far from equilibrium such as occurs when they become turbulent due to fast flow through a grid or tidal motions. In turbulent flows the distribution of vortex energy as a function of the inverse length scale [or wavenumber 'k'] of motion is proportional to 1/k(sup 5/3) which is the celebrated law of Kolmogorov. Although this law gives a good description of the average motion, fluctuations around the average are huge. This stands in contrast with thermally activated motion where large fluctuations around thermal equilibrium are highly unfavorable. The problem of turbulence is the problem of understanding why large fluctuations are so prevalent which is also called the problem of 'intermittency'. Turbulence is a remarkable problem in that its solution sits simultaneously at the forefront of physics, mathematics, engineering and computer science. A recent conference [March 2002] on 'Statistical Hydrodynamics' organized by the Los Alamos Laboratory Center for Nonlinear Studies brought together researchers in all of these fields. Although turbulence is generally thought to be described by the Navier-Stokes Equations of fluid mechanics the solution as well as its existence has eluded researchers for over 100 years. In fact proof of the existence of such a solution qualifies for a 1 M$ millennium prize. As part of our NASA funded research we have proposed building a bridge between vortex turbulence and wave turbulence. The latter occurs when high amplitude waves of various wavelengths are allowed to mutually interact in a fluid. In particular we have proposed measuring the interaction of ripples [capillary waves] that run around on the surface of a fluid sphere suspended in a microgravity environment. The problem of ripple turbulence poses similar mathematical challenges to the problem of vortex turbulence. The waves can have a high amplitude and a strong nonlinear interaction. Furthermore, the steady state distribution of energy again follows a Kolmogorov scaling law; in this case the ripple energy is distributed according to 1/k (sup 7/4). Again, in parallel with vortex turbulence ripple turbulence exhibits intermittency. The problem of ripple turbulence presents an experimental opportunity to generate data in a controlled, benchmarked system. In particular the surface of a sphere is an ideal environment to study ripple turbulence. Waves run around the sphere and interact with each other, and the effect of walls is eliminated. In microgravity this state can be realized for over 2 decades of frequency. Wave turbulence is a physically relevant problem in its own right. It has been studied on the surface of liquid hydrogen and its application to Alfven waves in space is a source of debate. Of course, application of wave turbulence perspectives to ocean waves has been a major success. The experiment which we plan to run in microgravity is conceptually straightforward. Ripples are excited on the surface of a spherical drop of fluid and then their amplitude is recorded with appropriate photography. A key challenge is posed by the need to stably position a 10cm diameter sphere of water in microgravity. Two methods are being developed. Orbitec is using controlled puffs of air from at least 6 independent directions to provided the positioning force. This approach has actually succeeded to position and stabilize a 4cm sphere during a KC 135 segment. Guigne International is using the radiation pressure of high frequency sound. These transducers have been organized into a device in the shape of a dodecahedron. This apparatus 'SPACE DRUMS' has already been approved for use for combustion synthesis experiments on the International Space Station. A key opportunity presented by the ripple turbulence data is its use in driving the development of codes to simulate its properties.

Putterman, Seth↗

"One-Stop Shopping" for Ocean Remote-Sensing and Model Data

OurOcean Portal 2.0 (http:// ourocean.jpl.nasa.gov) is a software system designed to enable users to easily gain access to ocean observation data, both remote-sensing and in-situ, configure and run an Ocean Model with observation data assimilated on a remote computer, and visualize both the observation data and the model outputs. At present, the observation data and models focus on the California coastal regions and Prince William Sound in Alaska. This system can be used to perform both real-time and retrospective analyses of remote-sensing data and model outputs. OurOcean Portal 2.0 incorporates state-of-the-art information technologies (IT) such as MySQL database, Java Web Server (Apache/Tomcat), Live Access Server (LAS), interactive graphics with Java Applet at the Client site and MatLab/GMT at the server site, and distributed computing. OurOcean currently serves over 20 real-time or historical ocean data products. The data are served in pre-generated plots or their native data format. For some of the datasets, users can choose different plotting parameters and produce customized graphics. OurOcean also serves 3D Ocean Model outputs generated by ROMS (Regional Ocean Model System) using LAS. The Live Access Server (LAS) software, developed by the Pacific Marine Environmental Laboratory (PMEL) of the National Oceanic and Atmospheric Administration (NOAA), is a configurable Web-server program designed to provide flexible access to geo-referenced scientific data. The model output can be views as plots in horizontal slices, depth profiles or time sequences, or can be downloaded as raw data in different data formats, such as NetCDF, ASCII, Binary, etc. The interactive visualization is provided by graphic software, Ferret, also developed by PMEL. In addition, OurOcean allows users with minimal computing resources to configure and run an Ocean Model with data assimilation on a remote computer. Users may select the forcing input, the data to be assimilated, the simulation period, and the output variables and submit the model to run on a backend parallel computer. When the run is complete, the output will be added to the LAS server for

Li, P. Peggy↗

An Efficient and Accurate Algorithm for Computing Grid-Averaged Solar Fluxes for Horizontally Inhomogeneous Clouds

A computationally efficient method is presented to account for the horizontal cloud inhomogeneity by using a radiatively equivalent plane parallel homogeneous (PPH) cloud. The algorithm can accurately match the calculations of the reference (rPPH) independent column approximation (ICA) results, but use only the same computational time required for a single plane parallel computation. The effective optical depth of this synthetic sPPH cloud is derived by exactly matching the direct transmission to that of the inhomogeneous ICA cloud. The ffective9 scattering asymmetry factor is found from a pre-calculated albedo inverse look-up-table that is allowed to vary over the range from -1.0 to 1.0. In the special cases of conservative scattering and total absorption, the synthetic method is exactly equivalent to the ICA, with only a small bias (about 0.2% in flux) relative to ICA due to imperfect interpolation in using the look-up tables. In principle, the ICA albedo can be approximated accurately regardless of cloud inhomogeneity. For a more complete comparison, the broadband shortwave albedo and transmission calculated from the synthetic sPPH cloud and averaged over all incident directions, have the RMS biases of 0.26% and 0.76%, respectively, for inhomogeneous clouds over a wide variation of particle size. The advantages of the synthetic PPH method are that (1) it is not required that all the cloud subcolumns have uniform microphysical characteristic, (2) it is applicable to any 1D radiative transfer scheme, and (3) it can handle arbitrary cloud optical depth distributions and an arbitrary number of cloud subcolumns with uniform computational efficiency.

cloud inhomogeneity↗

Multiphysics Demonstration of Temperature-Driven Assembly Bowing in SFRs using MOOSE-Based Codes

Core bowing is an important passive safety mechanism in liquid metal cooled fast reactors. When the core restraint system is properly designed, temperature and flux gradients influence assemblies in the core to bow into less reactive configurations during accident scenarios, resulting in negative reactivity feedback. Prediction of core bowing involves complex interplay of radiation transport, impacts of fluid flow and heat transfer on duct temperature, and mechanical responses to the induced temperature and flux gradients. Under the U.S. Department of Energy Office of Nuclear Energy’s Advanced Modeling and Simulation (NEAMS) Program [1], an integrated multiphysics approach is being developed to model the core bowing phenomena in liquid metal-cooled fast reactors with the Multiphysics Object Oriented Simulation Environment (MOOSE) [2]. In this methodology, the MOOSE-based reactor physics code Griffin [3] will solve the neutron transport equation and determine the power distribution. With the detailed power distribution from Griffin, the subchannel analysis codes MOOSE-Subchannel [4] and Pronghorn [5] are utilized to calculate the assembly temperature distribution. MOOSE’s Solid Mechanics [6] and Contact [7] Modules are leveraged to calculate the thermal expansion and duct bowing displacement with the duct wall temperature from thermal hydraulics calculation. In this work, an initial one-way coupling demonstration of the integrated multiphysics approach has been performed on a seven-assembly problem based on the sodium-cooled fast reactor ABR-1000 design [8]. The neutronics calculation with Griffin is not yet involved in the current simulation. MOOSE-Subchannel and Pronghorn evaluate fluid and solid temperature based on a fixed power distribution. In addition, one-way coupling is utilized in this coupled calculation, via Pronghorn passing the duct temperature data to the MOOSE Solid Mechanics calculation. An assessment of the Solid Mechanics module was performed in parallel to verify duct bowing behavior with duct-to-duct contact phenomenon [9]. The displacement from MOOSE Solid Mechanics is not yet transferred back and utilized in the Pronghorn and MOOSE-Subchannel calculation. This model will be available on the National Reactor Innovation Center (NRIC) Virtual Test Bed (VTB) repository [10]. Future stages of this work will involve solving problems of increasing complexity as well as adding more physics (e.g. reactor physics) to the integrated workflow to reach the end goal of modeling the core bowing phenomenon with an integrated multiphysics workflow.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

HARMONY: Large-Scale Architecture Search for Efficient Hybrid Language Models

As large language models scale to trillions of parameters, their computational and memory requirements present critical challenges for efficient training and deployment. While Mixture of Experts (MoE) architectures enable efficient scaling through sparse parameter activation, and state-space models like Mamba offer linear-time complexity, principled methods for combining these paradigms remain undeveloped. We introduce HARMONY (Hybrid Architecture Research for Mamba, Optimized with Neural efficiencY), a multi-objective evolutionary neural architecture search framework for discovering efficient hybrid language models that integrate Transformer attention mechanisms, Mixture-of-Experts routing, and Mamba state-space components. Through large-scale distributed search using 16,384 MI250X GPUs on the Frontier supercomputer, HARMONY explores a comprehensive design space encompassing six attention variants (MHA, MQA, GQA, MLA, SWA, and Mamba-2), variable MoE configurations with both routed and shared experts, and extensive Mamba hyperparameters. Our framework discovers heterogeneous architectures that balance training performance with computational efficiency through multi-objective optimization incorporating latency penalties and fitness-based selection. Analysis of discovered architectures reveals that optimal hybrid designs favor heterogeneous component mixing rather than homogeneous patterns, with Mamba-2 and Multi-Head Latent Attention (MLA) emerging as preferred mechanisms. Discovered architectures demonstrate superior training efficiency: our best configuration achieves a final perplexity of 1.0874 with 2.38B parameters while processing 4,320 tokens/second, outperforming significantly larger manually designed models. Full-scale evaluation shows HARMONY's top architectures achieve better loss trajectories than equivalently-sized models using state-of-the-art configurations including Mixtral, Jamba, and Samba. Additionally, we demonstrate 91% weak scaling efficiency when training discovered 36B-parameter models across 1,024 GPUs. HARMONY is released as an open framework with comprehensive tools for building and training hybrid models using expert-data-pipeline parallelism, democratizing access to automated architecture design for next-generation language models.

Herron, Emily [ORNL] (ORCID:0000000273008172)↗

Three Dimensional Atmospheric Radiative Transfer-Applications and Methods Comparison

We review applications of 3D radiative transfer in the atmosphere, emphasizing the wide spectrum of scales important to remote sensing and modeling of cloud fields, and the characteristic scales introduced into observed radiances and fluxes by the distribution of photon pathlengths at conservative and absorbing wavelengths. We define the "plane-parallel bias", which is a measure of the importance of 3D cloud structure in large-scale models, and the "independent pixel errors" that quantify the significance of 3D effects in remote sensing, and emphasize their relative magnitude and scale dependence. A variety of approaches in current use in 3D radiative transfer, and issues of speed, accuracy, and flexibility are summarized. We also describe a recently initiated "International Intercomparison of 3-Dimensional Radiation Codes", or I3RC. I3RC is a 3-phase effort that has as its goals to: (1) understand the errors and limits of 3D methods; (2) provide "baseline" cases for future 3D code development; (3) promote sharing of 3D tools; (4) derive guidelines for 3D tool selection; and (5) improve atmospheric science education in 3D radiative transfer. Selected results from Phases 1 and 2 of I3RC are discussed. These are taken from five cloud fields: a 1D field of bar clouds, a 2D radar-derived field, a 3D Landsat-derived field, a stratiform cloud from the model of C. Moeng, and a convective cloud from the model of B. Stevens. Computations have been carried out for three monochromatic wavelengths (one conservative, one absorptive, and one thermal) and two solar zenith angles (0, 60 degrees).

Cahalan, Robert F.↗

Efficacy of Code Optimization on Cache-based Processors

The current common wisdom in the U.S. is that the powerful, cost-effective supercomputers of tomorrow will be based on commodity (RISC) micro-processors with cache memories. Already, most distributed systems in the world use such hardware as building blocks. This shift away from vector supercomputers and towards cache-based systems has brought about a change in programming paradigm, even when ignoring issues of parallelism. Vector machines require inner-loop independence and regular, non-pathological memory strides (usually this means: non-power-of-two strides) to allow efficient vectorization of array operations. Cache-based systems require spatial and temporal locality of data, so that data once read from main memory and stored in high-speed cache memory is used optimally before being written back to main memory. This means that the most cache-friendly array operations are those that feature zero or unit stride, so that each unit of data read from main memory (a cache line) contains information for the next iteration in the loop. Moreover, loops ought to be 'fat', meaning that as many operations as possible are performed on cache data-provided instruction caches do not overflow and enough registers are available. If unit stride is not possible, for example because of some data dependency, then care must be taken to avoid pathological strides, just ads on vector computers. For cache-based systems the issues are more complex, due to the effects of associativity and of non-unit block (cache line) size. But there is more to the story. Most modern micro-processors are superscalar, which means that they can issue several (arithmetic) instructions per clock cycle, provided that there are enough independent instructions in the loop body. This is another argument for providing fat loop bodies. With these restrictions, it appears fairly straightforward to produce code that will run efficiently on any cache-based system. It can be argued that although some of the important computational algorithms employed at NASA Ames require different programming styles on vector machines and cache-based machines, respectively, neither architecture class appeared to be favored by particular algorithms in principle. Practice tells us that the situation is more complicated. This report presents observations and some analysis of performance tuning for cache-based systems. We point out several counterintuitive results that serve as a cautionary reminder that memory accesses are not the only factors that determine performance, and that within the class of cache-based systems, significant differences exist.

VanderWijngaart, Rob F.↗

Snow ALbedo eVOlution (SALVO) Campaign Broadband Albedo from April - June, 2024 in Utqiagivk, AK level a1

A field-portable broadband (285 – 2800 nm) albedometer was used to make spatially distributed albedo measurements on tundra and sea ice surfaces. The albedometer consists of paired upward-looking and downward-looking pyranometers, which were both connected to a data logger. The instrument was mounted approximately 1 m above the surface using a tripod and was placed on a 1.4 m-long boom to minimize the impacts of shading from the operator and to observe surfaces undisturbed by footprints (see Appendix for photos of measurement setup and uncertainty assessment). Albedo measurements were taken parallel to the 200-m albedo lines at 5-m increments (41 measurements) ~1.2 m south of the line. On the operator’s end of the boom, there was a bubble level that was aligned with the bubble level on the upward-looking pyranometer. To take a measurement, the operator first relocated the tripod to the measurement location, then leveled the instrument and held it level for at least twice the pyranometers’ response time (5 or 15 seconds, see below), and finally depressed a trigger on the data logger. The data logger recorded the instantaneous voltage on both pyranometers, the measurement number, and the time. The data logger also converted the voltages to irradiances, and from these computed the ratio (outgoing/incoming) for albedo, which could be checked in the field. The operator recorded in a field notebook the measurement number that corresponded with the locations on the line and any pertinent notes (e.g., invalid measurements). With this setup, a trained operator could measure a 200-m albedo line (41 measurements) in approximately 30 minutes. Measurements were made within 3 hours of solar noon. The data logger had sufficient storage capacity to record all measurements from the campaign, but data were downloaded to a computer after each measurement day.

54 ENVIRONMENTAL SCIENCES↗

A two-level GPU-accelerated incomplete LU preconditioner for general sparse linear systems

This paper presents a parallel preconditioning approach based on incomplete LU (ILU) factorizations in the framework of Domain Decomposition (DD) for general sparse linear systems. We focus on distributed memory parallel architectures, specifically, those that are equipped with graphic processing units (GPUs). In addition to block-Jacobi, we present general purpose two-level ILU Schur complement-based approaches, where different strategies are presented to solve the coarse-level reduced system. These strategies are combined with modified ILU methods in the construction of the coarse-level operator, in order to effectively remove smooth errors by targeting an algebraically smooth vector. We leverage available GPU-based sparse matrix kernels to accelerate the setup and the solve phases of the proposed ILU preconditioner. We evaluate the efficiency of the proposed methods as a smoother for algebraic multigrid (AMG) and as a preconditioner for Krylov subspace methods on challenging anisotropic diffusion problems and a collection of general sparse matrices.

97 MATHEMATICS AND COMPUTING↗

Computation of Gust-Cascade Interaction Using the CE/SE Method

The problem 2 in Category 3 of the 4th Computational Aeroacoustic(CAA) Workshop is solved using the space-time conservation element and solution element (CE/SE) method. This problem models rotor-stator interaction in a 2D cascade. It involves complex geometries and flow physics including vortex shedding and acoustic radiation. The parallel version of the 2D nonlinear Euler solver is used with an unstructured triangular mesh to solve this problem. The Giles approach is incorporated with the CE/SE method to handle non-equal pitches of the rotor and stator. Validation on the Giles approach is performed using Problem 3.1 in the 2nd CAA Workshop. The space-time CE/SE method is a finite volume method with second-order accuracy in both space and time. The flux conservation is enforced in both space and time instead of space only. It has low numerical dissipation and dispersion errors. It uses simple non-reflecting boundary conditions and is compatible with unstructured meshes. It is simple, flexible, and generate reasonably accurate solutions. The CE/SE method has been successfully applied to solve numerous practical problems, especially aeroacoustic problems. Some preliminary numerical results of the benchmark problem 3.2 of the 4th CAA Workshop are shown. The steady-state pressure contour is plotted. The mean pressure distribution on the blade surface is compared with Turbo solution showing a good agreement. The sound pressure level versus the rotor harmonic n at the six designated positions on the blade surface, three locations at inlet plane, and three locations at the outlet plane are plotted. It can be seen that the acoustic response exists only at the excitation frequencies (n = 1,2,3). On the blade surface, the acoustic wave at n = 1 is dominant, while at the inlet and outlet planes, the sound pressure level at n = 2 becomes the largest, which is similar to the results presented. The distribution of sound pressure level at different spatial modes along the z- direction is plotted for n = 1,2,3, respectively. It shows that the spatial modes m = -32 and 22 at n = 1 exponentially decay, and the spatial modes m = 10 at n = 2, m = -42 and 12 at n = 3 propagate both upstream and downstream, which agrees with the prediction based on the linearized theory. Some oscillations are observed, which needs to be investigated further. In the final paper, the numerical results will be compared with a frequency-domain solver LINFLUX solution if it is available.

Wang, X.-Y.↗

Downsampling Photodetector Array with Windowing

In a photon counting detector array, each pixel in the array produces an electrical pulse when an incident photon on that pixel is detected. Detection and demodulation of an optical communication signal that modulated the intensity of the optical signal requires counting the number of photon arrivals over a given interval. As the size of photon counting photodetector arrays increases, parallel processing of all the pixels exceeds the resources available in current application-specific integrated circuit (ASIC) and gate array (GA) technology; the desire for a high fill factor in avalanche photodiode (APD) detector arrays also precludes this. Through the use of downsampling and windowing portions of the detector array, the processing is distributed between the ASIC and GA. This allows demodulation of the optical communication signal incident on a large photon counting detector array, as well as providing architecture amenable to algorithmic changes. The detector array readout ASIC functions as a parallel-to-serial converter, serializing the photodetector array output for subsequent processing. Additional downsampling functionality for each pixel is added to this ASIC. Due to the large number of pixels in the array, the readout time of the entire photodetector is greater than the time between photon arrivals; therefore, a downsampling pre-processing step is done in order to increase the time allowed for the readout to occur. Each pixel drives a small counter that is incremented at every detected photon arrival or, equivalently, the charge in a storage capacitor is incremented. At the end of a user-configurable counting period (calculated independently from the ASIC), the counters are sampled and cleared. This downsampled photon count information is then sent one counter word at a time to the GA. For a large array, processing even the downsampled pixel counts exceeds the capabilities of the GA. Windowing of the array, whereby several subsets of pixels are designated for processing, is used to further reduce the computational requirements. The grouping of the designated pixel frame as the photon count information is sent one word at a time to the GA, the aggregation of the pixels in a window can be achieved by selecting only the designated pixel counts from the serial stream of photon counts, thereby obviating the need to store the entire frame of pixel count in the gate array. The pixel count se quence from each window can then be processed, forming lower-rate pixel statistics for each window. By having this processing occur in the GA rather than in the ASIC, future changes to the processing algorithm can be readily implemented. The high-bandwidth requirements of a photon counting array combined with the properties of the optical modulation being detected by the array present a unique problem that has not been addressed by current CCD or CMOS sensor array solutions.

Patawaran, Ferze D.↗

Dynamic Smagorinsky Modeled Large-Eddy Simulations of Turbulence Using Tetrahedral Meshes

Eddy-resolving numerical computations of turbulent flows are emerging as viable alternatives to Reynolds Averaged Navier-Stokes (RANS) calculations for flows with an intrinsically steady mean state due to the advances in large-scale parallel computing. In these computations, medium to large turbulent eddies are resolved by the numerics while the smaller or subgrid scales are either modeled or taken care of by the inherent numerical dissipation. To advance the state of the art of unstructured-mesh turbulence simulation capabilities, large eddy simulations (LES) using the dynamic Smagorinsky model (DSM) on tetrahedral meshes are carried out with the space-time conservation element, solution element (CESE) method. In contrast to what has been reported in the literature, the present implementation of dynamic models allows for active backscattering without any ad-hoc limiting of the eddy viscosity calculated from the subgrid-scale model. For the benchmark problems involving compressible isotropic turbulence decay as well as the shock/turbulent boundary layer interaction benchmark problems, no numerical instability associated with kinetic energy growth is observed and the volume percentage of the backscattering portion accounts for about 38-40% of the simulation domain. A slip-wall model in conjunction with the implemented DSM is used to simulate a relatively high Reynolds number Mach 2.85 turbulent boundary layer over a 30° ramp with several tetrahedral meshes and a wall-normal spacing of either Δγ& = 10 or Δγ& = 20. The computed mean wall pressure distribution, separation region size, mean velocity profiles, and Reynolds stress agree reasonably well with experimental data.

CESE method↗

Electrodynamic Dust Shield (EDS) Preparation for MISSE-11 Launch

The dusty surfaces of the Moon, Mars, and various asteroids present a significant challenge to NASA’s manned and unmanned space exploration efforts. The fine, electrostatically charged dust is difficult to remove and has resulted in vision obscuration, false instrument readings, contamination, elevated temperatures, performance reduction, and equipment failure. To alleviate these problems, NASA KSC’s Electrostatics and Surface Physics Laboratory (ESPL) is developing active dust mitigation systems for solar system exploration and in situ resource utilization (ISRU). Part of KSC’s Swamp Works system, the lab emphasizes fast-paced, hands-on, cost-effective, and collaborative innovation. The Electrodynamic Dust Shield (EDS) is an ESPL technology being developed to prevent dust accumulation on space components. It consists of a dielectric substrate embedded with parallel electrodes which, when applied with high voltage out-of-phase waveforms, produce a traveling electric field wave that scatters dust particles from its surface via dielectrophoretic and electrostatic forces. The EDS is currently fully functional and has been tested for scalability and endurance in reduced gravity/pressure flight. Next, it will be part of the May 2019 Materials International Space Station Experiment-11 (MISSE-11) payload, which will test small shield samples for performance durability with exposure to the harsh space environment. The summer objective was to test EDS functionality before active and passive dust shield samples launch on MISSE-11, in order to screen for good flight/control samples as well as to compare pre-flight and post-flight clearing efficiencies. I imaged glass and Kapton shields of 2-phase and 3-phase electrode configurations with the Keyence VHX-5000 digital microscope to check for damage before and after all testing for a qualitative pre-flight baseline; helped with vacuum testing of high-voltage shield breakdown; assembled shields and wires using Electrobond 004 silver epoxy; and helped with vacuum and cryogenic dust clearance testing of assembled EDS systems. I learned CAD (computer-aided design) and 3D printing skills to create parts for a shaker table/motor assembly, which will allow for repeatable experiments by evening out EDS dust distribution in a consistent manner prior to dust scattering testing. I will also be improving the interfacing with Alpha Space’s prototype Data Communications Unit (DCU), a payload component to control power to the EDS and store data, by mapping the DCU structure upon power-up and writing shell scripts to improve various processes. Finally, I also CADed and 3D-printed plastic supports for three different geometries of the Electrostatic Precipitator (ESP), a project that aims to use high voltage electrodes to charge and filter out dust from the Martian atmosphere drawn for ISRU devices. This work on dust mitigation technologies will help to support future robotic and human space exploration missions, including NASA’s Journey to Mars; the EDS has especially significant applications to spacecraft solar panels, thermal radiators, viewports, and astronaut visors. This opportunity has also allowed me to become more familiar with various lab equipment and hardware testing, develop key modeling skills in SolidWorks, explore past and present components critical to America’s space program, work with inspiring peers and mentors, and ascertain my dream of working permanently for NASA.

space↗

Performance-Aligned LLMs for Generating Fast HPC Code

Optimizing scientific software is a difficult task because codebases are often large and complex, and performance can depend upon several factors including the algorithm, its implementation, and hardware among others. Causes of poor performance can originate from disparate sources and be difficult to diagnose. Recent years have seen a multitude of work that use large language models (LLMs) to assist in software development tasks. However, these tools are trained to model the distribution of code as text, and are not specifically designed to understand performance aspects of code. In this work, we introduce a reinforcement learning based methodology to align the outputs of code LLMs with performance. This allows us to build upon the current code modeling capabilities of LLMs and extend them to generate better performing code. Here, we demonstrate that our fine-tuned model improves the expected speedup of generated code over base models for a set of benchmark tasks from 0.9 to 1.6 for serial code and 1.9 to 4.5 for OpenMP parallel code.

Computer science↗

Coupled Monte Carlo Probability Density Function/ SPRAY/CFD Code Developed for Modeling Gas-Turbine Combustor Flows

The success of any solution methodology for studying gas-turbine combustor flows depends a great deal on how well it can model various complex, rate-controlling processes associated with turbulent transport, mixing, chemical kinetics, evaporation and spreading rates of the spray, convective and radiative heat transfer, and other phenomena. These phenomena often strongly interact with each other at disparate time and length scales. In particular, turbulence plays an important role in determining the rates of mass and heat transfer, chemical reactions, and evaporation in many practical combustion devices. Turbulence manifests its influence in a diffusion flame in several forms depending on how turbulence interacts with various flame scales. These forms range from the so-called wrinkled, or stretched, flamelets regime, to the distributed combustion regime. Conventional turbulence closure models have difficulty in treating highly nonlinear reaction rates. A solution procedure based on the joint composition probability density function (PDF) approach holds the promise of modeling various important combustion phenomena relevant to practical combustion devices such as extinction, blowoff limits, and emissions predictions because it can handle the nonlinear chemical reaction rates without any approximation. In this approach, mean and turbulence gas-phase velocity fields are determined from a standard turbulence model; the joint composition field of species and enthalpy are determined from the solution of a modeled PDF transport equation; and a Lagrangian-based dilute spray model is used for the liquid-phase representation with appropriate consideration of the exchanges of mass, momentum, and energy between the two phases. The PDF transport equation is solved by a Monte Carlo method, and existing state-of-the-art numerical representations are used to solve the mean gasphase velocity and turbulence fields together with the liquid-phase equations. The joint composition PDF approach was extended in our previous work to the study of compressible reacting flows. The application of this method to several supersonic diffusion flames associated with scramjet combustor flow fields provided favorable comparisons with the available experimental data. A further extension of this approach to spray flames, three-dimensional computations, and parallel computing was reported in a recent paper. The recently developed PDF/SPRAY/computational fluid dynamics (CFD) module combines the novelty of the joint composition PDF approach with the ability to run on parallel architectures. This algorithm was implemented on the NASA Lewis Research Center's Cray T3D, a massively parallel computer with an aggregate of 64 processor elements. The calculation procedure was applied to predict the flow properties of both open and confined swirl-stabilized spray flames.

Source record↗