Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel Performance Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60

Investigation of Nucleate Boiling Mechanisms Under Microgravity Conditions

The present work is aimed at the experimental studies and numerical modeling of the bubble growth mechanisms of a single bubble attached to a heating surface and of a bubble sliding along an inclined heated plate. Single artificial cavity of 10 microns in diameter was made on the polished Silicon wafer which was electrically heated at the back side in order to control the surface nucleation superheat. Experiments with a sliding bubble were conducted at different inclination angles of the downward facing heated surface for the purpose of studying the effect of magnitude of components of gravity acting parallel to and normal to the heat transfer surface. Information on the bubble shape and size, the bubble induced liquid velocities as well as the surface temperature were obtained using the high speed imaging and hydrogen bubble techniques. Analytical/numerical models were developed to describe the heat transfer through the micro-macro layer underneath and around a bubble formed at a nucleation site. In the micro layer model the capillary and disjoining pressures were included. Evolution of the bubble-liquid interface along with induced liquid motion was modeled. As a follow-up to the studies at normal gravity, experiments are being conducted in the KC-135 aircraft to understand the bubble growth/detachment under low gravity conditions. Experiments have been defined to be performed under long duration of microgravity conditions in the space shuttle. The experiment in the space shuttle will provide bubble growth and detachment data at microgravity and will lead to validation of the nucleate boiling heat transfer model developed from the preceding studies conducted at normal and low gravity (KC-135) conditions.

Dhir, V. K.↗

A Summary of Test and Analysis Results from a Second Lift+Cruise Full-Scale Drop Test

The realization of advanced air mobility markets is enabling new forms of transportation to take shape in the United States and around the world. Though currently in development, as these markets mature, new types of vertical take-off and landing (VTOL) vehicles have been undergoing development for use. There are many factors which must be addressed prior to these types of vehicles becoming viable alternative forms of transportation in these markets. These factors include incorporation into the existing airspaces, the logistics of operating in urban environments, along with numerous factors associated with safety and reliability. To address some of the safety aspects associated with the development of these new types of vehicles, NASA has been conducting research into the performance of an example electric VTOL (eVTOL) aircraft as a part of the Revolutionary Vertical Lift Technology (RVLT) project. Over the course of this research, many aspects including the development of energy absorbing components, the evaluation of seating systems, the development of advanced finite element material model systems and the acquisition of full-scale vehicle impact data were investigated. The report will discuss aspects related to the acquisition of full-scale vehicle data which occurred in the form of a full-scale impact test conducted in the Summer of 2025. This test was on a NASA designed Lift+Cruise composite cabin test article and represented a partial capstone in the entirety of previous eVTOL research conducted for the project. In this test, a variety of experiments were included in order to investigate the effect of a full-scale environment on the experiment results. In parallel, the development of a computational impact model to simulate the full-scale test will be discussed in this report. A model of the Lift+Cruise test article was developed utilizing data collected from previous sub- and full-scale test data and then simulated in the current test environment. The model development, its use in pre-test predictions, and its use in post-test correlation will all be presented. This report will present the test data acquired from the Lift+Cruise test and document several of the results obtained. One intended result is to determine the effect of a complex full-scale crash impact on the identification of occupant injury risk within seat and vehicle designs. A second intended result is to determine whether high-fidelity models can be used with some confidence in the prediction of test events and can allow for additional test cases to be simulated without the need of having to conduct additional tests. The overall goal of the test is to provide the community with data that can be used for design, development or certification efforts, along with providing data on what an example eVTOL crash incident could entail.

energy storage systems↗

HARMONY: Large-Scale Architecture Search for Efficient Hybrid Language Models

As large language models scale to trillions of parameters, their computational and memory requirements present critical challenges for efficient training and deployment. While Mixture of Experts (MoE) architectures enable efficient scaling through sparse parameter activation, and state-space models like Mamba offer linear-time complexity, principled methods for combining these paradigms remain undeveloped. We introduce HARMONY (Hybrid Architecture Research for Mamba, Optimized with Neural efficiencY), a multi-objective evolutionary neural architecture search framework for discovering efficient hybrid language models that integrate Transformer attention mechanisms, Mixture-of-Experts routing, and Mamba state-space components. Through large-scale distributed search using 16,384 MI250X GPUs on the Frontier supercomputer, HARMONY explores a comprehensive design space encompassing six attention variants (MHA, MQA, GQA, MLA, SWA, and Mamba-2), variable MoE configurations with both routed and shared experts, and extensive Mamba hyperparameters. Our framework discovers heterogeneous architectures that balance training performance with computational efficiency through multi-objective optimization incorporating latency penalties and fitness-based selection. Analysis of discovered architectures reveals that optimal hybrid designs favor heterogeneous component mixing rather than homogeneous patterns, with Mamba-2 and Multi-Head Latent Attention (MLA) emerging as preferred mechanisms. Discovered architectures demonstrate superior training efficiency: our best configuration achieves a final perplexity of 1.0874 with 2.38B parameters while processing 4,320 tokens/second, outperforming significantly larger manually designed models. Full-scale evaluation shows HARMONY's top architectures achieve better loss trajectories than equivalently-sized models using state-of-the-art configurations including Mixtral, Jamba, and Samba. Additionally, we demonstrate 91% weak scaling efficiency when training discovered 36B-parameter models across 1,024 GPUs. HARMONY is released as an open framework with comprehensive tools for building and training hybrid models using expert-data-pipeline parallelism, democratizing access to automated architecture design for next-generation language models.

Herron, Emily [ORNL] (ORCID:0000000273008172)↗

Optimal design of structures with multiple design variables per group and multiple loading conditions on the personal computer

A finite element based programming system for minimum weight design of a truss-type structure subjected to displacement, stress, and lower and upper bounds on design variables is presented. The programming system consists of a number of independent processors, each performing a specific task. These processors, however, are interfaced through a well-organized data base, thus making the tasks of modifying, updating, or expanding the programming system much easier in a friendly environment provided by many inexpensive personal computers. The proposed software can be viewed as an important step in achieving a 'dummy' finite element for optimization. The programming system has been implemented on both large and small computers (such as VAX, CYBER, IBM-PC, and APPLE) although the focus is on the latter. Examples are presented to demonstrate the capabilities of the code. The present programming system can be used stand-alone or as part of the multilevel decomposition procedure to obtain optimum design for very large scale structural systems. Furthermore, other related research areas such as developing optimization algorithms (or in the larger level: a structural synthesis program) for future trends in using parallel computers may also benefit from this study.

Nguyen, D. T.↗

Incorporation of a progressive failure analysis method in the CSM testbed software system

Analysis of the postbuckling behavior of composite shell structures pose many difficult and challenging problems in the field of structural mechanics. Current analysis methods perform well for most cases in predicting the postbuckling response of undamaged components. To predict component behavior accurately at higher load levels, the analysis must include the effects of local material failures. The CSM testbed software system is a highly modular structural analysis system currently under development at Langley Research Center. One of the primary goals of the CSM testbed is to provide a software environment for the development of advanced structural analysis methods and modern numerical methods which will exploit advanced computer architecture such as parallel-vector processors. Development of a progressive failure analysis method consists of the design and implementation of a processor which will perform the ply-level progressive failure analysis and the development of a geometrically nonlinear analysis procedure which incorporates the progressive failure processor. Regarding the development of the progressive failure processor, two components are required: failure criteria and a degradation model. For the initial implementation, the failure criteria of Hashin will be used. For a matrix failure which typically indicates the development of transverse matrix cracks, the ply properties will be degraded. Work to date includes the design of the progressive failure analysis processor and initial plans for the controlling geometrically nonlinear analysis procedure. The implementation of the progressive failure analysis has begun. Access to the model database and the Hashin failure criteria are completed. Work is in progress on the input/output operations for the processor related data and the finite element model updating procedures. In total the progressive failure processor is approximately one-third complete.

Arenburg, Robert T.↗

Fast Multipole Methods for Three-Dimensional N-body Problems

We are developing computational tools for the simulations of three-dimensional flows past bodies undergoing arbitrary motions. High resolution viscous vortex methods have been developed that allow for extended simulations of two-dimensional configurations such as vortex generators. Our objective is to extend this methodology to three dimensions and develop a robust computational scheme for the simulation of such flows. A fundamental issue in the use of vortex methods is the ability of employing efficiently large numbers of computational elements to resolve the large range of scales that exist in complex flows. The traditional cost of the method scales as Omicron (N(sup 2)) as the N computational elements/particles induce velocities at each other, making the method unacceptable for simulations involving more than a few tens of thousands of particles. In the last decade fast methods have been developed that have operation counts of Omicron (N log N) or Omicron (N) (referred to as BH and GR respectively) depending on the details of the algorithm. These methods are based on the observation that the effect of a cluster of particles at a certain distance may be approximated by a finite series expansion. In order to exploit this observation we need to decompose the element population spatially into clusters of particles and build a hierarchy of clusters (a tree data structure) - smaller neighboring clusters combine to form a cluster of the next size up in the hierarchy and so on. This hierarchy of clusters allows one to determine efficiently when the approximation is valid. This algorithm is an N-body solver that appears in many fields of engineering and science. Some examples of its diverse use are in astrophysics, molecular dynamics, micro-magnetics, boundary element simulations of electromagnetic problems, and computer animation. More recently these N-body solvers have been implemented and applied in simulations involving vortex methods. Koumoutsakos and Leonard (1995) implemented the GR scheme in two dimensions for vector computer architectures allowing for simulations of bluff body flows using millions of particles. Winckelmans presented three-dimensional, viscous simulations of interacting vortex rings, using vortons and an implementation of a BH scheme for parallel computer architectures. Bhatt presented a vortex filament method to perform inviscid vortex ring interactions, with an alternative implementation of a BH scheme for a Connection Machine parallel computer architecture.

Koumoutsakos, P.↗

Spatially Accelerated Winding Numbers for Curved Geometry

The generalized winding number (GWN) is a scalar field that supports robust containment queries on curved geometry, including non-watertight, overlapping, and nested boundary representations. While queries can be easily parallelized over samples, direct evaluation on parametric curves and surfaces remains costly for large and complex models. Fast, state-of-the-art GWN approaches leverage a spatial index to approximate the GWN, typically coupled with a Taylor expansion which approximates the GWN contribution for far clusters of geometric primitives. However, such methods operate only on discrete inputs such as triangle meshes and point clouds, and would introduce containment errors near boundaries if applied to curved input. We extend support for fast GWN evaluation over arbitrary collections of NURBS curves in 2D and trimmed NURBS patches in 3D via a Bounding Volume Hierarchy that stores efficiently precomputed moment data in the hierarchy nodes. When querying the hierarchy, approximations for far clusters are used alongside direct evaluation for nearby NURBS primitives, achieving sub-linear complexity while preserving the geometric features in the vicinity of the query point. Central to our performance improvements is an adaptive subdivision strategy for NURBS primitives during a preprocessing phase, creating better spatial partitions while retaining the same accuracy for containment decisions as a direct evaluation. We demonstrate the performance and accuracy of our approach across a large collection of 2D and 3D datasets.

Computer science↗

Rao-Blackwellization for Adaptive Gaussian Sum Nonlinear Model Propagation

When dealing with imperfect data and general models of dynamic systems, the best estimate is always sought in the presence of uncertainty or unknown parameters. In many cases, as the first attempt, the Extended Kalman filter (EKF) provides sufficient solutions to handling issues arising from nonlinear and non-Gaussian estimation problems. But these issues may lead unacceptable performance and even divergence. In order to accurately capture the nonlinearities of most real-world dynamic systems, advanced filtering methods have been created to reduce filter divergence while enhancing performance. Approaches, such as Gaussian sum filtering, grid based Bayesian methods and particle filters are well-known examples of advanced methods used to represent and recursively reproduce an approximation to the state probability density function (pdf). Some of these filtering methods were conceptually developed years before their widespread uses were realized. Advanced nonlinear filtering methods currently benefit from the computing advancements in computational speeds, memory, and parallel processing. Grid based methods, multiple-model approaches and Gaussian sum filtering are numerical solutions that take advantage of different state coordinates or multiple-model methods that reduced the amount of approximations used. Choosing an efficient grid is very difficult for multi-dimensional state spaces, and oftentimes expensive computations must be done at each point. For the original Gaussian sum filter, a weighted sum of Gaussian density functions approximates the pdf but suffers at the update step for the individual component weight selections. In order to improve upon the original Gaussian sum filter, Ref. [2] introduces a weight update approach at the filter propagation stage instead of the measurement update stage. This weight update is performed by minimizing the integral square difference between the true forecast pdf and its Gaussian sum approximation. By adaptively updating each component weight during the nonlinear propagation stage an approximation of the true pdf can be successfully reconstructed. Particle filtering (PF) methods have gained popularity recently for solving nonlinear estimation problems due to their straightforward approach and the processing capabilities mentioned above. The basic concept behind PF is to represent any pdf as a set of random samples. As the number of samples increases, they will theoretically converge to the exact, equivalent representation of the desired pdf. When the estimated qth moment is needed, the samples are used for its construction allowing further analysis of the pdf characteristics. However, filter performance deteriorates as the dimension of the state vector increases. To overcome this problem Ref. [5] applies a marginalization technique for PF methods, decreasing complexity of the system to one linear and another nonlinear state estimation problem. The marginalization theory was originally developed by Rao and Blackwell independently. According to Ref. [6] it improves any given estimator under every convex loss function. The improvement comes from calculating a conditional expected value, often involving integrating out a supportive statistic. In other words, Rao-Blackwellization allows for smaller but separate computations to be carried out while reaching the main objective of the estimator. In the case of improving an estimator's variance, any supporting statistic can be removed and its variance determined. Next, any other information that dependents on the supporting statistic is found along with its respective variance. A new approach is developed here by utilizing the strengths of the adaptive Gaussian sum propagation in Ref. [2] and a marginalization approach used for PF methods found in Ref. [7]. In the following sections a modified filtering approach is presented based on a special state-space model within nonlinear systems to reduce the dimensionality of the optimization problem in Ref. [2]. First, the adaptive Gaussian sum propagation is explained and then the new marginalized adaptive Gaussian sum propagation is derived. Finally, an example simulation is presented.

state estimation↗

Measurement and Computation of Supersonic Flow in a Lobed Diffuser-Mixer for Trapped Vortex Combustors

The trapped vortex combustor (TVC) pioneered by Air Force Research Laboratories (AFRL) is under consideration as an alternative to conventional gas turbine combustors. The TVC has demonstrated excellent operational characteristics such as high combustion efficiency, low NO(x) emissions, effective flame stabilization, excellent high-altitude relight capability, and operation in the lean-burn or rich burn-quick quench-lean burn (RQL) modes of combustion. It also has excellent potential for lowering the engine combustor weight. This performance at low to moderate combustor mach numbers has stimulated interest in its ability to operate at higher combustion mach number, and for aerospace, this implies potentially higher flight mach numbers. To this end, a lobed diffuser-mixer that enhances the fuel-air mixing in the TVC combustor core was designed and evaluated, with special attention paid to the potential shock system entering the combustor core. For the present investigation, the lobed diffuser-mixer combustor rig is in a full annular configuration featuring sixfold symmetry among the lobes, symmetry within each lobe, and plain parallel, symmetric incident flow. During hardware cold-flow testing, significant discrepancies were found between computed and measured values for the pitot-probe-averaged static pressure profiles at the lobe exit plane. Computational fluid dynamics (CFD) simulations were initiated to determine whether the static pressure probe was causing high local flow-field disturbances in the supersonic flow exiting the diffuser-mixer and whether shock wave impingement on the pitot probe tip, pressure ports, or surface was the cause of the discrepancies. Simulations were performed with and without the pitot probe present in the modeling. A comparison of static pressure profiles without the probe showed that static pressure was off by nearly a factor of 2 over much of the radial profile, even when taking into account potential axial displacement of the probe by up to 0.25 in. (0.64 cm). Including the pitot probe in the CFD modeling and data interpretation lead to good agreement between measurement and prediction. Graphical inspection of the results showed that the shock waves impinging on the probe surface were highly nonuniform, with static pressure varying circumferentially among the pressure ports by over 10 percent in some cases. As part of the measurement methodology, such measurements should be routinely supplemented with CFD analyses that include the pitot probe as part of the flow-path geometry.

Brankovic, Andreja↗

Range-Measuring Video Sensors

Optoelectronic sensors of a proposed type would perform the functions of both electronic cameras and triangulation- type laser range finders. That is to say, these sensors would both (1) generate ordinary video or snapshot digital images and (2) measure the distances to selected spots in the images. These sensors would be well suited to use on robots that are required to measure distances to targets in their work spaces. In addition, these sensors could be used for all the purposes for which electronic cameras have been used heretofore. The simplest sensor of this type, illustrated schematically in the upper part of the figure, would include a laser, an electronic camera (either video or snapshot), a frame-grabber/image-capturing circuit, an image-data-storage memory circuit, and an image-data processor. There would be no moving parts. The laser would be positioned at a lateral distance d to one side of the camera and would be aimed parallel to the optical axis of the camera. When the range of a target in the field of view of the camera was required, the laser would be turned on and an image of the target would be stored and preprocessed to locate the angle (a) between the optical axis and the line of sight to the centroid of the laser spot.

Howard, Richard T.↗

High Yield Xray Imager Final Design Review

The High Yield Xray Imager (HYXI) is a new NIF target diagnostic system currently under development. The goal of HYXI is to provide high-fidelity, high temporal resolution x-ray imaging capability on high yield NIF implosions at 10MJ and above. The HYXI instrument design concept is based on the combination of two technologies that have been successfully utilized at the NIF on previous instruments, electron pulse-dilation and hybrid-CMOS sensor imaging. The combination of these two techniques will give HYXI sufficient data quality to ascertain differences in hot spot formation dynamics between high and low yield implosions. This information will highlight the critical hot spot conditions needed for ignition and burn. The HYXI design leverages the successful operation of the PDIXI x-ray imager at the NIF on multi MJ yield shots. A new radiation tolerant CMOS imaging array (HYPERION) is being developed to eliminate the significant background noise which limits the data quality of PDIXI. We successfully placed the contract with Advanced hCMOS Systems (AHS) to develop the HYPERION sensor, which fulfils our criteria to place long lead time item procurements by end of FY24. The HYXI Final Design Review was completed at the end of Q4 FY24 (Sep 24 th and Sep 30 th ). The HYXI project is a multi-year effort with a phased approach to be bring up system functionality over time in parallel with the development and fabrication effort of the HYPERION CMOS imaging array. In Phase 1, time-integrated x-ray images on NIF DT experiments will be collected starting in Q3 FY25. In Phase 2 of the project, time-resolved imaging with HYXI utilizing a spare microchannel plate detector back-end will begin in Q3 FY26. Phase 3 concludes the project with the installation of the HYPERION sensor array and the final performance qualification of the HYXI instrument which is scheduled for Q3 FY27 as discussed in the PDR and MRT report on this project in FY23.

42 ENGINEERING↗

Simultaneous measurements of waves and precipitating electrons near the equator in the outer radiation belt

An investigation of wave-particle interactions is made using several simultaneous electron and wave measurements performed at near-equatorial positions from the Combined Release and Radiation Effects Satellite (CRRES) satellite. Bursts of electron precipitation were observed, most frequently at local times near dawn. Examples of bursts are presented in which the fluxes of the precipitating electrons and the wave intensities are correlated with coefficients as high as 0.7. During bursts the frequencies of the enhanced waves spanned a wide range from 311 Hz to 3.11 kHz, and the energies of the enhanced electrons were in the range 1.7 keV to 288 keV. The changes of the precipitating fluxes were generally less pronounced at the lowest energies. On the basis of electron-cyclotron resonant calculations using the cold plasma densities and ambient magnetic fields taken from the CRRES measurements it was found that the wave frequencies and precipitating electron energies were generally consistent with those expected from electron resonance with parallel propagating whistler waves. The electron data of principal concern here were acquired in and about the loss cone with narrow angular resolution spectrometers covering the energy range 340 eV to 5 MeV. The wave data included electric field measurements spanning frequencies from 5 Hz to 400 kHz and magnetic field measurements from 5 Hz to 10 kHz.

Imhof, W. L.↗

Cooperative Three-Robot System for Traversing Steep Slopes

Teamed Robots for Exploration and Science in Steep Areas (TRESSA) is a system of three autonomous mobile robots that cooperate with each other to enable scientific exploration of steep terrain (slope angles up to 90 ). Originally intended for use in exploring steep slopes on Mars that are not accessible to lone wheeled robots (Mars Exploration Rovers), TRESSA and systems like TRESSA could also be used on Earth for performing rescues on steep slopes and for exploring steep slopes that are too remote or too dangerous to be explored by humans. TRESSA is modeled on safe human climbing of steep slopes, two key features of which are teamwork and safety tethers. Two of the autonomous robots, denoted Anchorbots, remain at the top of a slope; the third robot, denoted the Cliffbot, traverses the slope. The Cliffbot drives over the cliff edge supported by tethers, which are payed out from the Anchorbots (see figure). The Anchorbots autonomously control the tension in the tethers to counter the gravitational force on the Cliffbot. The tethers are payed out and reeled in as needed, keeping the body of the Cliffbot oriented approximately parallel to the local terrain surface and preventing wheel slip by controlling the speed of descent or ascent, thereby enabling the Cliffbot to drive freely up, down, or across the slope. Due to the interactive nature of the three-robot system, the robots must be very tightly coupled. To provide for this tight coupling, the TRESSA software architecture is built on a combination of (1) the multi-robot layered behavior-coordination architecture reported in "An Architecture for Controlling Multiple Robots" (NPO-30345), NASA Tech Briefs, Vol. 28, No. 10 (October 2004), page 65, and (2) the real-time control architecture reported in "Robot Electronics Architecture" (NPO-41784), NASA Tech Briefs, Vol. 32, No. 1 (January 2008), page 28. The combination architecture makes it possible to keep the three robots synchronized and coordinated, to use data from all three robots for decision- making at each step, and to control the physical connections among the robots. In addition, TRESSA (as in prior systems that have utilized this architecture) , incorporates a capability for deterministic response to unanticipated situations from yet another architecture reported in Control Architecture for Robotic Agent Command and Sensing (NPO-43635), NASA Tech Briefs, Vol. 32, No. 10 (October 2008), page 40. Tether tension control is a major consideration in the design and operation of TRESSA. Tension is measured by force sensors connected to each tether at the Cliffbot. The direction of the tension (both azimuth and elevation) is also measured. The tension controller combines a controller to counter gravitational force and an optional velocity controller that anticipates the motion of the Cliffbot. The gravity controller estimates the slope angle from the inclination of the tethers. This angle and the weight of the Cliffbot determine the total tension needed to counteract the weight of the Cliffbot. The total needed tension is broken into components for each Anchorbot. The difference between this needed tension and the tension measured at the Cliffbot constitutes an error signal that is provided to the gravity controller. The velocity controller computes the tether speed needed to produce the desired motion of the Cliffbot. Another major consideration in the design and operation of TRESSA is detection of faults. Each robot in the TRESSA system monitors its own performance and the performance of its teammates in order to detect any system faults and prevent unsafe conditions. At startup, communication links are tested and if any robot is not communicating, the system refuses to execute any motion commands. Prior to motion, the Anchorbots attempt to set tensions in the tethers at optimal levels for counteracting the weight of the Cliffbot; if either Anchorbot fails to reach its optimal tension level within a specified time, it sends message to the other robots and the commanded motion is not executed. If any mechanical error (e.g., stalling of a motor) is detected, the affected robot sends a message triggering stoppage of the current motion. Lastly, messages are passed among the robots at each time step (10 Hz) to share sensor information during operations. If messages from any robot cease for more than an allowable time interval, the other robots detect the communication loss and initiate stoppage.

Stroupe, Ashley↗

CG-Kit: Code Generation Toolkit for performant and maintainable variants of source code applied to Flash-X hydrodynamics simulations

CG-Kit is a new Code Generation tool-Kit that we have developed as a part of the solution for portability and maintainability for multiphysics computing applications. The development of CG-Kit is rooted in the urgent need created by the shifting landscape of high-performance computing platforms and the algorithmic complexities of a particular large-scale multiphysics application: Flash-X. To efficiently use computing resources on a heterogeneous node, an application must have a map of computation to resources and a mechanism to move the data and computation to the resources according to the map. Most existing performance portability solutions are focussed on abstracting the expression of computations so that a unified source code can be specialized to run on different resources. However, such an approach is insufficient for a code like Flash-X, which has a multitude of code components that can be assembled in various permutations and combinations to form different instances of applications. Similar challenges apply to any code that has composability, where a single specified way of apportioning work among devices may not be optimal. Additionally, use cases arise where the optimal control flow of computation may differ for different devices while the underlying numerics remain identical. This combination leads to unique challenges including handling an existing large code base in Fortran and/or C/C++, subdivision of code into a great variety of units supporting a wide range of physics and numerical methods, different parallelization techniques for distributed and shared memory systems and accelerator devices, and heterogeneity of computing platforms requiring coexisting variants of parallel algorithms. All of these challenges demand that scientific software developers apply existing knowledge about domain applications, algorithms, and computing platforms to determine custom abstractions and granularity for code generation. There is a critical lack of tools to tackle those problems. CG-Kit is designed to fill this gap by providing a user with the ability to express their desired control flow and computation-to-resource map in the form a pseudocode-like recipe. It consists of standalone tools that can be combined into highly specific and, we argue, highly effective portability and maintainability toolchains. Here we present the design of our new tools: parametrized source trees, control flow graphs, and recipes. The tools are implemented in Python. They are agnostic to the programming language of the source code targeted for code generation. In conclusion, we demonstrate the capabilities of the toolkit with two examples, first, multithreaded variants of the basic AXPY operation, and second, variants of parallel algorithms within a hydrodynamics solver, called Spark, from Flash-X that operates on block-structured adaptive meshes.

Algorithmic portability↗

Experimental Validation of Numerical Simulations for an Acoustic Liner in Grazing Flow

A coordinated experimental and numerical simulation effort is carried out to improve our understanding of the physics of acoustic liners in a grazing flow as well our computational aeroacoustics (CAA) method prediction capability. A numerical simulation code based on advanced CAA methods is developed. In a parallel effort, experiments are performed using the Grazing Flow Impedance Tube at the NASA Langley Research Center. In the experiment, a liner is installed in the upper wall of a rectangular flow duct with a 2 inch by 2.5 inch cross section. Spatial distribution of sound pressure levels and relative phases are measured on the wall opposite the liner in the presence of a Mach 0.3 grazing flow. The computer code is validated by comparing computed results with experimental measurements. Good agreements are found. The numerical simulation code is then used to investigate the physical properties of the acoustic liner. It is shown that an acoustic liner can produce self-noise in the presence of a grazing flow and that a feedback acoustic resonance mechanism is responsible for the generation of this liner self-noise. In addition, the same mechanism also creates additional liner drag. An estimate, based on numerical simulation data, indicates that for a resonant liner with a 10% open area ratio, the drag increase would be about 4% of the turbulent boundary layer drag over a flat wall.

Tam, Christopher K. W.↗

Predicting Ares I Reaction Control System Performance by Utilizing Analysis Anchored with Development Test Data

The Ares I launch vehicle is an integral part of NASA s Constellation Program, providing a foundation for a new era of space access. The Ares I is designed to lift the Orion Crew Module and will enable humans to return to the Moon as well as explore Mars.1 The Ares I is comprised of two inline stages: a Space Shuttle-derived five-segment Solid Rocket Booster (SRB) First Stage (FS) and an Upper Stage (US) powered by a Saturn V-derived J-2X engine. A dedicated Roll Control System (RoCS) located on the connecting interstage provides roll control prior to FS separation. Induced yaw and pitch moments are handled by the SRB nozzle vectoring. The FS SRB operates for approximately two minutes after which the US separates from the vehicle and the US Reaction Control System (ReCS) continues to provide reaction control for the remainder of the mission. A representation of the Ares I launch vehicle in the stacked configuration and including the Orion Crew Exploration Vehicle (CEV) is shown in Figure 1. Each Reaction Control System (RCS) design incorporates a Gaseous Helium (GHe) pressurization system combined with a monopropellant Hydrazine (N2H4) propulsion system. Both systems have two diametrically opposed thruster modules. This architecture provides one failure tolerance for function and prevention of catastrophic hazards such as inadvertent thruster firing, bulk propellant leakage, and over-pressurization. The pressurization system on the RoCS includes two ambient pressure-referenced regulators on parallel strings in order to attain the required system level single Fault Tolerant (FT) design for function while the ReCS utilizes a blow-down approach. A single burst disk and relief valve assembly is also included on the RoCS to ensure single failure tolerance for must-not-occur catastrophic hazards. The Reaction Control Systems are designed to support simultaneously firing multiple thrusters as required

Stein, William B.↗

BM3DORNL

BM3DORNL is a high-performance, open-source library for removing streak and ring artifacts from computed-tomography (CT) data, developed for neutron imaging at Oak Ridge National Laboratory's Spallation Neutron Source (VENUS beamline) and applicable to X-ray CT as well. Ring artifacts — concentric rings in reconstructed slices caused by detector pixel-to-pixel response non-uniformities — appear as vertical streaks in the sinogram and degrade both image quality and quantitative analysis. BM3DORNL operates in the sinogram domain using an adaptation of the BM3D (block-matching and 3D collaborative filtering) algorithm (Dabov et al., 2007). It provides a dedicated streak-removal mode, a true multi-scale BM3D variant (after Mäkinen et al., 2021) that suppresses wide streaks single-scale methods miss, and an alternative Fourier–SVD method (~2.6× faster) combining FFT-based energy detection with rank-1 SVD. The computationally intensive core is implemented in Rust with parallel (Rayon) block matching, integral-image pre-screening, and optimized transforms, and is exposed through a simple Python API (with an optional GUI) so it integrates directly into existing tomography reconstruction pipelines. It processes both 2D sinograms and 3D sinogram stacks, is pip-installable for Linux and macOS, and is documented at https://bm3dornl.readthedocs.io.

Zhang, Chen [Oak Ridge National Laboratory (ORNL),↗

Cloud Thermodynamic Phase Detection with Polarimetrically Sensitive Passive Sky Radiometers

The primary goal of this project has been to investigate if ground-based visible and near-infrared passive radiometers that have polarization sensitivity can determine the thermodynamic phase of overlying clouds, i.e. if they are comprised of liquid droplets or ice particles. While this knowledge is important by itself for our understanding of the global climate, it can also help improve cloud property retrieval algorithms that use total (unpolarized) radiance to determine Cloud Optical Depth (COD). This is a potentially unexploited capability of some instruments in the NASA Aerosol Robotic Network (AERONET), which, if practical, could expand the products of that global instrument network at minimal additional cost. We performed simulations that found, for zenith observations, cloud thermodynamic phase is often expressed in the sign of the Q component of the Stokes polarization vector. We chose our reference frame as the plane containing solar and observation vectors, so the sign of Q indicates the polarization direction, parallel (negative) or perpendicular (positive) to that plane. Since the quantity of polarization is inversely proportional to COD, optically thin clouds are most likely to create a signal greater than instrument noise. Besides COD and instrument accuracy, other important factors for the determination of cloud thermodynamic phase are the solar and observation geometry (scattering angles between 40 and 60 degrees are best), and the properties of ice particles (pristine particles may have halos or other features that make them difficult to distinguish from water droplets at specific scattering angles, while extreme ice crystal aspect ratios polarize more than compact particles). We tested the conclusions of our simulations using data from polarimetrically sensitive versions of the Cimel 318 sun photometerradiometer that comprise AERONET. Most algorithms that exploit Cimel polarized observations use the Degree of Linear Polarization (DoLP), not the individual Stokes vector elements (such as Q). For this reason, we had no information about the accuracy of Cimel observed Q and the potential for cloud phase determination. Indeed, comparisons to ceilometer observations with a single polarized spectral channel version of the Cimel at a site in the Netherlands showed little correlation. Comparisons to Lidar observations with a more recently developed, multi-wavelength polarized Cimel in Maryland, USA, show more promise. This divergence between simulations and observations has prompted us to begin the development of a small test instrument called the Sky Polarization Radiometric Instrument for Test and Evaluation (SPRITE). This instrument is specifically devoted to the accurate observation of Q, and the testing of calibration and uncertainty assessment techniques, with the ultimate goal of understanding the practical feasibility of these measurements.

Knobelspiesse, Kirk D.↗