Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel Performance Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48

Embedded Sensors for Measuring Surface Regression

The development and evaluation of new hybrid and solid rocket motors requires accurate characterization of the propellant surface regression as a function of key operational parameters. These characteristics establish the propellant flow rate and are prime design drivers affecting the propulsion system geometry, size, and overall performance. There is a similar need for the development of advanced ablative materials, and the use of conventional ablatives exposed to new operational environments. The Miniature Surface Regression Sensor (MSRS) was developed to serve these applications. It is designed to be cast or embedded in the material of interest and regresses along with it. During this process, the resistance of the sensor is related to its instantaneous length, allowing the real-time thickness of the host material to be established. The time derivative of this data reveals the instantaneous surface regression rate. The MSRS could also be adapted to perform similar measurements for a variety of other host materials when it is desired to monitor thicknesses and/or regression rate for purposes of safety, operational control, or research. For example, the sensor could be used to monitor the thicknesses of brake linings or racecar tires and indicate when they need to be replaced. At the time of this reporting, over 200 of these sensors have been installed into a variety of host materials. An MSRS can be made in either of two configurations, denoted ladder and continuous (see Figure 1). A ladder MSRS includes two highly electrically conductive legs, across which narrow strips of electrically resistive material are placed at small increments of length. These strips resemble the rungs of a ladder and are electrically equivalent to many tiny resistors connected in parallel. A substrate material provides structural support for the legs and rungs. The instantaneous sensor resistance is read by an external signal conditioner via wires attached to the conductive legs on the non-eroding end of the sensor. The sensor signal can be transmitted from inside a high-pressure chamber to the ambient environment, using commercially available feedthrough connectors. Miniaturized internal recorders or wireless data transmission could also potentially be employed to eliminate the need for producing penetrations in the chamber case. The rungs are designed so that as each successive rung is eroded away, the resistance changes by an amount that yields a readily measurable signal larger than the background noise. (In addition, signal-conditioning techniques are used in processing the resistance readings to mitigate the effect of noise.) Hence, each discrete change of resistance serves to indicate the arrival of the regressing host material front at the known depth of the affected resistor rung. The average rate of regression between two adjacent resistors can be calculated simply as the distance between the resistors divided by the time interval between their resistance jumps. Advanced data reduction techniques have also been developed to establish the instantaneous surface position and regression rate when the regressing front is between rungs.

Gramer, Daniel J.↗

Effect of Anoxic Iron Corrosion on WIPP Brine Geochemistry FY23 Final Report (U)

A 280-day study was completed to evaluate the effect of zero-valent iron (Fe 0 ) on the Waste Isolation Pilot Plant (WIPP) brine geochemistry under anticipated reducing conditions. Hydrogen (H 2 ) gas is expected to be present in the repository after closure due to the anoxic corrosion of a vast quantity of iron contained in the waste forms disposed at WIPP; therefore, a background argon atmosphere containing H 2 was chosen for this study. WIPP groundwater brine pH and E h will impact the mobility and fate of plutonium within the repository. Modeling and laboratory results for Castile WIPP brine indicate that equilibrium fa values relative to the standard hydrogen electrode (SHE) are 40 mV more reducing (i.e., more negative) than those for Salado WIPP brine (-480 mV vs. -440 mV, respectively) because of the higher pH of the Castile brine (pH 9 .3 for Castile vs. pH 8.8 for Salado). The E h and pH data were corrected for the effects of high ionic strength. The experimental results for both brines are consistent with thermodynamic predictions using OLI Systems' Mixed Solvent Electrolyte chemical equilibrium model. The measured and corrected pH and E h data from this study are provided in Table ES-I and Table ES-2, respectively. The experimental study, with four test conditions in triplicate, was performed in a dual glovebox with a nominally 3 vol.% H 2 in argon atmosphere (target H 2 range: 3 ± I vol.%). Simulants containing MgO only ( experimental control) and MgO+Fe 0 (WIPP base case) were prepared for both the Salado and Castile brines. MgO was included in all simulants to account for the use of bulk magnesium oxide in the WIPP repository. Fe 0 was included in some simulants to incorporate the effects of the anoxic corrosion of iron and in-situ hydrogen generation in the study. The brine compositions were developed by Sandia National Laboratory (SNL; Xiong, 2008) and have been used in previous WIPP evaluations. The test method (agitation, etc.) is partially based on ASTM D3987-12. Twelve rounds of periodic measurements of pH and E h were performed over the course of the study. Chemical analysis results for liquids and solids (ICP-MS, ICP-ES, IC Anion, TIC, SEM-EDX) are consistent with the pH, E h , and thermodynamic modeling results. This study included the following conditions that deviate from anticipated post-closure conditions following brine intrusion, but were selected to facilitate bench-scale testing to validate modeling of pH and E h for the post-closure WIP P repository: an anoxic glove box atmosphere containing ≤ 4 vol. % H 2 vs. substantially higher H 2 gas concentrations assumed in the WIPP Performance Assessment (PA); a significantly higher liquid-to-solid test ratio compared to the much lower phase ratio anticipated in the WIP P repository; agitation of the simulant bottles to maximize mass transfer; and finally the use of Fe 0 reagents having a much greater surface area than expected in the WIP P repository. Non-representative conditions were chosen for various reasons such as: to provide bounding conservative results, to provide a margin of safety for testing, or to facilitate simulant sub-sampling and analysis. In a parallel effort, aqueous electrolyte thermodynamic models were developed for the synthetic Salado and Castile brines to inform the experimental design, facilitate laboratory data interpretation, and allow extension of evaluations beyond the parameters tested. Thermodynamic modeling simulations including the MgO and Fe 0 additives that are directly relevant to the experimental measurements (e.g., pH calibration curve, ORP corrections) are included in this report. The measured fa of the simulants was close to the OLI model predictions for both brines and was largely controlled by the background H 2 partial pressure in the vapor phase as well as H 2 generated in situ in the aqueous phase by the Fe 0 corrosion. The H 2 gas-phase concentration tested and thermodynamically evaluated was much lower than is assumed in the WIPP PA; however, H 2 (g) concentrations significantly below this level are still predicted to result in very reducing conditions. In conclusion: • The experimental results are consistent with thermodynamic model predictions for fa, pH, and the effects of high ionic strength. • Evidence to date suggests that the H2 concentration in the glovebox atmosphere ultimately determined the final E h values of the simulants and resulted in highly reducing conditions. As a result, little difference was observed between the control simulants containing only MgO and the WIPP base-case simulants that contained MgO and Fe 0 . • This test methodology is recommended for future studies evaluating WIPP repository conditions. The methodology includes: (1) background H 2 in argon with agitation ( or could alternatively include in-situ-generated H 2 in sealed bottles); (2) carefully measured and corrected ORP data ( with much effort focused on allowing the probes to fully stabilize); and (3) ionic-strength-corrected pH data. Other best practices, such as simulant sparging/handling, ORP probe replacement, etc., should also be considered. • The coupling of experimental studies and thermodynamic modeling is also highly recommended because these methods inform and direct one another leading to greater confidence in and understanding of the results.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

ELF waves and ion resonances produced by an electron beam emitting rocket in the ionosphere

Results are reported from the ECHO-6 electron-beam-injection experiment, performed in the auroral-zone ionosphere on March 30, 1983 using a sounding rocket equipped with two electron guns and a free-flying plasma-diagnostics instrument package. The data are presented in extensive graphs and diagrams and characterized in detail. Large ELF wave variations, superposed on the strong beam-sector-directed quasi-dc component, are observed in the 100-eV beam-induced plasma when the beam is injected in a transverse spiral, but not when it is injected upward parallel to the magnetic-field line. ELF activity is found to be suppressed whenever the rocket passed through field lines with auroral activity, suggesting that the waves are produced by the interaction of the beam potentials, plasma currents, and return currents neutralizing the accelerator payload.

Winckler, J. R.↗

A Feasibility Study for Perioperative Ventricular Tachycardia Prognosis and Detection and Noise Detection Using a Neural Network and Predictive Linear Operators

To locate the accessory pathway(s) in preexicitation syndromes, epicardial and endocardial ventricular mapping is performed during anterograde ventricular activation via accessory pathway(s) from data originally received in signal form. As the number of channels increases, it is pertinent that more automated detection of coherent/incoherent signals is achieved as well as the prediction and prognosis of ventricular tachywardia (VT). Today's computers and computer program algorithms are not good in simple perceptual tasks such as recognizing a pattern or identifying a sound. This discrepancy, among other things, has been a major motivating factor in developing brain-based, massively parallel computing architectures. Neural net paradigms have proven to be effective at pattern recognition tasks. In signal processing, the picking of coherent/incoherent signals represents a pattern recognition task for computer systems. The picking of signals representing the onset ot VT also represents such a computer task. We attacked this problem by defining four signal attributes for each potential first maximal arrival peak and one signal attribute over the entire signal as input to a back propagation neural network. One attribute was the predicted amplitude value after the maximum amplitude over a data window. Then, by using a set of known (user selected) coherent/incoherent signals, and signals representing the onset of VT, we trained the back propagation network to recognize coherent/incoherent signals, and signals indicating the onset of VT. Since our output scheme involves a true or false decision, and since the output unit computes values between 0 and 1, we used a Fuzzy Arithmetic approach to classify data as coherent/incoherent signals. Furthermore, a Mean-Square Error Analysis was used to determine system stability. The neural net based picking coherent/incoherent signal system achieved high accuracy on picking coherent/incoherent signals on different patients. The system also achieved a high accuracy of picking signals which represent the onset of VT, that is, VT immediately followed these signals. A special binary representation of the input and output data allowed the neural network to train very rapidly as compared to another standard decimal or normalized representations of the data.

Moebes, T. A.↗

Concurrent Probabilistic Simulation of High Temperature Composite Structural Response

A computational structural/material analysis and design tool which would meet industry's future demand for expedience and reduced cost is presented. This unique software 'GENOA' is dedicated to parallel and high speed analysis to perform probabilistic evaluation of high temperature composite response of aerospace systems. The development is based on detailed integration and modification of diverse fields of specialized analysis techniques and mathematical models to combine their latest innovative capabilities into a commercially viable software package. The technique is specifically designed to exploit the availability of processors to perform computationally intense probabilistic analysis assessing uncertainties in structural reliability analysis and composite micromechanics. The primary objectives which were achieved in performing the development were: (1) Utilization of the power of parallel processing and static/dynamic load balancing optimization to make the complex simulation of structure, material and processing of high temperature composite affordable; (2) Computational integration and synchronization of probabilistic mathematics, structural/material mechanics and parallel computing; (3) Implementation of an innovative multi-level domain decomposition technique to identify the inherent parallelism, and increasing convergence rates through high- and low-level processor assignment; (4) Creating the framework for Portable Paralleled architecture for the machine independent Multi Instruction Multi Data, (MIMD), Single Instruction Multi Data (SIMD), hybrid and distributed workstation type of computers; and (5) Market evaluation. The results of Phase-2 effort provides a good basis for continuation and warrants Phase-3 government, and industry partnership.

Abdi, Frank↗

Code Development for Unsteady Inlet Flows using Parallel Processing, Iced Airfoils

As a part of an effort for the development of an "Integrated Solution Process for Iced Airfoils" which combines a CFD (Computational Fluid Dynamics) code and an ice accretion code for accurate prediction of ice growth and performance degradation, a study on the effect of iced-geometry-smoothing was initiated. As a first step in this study, the degree of smoothing was defined by the number of control points for the given iced airfoil geometry. Then, reducing these number of control points in a systematical way provided various degrees of grid generation. This study will be continued by comparing CFD computed data such as pressure, lift, drag, flow separation, and wake flow patterns between the iced airfoils against any existing experimental data.

Chung, Joongkee↗

Automatic Data Traffic Control on DSM Architecture

We study data traffic on distributed shared memory machines and conclude that data placement and grouping improve performance of scientific codes. We present several methods which user can employ to improve data traffic in his code. We report on implementation of a tool which detects the code fragments causing data congestions and advises user on improvements of data routing in these fragments. The capabilities of the tool include deduction of data alignment and affinity from the source code; detection of the code constructs having abnormally high cache or TLB misses; generation of data placement constructs. We demonstrate the capabilities of the tool on experiments with NAS parallel benchmarks and with a simple computational fluid dynamics application ARC3D.

Frumkin, Michael↗

NASA Tech Briefs, March 2013

Topics covered include: Remote Data Access with IDL Data Compression Algorithm Architecture for Large Depth-of-Field Particle Image Velocimeters Vectorized Rebinning Algorithm for Fast Data Down-Sampling Display Provides Pilots with Real-Time Sonic-Boom Information Onboard Algorithms for Data Prioritization and Summarization of Aerial Imagery Monitoring and Acquisition Real-time System (MARS) Analog Signal Correlating Using an Analog-Based Signal Conditioning Front End Micro-Textured Black Silicon Wick for Silicon Heat Pipe Array Robust Multivariable Optimization and Performance Simulation for ASIC Design; Castable Amorphous Metal Mirrors and Mirror Assemblies; Sandwich Core Heat-Pipe Radiator for Power and Propulsion Systems; Apparatus for Pumping a Fluid; Cobra Fiber-Optic Positioner Upgrade; Improved Wide Operating Temperature Range of Li-Ion Cells; Non-Toxic, Non-Flammable, -80 C Phase Change Materials; Soft-Bake Purification of SWCNTs Produced by Pulsed Laser Vaporization; Improved Cell Culture Method for Growing Contracting Skeletal Muscle Models; Hand-Based Biometric Analysis; The Next Generation of Cold Immersion Dry Suit Design Evolution for Hypothermia Prevention; Integrated Lunar Information Architecture for Decision Support Version 3.0 (ILIADS 3.0); Relay Forward-Link File Management Services (MaROS Phase 2); Two Mechanisms to Avoid Control Conflicts Resulting from Uncoordinated Intent; XTCE GOVSAT Tool Suite 1.0; Determining Temperature Differential to Prevent Hardware Cross-Contamination in a Vacuum Chamber; SequenceL: Automated Parallel Algorithms Derived from CSP-NT Computational Laws; Remote Data Exploration with the Interactive Data Language (IDL); Mixture-Tuned, Clutter Matched Filter for Remote Detection of Subpixel Spectral Signals; Partitioned-Interval Quantum Optical Communications Receiver; and Practical UAV Optical Sensor Bench with Minimal Adjustability.

Source record↗

Influence of the preshock temperature on shock effects in quartz

Shock metamorphic features are the prime indicators for recognizing impact phenomena on Earth and other planetary bodies. Although the pressure dependence of shock features is well known, information about the influence of the preshock temperature is almost lacking. Especially in the case of large-scale impacts like Sudbury, it is expected that deep-seated crustal rocks were subjected to shock at elevated temperatures. Therefore, we continued to perform shock experiments at elevated temperatures on less than 0.5-mm thin disks of single crystal quartz cut parallel to the (1010) face. All recovered quartz samples were investigated by universal stage, spindle stage, and a newly developed density gradient technique. Errors of refractive index and density measurements are +/- 0.0005 and +/- 0.002 g/cu cm respectively. Our investigations indicate that shock metamorphic features are strongly dependent on the preshock temperature. This statement has far-reaching implications with respect to shock wave barometry that is based on data from recovery experiments at room temperature. These datasets might be applicable only to low-temperature target rocks. Moreover, this study demonstrates that shock recovery experiments are definitely required for understanding the complete pressure-temperature regime of shock metamorphism on planetary bodies.

Langenhorst, F.↗

Buoyancy Effects on Flow Structure and Instability of Low-Density Gas Jets

A low-density gas jet injected into a high-density ambient gas is known to exhibit self-excited global oscillations accompanied by large vortical structures interacting with the flow field. The primary objective of the proposed research is to study buoyancy effects on the origin and nature of the flow instability and structure in the near-field of low-density gas jets. Quantitative rainbow schlieren deflectometry, Computational fluid dynamics (CFD) and Linear stability analysis were the techniques employed to scale the buoyancy effects. The formation and evolution of vortices and scalar structure of the flow field are investigated in buoyant helium jets discharged from a vertical tube into quiescent air. Oscillations at identical frequency were observed throughout the flow field. The evolving flow structure is described by helium mole percentage contours during an oscillation cycle. Instantaneous, mean, and RMS concentration profiles are presented to describe interactions of the vortex with the jet flow. Oscillations in a narrow wake region near the jet exit are shown to spread through the jet core near the downstream location of the vortex formation. The effects of jet Richardson number on characteristics of vortex and flow field are investigated and discussed. The laminar, axisymmetric, unsteady jet flow of helium injected into air was simulated using CFD. Global oscillations were observed in the flow field. The computed oscillation frequency agreed qualitatively with the experimentally measured frequency. Contours of helium concentration, vorticity and velocity provided information about the evolution and propagation of vortices in the oscillating flow field. Buoyancy effects on the instability mode were evaluated by rainbow schlieren flow visualization and concentration measurements in the near-field of self-excited helium jets undergoing gravitational change in the microgravity environment of 2.2s drop tower at NASA John H. Glenn Research Center. The jet Reynolds number was varied from 200 to 1500 and jet Richardson number was varied from 0.72 to 0.002. Power spectra plots generated from Fast Fourier Transform (FFT) analysis of angular deflection data acquired at a temporal resolution of 1000Hz reveal substantial damping of the oscillation amplitude in microgravity at low Richardson numbers (~0.002). Quantitative concentration data in the form of spatial and temporal evolutions of the instability data in Earth gravity and microgravity reveal significant variations in the jet flow structure upon removal of buoyancy forces. Radial variation of the frequency spectra and time traces of helium concentration revealed the importance of gravitational effects in the jet shear layer region. Linear temporal and spatio-temporal stability analyses of a low-density round gas jet injected into a high-density ambient gas were performed by assuming hyper-tan mean velocity and density profiles. The flow was assumed to be non parallel. Viscous and diffusive effects were ignored. The mean flow parameters were represented as the sum of the mean value and a small normal-mode fluctuation. A second order differential equation governing the pressure disturbance amplitude was derived from the basic conservation equations. The effects of the inhomogeneous shear layer and the Froude number (signifying the effects of gravity) on the temporal and spatio-temporal results were delineated. A decrease in the density ratio (ratio of the density of the jet to the density of the ambient gas) resulted in an increase in the temporal amplification rate of the disturbances. The temporal growth rate of the disturbances increased as the Froude number was reduced. The spatio-temporal analysis performed to determine the absolute instability characteristics of the jet yield positive absolute temporal growth rates at all Fr and different axial locations. As buoyancy was removed (Fr . 8), the previously existing absolute instability disappeared at all locations establhing buoyancy as the primary instability mechanism in self-excited low-density jets.

Pasumarthi, Kasyap Sriramachandra↗

Characterization of Structural Vibration and Acoustic Radiation with a Beam-Array Doppler Vibrometer

This paper discusses the operational principles of an assembly of two opto-electronic modules with their configurations based on a Laser-Array Doppler Vibrometer (LADV) and a Shack-Hartmann Wavefront Sensor (SHWFS). Both sensors can operate concurrently providing complementary data. While the SHWFS is a well-established system that enables low-bandwidth spatio-temporal characterization of aerodynamic turbulence, the LADV technology allows real-time detection of structural vibrations and associated acoustic radiation fields. Practical examples are shown of the instruments’ capabilities, focusing on the ability to simultaneously capture, visualize and quantitatively characterize full-field non-stationary structural dynamics and unsteady sound fields or transient flow fields around ground test facility airframe models or other structures of interest. The parallel multi-channel architecture of the LADV system enables synchronous detection and temporal correlation of indicators related to the spatio-temporal vibration of the structure under inspection. This unique feature is essential for real time detection and categorization of the structural and acoustic dynamics of transient events. For the present study, the LADV data have been successfully compared with measurements obtained with the SHWFS in the wake of a subsonic scale-model airfoil. The ability of the LADV-SHWFS coupled measurements to perform real time non-intrusive evaluation and characterization of dynamic processes at operationally relevant bandwidths should provide a deeper insight into the complex structural dynamics that contribute to radiated sound fields.

Vladimir B Markov↗

Parallel High-Order Anisotropic Meshing Using Discrete Metric Tensors

This paper presents a metric-aligned meshing algorithm that relies on the Lp-Centroidal Voronoi Tesselation approach. A prototype of this algorithm was first presented at the Scitech conference of 2018 and this work is an extension to that paper. At the end of the previously presented work, a set of problems were mentioned which we are trying to address in this paper. First, we show a significant improvement in code performance since we were limited to present relatively benign (analytical) test cases. Second, we demonstrate here that we are able to rely on discrete metric data that is delivered by a Computational Fluid Dynamics (CFD) solver. Third, we demonstrate how to generate high-order curved elements that are aligned with the underlying discrete metric field.

Ekelschot, Dirk↗

Flowfield Analysis of a Small Entry Probe (SPRITE) Tested in an Arc Jet

A novel concept of small size (diameter less than 15 inches) entry probes named SPRITE (Small Probe Re-entry Investigation for TPS Engineering) has been developed at NASA Ames Research Center (ARC). These flight probes have on-board data acquisition systems that have also been developed in parallel at NASA ARC by Greg Swanson1. Flight probes of this size facilitate testing over a wide range of conditions in arc jets available at NASA ARC, thereby fulfilling a 'test what you fly' paradigm. As indicated by the acronym, these probes, with suitably tailored trajectories, are primarily meant to be robotic flight test beds for TPS materials, although the design is flexible enough to accommodate additional objectives of flight-testing other vehicle subsystems. A first step towards establishing the feasibility of the SPRITE concept is to arc-jet test fully instrumented models at flight scale. In a follow-on to the Large-Scale Article Tests (LSAT2) performed in the 60 MW Interaction Heating Facility (IHF) in late 2008/early 2009, a full-scale model of Deep Space-2 (DS23) made of red oak was tested in the 20 MW Aerodynamic Heating Facility (AHF). There were no issues with mass capture by the diffuser for blunt bodies of roughly 15 inches diameter tested in the 18-inch nozzle of the AHF. Building on this initial success, two identical test articles - SPRITE-T1-1 and SPRITE-T1-2 (T1 indicating the choice of back shell geometry) - were fabricated, and one of them, SPRITE-T1-1, was tested in the AHF recently. Both these test articles, 14 inches in diameter, have a 45deg sphere-cone (like DS2) made of PICA bonded on to a 1/8th inch thick aluminum shell using RTV. The aft portion of the test article is a conical frustum (15deg cone angle) with LI-2200 bonded on to the aluminum shell. Each model is fully instrumented with: (a) thermocouples imbedded in plugs in the heat shield, (b) thermocouples bonded to the aluminum substructure; the thermocouples are distributed over the entire shell, and (c) a few strain gages. Data from some of the thermocouples and gages are acquired by the on-board data acquisition system (DAS), while data from the others are routed to the facility-provided DAS, thereby enabling a cross check on the in situ measurement capability. as inputs to v2.6.1 of the in-house materials thermal response code, FIAT

atmospheric entry↗

Photon detection with parallel asynchronous processing

An approach to photon detection with a parallel asynchronous signal processor is described. The visible or IR photon-detection capability of the silicon p(+)-n-n(+) detectors and the parallel asynchronous processing are addressed separately. This approach would permit an independent analog processing channel to be dedicated to every pixel. A laminar architecture consisting of a stack of planar arrays of the devices would form a 2D array processor with a 2D array of inputs located directly behind a focal-plane detector array. A 2D image data stream would propagate in neuronlike asynchronous pulse-coded form through the laminar processor. Such systems can integrate image acquisition and image processing. Acquisition and processing would be performed concurrently as in natural vision systems. The possibility of multispectral image processing is addressed.

Coon, D. D.↗

Efficient Load Balancing and Data Remapping for Adaptive Grid Calculations

Mesh adaption is a powerful tool for efficient unstructured- grid computations but causes load imbalance among processors on a parallel machine. We present a novel method to dynamically balance the processor workloads with a global view. This paper presents, for the first time, the implementation and integration of all major components within our dynamic load balancing strategy for adaptive grid calculations. Mesh adaption, repartitioning, processor assignment, and remapping are critical components of the framework that must be accomplished rapidly and efficiently so as not to cause a significant overhead to the numerical simulation. Previous results indicated that mesh repartitioning and data remapping are potential bottlenecks for performing large-scale scientific calculations. We resolve these issues and demonstrate that our framework remains viable on a large number of processors.

Oliker, Leonid↗

torch-einshard v1.0

torch-einshard is a Python library for describing local and distributed PyTorch tensor computations with compact, einsum-like notation. Its expressions name logical axes, specify how they are sharded across a PyTorch DeviceMesh, and represent partial reductions. The library automatically performs contractions, permutations, reshaping, splitting, gathering, reduction, reduce-scatter, and repartitioning while preserving autograd. Additional features include sharding-aware FFTs, tensor rolls, halo exchange, sliding windows, 1D–3D convolutions, uneven-shard handling, parameter initialization and gradient management, and cost-based execution planning. It is designed for scientific machine learning and large-model workloads, including tensor-, sequence-, and spatial-parallel MLPs, attention, convolutions, and spectral operations. Compared with manually combining torch.einsum and distributed collectives, torch-einshard expresses both the mathematical operation and data placement in one readable formula. This reduces boilerplate and synchronization errors, keeps forward and backward communication consistent, and allows the library to select optimized collective strategies without changing model code.

Morozov, Dmitriy [Lawrence Berkeley National Labor↗

Machine Learning-Driven Conservative-to-Primitive Conversion in Hybrid Piecewise Polytropic and Tabulated Equations of State

We present a novel machine learning (ML)-based method to accelerate conservative-to-primitive inversion, focusing on hybrid piecewise polytropic and tabulated equations of state. Traditional root-finding techniques are computationally expensive, particularly for large-scale relativistic hydrodynamics simulations. To address this, we employ feedforward neural networks (NNC2PS and NNC2PL), trained in PyTorch (2.0+) and optimized for GPU inference using NVIDIA TensorRT (8.4.1), achieving significant speedups with minimal accuracy loss. The NNC2PS model achieves 𝐿 1 and 𝐿 ∞ errors of 4.54 × 10 −7 and 3.44 × 10−6, respectively, while the NNC2PL model exhibits even lower error values. TensorRT optimization with mixed-precision deployment substantially accelerates performance compared to traditional root-finding methods. Specifically, the mixed-precision TensorRT engine for NNC2PS achieves inference speeds approximately 400 times faster than a traditional single-threaded CPU implementation for a dataset size of 1,000,000 points. Ideal parallelization across an entire compute node in the Delta supercomputer (dual AMD 64-core 2.45 GHz Milan processors and 8 NVIDIA A100 GPUs with 40 GB HBM2 RAM and NVLink) predicts a 25-fold speedup for TensorRT over an optimally parallelized numerical method when processing 8 million data points. Moreover, the ML method exhibits sub-linear scaling with increasing dataset sizes. We release the scientific software developed, enabling further validation and extension of our findings. By exploiting the underlying symmetries within the equation of state, these findings highlight the potential of ML, combined with GPU optimization and model quantization, to accelerate conservative-to-primitive inversion in relativistic hydrodynamics simulations.

conservative-to-primitive conversion↗

Multiscale and Multifidelity Modeling of a 3D Woven Composite Thermal Protection System

Complex three-dimensional (3D) woven composites have been considered by multiple NASA projects in recent years as a means of offering improved mechanical and thermal performance over traditional laminated composite systems. Parallel efforts have focused on developing simulation capabilities for these systems, which have traditionally and heavily relied on experimental testing to evaluate composite performance. One system is the Heatshield for Extreme Entry Environment Technology (HEEET), which is being considered for the thermal protection system on reentry spacecraft. Optical microscopy was used to characterize the blended carbon and phenolic fiber tows. A section of HEEET insulation layer was imaged with high-resolution micro-computed tomography (microCT) and segmented to separate individual tows, porous matrix, and voids. These data were used to develop multiscale thermomechanical computational models within the NASA Multiscale Analysis Tool (NASMAT). Two NASMAT modeling approaches were considered to capture the details of the 3D woven architecture: a coarse model appropriate for inclusion in multiscale structural analyses and a high-fidelity model created by downsampling the microCT data. Both elastic and thermal properties were computed and compared. The feasibility and challenges associated with modeling complex, hybrid 3D woven composites were also addressed.

NASMAT↗