Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel hybrid”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Classic and Quantum Task-Based Intelligent Runtime for QIRs Running on Multiple QPUs

High-performance computing systems are rapidly evolving into heterogeneous platforms that fuse quantum accelerators with traditional classical processing units (CPUs) and graphical processing units (GPUs). This convergence calls for runtimes capable of managing both classical and quantum workloads in a unified manner. We introduce an intelligent, task-based runtime that marries the Intelligent RuntIme System (IRIS) asynchronous scheduler with a quantum programming stack through the Quantum Intermediate Representation Execution Engine (QIR-EE). Our design allows programs written in the quantum intermediate representation (QIR) to be dispatched concurrently to a variety of back-ends, including multiple quantum simulators and nascent quantum processors, enabling genuine hybrid execution on a single node. To illustrate its practicality, we partition a 4-qubit and 20-qubit circuit into three sub-circuits using quantum circuit cutting via the QCut library. Each sub-circuit is simulated independently by the QIR-EE driver within IRIS, after which a classical post-processing step merges the simulation results to recover the outcome of the original full-circuit computation. This case study demonstrates how finer task granularity can enable the parallel execution and lower the simulation burden per quantum task while preserving overall accuracy, highlighting the feasibility of our hybrid approach.

Miniskar, Narasinga Rao [ORNL] (ORCID:000000018259↗

Comparing the OpenMP, MPI, and Hybrid Programming Paradigm on an SMP Cluster

Clusters of SMP (Symmetric Multi-Processors) nodes provide support for a wide range of parallel programming paradigms. The shared address space within each node is suitable for OpenMP parallelization. Message passing can be employed within and across the nodes of a cluster. Multiple levels of parallelism can be achieved by combining message passing and OpenMP parallelization. Which programming paradigm is the best will depend on the nature of the given problem, the hardware components of the cluster, the network, and the available software. In this study we compare the performance of different implementations of the same CFD benchmark application, using the same numerical algorithm but employing different programming paradigms.

Jost, Gabriele↗

Concurrent Probabilistic Simulation of High Temperature Composite Structural Response

A computational structural/material analysis and design tool which would meet industry's future demand for expedience and reduced cost is presented. This unique software 'GENOA' is dedicated to parallel and high speed analysis to perform probabilistic evaluation of high temperature composite response of aerospace systems. The development is based on detailed integration and modification of diverse fields of specialized analysis techniques and mathematical models to combine their latest innovative capabilities into a commercially viable software package. The technique is specifically designed to exploit the availability of processors to perform computationally intense probabilistic analysis assessing uncertainties in structural reliability analysis and composite micromechanics. The primary objectives which were achieved in performing the development were: (1) Utilization of the power of parallel processing and static/dynamic load balancing optimization to make the complex simulation of structure, material and processing of high temperature composite affordable; (2) Computational integration and synchronization of probabilistic mathematics, structural/material mechanics and parallel computing; (3) Implementation of an innovative multi-level domain decomposition technique to identify the inherent parallelism, and increasing convergence rates through high- and low-level processor assignment; (4) Creating the framework for Portable Paralleled architecture for the machine independent Multi Instruction Multi Data, (MIMD), Single Instruction Multi Data (SIMD), hybrid and distributed workstation type of computers; and (5) Market evaluation. The results of Phase-2 effort provides a good basis for continuation and warrants Phase-3 government, and industry partnership.

Abdi, Frank↗

A Comparison of Lifting-Line and CFD Methods with Flight Test Data from a Research Puma Helicopter

Four lifting-line methods were compared with flight test data from a research Puma helicopter and the accuracy assessed over a wide range of flight speeds. Hybrid Computational Fluid Dynamics (CFD) methods were also examined for two high-speed conditions. A parallel analytical effort was performed with the lifting-line methods to assess the effects of modeling assumptions and this provided insight into the adequacy of these methods for load predictions.

Bousman, William G.↗

Message Passing and Shared Address Space Parallelism on an SMP Cluster

Currently, message passing (MP) and shared address space (SAS) are the two leading parallel programming paradigms. MP has been standardized with MPI, and is the more common and mature approach; however, code development can be extremely difficult, especially for irregularly structured computations. SAS offers substantial ease of programming, but may suffer from performance limitations due to poor spatial locality and high protocol overhead. In this paper, we compare the performance of and the programming effort required for six applications under both programming models on a 32-processor PC-SMP cluster, a platform that is becoming increasingly attractive for high-end scientific computing. Our application suite consists of codes that typically do not exhibit scalable performance under shared-memory programming due to their high communication-to-computation ratios and/or complex communication patterns. Results indicate that SAS can achieve about half the parallel efficiency of MPI for most of our applications, while being competitive for the others. A hybrid MPI+SAS strategy shows only a small performance advantage over pure MPI in some cases. Finally, improved implementations of two MPI collective operations on PC-SMP clusters are presented.

Shan, Hongzhang↗

Large-Scale Numerical Simulations of Human Motion

This paper examines the feasibility of using massively-parallel and vector-processing supercomputers to solve large-scale optimal control problems for human movement. Specifically, we compare the computational expense of determining the optimal controls for the single support phase of walking using a conventional serial machine (a Silicon Graphics Personal Iris 4D25 workstation), a MIMD parallel machine (an Intel iPSC/860 comprising 128 processors), and a parallel-vector-processing machine (a Cray Y-MP 8/864). With the human body modeled as a 14 degree-of-freedom linkage actuated by 46 musculotendinous units, computation of the optimal controls for walking could take up to 3 months of CPU time on the Iris. Both the Cray Y-MP and the Intel iPSC/860 are able to reduce this time to practical levels. The optimal control solution for walking can be found with about 77 hours of CPU time on the Cray, and with about 88 hours of CPU time on the Intel. Although the overall speeds of the Cray and the Intel were found to be similar, the unique capabilities of each machine are best suited to different parts of the optimal control algorithm used. The Intel performed best in the calculation of the derivatives of the performance criterion and the constraints. In contrast, the Cray performed best during parameter optimization of the controls. These results suggest that the ideal computer architecture for solving very large-scale optimal control problems is a hybrid system in which a vector-processing machine is integrated into the communication network of a MIMD parallel machine.

Anderson, Frank C.↗

Validation study of RWM stability in DIII-D high- β N plasmas

The n = 1 (n is the toroidal mode number) resistive wall mode (RWM) stability is numerically investigated for two DIII-D high-β N discharges 176440 and 172461, utilizing the MARS-F (Liu et al 2000 Phys. Plasmas 7 3681) and MARS-K (Liu et al 2008 Phys. Plasmas 15 112503) codes. Systematic validation efforts are attempted, for the first time, for discharges with very slow or vanishing toroidal flow for a large fraction of the plasma volume. While gaining physics insights in accessing stable operation regime at β N exceeding the Troyon no-wall limit in these slow-rotation experiments, the predictive capability of fluid and non-perturbative magnetohydrodynamic-kinetic hybrid models for the RWM is further confirmed. The MARS-F fluid model, with a strong but numerically tunable viscosity mimicking ion Landau damping of parallel sound waves, finds complete stabilization of the n = 1 RWM in the considered DIII-D plasmas under the experimental flow conditions. Similarly, either full stabilization (for discharge 176440) or marginal stability (for discharge 172461) of the mode is computed by the MARS-K hybrid model, which is first-principle based without free model parameters. In particular, all drift kinetic resonances, including those of thermal and energetic particles, are found to synergistically act to marginally stabilize the RWM in discharge 172461. These MARS-F/K modeling results explain the experimentally observed stable operational regime in DIII-D, as far as the RWM stability is concerned. Extensive numerical sensitivity studies, with respect to the plasma toroidal flow speed as well as the radial location of the resistive wall, are also carried out to further support the validation study.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Sequence length scaling in vision transformers for scientific images on frontier

Vision Transformers (ViTs) are pivotal for foundational models in scientific imagery, including Earth science applications, due to their capability to process large sequence lengths. While transformers for text have inspired scaling sequence lengths in ViTs, adapting these for ViTs introduces unique challenges. We develop distributed sequence parallelism for ViTs, enabling them to handle up to 1M tokens. Our approach, leveraging DeepSpeed-Ulysses and Long-Sequence-Segmentation with model sharding, is the first to apply sequence parallelism in ViT training, achieving a 94% batch scaling efficiency on 2,048 AMD-MI250X GPUs. Evaluating sequence parallelism in ViTs, particularly in models up to 10B parameters, highlighted substantial bottlenecks. We countered these with hybrid sequence, pipeline, and flash attention strategies, to scale beyond single GPU memory limits. Our method significantly enhances climate modeling accuracy by 20% in temperature predictions, marking the first training of a vision transformer model to convergence with a sequence length of 188K tokens, using full self-attention.

Tsaris, Aristeidis (aris) [ORNL] (ORCID:0000000277↗

Average electric wave spectra across the plasma sheet and their relation to ion bulk speed

Using 4 months of tail data obtained by the ELF/MF spectrum analyzer of the wave experiment and the three-dimensional plasma instrument on board the AMPTE/IRM satellite, a statistical survey on the electric wave spectral density in the earth's plasma sheet has been conducted. More than 50,000 10-s-averaged electric wave spectra were analyzed with respect to differences between their values in the inner and outer central plasma sheet and the plasma sheet boundary layer as well as their dependence on radial distance and ion bulk speed. High-speed flows are dominated by broadband electrostatic noise with highest spectral densities in the plasma sheet boundary, where broadband electrostatic noise also exists during periods of low-speed flows. The broadband electrostatic noise has a typical spectral index of about -2. During low-speed flows the spectra in the central plasma sheet show distinct emissions at the electron cyclotron odd half-harmonic and upper hybrid frequency. Wave intensities during episodes of fast perpendicular flows are higher than those associated with fast parallel flows.

Baumjohann, W.↗

Impacts of Spontaneous Hot Flow Anomalies on the Magnetosheath and Magnetopause

Spacecraft observations and global hybrid (kinetic ions and fluid electrons) simulations have demonstrated that ion dissipation processes at the quasi-parallel bow shock are associated with the formation of structures called spontaneous hot flow anomalies (SHFAs). Previous simulations and recent spacecraft observations have also established that SHFAs result in the formation of magnetosheath filamentary structures(MFS). In this paper we demonstrate that in addition to MFS, SHFAs also result in the formation of magnetos heath cavities that are associated with decreases in density, velocity, and magnetic field and enhancements in temperature. We use the results of a global MHD run to determine the change in the magnetosheath properties associated with cavities due to ion kinetic effects. The results also show the formation of regions of high flow speed called magnetosheath jets whose properties as a function of solar wind Mach number are described in this study. Comparing the properties of the simulated magnetosheath cavities and jets to past spacecraft observations provides good agreement in both cases. We also demonstrate that pressure variations associated with cavities and SHFAs in the sheath result in a continuous sunward and anti sunward magnetopause motion. This result is consistent with previous suggestions that SHFAs may be responsible for the generation of ion cyclotron waves and precipitation of ring current protons in the outer magnetosphere.

Omidi, N.↗

Interpretable Models for Workflow Differentiation in High-Performance Scientific Networks

Scientific workflows in high-performance networks spawn hundreds of interdependent flows that must be managed collectively—yet existing network classifiers treat each flow in isolation, leading to fragmented QoS decisions and missed interflow patterns. We present a novel traffic classification solution that operates at the workflow level, distinguishing entire filetransfer operations from streaming analytics by capturing how concurrent flows interact and burst together. We introduce a workflow identification window (WIW) that ingests raw packet headers from parallel flows into unified tensors, preserving the spatial-temporal patterns that differentiate scientific workflows. This approach achieves 98.7% accuracy using CNN, LSTM, and hybrid architectures, while maintaining 84% accuracy on production traffic collected a week later—demonstrating robustness to temporal drift. By integrating SHAP and GradCAM explainability, we reveal that early-packet timing patterns and cross-flow correlations drive classification decisions, providing operators with interpretable insights. Our system enables coherent workflow-level QoS enforcement and dynamic bandwidth allocation in scientific networks, eliminating manual per-flow configuration while maintaining classification latency at millisecond level.

Giannakou, Anna [LBL, Berkeley]↗

A Hybrid Mars Ascent Vehicle Design and FY 2016 Technology Development

Hybrid propulsion is currently favored for a Mars Ascent Vehicle (MAV) concept from a thermal performance and Gross Lift Off Mass standpoint. However, it is at a relatively low level of maturity compared to conventional propulsion options. Technology development efforts are currently underway to bring hybrid propulsion to a technology readiness level that would enable its infusion into potential Mars Sample Return. A new propellant combination is being considered for this design that has excellent low temperature behavior. Preliminary results of two ground test campaigns are currently underway to characterize this propellant combination. Hotfire testing is being carried out in parallel at Parabilis Space Technologies and Space Propulsion Group. In addition to the new propellant combination, several other technologies are being pursued for a potential hybrid MAV: hypergolic ignition and Liquid Injection Thrust Vector Control. Both of these technologies have been applied in other rocket applications, e.g. liquid propulsion commonly uses hypergolic propellants and missiles, such as the Minuteman II, have used LITVC in the past. Hypergolic ignition, when oxidizer and fuel combust upon contact, is highly desirable for multiple starts required by the MAV concept. Therefore, testing at Penn State and Purdue is being completed in this area. An updated hybrid propulsion system design for a Mars Ascent Vehicle concept based on JPL’s current understanding of potential Mars Sample Return requirements will be presented, leveraging the advances in technology development as well as updated understanding of how requirements may evolve.

Karp, Ashley C.↗

Analysis and Simulation of the Simplified Aircraft-Based Paired Approach Concept With the ALAS Alerting Algorithm in Conjunction With Echelon and Offset Strategies

This report presents analytical and simulation results of an investigation into proposed operational concepts for closely spaced parallel runways, including the Simplified Aircraft-based Paired Approach (SAPA) with alerting and an escape maneuver, MITRE?s echelon spacing and no escape maneuver, and a hybrid concept aimed at lowering the visibility minima. We found that the SAPA procedure can be used at 950 ft separations or higher with next-generation avionics and that 1150 ft separations or higher is feasible with current-rule compliant ADS-B OUT. An additional 50 ft reduction in runway separation for the SAPA procedure is possible if different glideslopes are used. For the echelon concept we determined that current generation aircraft cannot conduct paired approaches on parallel paths using echelon spacing on runways less than 1400 ft apart and next-generation aircraft will not be able to conduct paired approach on runways less than 1050 ft apart. The hybrid concept added alerting and an escape maneuver starting 1 NM from the threshold when flying the echelon concept. This combination was found to be effective, but the probability of a collision can be seriously impacted if the turn component of the escape maneuver has to be disengaged near the ground (e.g. 300 ft or below) due to airport buildings and surrounding terrain. We also found that stabilizing the approach path in the straight-in segment was only possible if the merge point was at least 1.5 to 2 NM from the threshold unless the total system error can be sufficiently constrained on the offset path and final turn.

Torres-Pomales, Wilfredo↗

Particle-based modelling of axisymmetric tandem mirror devices

In this work, we describe the use of a 1D-2V quasi-neutral hybrid electrostatic PIC with Monte-Carlo Coulomb collisions and non-uniform magnetic field to model the parallel transport and confinement in an axisymmetric tandem mirror device. End-plugs, based on simple-mirrors, are positioned at each end of the device and fueled with neutral beams (25 and 100 keV) to produce a sloshing ion population and increase the density of the end-plugs relative to the central cell. Results show the formation of a potential difference barrier between the central cell and the end-plugs. This potential confines a large fraction of the low energy thermal ions in the central cell which would otherwise be lost in a simple mirror, demonstrating the advantage of the beam-driven tandem mirror configuration relative to simple mirrors. In addition, we explore the effect of end-plug electron temperature on the confinement time of the device and compare it with theoretical estimates. Finally, we discuss the limitations of the code in its present form and describe the next logical steps to improve its predictive capability such as a fully nonlinear Fokker–Planck collision operator, multiply nested flux surface solutions and modeling the exhaust region up to the wall.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Ferroelectric/Optoelectronic Memory/Processor

Proposed hybrid optoelectronic nonvolatile analog memory and data processor comprises planar array of microscopic photosensitive ferroelectric capacitors performing massively parallel analog computations. Processors overcome electronic crosstalk and limitations on number of input/output contacts inherent in electronic implementations of large interconnection arrays. Used in general optical computing, recognition of patterns, and artificial neural networks.

Thakoor, Sarita↗

Effects of magnetospheric electrons on polar plasma outflow - A semikinetic model

The effect or hot magnetospheric electrons on the polar-plasma outflow was investigated, using a semikinetic model developed by Wilson et al. (1990) and Ho et al. (1991) to simulate the effect. The model is based on a hybrid particle-in-cell approach, in which the H(+) and O(+) ions are treated as adiabatic parallel-drifting gyrocenters injected as the upgoing portions of drifting bi-Maxwellian distributions at 1.6 R(E), while the electrons are treated as a massless neutralizing fluid. The results show that, in order to simulate the polar outflow under the influence of hot magnetospheric electrons, it is necessary to consider the effect of the electron temperature gradient.

Ho, C. W.↗

Comparison of weakly and strongly nonlinear wave evolution in a dispersive plasma

An investigation of the evolution of strongly nonlinear, low frequency (ion gyrofrequency), parallel propagating wave packets in a dispersive, collisionless, and low beta(= 8piP/B-squared = 0.3) plasma is undertaken using a hybrid numerical code. These strongly nonlinear wave packets have a transverse magnetic field strength, or wave amplitude, which is of order or greater the field strength along the direction of propagation, and their evolution can differ qualitatively from that of weakly nonlinear packets. The development of spreading fast wave (right helicity) and rarefraction regions competes strongly with steepening, and leads to a long time waveform which differs greatly from that for weak nonlinearity. Results are used to suggest that strongly nonlinear wave evolution occurs frequently in the earth's foreshock.

Vasquez, Bernard J.↗

Identification of genes regulated during mechanical load-induced cardiac hypertrophy

Cardiac hypertrophy is associated with both adaptive and adverse changes in gene expression. To identify genes regulated by pressure overload, we performed suppressive subtractive hybridization between cDNA from the hearts of aortic-banded (7-day) and sham-operated mice. In parallel, we performed a subtraction between an adult and a neonatal heart, for the purpose of comparing different forms of cardiac hypertrophy. Sequencing more than 100 clones led to the identification of an array of functionally known (70%) and unknown genes (30%) that are upregulated during cardiac growth. At least nine of those genes were preferentially expressed in both the neonatal and pressure over-load hearts alike. Using Northern blot analysis to investigate whether some of the identified genes were upregulated in the load-independent calcineurin-induced cardiac hypertrophy mouse model, revealed its incomplete similarity with the former models of cardiac growth. Copyright 2000 Academic Press.

NASA Discipline Cardiopulmonary↗