Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallelism”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60

Parallel Grand Canonical Monte Carlo (ParaGrandMC) Simulation Code

This report provides an overview of the Parallel Grand Canonical Monte Carlo (ParaGrandMC) simulation code. This is a highly scalable parallel FORTRAN code for simulating the thermodynamic evolution of metal alloy systems at the atomic level, and predicting the thermodynamic state, phase diagram, chemical composition and mechanical properties. The code is designed to simulate multi-component alloy systems, predict solid-state phase transformations such as austenite-martensite transformations, precipitate formation, recrystallization, capillary effects at interfaces, surface absorption, etc., which can aid the design of novel metallic alloys. While the software is mainly tailored for modeling metal alloys, it can also be used for other types of solid-state systems, and to some degree for liquid or gaseous systems, including multiphase systems forming solid-liquid-gas interfaces.

Vesselin I Yamakov↗

Enhancing Application Performance Using Mini-Apps: Comparison of Hybrid Parallel Programming Paradigms

In many fields, real-world applications for High Performance Computing have already been developed. For these applications to stay up-to-date, new parallel strategies must be explored to yield the best performance; however, restructuring or modifying a real-world application may be daunting depending on the size of the code. In this case, a mini-app may be employed to quickly explore such options without modifying the entire code. In this work, several mini-apps have been created to enhance a real-world application performance, namely the VULCAN code for complex flow analysis developed at the NASA Langley Research Center. These mini-apps explore hybrid parallel programming paradigms with Message Passing Interface (MPI) for distributed memory access and either Shared MPI (SMPI) or OpenMP for shared memory accesses. Performance testing shows that MPI+SMPI yields the best execution performance, while requiring the largest number of code changes. A maximum speedup of 23 was measured for MPI+SMPI, but only 11 was measured for MPI+OpenMP.

Lawson, Gary↗

Position Paper - pFLogger: The Parallel Fortran Logging framework for HPC Applications

In the context of high performance computing (HPC), software investments in support of text-based diagnostics, which monitor a running application, are typically limited compared to those for other types of IO. Examples of such diagnostics include reiteration of configuration parameters, progress indicators, simple metrics (e.g., mass conservation, convergence of solvers, etc.), and timers. To some degree, this difference in priority is justifiable as other forms of output are the primary products of a scientific model and, due to their large data volume, much more likely to be a significant performance concern. In contrast, text-based diagnostic content is generally not shared beyond the individual or group running an application and is most often used to troubleshoot when something goes wrong. We suggest that a more systematic approach enabled by a logging facility (or logger) similar to those routinely used by many communities would provide significant value to complex scientific applications. In the context of high-performance computing, an appropriate logger would provide specialized support for distributed and shared-memory parallelism and have low performance overhead. In this paper, we present our prototype implementation of pFlogger a parallel Fortran-based logging framework, and assess its suitability for use in a complex scientific application.

Fortran↗

POSITION PAPER - pFLogger: The Parallel Fortran Logging Framework for HPC Applications

In the context of high performance computing (HPC), software investments in support of text-based diagnostics, which monitor a running application, are typically limited compared to those for other types of IO. Examples of such diagnostics include reiteration of configuration parameters, progress indicators, simple metrics (e.g., mass conservation, convergence of solvers, etc.), and timers. To some degree, this difference in priority is justifiable as other forms of output are the primary products of a scientific model and, due to their large data volume, much more likely to be a significant performance concern. In contrast, text-based diagnostic content is generally not shared beyond the individual or group running an application and is most often used to troubleshoot when something goes wrong. We suggest that a more systematic approach enabled by a logging facility (or 'logger') similar to those routinely used by many communities would provide significant value to complex scientific applications. In the context of high-performance computing, an appropriate logger would provide specialized support for distributed and shared-memory parallelism and have low performance overhead. In this paper, we present our prototype implementation of pFlogger - a parallel Fortran-based logging framework, and assess its suitability for use in a complex scientific application.

Clune, Thomas L.↗

pFlogger: The Parallel Fortran Logging Utility

In the context of high performance computing (HPC), software investments in support of text-based diagnostics, which monitor a running application, are typically limited compared to those for other types of IO. Examples of such diagnostics include reiteration of configuration parameters, progress indicators, simple metrics (e.g., mass conservation, convergence of solvers, etc.), and timers. To some degree, this difference in priority is justifiable as other forms of output are the primary products of a scientific model and, due to their large data volume, much more likely to be a significant performance concern. In contrast, text-based diagnostic content is generally not shared beyond the individual or group running an application and is most often used to troubleshoot when something goes wrong. We suggest that a more systematic approach enabled by a logging facility (or 'logger)' similar to those routinely used by many communities would provide significant value to complex scientific applications. In the context of high-performance computing, an appropriate logger would provide specialized support for distributed and shared-memory parallelism and have low performance overhead. In this paper, we present our prototype implementation of pFlogger - a parallel Fortran-based logging framework, and assess its suitability for use in a complex scientific application.

Clune, Tom↗

Dynamic Analysis of the hFan, a Parallel Hybrid Electric Turbofan Engine

NASA and a variety of aerospace industry stakeholders are investing in conceptual studies of electrified aircraft, including parallel hybrid electric aircraft such as the Subsonic Ultra Green Aircraft Research (SUGAR) Volt. At this point, little of the work published in the literature has examined the transient behavior of the turbomachinery in these systems. This paper describes a control system built around the hFan, the parallel hybrid electric turbofan engine designed for the SUGAR Volt concept aircraft. This control system is used to show that the hFan, running with its baseline concept of operations, is capable of transient operation throughout the envelope. The design parameters of this controller are varied to assess the amount of operability margin built into the engine design, and whether this margin can be reduced to enable more aggressive designs, that may feature better fuel economy. Further, studies are performed as parameters for the hFan electric motor are varied to determine how the motor impacts the engine's need for transient operability margin. The studies suggest that the engine may be redesigned with as much as a 3% reduction in high pressure compressor stall margin. It was also demonstrated that appropriate design and control of the electric motor may be able to buy an additional 0.5% stall margin reduction or a turbine inlet temperature reduction of 35 degR, as tested at the sea-level static condition.

turboelectric↗

Dynamic Analysis of the hFan, a Parallel Hybrid Electric Turbofan Engine

NASA and a variety of aerospace industry stakeholders are investing in conceptual studies of electrified aircraft, including parallel hybrid electric aircraft such as the Subsonic Ultra Green Aircraft Research (SUGAR) Volt. At this point, little of the work published in the literature has examined the transient behavior of the turbomachinery in these systems. This paper describes a control system built around the hFan, the parallel hybrid electric turbofan engine designed for the SUGAR Volt concept aircraft. This control system is used to show that the hFan, running with its baseline concept of operations, is capable of transient operation throughout the envelope. The design parameters of this controller are varied to assess the amount of operability margin built into the engine design, and whether this margin can be reduced to enable more aggressive designs, that may feature better fuel economy. Further, studies are performed as parameters for the hFan electric motor are varied to determine how the motor impacts the engine's need for transient operability margin. The studies suggest that the engine may be redesigned with as much as a 3% reduction in high pressure compressor stall margin. It was also demonstrated that appropriate design and control of the electric motor may be able to buy an additional 0.5% stall margin reduction or a turbine inlet temperature reduction of 35 R, as tested at the sea-level static condition.

SUGAR Volt↗

Parallel Grand-Canonical Monte Carlo (ParaGrandMC) User’s Manual Version 2.0

This manual describes the commands and command line options for the Parallel Grand Canonical Monte Carlo version 2.0 (ParaGrandMC.2.0) simulation code. This is a highly scalable parallel FORTRAN 2003 code for simulating the thermodynamic evolution of materials at the atomic level, and predicting their thermodynamic state, phase diagram, chemical composition and mechanical properties. The code is specifically designed to simulate multi-component alloy systems, predict solid-state phase transformations such as austenite-martensite transformations, precipitate formation, recrystallization, capillary effects at interfaces, surface absorption, etc., which can aid the design of novel metallic alloys. While the software is mainly tailored for modeling metal alloys, it can also be used for other types of solid-state systems, and to some degree for liquid or gaseous systems, including multiphase systems forming solid-liquid-gas interfaces. In addition to performing Monte Carlo (MC) simulations, the code can also perform Molecular Dynamics (MD) and Langevin Dynamics (LD) simulations, which can be combined and interchanged with MC for faster and more efficient system evolution. A detailed description of the MC part of the code is provided in the NASA ParaGrandMC report: NASA/CR–2016-219202; http://www.sti.nasa.gov.

High performance computing↗

Enabling Thread Safety and Parallelism in the Program to Optimize Simulated Trajectories II

Development of the Program to Optimize Simulated Trajectories (POST) began in the 1970s. Since then, it has become widely utilized across NASA, industry, and academia to solve a variety of atmospheric ascent and entry problems. Its successor, POST2, has undergone many upgrades since its release in the 1990s. Recently, there has been an increasing desire to take advantage of the advances in parallel computing for both offline and online systems. Thus, modifications were made to allow POST2 to simulate multiple trajectories simultaneously without adversely affecting results. This capability is leveraged to calculate optimization solutions in parallel as opposed to sequentially. A demonstration of the benefits is presented using a small set of POST2 regression tests, as well as a project simulating a human-scale Lunar lander.

R. Anthony Williams↗

Enabling Thread Safety and Parallelism in the Program to Optimize Simulated Trajectories II

Development of the Program to Optimize Simulated Trajectories (POST) began in the 1970s. Since then, it has become widely utilized across NASA, industry, and academia to solve a variety of atmospheric ascent and entry problems. Its successor, POST2, has undergone many upgrades since its release in the 1990s. Recently, there has been an increasing desire to take advantage of the advances in parallel computing for both offline and online systems. Thus, modifications were made to allow POST2 to simulate multiple trajectories simultaneously without adversely affecting results. This capability is leveraged to calculate optimization solutions in parallel as opposed to sequentially. A demonstration of the benefits is presented using a small set of POST2 regression tests, as well as a project simulating a human-scale Lunar lander.

Anthony Williams↗

Electron Acceleration and Heating during Magnetic Reconnection in the Earth's Quasi-parallel Bow Shock

We perform a 2.5-dimensional particle-in-cell simulation of a quasi-parallel shock, using parameters for the Earth's bow shock, to examine electron acceleration and heating due to magnetic reconnection. The shock transition region evolves from the ion-coupled reconnection dominant stage to the electron-only reconnection dominant stage, as time elapses. The electron temperature enhances locally in each reconnection site, and ion-scale magnetic islands generated by ion-coupled reconnection show the most significant enhancement of the electron temperature. The electron energy spectrum shows a power law, with a power-law index around 6. We perform electron trajectory tracing to understand how they are energized. Some electrons interact with multiple electron-only reconnection sties, and Fermi acceleration occurs during multiple reflections. Electrons trapped in ion-scale magnetic islands can be accelerated in another mechanism. Islands move in the shock transition region, and electrons can obtain larger energy from the in-plane electric field than the electric potential in those islands. These newly found energization mechanisms in magnetic islands in the shock can accelerate electrons to energies larger than the achievable energies by the conventional energization due to the parallel electric field and shock drift acceleration. This study based on the selected particle analysis indicates that the maximum energy in the nonthermal electrons is achieved through acceleration in ion-scale islands, and electron-only reconnection accounts for no more than half of the maximum energy, as the lifetime of sub-ion-scale islands produced by electron-only reconnection is several times shorter than that of ion-scale islands.

Solar magnetic reconnection↗

Parametric Modeling and Mission Performance Analysis of a True Parallel Hybrid Turboprop Aircraft for Freighter Operations

Hybrid-electric propulsion systems for short-haul, cargo carrying aircraft have emerged as promising solutions for an environmentally sustainable future for commercial aviation. The novel propulsion architecture presents significant complexity and requires the development of new methodologies to account for the unique coupling and integration between critical design variables. This paper presents a comprehensive study on the modeling, performance assessment, and design space exploration of a C-130H freighter aircraft retrofitted with a parallel hybrid electric powertrain. Details on the development of parametric models for both the baseline and true parallel hybrid (TPH) aircraft and the integrated electrified aircraft propulsion (EAP) system sizing approach are presented along with detailed performance analyses of payload/range capabilities for short-haul cargo missions. For 2030, 2040, and 2050 EAP technology levels, the TPH C-130H configuration with a ~2.12 MW class EAP system has a range capability of 485-1,028 nautical miles and block fuel savings of 27-44%.

efficiency↗

Oblique Instability of Quasi-Parallel Whistler Waves in the Presence of Cold and Warm Electron Populations

Whistler waves propagating nearly parallel to the ambient magnetic field experience a nonlinear instability due to transverse currents when the background plasma has a population of sufficiently low energy electrons. Intriguingly, this nonlinear process may generate oblique electrostatic waves, including whistlers near the resonance cone with properties resembling oblique chorus waves in the Earth’s magnetosphere. Focusing on the generation of oblique whistlers, earlier analysis of the instability is extended here to the case where low-energy background plasma consists of both a “cold” population with energy of a few eV and a “warm” electron component with energy of the order of 100 eV. This is motivated by spacecraft observations in the Earth’s magnetosphere where oblique chorus waves were shown to interact resonantly with the warm electrons. The main new results are: 1) the instability producing oblique electrostatic waves is sensitive to the shape of the electron distribution at low energies. In the whistler range of frequencies, two distinct peaks in the growth rate are typically present for the model considered: a peak associated with the warm electron population at relatively low wavenumbers and a peak associated with the cold electron population at relatively high wavenumbers; 2) overall, the instability producing oblique whistler waves near the resonance cone persists (with a reduced growth rate) even in the cases where the temperature of the cold population is relatively high, including cases where cold population is absent and only the warm population is included; 3) particle-in-cell simulations show that the instability leads to heating of the background plasma and formation of characteristic plateau and beam features in the parallel electron distribution function in the range of energies resonant with the instability. The plateau/beam features have been previously detected in spacecraft observations of oblique chorus waves. However, they have been attributed to external sources and have been proposed to be the mechanism generating oblique chorus. In the present scenario, the causality link is reversed and the instability generating oblique whistler waves is shown to be a possible mechanism for formation of the plateau and beam features.

Vadim Roytershteyn↗

Plastic Parallel Pathways Platform - 4P Model

Global momentum is building towards a circular economy capable of keeping plastics in use and out of waste streams. Given that 79% of all plastic produced since 1950 has accumulated in landfills or the natural environment,rapid implementation of various end-of-life (EoL) management technologies will be needed to reach this target. However, it can be challenging to develop an effective plastic EoL strategy when the available options - chemical or molecular recycling, energy recovery, upcycling, downcycling, closed-loop (plastic-to-plastic) or open-loop (plastic-to-x) recycling, among others - can generate products ranging from low-grade to virgin-quality plastic and from fuels to value-added chemicals. We present a flexible material flow model capable of analyzing the effects of both plastic-to-plastic and plastic-to-x EoL management strategies on the U.S. PET economy. This Plastic Parallel Pathways Platform (4P) assesses the environmental impacts, costs, and circularity of a PET system in which waste is managed through six potential EoL pathways: landfill, incineration with energy recovery, pyrolysis to fuel oil, upcycling to glass fiber reinforced plastic (GFRP), mechanical recycling to low-grade PET, and chemical recycling (glycolysis) to bottle-grade PET. We compare the pathways across multiple metrics using multi-criteria decision analysis (MCDA) and then use a brute force algorithm to predict an optimal combination of EoL pathways to minimize greenhouse gas (GHG) emissions and costs and maximize circularity. This work highlights the need to implement a diverse portfolio of EoL strategies in parallel to enable a PET economy that meets environmental, economic, and circularity requirements simultaneously.

downcycling↗

Parallel-in-Time Solution of Allen-Cahn Equations by Integrating Operator Learning into the Parareal Method

While recent advances in deep learning have shown promising efficiency gains in solving time-dependent partial differential equations (PDEs), matching the accuracy of conventional numerical solvers still remains a challenge. One strategy to improve the accuracy of deep learning-based solutions for time-dependent PDEs is to use the learned model as the coarse propagator in the Parareal method and a traditional numerical method as the fine solver. However, successful integration of deep learning into the Parareal method requires consistency between the coarse and fine solvers, particularly for PDEs exhibiting rapid changes such as sharp transitions. Here, to ensure this consistency, we propose using convolutional neural networks (CNNs) to learn the fully discrete time-stepping operator defined by the same numerical scheme employed as the fine solver. We demonstrate the effectiveness of the proposed method in solving the classical and mass-conservative Allen–Cahn (AC) equations. Through iterative updates in the Parareal algorithm, our approach achieves a significant computational speedup compared to traditional fine solvers while converging to high-accuracy solutions. Our results highlight that the proposed hybrid Parareal algorithm effectively accelerates simulations, particularly when implemented on multiple GPUs, and converges to the desired accuracy in only a few iterations. Another advantage of our method is that the CNN model is trained on trajectory-based data generated from random initial conditions, such that the trained model can be used to solve the AC equations with various initial conditions without retraining. This work demonstrates the potential of integrating neural network methods into parallel-in-time frameworks for efficient and accurate simulations of time-dependent PDEs.

97 MATHEMATICS AND COMPUTING↗

QRCODE: Massively parallelized real-time time-dependent density functional theory for periodic systems

We present a new software module, QRCODE (Quantum Research for Calculating Optically Driven Excitations), for massively parallelized real-time time-dependent density functional theory (RT-TDDFT) calculations of periodic systems in the open-source Qbox software package. Our approach utilizes a custom implementation of a fast Fourier transformation scheme that significantly reduces inter-node message passing interface (MPI) communication of the major computational kernel and shows impressive scaling up to 16,344 CPU cores. In addition to improving computational performance, QRCODE contains a suite of various time propagators for accurate RT-TDDFT calculations. As benchmark applications of QRCODE, we calculate the current density and optical absorption spectra of hexagonal boron nitride (h-BN) and photo-driven reaction dynamics of the ozone-oxygen reaction. We also calculate the second and higher harmonic generation of monolayer and multi-layer boron nitride structures as examples of large material systems. Our optimized implementation of RT-TDDFT in QRCODE enables large-scale calculations of real-time electron dynamics of chemical and material systems with enhanced computational performance and impressive scaling across several thousand CPU cores.

97 MATHEMATICS AND COMPUTING↗

Parallel computing for power system climate resiliency: Solving a large-scale stochastic capacity expansion problem with mpi-sppy

Here we propose a nodal stochastic generation and transmission expansion planning model that incorporates the output from high-resolution global climate models through load and generation availability scenarios. We implement our model in Pyomo and perform computational studies on a realistically-sized test case of the California electric grid in a high performance computing environment. We propose model reformulations and algorithm tuning to efficiently solve this large problem using a variant of the Progressive Hedging Algorithm. We utilize the parallelization capabilities and overall versatility of mpi-sppy, exploiting its hub-and-spoke architecture to concurrently obtain inner and outer bounds on an optimal expansion plan. Initial results show that instances with 360 representative days on a system with over 8,000 buses can be solved to within 5% of optimality in under 4 h of wall clock time, a first step towards solving a large-scale power system expansion planning problem across a wide range of climate-informed operational scenarios.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A time-parallel multiple-shooting method for large-scale quantum optimal control

Quantum optimal control plays a crucial role in quantum computing by providing the interface between compiler and hardware. Solving the optimal control problem is particularly challenging for multi-qubit gates, due to the exponential growth in computational complexity with the system's dimensionality and the deterioration of optimization convergence. To ameliorate the computational complexity of time-integration, this paper introduces a multiple-shooting approach in which the time domain is divided into multiple windows and the intermediate states at window boundaries are treated as additional optimization variables. Further, this enables parallel computation of state evolution across time-windows, significantly accelerating objective function and gradient evaluations. Since the initial state matrix in each window is only guaranteed to be unitary upon convergence of the optimization algorithm, the conventional gate trace infidelity is replaced by a generalized infidelity that is convex for non-unitary state matrices. Continuity of the state across window boundaries is enforced by equality constraints. A quadratic penalty optimization method is used to solve the constrained optimal control problem, and an efficient adjoint technique is employed to calculate the gradients in each iteration. We demonstrate the effectiveness of the proposed method through numerical experiments on quantum Fourier transform gates in systems with 2, 3, and 4 qubits, noting a speedup of 80x for evaluating the gradient in the 4-qubit case, highlighting the method's potential for optimizing control pulses in multi-qubit quantum systems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗