Search NASA⌕ Search

SEARCH · Search NASA

Results for “MPI”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Framework for Extensible, Asynchronous Task Scheduling (FEATS) in Fortran

Most parallel scientific programs contain compiler directives (pragmas) such as those from OpenMP, explicit calls to runtime library procedures such as those implementing the Message Passing Interface (MPI), or compiler-specific language extensions such as those provided by CUDA. By contrast, the recent Fortran standards empower developers to express parallel algorithms without directly referencing lower-level parallel programming models. Fortran’s parallel features place the language within the Partitioned Global Address Space (PGAS) class of programming models. When writing programs that exploit data-parallelism, application developers often find it straightforward to develop custom parallel algorithms. Problems involving complex, heterogeneous, staged calculations, however, pose much greater challenges. Such applications require careful coordination of tasks in a manner that respects dependencies prescribed by a directed acyclic graph. When rolling one’s own solution proves difficult, extending a customizable framework becomes attractive. The paper presents the design, implementation, and use of the Framework for Extensible Asynchronous Task Scheduling (FEATS), which we believe to be the first task-scheduling tool written in modern Fortran. We describe the benefits and compromises associated with choosing Fortran as the implementation language, and we propose ways in which future Fortran standards can best support the use case in this paper.

Modern Fortran↗

Radiative Heat Transfer Capability Implemented in OpenNCC for Conjugate Heat Transfer Applications

Thermal efficiency of gas turbine engine increases as the temperature and pressure at the combustor increases. Consequently, the materials used inside a combustor must survive an increasingly challenging environment. For this reason, accurate assessment of heat transfer is crucial for combustor design. While all three modes of heat transfer are present inside a combustor, the focus of this paper is the thermal radiation. Radiative heat transfer in a gas turbine combustors are particularly interesting from three reasons. Firstly, the radiative heat loss from the combustion region may affect the emission performance. Secondly, the cooling air will protect the liner from convection but not necessary from radiation. Finally, it is less frequently incorporated in CFD analysis than other forms of heat transfer. In this work, radiative heat transfer using discrete ordinate method has been incorporated in OpenNCC (a publicly releasable version of the National Combustion Code) developed at NASA Glenn Research Center. Aside from massively parallel computation capability using MPI and the ability to utilize unstructured mesh, the current implementation includes two types of spectral models, namely, the weighted some of gray gas model and the full spectrum correlated k-distribution model. After presenting the theory and the strategy of implementation, results of validation cases for gray gas and spectral models will be presented. While the implementation of the radiation solver is intended for gas turbine application, the radiation solver can run independently from the convection/combustion solver and the same theory can be applied to other application.

OpenNCC↗

Radiative Heat Transfer Capability Implemented in OpenNCC for Conjugate Heat Transfer Applications

Thermal efficiency of gas turbine engine increases as the temperature and pressure at the combustor increases. Consequently, the materials used inside a combustor must survive an increasingly challenging environment. For this reason, accurate assessment of heat transfer is crucial for combustor design. While all three modes of heat transfer are present inside a combustor, the focus of this paper is the thermal radiation. Radiative heat transfer in a gas turbine combustors are particularly interesting from three reasons. Firstly, the radiative heat loss from the combustion region may affect the emission performance. Secondly, the cooling air will protect the liner from convection but not necessary from radiation. Finally, it is less frequently incorporated in CFD analysis than other forms of heat transfer. In this work, radiative heat transfer using discrete ordinate method has been incorporated in OpenNCC (a publicly releasable version of the National Combustion Code) developed at NASA Glenn Research Center. Aside from massively parallel computation capability using MPI and the ability to utilize unstructured mesh, the current implementation includes two types of spectral models, namely, the weighted some of gray gas model and the full spectrum correlated k-distribution model. After presenting the theory and the strategy of implementation, results of validation cases for gray gas and spectral models will be presented. While the implementation of the radiation solver is intended for gas turbine application, the radiation solver can run independently from the convection/combustion solver and the same theory can be applied to other application.

OpenNCC↗

Containerized GEOS: Toward a Portable Climate Model

The NASA Goddard Earth Observing System (GEOS) is an Earth system model used for weather, climate, and other scientific applications. GEOS consists of linked components that can run in various configurations such as atmosphere-only and coupled atmosphere-ocean. Running this model on any new supercomputing system depends on operating systems, compilers, MPI stacks, and libraries being present and correctly configured. To remove that burden from users, our project explores building and running GEOS using Singularity containers – files containing all the needed software dependencies – on both NASA high-end computing systems and commercial cloud computing environments. Ultimately, the goal is for containerized GEOS to make it easier for users outside of NASA to deploy and run the model on any machine.

Matthew Thompson↗

LAURA Users Manual: 5.7

This users manual provides in-depth information concerning installation and execution of Laura, version 5. Laura is a structured, multi-block, compu- tational aerothermodynamic simulation code. Version 5 represents a major refactoring of the original Fortran 77 Laura code toward a modular structure afforded by Fortran 2003. The refactoring improved usability and maintain- ability by eliminating the requirement for problem-dependent re-compilations, providing more intuitive distribution of functionality, and simplifying inter- faces required for multi-physics coupling. As a result, Laura now shares gas-physics modules, MPI modules, and other low-level modules with the Fun3D unstructured-grid code. In addition to internal refactoring, several new features and capabilities have been added, e.g., a GNU-standard instal- lation process, parallel load balancing, automatic trajectory point sequencing, free-energy minimization, and coupled ablation and flowfield radiation.

CFD hypersonics reentry↗

The Profiled Feldman-Cousins Method for Confidence Interval Construction for the Nova 3-Flavor Oscillation Analysis

The small interaction cross-section of neutrinos makes experimental neutrino physics particularly responsive to technological advancements. A significant development leveraged by the NOvA experiment is large-scale parallel processing, enabling novel computational approaches to longstanding experimental challenges. Central to managing the resulting high-throughput data is NOvA’s implementation of the Freight Train model, designed for efficient data production and handling.This dissertation details the methodology and execution of the NOvA 2024 3-Flavor Oscillation Analysis, supported by a comprehensive dataset spanning ten years. It emphasizes frequentist results refined through the Feldman-Cousins (FC) technique, specifically addressing confidence interval corrections in parameter estimation. The computational intensity associated with Feldman-Cousins arises from extensive Monte Carlo simulations, which were substantially mitigated through parallel computing on the Perlmutter supercomputer at the National Energy Research Scientific Computing Center (NERSC), employing the MPI framework.To further enhance computational efficiency, an Importance Sampling method is introduced and evaluated, demonstrating significant potential to reduce complexity, particularly in exploring extreme parameter space regions. This thesis presents both the successful application of advanced computational resources and the development of sophisticated statistical techniques, aiming to enhance the precision and scope of neutrino oscillation analyses.

Dye ajdye11190@gmail.com, Andrew Joseph [Mississip↗

Performance Evaluation of Different Parallel Programming Models in SCALE-Shift Sequences for Criticality and Shielding Applications [Abstract]

The SCALE code system has been widely used for nuclear criticality safety, reactor physics, radiation shielding, source term generation, and inventory analyses by researchers, industry, and regulatory bodies. Although limited support for shared- and distributed-memory parallel processing was introduced via C++ threading, OpenMP, and MPI, a hybrid parallel programming model with both distributed- and shared-memory parallelism has not been fully supported in the SCALE code system.

Nuclear Criticality Safety Program (NCSP)↗

Transient Triplet Metallopnictinidenes M–Pn (M = Pd II , Pt II ; Pn = P, As, Sb): Characterization and Dimerization

Nitrenes (R–N) have been subject to a large body of experimental and theoretical studies. The fundamental reactivity of this important class of transient intermediates has been attributed to their electronic structures, particularly the accessibility of triplet vs singlet states. In contrast, electronic structure trends along the heavier pnictinidene analogues (R–Pn; Pn = P–Bi) are much less systematically explored. We here report the synthesis of a series of metallodipnictenes, {M–Pn=Pn–M} (M = Pd II , Pt II ; Pn = P, As, Sb, Bi) and the characterization of the transient metallopnictinidene intermediates, {M–Pn} for Pn = P, As, Sb. Structural, spectroscopic, and computational analysis revealed spin triplet ground states for the metallopnictinidenes with characteristic electronic structure trends along the series. In comparison to the nitrene, the heavier pnictinidenes exhibit lower-lying ground state SOMOs and singlet excited states, thus suggesting increased electrophilic reactivity. Furthermore, the splitting of the triplet magnetic microstates is beyond the phosphinidenes {M–P} dominated by heavy pnictogen atom induced spin–orbit coupling.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Enhanced Interfacial Strength in Carbon Fiber Composites via Mussel‐Inspired Sizing Polymers

Composite materials possess a high strength-to-weight ratio. A key determinant of their mechanical performance is the interfacial strength between the fibers and the matrix. Sizing agents are commonly used to improve this interface by promoting better adhesion, though optimizing this interaction remains a significant challenge. Here, this study evaluates the use of poly(catechol-styrene) (PCS), a mussel-inspired sizing agent, to enhance fiber–matrix bonding in carbon fiber composites. Woven carbon fiber laminates were dip-coated with varying concentrations of PCS (0.05 and 0.1 wt%) and subsequently fabricated using vacuum-assisted resin transfer molding followed by compression molding. Interlaminar shear strength (ILSS) tests showed improvements of 4% and 8% for the 0.05% and 0.1% PCS treatments, respectively. These results indicate that PCS is effective in reinforcing interfacial adhesion, thereby improving the mechanical integrity of carbon fiber-reinforced composites.

Carbon fiber composites↗

A time-parallel multiple-shooting method for large-scale quantum optimal control

Quantum optimal control plays a crucial role in quantum computing by providing the interface between compiler and hardware. Solving the optimal control problem is particularly challenging for multi-qubit gates, due to the exponential growth in computational complexity with the system's dimensionality and the deterioration of optimization convergence. To ameliorate the computational complexity of time-integration, this paper introduces a multiple-shooting approach in which the time domain is divided into multiple windows and the intermediate states at window boundaries are treated as additional optimization variables. Further, this enables parallel computation of state evolution across time-windows, significantly accelerating objective function and gradient evaluations. Since the initial state matrix in each window is only guaranteed to be unitary upon convergence of the optimization algorithm, the conventional gate trace infidelity is replaced by a generalized infidelity that is convex for non-unitary state matrices. Continuity of the state across window boundaries is enforced by equality constraints. A quadratic penalty optimization method is used to solve the constrained optimal control problem, and an efficient adjoint technique is employed to calculate the gradients in each iteration. We demonstrate the effectiveness of the proposed method through numerical experiments on quantum Fourier transform gates in systems with 2, 3, and 4 qubits, noting a speedup of 80x for evaluating the gradient in the 4-qubit case, highlighting the method's potential for optimizing control pulses in multi-qubit quantum systems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Computing Sparse Tensor Decompositions via Chapel and C++/MPI Interoperability without Intermediate I/O

We extend an existing approach for efficient use of shared mapped memory across Chapel and C++ for graph data stored as 1-D arrays to sparse tensor data stored using a combination of 2-D and 1-D arrays. We describe the specific extensions that provide use of shared mapped memory tensor data for a particular C++ tensor decomposition tool called GentenMPI. We then demonstrate our approach on several real-world datasets, providing timing results that illustrate minimal overhead incurred using this approach. Finally, we extend our work to improve memory usage and provide convenient random access to sparse shared mapped memory tensor elements in Chapel, while still being capable of leveraging high performance implementations of tensor algorithms in C++.

97 MATHEMATICS AND COMPUTING↗

Parthenon

Explore the source record for details and available documents.

MPI, GPU, parallelism↗

Uniqueness of MHV Gravity Amplitudes

We investigate MHV tree-level gravity amplitudes as defined on the spinor-helicity variety. Unlike their gluon counterparts, the gravity amplitudes do not have logarithmic singularities and do not admit Amplituhedron-like construction. Importantly, they are not determined just by their singularities, but rather their numerators have interesting zeroes. We make a conjecture about the uniqueness of the numerator and explore this feature from a more mathematical perspective. This leads us to a new approach for examining adjoints. We outline steps of our proposed proof and provide computational evidence for its validity in specific cases.

adjoints↗

Position Paper - pFLogger: The Parallel Fortran Logging framework for HPC Applications

In the context of high performance computing (HPC), software investments in support of text-based diagnostics, which monitor a running application, are typically limited compared to those for other types of IO. Examples of such diagnostics include reiteration of configuration parameters, progress indicators, simple metrics (e.g., mass conservation, convergence of solvers, etc.), and timers. To some degree, this difference in priority is justifiable as other forms of output are the primary products of a scientific model and, due to their large data volume, much more likely to be a significant performance concern. In contrast, text-based diagnostic content is generally not shared beyond the individual or group running an application and is most often used to troubleshoot when something goes wrong. We suggest that a more systematic approach enabled by a logging facility (or logger) similar to those routinely used by many communities would provide significant value to complex scientific applications. In the context of high-performance computing, an appropriate logger would provide specialized support for distributed and shared-memory parallelism and have low performance overhead. In this paper, we present our prototype implementation of pFlogger a parallel Fortran-based logging framework, and assess its suitability for use in a complex scientific application.

Fortran↗