Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel processing (computers)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56

Spaceflight optical disk recorder development

Mass memory systems based on rewriteable optical disk media are expected to play an important role in meeting the data system requirements for future NASA spaceflight missions. NASA has established a program to develop a high performance (high rate, large capacity) optical disk recorder focused on use aboard unmanned Earth orbiting platforms. An expandable, adaptable system concept is proposed based on disk drive modules and a modular controller. Drive performance goals are 10 gigabyte capacity, 300 megabit/s transfer rate, 10 exp -12 corrected bit error rate, and 150 millisec access time. This performance is achieved by writing eight data tracks in parallel on both sides of a 14 in. optical disk using two independent heads. System goals are 160 gigabyte capacity, 1.2 gigabits/s data rate with concurrent I/O, 250 millisec access time, and two to five year operating life on orbit. The system can be configured to meet various applications. This versatility is provided by the controller. The controller provides command processing, multiple drive synchronization, data buffering, basic file management, error processing, and status reporting. Technology developments, design concepts, current status including a computer model of the system and a Controller breadboard, and future plans for the Drive and Controller are presented.

Jurczyk, Stephen G.↗

Optimal Control within the Context of Multidisciplinary Design, Analysis, and Optimization

Multidisciplinary design, analysis and optimization involves modeling the interactions of complex systems across a variety of disciplines. The optimization of such systems can be a computationally expensive exercise with multiple levels of nested nonlinear solvers running under an optimizer.The application of optimal control in project development often involves performing trajectory optimization for fixed vehicle designs or parametric sweeps across some key vehicle properties.This information is then relayed to the subsystem design teams who update their designs and relay some bulk characteristics back to the trajectory optimization procedure.This iteration is then repeated until the design closes.However, with increasing interest in more tightly coupled systems, such as electric and hybrid-electric aircraft propulsion and boundary layer ingestion, this process is prone to ignore subtle coupling between vehicle subsystem designs and vehicle operation on a given mission.Integrating trajectory optimization into a tightly coupled multidisciplinary design procedure can be computationally prohibitive, depending on the complexity of the subsystem analyses and the optimal control technique applied.To address these issues a new optimal control software tool, Dymos, has been developed.Dymos is built upon NASA's OpenMDAO software and can leverage its capabilities to efficiently compute gradients for the optimization and optimize complex models in parallel on distributed memory systems.This report provides some explanation into the numerical methods employed in Dymos and provides several use cases that demonstrate its performance on traditional optimal control problems and improvements ino techniques have been used extensively in recent decades to solve a variety of optimal control problems, typically in the form of aerospace vehicle trajectory optimization.

pseudospectral↗

PyHydroGeophysX: An extensible open-source platform for integrating hydrological models with geophysical measurements

Hydrological models and geophysical measurements are widely used tools for understanding subsurface hydrological processes relevant to water resource management, yet they typically remain disconnected due to technical barriers. We present PyHydroGeophysX, an open-source Python platform bridging this gap by providing standardized interfaces between hydrological modeling software (MODFLOW, ParFlow) and geophysical simulation tools (PyGIMLi, SimPEG). The platform implements bidirectional workflows: translating hydrological outputs into simulated geophysical responses through petrophysical models, and extracting hydrological information from geophysical inversions. Key features include bidirectional workflow modules, configurable petrophysical models, time-lapse inversion with temporal regularization, parallel computing, and mesh utilities for property transfer between geophysical and hydrological grids. The modular architecture of PyHydroGeophysX enables researchers to incorporate additional models and methods, fostering broader adoption of integrated hydrogeophysical approaches. The software is freely available on GitHub and is intended for researchers and practitioners working at the intersection of hydrology and geophysics.

Hydrogeophysics↗

A system for routing arbitrary directed graphs on SIMD architectures

There are many problems which can be described in terms of directed graphs that contain a large number of vertices where simple computations occur using data from connecting vertices. A method is given for parallelizing such problems on an SIMD machine model that is bit-serial and uses only nearest neighbor connections for communication. Each vertex of the graph will be assigned to a processor in the machine. Algorithms are given that will be used to implement movement of data along the arcs of the graph. This architecture and algorithms define a system that is relatively simple to build and can do graph processing. All arcs can be transversed in parallel in time O(T), where T is empirically proportional to the diameter of the interconnection network times the average degree of the graph. Modifying or adding a new arc takes the same time as parallel traversal.

Tomboulian, Sherryl↗

Near-Body Grid Adaption for Overset Grids

A solution adaption capability for curvilinear near-body grids has been implemented in the OVERFLOW overset grid computational fluid dynamics code. The approach follows closely that used for the Cartesian off-body grids, but inserts refined grids in the computational space of original near-body grids. Refined curvilinear grids are generated using parametric cubic interpolation, with one-sided biasing based on curvature and stretching ratio of the original grid. Sensor functions, grid marking, and solution interpolation tasks are implemented in the same fashion as for off-body grids. A goal-oriented procedure, based on largest error first, is included for controlling growth rate and maximum size of the adapted grid system. The adaption process is almost entirely parallelized using MPI, resulting in a capability suitable for viscous, moving body simulations. Two- and three-dimensional examples are presented.

Buning, Pieter G.↗

Implementing Atmospheric Infrared Sounder (AIRS) and Cross-Track Infrared Sounder (CrIS) Cloud-Clearing Algorithm into the NASA GEOS: Focus on the 2017 Atlantic Tropical Cyclone Season

Numerical Weather Prediction (NWP) centers assimilate cloud-free infrared (IR) radiances because the assimilation of all-sky IR radiances is not yet operationally achievable. The cloud-clearing procedure offers a simpler, but effective strategy that produces cloud-affected radiances suitable for assimilation in partially cloudy regions. Several studies conducted by this team have demonstrated that IR Cloud-Cleared Radiances (CCRs), if thinned more aggressively than clear-sky radiances, can improve analysis and forecasts, particularly in meteorologically active areas. However, CCRs are not used by operational centers due partly to the thought that the process of cloud-clearing may affect latency and introduce difficult-to-control external dependencies. This study presents the results of implementing an Atmospheric Infrared Sounder (AIRS) and Cross-Track Infrared Sounder (CrIS) cloud-clearing procedure into the NASA Goddard Earth Observing System (GEOS) to demonstrate the portability of the procedure. The AIRS and CrIS cloud-clearing algorithms have been deprived of external dependencies, made customizable to any specific model, and the computational efficiency has been improved via parallelization. The revised AIRS and CrIS cloud-clearing algorithms allow a customized choice of channel selection, the use of a user-specified model's fields as first guess, and can perform in real time. Data assimilation experiments with the hybrid 4DEnVar GEOS system were successfully performed for the 2017 tropical cyclones (TC) season with a focus on three major hurricanes (Harvey, Irma, and Maria). This study shows that assimilation of locally-generated CCRs have a positive impact on both global skill and TC representation, compared to the assimilation of AIRS and CrIS clear-sky radiances, and a comparable or slightly improved impact compared to assimilation of CCRs produced by external sources, such as NASA's Distributed Active Archive Centers and NOAA’s Comprehensive Large Array-data Stewardship System. The customization and computational efficiency of the revised procedure would enable its usability in a real-time forecast context.

Niama Boukachaba↗

Tailoring microstructures with mild magnetic-field processing: A case study of CuNiFe alloys

Combined experimental and computational investigations of the CuNiFe spinodal system confirm that application of a mild magnetic field during thermal treatment alters elemental redistribution and the resulting microstructure, relative to that obtained from zero-field annealing. Spinodal decomposition of a Cu 40 Ni 42 Fe 18 alloy was initiated during thermal treatment at 773 K, conducted either under zero field or modest (60 mT) magnetic f ield conditions for up to 200 h. Periodic (~10 nm) chemical modulations into Cu-rich and NiFe-rich regions were observed under both conditions, with the amplitude and wavelength of the segregated regions increasing with treatment time. However, magnetic field annealing resulted in a more than twofold increase in the amplitude of elemental modulations relative to zero-field conditions – consistent with enhanced diffusional f luxes during spinodal decomposition – while the modulation wavelength remained largely unaffected. These microstructural differences are reflected in various extrinsic magnetic properties. In parallel, first-principles DFT calculations indicate that long-range ferromagnetic order, as induced by an applied magnetic field, substantially alters the strength and nature of atomic interactions, enhancing the thermodynamic instability of the CuNiFe solid solution. Collectively, these results suggest that incorporating a mild (millitesla-level) magnetic field – distinct from the strong (tesla-level) fields commonly used in prior studies – during thermal processing has the potential to deliver enhanced control of microstructures for targeted engineering outcomes.

36 MATERIALS SCIENCE↗

HPDR: High-Performance Portable Scientific Data Reduction Framework

The rapid growth in scientific data generation is outpacing advancements in computing systems necessary for efficient storage, transfer, and analysis, particularly in the context of exascale computing. With the deployment of first-generation exascale computing systems and next-generation experimental facilities, this gap is widening and necessitates effective data reduction techniques to manage enormous data volumes. Over the past decade, various data reduction methods, including lossless compression, error-controlled lossy compression, and data refactoring, have been developed to accelerate I/O in scientific workflows. Despite significant reductions in data volume, these methods introduce considerable computational overhead, which can become the new bottleneck in data processing. To mitigate this, GPU-accelerated data reduction algorithms have been introduced. However, challenges remain in their integration into exascale workflows, including limited portability across different GPU architectures, substantial memory transfer overhead, and reduced scalability on dense multi-GPU systems. To address these challenges, we propose HPDR, a high-performance and portable data reduction framework. HPDR is designed to enable the execution of state-of-the-art reduction algorithms across diverse processor architectures while reducing memory transfer overhead to 2.3 % of the original, resulting in up to 3.5× faster throughput compared to existing solutions. It also achieves up to 96% of the theoretical speedup in multi-GPU settings. In addition, evaluations on accelerating I/O operations at scale up to 1,024 nodes of the Frontier supercomputer demonstrate that HPDR can achieve up to 103 TB/s reduction throughput, providing up to 4× acceleration in parallel I/O performance compared to existing data reduction routines. This work highlights the potential of HPDR to significantly enhance data reduction efficiency in exascale computing environments.

Chen, Jieyang [University of Oregon]↗

Poplar: a phylogenomics pipeline

Motivation Generating phylogenomic trees from the genomic data is essential in understanding biological systems. Each step of this complex process has received extensive attention and has been significantly streamlined over the years. Given the public availability of data, obtaining genomes for a wide selection of species is straightforward. However, analyzing that data to generate a phylogenomic tree is a multistep process with legitimate scientific and technical challenges, often requiring a significant input from a domain-area scientist. Results We present Poplar, a new, streamlined computational pipeline, to address the computational logistical issues that arise when constructing the phylogenomic trees. It provides a framework that runs state-of-the-art software for essential steps in the phylogenomic pipeline, beginning from a genome with or without an annotation, and resulting in a species tree. Running Poplar requires no external databases. In the execution, it enables parallelism for execution for clusters and cloud computing. The trees generated by Poplar match closely with state-of-the-art published trees. The usage and performance of Poplar is far simpler and quicker than manually running a phylogenomic pipeline. Availability and implementation Freely available on GitHub at https://github.com/sandialabs/poplar. Implemented using Python and supported on Linux.

Koning, Elizabeth [Sandia National Laboratories (S↗

Regional-scale fault-to-structure earthquake simulations with the EQSIM framework: Workflow maturation and computational performance on GPU-accelerated exascale platforms

Continuous advancements in scientific and engineering understanding of earthquake phenomena, combined with the associated development of representative physics-based models, is providing a foundation for high-performance, fault-to-structure earthquake simulations. However, regional-scale applications of high-performance models have been challenged by the computational requirements at the resolutions required for engineering risk assessments. The EarthQuake SIMulation (EQSIM) framework, a software application development under the US Department of Energy (DOE) Exascale Computing Project, is focused on overcoming the existing computational barriers and enabling routine regional-scale simulations at resolutions relevant to a breadth of engineered systems. This multidisciplinary software development—drawing upon expertise in geophysics, engineering, applied math and computer science—is preparing the advanced computational workflow necessary to fully exploit the DOE’s exaflop computer platforms coming online in the 2023 to 2024 timeframe. Achievement of the computational performance required for high-resolution regional models containing upward of hundreds of billions to trillions of model grid points requires numerical efficiency in every phase of a regional simulation. This includes run time start-up and regional model generation, effective distribution of the computational workload across thousands of computer nodes, efficient coupling of regional geophysics and local engineering models, and application-tailored highly efficient transfer, storage, and interrogation of very large volumes of simulation data. This article summarizes the most recent advancements and refinements incorporated in the workflow design for the EQSIM integrated fault-to-structure framework, which are based on extensive numerical testing across multiple graphics processing unit (GPU)-accelerated platforms, and demonstrates the computational performance achieved on the world’s first exaflop computer platform through representative regional-scale earthquake simulations for the San Francisco Bay Area in California, USA.

58 GEOSCIENCES↗

Fiats: Functional inference and training for surrogates

Fiats provides a platform for research on the training and deployment of neural-network surrogate models for computational science. Fiats also supports exploring, advancing, and combining functional, object-oriented, and parallel programming patterns in Fortran 2023. As such, the Fiats name has dual expansions: “Functional Inference And Training for Surrogates” or “Fortran Inference And Training for Science.” Fiats inference and training procedures are pure and therefore satisfy a language constraint imposed on procedure invocations inside Fortran’s parallel loop construct: do concurrent. Furthermore, the Fiats training procedures are built around a do concurrent parallel reduction. Several compilers can automatically parallelize do concurrent on Central Processing Units (CPUs) or Graphics Processing Units (GPUs). Fiats thus aims to achieve performance portability through standard language mechanisms.

Rouson, Damian [Lawrence Berkeley National Laborat↗

High performance flight simulation at NASA Langley

The use of real-time simulation at the NASA facility is reviewed specifically with regard to hardware, software, and the use of a fiberoptic-based digital simulation network. The network hardware includes supercomputers that support 32- and 64-bit scalar, vector, and parallel processing technologies. The software include drivers, real-time supervisors, and routines for site-configuration management and scheduling. Performance specifications include: (1) benchmark solution at 165 sec for a single CPU; (2) a transfer rate of 24 million bits/s; and (3) time-critical system responsiveness of less than 35 msec. Simulation applications include the Differential Maneuvering Simulator, Transport Systems Research Vehicle simulations, and the Visual Motion Simulator. NASA is shown to be in the final stages of developing a high-performance computing system for the real-time simulation of complex high-performance aircraft.

Cleveland, Jeff I., II↗

A Microfabricated Involute-Foil Regenerator for Stirling Engines

A segmented involute-foil regenerator has been designed, microfabricated and tested in an oscillating-flow rig with excellent results. During the Phase I effort, several approximations of parallel-plate regenerator geometry were chosen as potential candidates for a new microfabrication concept. Potential manufacturers and processes were surveyed. The selected concept consisted of stacked segmented-involute-foil disks (or annular portions of disks), originally to be microfabricated from stainless-steel via the LiGA (lithography, electroplating, and molding) process and EDM (electric discharge machining). During Phase II, re-planning of the effort led to test plans based on nickel disks, microfabricated via the LiGA process, only. A stack of nickel segmented-involute-foil disks was tested in an oscillating-flow test rig. These test results yielded a performance figure of merit (roughly the ratio of heat transfer to pressure drop) of about twice that of the 90% random fiber currently used in small ~ 100 W Stirling space-power convertors in the Reynolds Number range of interest (50-100). A Phase III effort is now underway to fabricate and test a segmented-involute-foil regenerator in a Stirling convertor. Though funding limitations prevent optimization of the Stirling engine geometry for use with this regenerator, the Sage computer code will be used to help evaluate the engine test results. Previous Sage Stirling model projections have indicated that a segmented-involute-foil regenerator is capable of improving the performance of an optimized involute-foil engine by 6-9%; it is also anticipated that such involute-foil geometries will be more reliable and easier to manufacture with tight-tolerance characteristics, than random-fiber or wire-screen regenerators. Beyond the near-term Phase III regenerator fabrication and engine testing, other goals are (1) fabrication from a material suitable for high temperature Stirling operation (up to 850 C for current engines; up to 1200 C for a potential engine-cooler for a Venus mission), and (2) reduction of the cost of the fabrication process to make it more suitable for terrestrial applications of segmented involute foils. Past attempts have been made to use wrapped foils to approximate the large theoretical figures of merit projected for parallel plates. Such metal wrapped foils have never proved very successful, apparently due to the difficulties of fabricating wrapped-foils with uniform gaps and maintaining the gaps under the stress of time-varying temperature gradients during start-up and shut-down, and relatively-steady temperature gradients during normal operation. In contrast, stacks of involute-foil disks, with each disk consisting of multiple involute-foil segments held between concentric circular ribs, have relatively robust structures. The oscillating-flow rig tests of the segmented-involute-foil regenerator have demonstrated a shift in regenerator performance strongly in the direction of the theoretical performance of ideal parallel-plate regenerators.

Tew, Roy↗

Attenuation of empennage buffet response through active control of damping using piezoelectric material

Dynamic response and damping data obtained from buffet studies conducted in a low-speed wind tunnel by using a simple, rigid model attached to spring supports are presented. The two parallel leaf spring supports provided a means for the model to respond in a vertical translation mode, thus simulating response in an elastic first bending mode. Wake-induced buffeting flow was created by placing an airfoil upstream of the model of that the wake of the airfoil impinged on the model. Model response was sensed by a strain gage mounted on one of the springs. The output signal from the strain gage was fed back through a control law implemented on a desktop computer. The processed signals were used to 'actuate' a piezoelectric bending actuator bonded to the other spring in such a way as to add damping as the model responded. The results of this 'proof-of-concept' study show that the piezoelectric actuator was effective in attenuating the wake-induced buffet response over the range of parameters investigated.

Heeg, Jennifer↗

Minimizing distortion in truss structures -- a Hopfield network solution

Distortions in truss structures can result from random errors in elemental lengths that are typical of a manufacturing process. These distortions may be minimized by an optimal selection of elements from those available for placement between the prescribed nodes -- a combinatorial optimization problem requiring significant investment of computational resource for all but the smallest problems. The present paper describes a formulation in which near-optimal element assignments are obtained as minimum energy, stable states, of an analogous Hopfield neural network. This requires mapping of the optimization problem into an energy function of the appropriate Lyapunov form. The computational architecture is ideally suited to a parallel processor implementation and offers significant savings in computational effort. A numerical implementation of the approach is discussed with reference to planar truss problems.

Fu, B.↗

A Microfabricated Involute-Foil Regenerator for Stirling Engines

A segmented involute-foil regenerator has been designed, microfabricated and tested in an oscillating-flow rig with excellent results. During the Phase I effort, several approximations of parallel-plate regenerator geometry were chosen as potential candidates for a new microfabrication concept. Potential manufacturers and processes were surveyed. The selected concept consisted of stacked segmented-involute-foil disks (or annular portions of disks), originally to be microfabricated from stainless-steel via the LiGA (lithography, electroplating, and molding) process and EDM. During Phase II, re-planning of the effort led to test plans based on nickel disks, microfabricated via the LiGA process, only. A stack of nickel segmented-involute-foil disks was tested in an oscillating-flow test rig. These test results yielded a performance figure of merit (roughly the ratio of heat transfer to pressure drop) of about twice that of the 90 percent random fiber currently used in small approx.100 W Stirling space-power convertors-in the Reynolds Number range of interest (50 to 100). A Phase III effort is now underway to fabricate and test a segmented-involute-foil regenerator in a Stirling convertor. Though funding limitations prevent optimization of the Stirling engine geometry for use with this regenerator, the Sage computer code will be used to help evaluate the engine test results. Previous Sage Stirling model projections have indicated that a segmented-involute-foil regenerator is capable of improving the performance of an optimized involute-foil engine by 6 to 9 percent; it is also anticipated that such involute-foil geometries will be more reliable and easier to manufacture with tight-tolerance characteristics, than random-fiber or wire-screen regenerators. Beyond the near-term Phase III regenerator fabrication and engine testing, other goals are (1) fabrication from a material suitable for high temperature Stirling operation (up to 850 C for current engines; up to 1200 C for a potential engine-cooler for a Venus mission), and (2) reduction of the cost of the fabrication process to make it more suitable for terrestrial applications of segmented involute foils. Past attempts have been made to use wrapped foils to approximate the large theoretical figures of merit projected for parallel plates. Such metal wrapped foils have never proved very successful, apparently due to the difficulties of fabricating wrapped-foils with uniform gaps and maintaining the gaps under the stress of time-varying temperature gradients during start-up and shut-down, and relatively-steady temperature gradients during normal operation. In contrast, stacks of involute-foil disks, with each disk consisting of multiple involute-foil segments held between concentric circular ribs, have relatively robust structures. The oscillating-flow rig tests of the segmented-involute-foil regenerator have demonstrated a shift in regenerator performance strongly in the direction of the theoretical performance of ideal parallel-plate regenerators.

Tew, Roy↗

Minimizing distortion in truss structures - A Hopfield network solution

Distortions in truss structures can result from random errors in element lengths that are typical of a manufacturing process. These distortions may be minimized by an optimal selection of elements from those available for placement between the prescribed nodes - a combinatorial optimization problem requiring significant investment of computational resource for all but the smallest problems. The present paper describes a formulation in which near-optimal element assignments are obtained as minimum-energy stable states, of an analogous Hopfield neural network. This requires mapping of the optimization problem into an energy function of the appropriate Liapunov form. The computational architecture is ideally suited to a parallel processor implementation and offers significant savings in computational effort. A numerical implementation of the approach is discussed with reference to planar truss problems.

Fu, B.↗

Checkpoint-based forward recovery using lookahead execution and rollback validation in parallel and distributed systems

This thesis studies a forward recovery strategy using checkpointing and optimistic execution in parallel and distributed systems. The approach uses replicated tasks executing on different processors for forwared recovery and checkpoint comparison for error detection. To reduce overall redundancy, this approach employs a lower static redundancy in the common error-free situation to detect error than the standard N Module Redundancy scheme (NMR) does to mask off errors. For the rare occurrence of an error, this approach uses some extra redundancy for recovery. To reduce the run-time recovery overhead, look-ahead processes are used to advance computation speculatively and a rollback process is used to produce a diagnosis for correct look-ahead processes without rollback of the whole system. Both analytical and experimental evaluation have shown that this strategy can provide a nearly error-free execution time even under faults with a lower average redundancy than NMR.

Long, Junsheng↗