Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel programming”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,153 records · Page 64

A Module Experimental Process System Development Unit (MEPSDU)

Restructuring research objectives from a technical readiness demonstration program to an investigation of high risk, high payoff activities associated with producing photovoltaic modules using non-CZ sheet material is reported. Deletion of the module frame in favor of a frameless design, and modification in cell series parallel electrical interconnect configuration are reviewed. A baseline process sequence was identified for the fabrication of modules using the selected dendritic web sheet material, and economic evaluations of the sequence were completed.

Source record↗

DeMAID: A Design Manager's Aide for Intelligent Decomposition user's guide

A design problem is viewed as a complex system divisible into modules. Before the design of a complex system can begin, the couplings among modules and the presence of iterative loops is determined. This is important because the design manager must know how to group the modules into subsystems and how to assign subsystems to design teams so that changes in one subsystem will have predictable effects on other subsystems. Determining these subsystems is not an easy, straightforward process and often important couplings are overlooked. Moreover, the planning task must be repeated as new information become available or as the design specifications change. The purpose of this research is to develop a knowledge-based tool called the Design Manager's Aide for Intelligent Decomposition (DeMAID) to act as an intelligent advisor for the design manager. DeMaid identifies the subsystems of a complex design problem, orders them into a well-structured format, and marks the couplings among the subsystems to facilitate the use of multilevel tools. DeMAID also provides the design manager with the capability of examining the trade-offs between sequential and parallel processing. This type of approach could lead to a substantial savings or organizing and displaying a complex problem as a sequence of subsystems easily divisible among design teams. This report serves as a User's Guide for the program.

Rogers, James L.↗

Design of object-oriented distributed simulation classes

Distributed simulation of aircraft engines as part of a computer aided design package is being developed by NASA Lewis Research Center for the aircraft industry. The project is called NPSS, an acronym for 'Numerical Propulsion Simulation System'. NPSS is a flexible object-oriented simulation of aircraft engines requiring high computing speed. It is desirable to run the simulation on a distributed computer system with multiple processors executing portions of the simulation in parallel. The purpose of this research was to investigate object-oriented structures such that individual objects could be distributed. The set of classes used in the simulation must be designed to facilitate parallel computation. Since the portions of the simulation carried out in parallel are not independent of one another, there is the need for communication among the parallel executing processors which in turn implies need for their synchronization. Communication and synchronization can lead to decreased throughput as parallel processors wait for data or synchronization signals from other processors. As a result of this research, the following have been accomplished. The design and implementation of a set of simulation classes which result in a distributed simulation control program have been completed. The design is based upon MIT 'Actor' model of a concurrent object and uses 'connectors' to structure dynamic connections between simulation components. Connectors may be dynamically created according to the distribution of objects among machines at execution time without any programming changes. Measurements of the basic performance have been carried out with the result that communication overhead of the distributed design is swamped by the computation time of modules unless modules have very short execution times per iteration or time step. An analytical performance model based upon queuing network theory has been designed and implemented. Its application to realistic configurations has not been carried out.

Schoeffler, James D.↗

Design of Object-Oriented Distributed Simulation Classes

Distributed simulation of aircraft engines as part of a computer aided design package being developed by NASA Lewis Research Center for the aircraft industry. The project is called NPSS, an acronym for "Numerical Propulsion Simulation System". NPSS is a flexible object-oriented simulation of aircraft engines requiring high computing speed. It is desirable to run the simulation on a distributed computer system with multiple processors executing portions of the simulation in parallel. The purpose of this research was to investigate object-oriented structures such that individual objects could be distributed. The set of classes used in the simulation must be designed to facilitate parallel computation. Since the portions of the simulation carried out in parallel are not independent of one another, there is the need for communication among the parallel executing processors which in turn implies need for their synchronization. Communication and synchronization can lead to decreased throughput as parallel processors wait for data or synchronization signals from other processors. As a result of this research, the following have been accomplished. The design and implementation of a set of simulation classes which result in a distributed simulation control program have been completed. The design is based upon MIT "Actor" model of a concurrent object and uses "connectors" to structure dynamic connections between simulation components. Connectors may be dynamically created according to the distribution of objects among machines at execution time without any programming changes. Measurements of the basic performance have been carried out with the result that communication overhead of the distributed design is swamped by the computation time of modules unless modules have very short execution times per iteration or time step. An analytical performance model based upon queuing network theory has been designed and implemented. Its application to realistic configurations has not been carried out.

Schoeffler, James D.↗

Remediation Technology Collaboration Development - A Compendium

During its multi-year period of performance, the Remediation Technology Collaboration Development (RTCD) task orders initial goals were to enhance the capability to specifically target reductions in the long-term liabilities associated with NASAs most challenging remediation sites. This was accomplished by identifying existing remediation processes and conditions, researching site-specific technologies (both past and present) while simultaneously looking for parallel situations where these technologies could be applied. In addition, the most promising of these solutions were developed from comprehensive research and bench studies into pilot studies or demonstration projects, which contributed significantly to the success of the RTCD program.

Olsen, W.↗

SAIGE-GPU: accelerating genome- and phenome-wide association studies using GPUs

Genome-wide association studies (GWAS) at biobank scale are computationally intensive, especially for admixed populations requiring robust statistical models. SAIGE is a widely used method for generalized linear mixed-model GWAS but is limited by its CPU-based implementation, making phenome-wide association studies impractical for many research groups. We developed SAIGE-GPU, a GPU-accelerated version of SAIGE that replaces CPU-intensive matrix operations with GPU-optimized kernels. The core innovation is distributing genetic relationship matrix calculations across GPUs and communication layers. Applied to 2068 phenotypes from 635 969 participants in the Million Veteran Program, including diverse and admixed populations, SAIGE-GPU achieved a 5-fold speedup in mixed model fitting on supercomputing infrastructure and cloud platforms. We further optimized the variant association testing step through multi-core and multi-trait parallelization. Deployed on Google Cloud Platform and Azure, the method provided substantial cost and time savings. Source code and binaries are available for download at https://github.com/saigegit/SAIGE/tree/SAIGE-GPU-1.3.3. A code snapshot is archived at Zenodo for reproducibility (DOI: [10.5281/zenodo.17642591]). SAIGE-GPU is available in a containerized format for use across HPC and cloud environments and is implemented in R/C++ and runs on Linux systems.

Rodriguez, Alex [Argonne National Laboratory (ANL)↗

The correction of Landsat data for the effects of haze, sun angle, and background reflectance

A technique has been developed for simulating the effects of haze, sun angle, and background reflectance in Landsat data and correcting for them. The atmospheric model assumes a two-layer atmosphere: a Rayleigh scattering molecular layer and a Mie scattering haze layer next to the earth's surface. Reflection and transmission matrices describe the reflection and transmission properties of the plane parallel scattering layers. The multispectral scanner response is computed for various values of the parameters under evaluation. This yields expressions for Landsat gray-scale levels used for determining the effect of changes in any parameter. The Atmospheric Correction computer program is used to determine the haze level from the data, to compute the reflectance, and to interpolate in order to find the correction coefficients necessary to make the desired correction.

Potter, J. F.↗

Enhancements to Program LAURA for computation of three-dimensional hypersonic flow

Changes to Program Laura (Langley Aerothermodynamic Upwind Relaxation Algorithm) are presented which enhance both stability and accuracy of the algorithm. A discussion of iteration/sweeping strategies and their relation to computer architectures is included to best exploit the capabilities of serial, vector, and parallel processor machines. Test cases for Mach 10 perfect gas flow and Mach 32 real gas flow in chemical nonequilibrium over a blunt, raked elliptic cone using the thin-layer Navier-Stokes equations are presented in order to demonstrate the current improved capabilities. Algorithm changes include the use of volume averaging, application of a symmetric total variation diminishing (TVD) scheme, and stronger interaction between the grid/shock alignment routine and the relaxation algorithm. Good comparisons with heat transfer and pitching moment data at three different angles of attack for the Mach 10 tests serve to further validate the present algorithm. Parameters are defined which control the coupling of the specie continuity equations with the solution of the mixture conservation equations. A discussion of the consequences involved in the choice of strong versus weak coupling is presented, and a sample nonequilibrium calculation on a fine grid over a full scale model of the Aeroassist Flight Experiment (AFE) demonstrates current capabilities.

Gnoffo, Peter A.↗

The utilization of parallel processing in solving the inviscid form of the average-passage equation system for multistage turbomachinery

A procedure is outlined which utilizes parallel processing to solve the inviscid form of the average-passage equation system for multistage turbomachinery along with a description of its implementation in a FORTRAN computer code, MSTAGE. A scheme to reduce the central memory requirements of the program is also detailed. Both the multitasking and I/O routines referred to in this paper are specific to the Cray X-MP line of computers and its associated SSD (Solid-state Storage Device). Results are presented for a simulation of a two-stage rocket engine fuel pump turbine.

Mulac, Richard A.↗

Utilization of parallel processing in solving the inviscid form of the average-passage equation system for multistage turbomachinery

A procedure is outlined which utilizes parallel processing to solve the inviscid form of the average-passage equation system for multistage turbomachinery along with a description of its implementation in a FORTRAN computer code, MSTAGE. A scheme to reduce the central memory requirements of the program is also detailed. Both the multitasking and I/O routines referred to are specific to the Cray X-MP line of computers and its associated SSD (Solid-State Disk). Results are presented for a simulation of a two-stage rocket engine fuel pump turbine.

Mulac, Richard A.↗

In Situ Resource Utilization Technologies for Enhancing and Expanding Mars Scientific and Exploration Missions

The primary objectives of the Mars exploration program are to collect data for planetary science in a quest to answer questions related to Origins, to search for evidence of extinct and extant life, and to expand the human presence in the solar system. The public and political engagement that is critical for support of a Mars exploration program is based on all of these objectives. In order to retain and to build public and political support, it is important for NASA to have an integrated Mars exploration plan, not separate robotic and human plans that exist in parallel or in sequence. The resolutions stemming from the current architectural review and prioritization of payloads may be pivotal in determining whether NASA will have such a unified plan and retain public support. There are several potential scientific and technological links between the robotic-only missions that have been flown and planned to date, and the combined robotic and human missions that will come in the future. Taking advantage of and leveraging those links are central to the idea of a unified Mars exploration plan. One such link is in situ resource utilization (ISRU) as an enabling technology to provide consumables such as fuels, oxygen, sweep and utility gases from the Mars atmosphere.

Sridhar, K. R.↗

ISRU Technologies for Mars Life Support

The primary objectives of the Mars Exploration program are to collect data for planetary science in a quest to answer questions related to Origins, to search for evidence of extinct and extant life, and to expand the human presence in the solar system. The public and political engagement that is critical for support of a Mars exploration program is based on all of these objectives. In order to retain and to build public and political support, it is important for NASA to have an integrated Mars exploration plan, not separate robotic and human plans that exist in parallel or in sequence. The resolution stemming from the current architectural review and prioritization of payloads may be pivotal in determining whether NASA will have such a unified plan and retain public support. There are several potential scientific and technological links between the robotic-only missions that have been flown and planned to date, and the robotic + human missions that will come in the future. Taking advantage of and leveraging those links are central to the idea of a unified Mars exploration plan. One such link is in situ resource utilization (ISRU) as an enabling technology to provide consumables such as fuels, oxygen, sweep and utility gases from the Mars atmosphere. ISRU for propellant production and for generation of life support consumables is a key element of human exploration mission plans because of the tremendous savings that can be realized in terms of launch costs and reduction in overall risk to the mission. The Human Exploration and Development of Space (HEDS) Enterprise has supported ISRU technology development for several years, and is funding the MIP and PROMISE payloads that will serve as the first demonstrations of ISRU technology for Mars. In our discussion and presentation at the workshop, we will highlight how the PROMISE ISRU experiment that has been selected by HEDS for a future Mars flight opportunity can extend and enhance the science experiments on board.

Finn, John E.↗

LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages

The rapid evolution of large language models (LLMs) has opened new possibilities for automating various tasks in software development. This paper evaluates the capabilities of the LLaMA 2-70B model in automating these tasks for scientific applications written in commonly used programming languages. Using representative test problems, we assess the model's capacity to generate code, documentation, and unit tests, as well as its ability to translate existing code between commonly used programming languages. Our comprehensive analysis evaluates the compilation, runtime behavior, and correctness of the generated and translated code. Additionally, we assess the quality of automatically generated code, documentation, and unit tests. Here, our results indicate that while LLaMA 2-70B frequently generates syntactically correct and functional code for simpler numerical tasks, it encounters substantial difficulties with more complex, parallelized, or distributed computations, requiring considerable manual corrections. We identify key limitations and suggest areas for future improvements to better leverage AI-driven automation in scientific computing workflows.

97 MATHEMATICS AND COMPUTING↗

Trade-off results and preliminary designs of Near-Term Hybrid Vehicles

Phase I of the Near-Term Hybrid Vehicle Program involved the development of preliminary designs of electric/heat engine hybrid passenger vehicles. The preliminary designs were developed on the basis of mission analysis, performance specification, and design trade-off studies conducted independently by four contractors. THe resulting designs involve parallel hybrid (heat engine/electric) propulsion systems with significant variation in component selection, power train layout, and control strategy. Each of the four designs is projected by its developer as having the potential to substitute electrical energy for 40% to 70% of the petroleum fuel consumed annually by its conventional counterpart.

Sandberg, J. J.↗

A high-order language for a system of closely coupled processing elements

The research reported in this paper was occasioned by the requirements on part of the Real-Time Digital Simulator (RTDS) project under way at NASA Lewis Research Center. The RTDS simulation scheme employs a network of CPUs running lock-step cycles in the parallel computations of jet airplane simulations. Their need for a high order language (HOL) that would allow non-experts to write simulation applications and that could be implemented on a possibly varying network can best be fulfilled by using the programming language Ada. We describe how the simulation problems can be modeled in Ada, how to map a single, multi-processing Ada program into code for individual processors, regardless of network reconfiguration, and why some Ada language features are particulary well-suited to network simulations.

Feyock, S.↗

Instructions to working groups

The key to the success of this workshop is your active participation in the working group process. The goals of this workshop are to address four major questions regarding Cockpit Resource Management (CRM) Training. To some extent the working group topic areas parallel these issues, but in some cases they do not. However, it is important for all of the working groups to keep these general questions in mind during their deliberations: (1) What are the essential elements of an optimal CRM Training program; (2) What are the strengths and weaknesses of current approaches to CRM Training; (3) How can CRM Training best be implemented, and what barriers exist; and (4) Is CRM Training effective, do we know, and if not, how can we find out.

Foushee, H. Clayton↗

DeepSAT: A Deep Learning Approach to Tree-Cover Delineation in 1-m NAIP Imagery for the Continental United States

High resolution tree cover classification maps are needed to increase the accuracy of current land ecosystem and climate model outputs. Limited studies are in place that demonstrates the state-of-the-art in deriving very high resolution (VHR) tree cover products. In addition, most methods heavily rely on commercial softwares that are difficult to scale given the region of study (e.g. continents to globe). Complexities in present approaches relate to (a) scalability of the algorithm, (b) large image data processing (compute and memory intensive), (c) computational cost, (d) massively parallel architecture, and (e) machine learning automation. In addition, VHR satellite datasets are of the order of terabytes and features extracted from these datasets are of the order of petabytes. In our present study, we have acquired the National Agriculture Imagery Program (NAIP) dataset for the Continental United States at a spatial resolution of 1-m. This data comes as image tiles (a total of quarter million image scenes with ~60 million pixels) and has a total size of ~65 terabytes for a single acquisition. Features extracted from the entire dataset would amount to ~8-10 petabytes. In our proposed approach, we have implemented a novel semi-automated machine learning algorithm rooted on the principles of "deep learning" to delineate the percentage of tree cover. Using the NASA Earth Exchange (NEX) initiative, we have developed an end-to-end architecture by integrating a segmentation module based on Statistical Region Merging, a classification algorithm using Deep Belief Network and a structured prediction algorithm using Conditional Random Fields to integrate the results from the segmentation and classification modules to create per-pixel class labels. The training process is scaled up using the power of GPUs and the prediction is scaled to quarter million NAIP tiles spanning the whole of Continental United States using the NEX HPC supercomputing cluster. An initial pilot over the state of California spanning a total of 11,095 NAIP tiles covering a total geographical area of 163,696 sq. miles has produced true positive rates of around 88 percent for fragmented forests and 74 percent for urban tree cover areas, with false positive rates lower than 2 percent for both landscapes.

Imagery↗

Status of the Delco Systems Operations forward looking windshear detection program

Delco Systems Operations, a division of General Motors Hughes Electronics Corporation, is developing a Forward Looking Windshear Detection System based on the integration of infrared remote sensing and accelerometer reactive sensing technologies. The infrared sensor is a multi-spectral, scanning radiometer operating in the 8 to 14 micron region. A 2 x 5 detector array with parallel-serial scanning produces 60 degrees horizontal and 10 degrees vertical-fields of view. Using multiple wavelength signals, azimuth temperature gradients are analyzed for characteristic signatures of thermally induced windshear phenomena. Elevation temperature gradients are processed through an atmosphere model to continuously compute a stability index for arming microburst detection criteria. The atmosphere model and proprietary computer processing algorithms combine to generate coarse estimates of disturbance ranges based on multiple wavelength radiance data with different extinction coefficients. Computer outputs of atmospheric stability, disturbance intensity, and azimuth and range information provide a situation display capability. A ground operated, experimental radiometer has been developed and is being used to verify the detection and discrimination concepts at an atmospheric and simulated rain test facility in Milwaukee. A prototype airborne radiometer is being developed for flight test evaluation during the summer of 1989.

Gallagher, Brian J.↗