Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel Performance Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38

Turbine Electrified Energy Management for Single Aisle Aircraft

Electrified aircraft propulsion technology is being developed to reduce the environmental impacts of the aviation industry. This is prompting the exploration of potential uses and benefits of hybrid systems in which electric powertrains are integrated with more traditional gas turbine propulsion systems. Turbine Electrified Energy Management (TEEM) is an energy management approach for hybrid-electric architectures in which electric machines are connected to the turbofan shafts and used to suppress the off-design operation naturally associated with engine transients. This reduces the need to maintain a large amount of compressor operability margin, thus allowing further exploration of the engine design space. In this study, a 19,000 lbf engine within a parallel hybrid propulsion system is considered along with a 30,000 lbf standalone engine. Data from prior TEEM applications are used to approximate the electric machine sizing required to achieve operability benefits. The TEEM controller is shown to improve operability during transients through the reduction of stall margin undershoots and the decrease of transient variations in component performance maps by over 29%.

Controls↗

Turbine Electrified Energy Management for Single Aisle Aircraft

Electrified aircraft propulsion technology is being developed to reduce the environmental impacts of the aviation industry. This is prompting the exploration of potential uses and benefits of hybrid systems in which electric powertrains are integrated with more traditional gas turbine propulsion systems. Turbine Electrified Energy Management (TEEM) is an energy management approach for hybrid-electric architectures in which electric machines are connected to the turbofan shafts and used to suppress the off-design operation naturally associated with engine transients. This reduces the need to maintain a large amount of compressor operability margin, thus allowing further exploration of the engine design space. In this study, a 19,000 lbf engine within a parallel hybrid propulsion system is considered along with a 30,000 lbf standalone engine. Data from prior TEEM applications are used to approximate the electric machine sizing required to achieve operability benefits. The TEEM controller is shown to improve operability during transients through the reduction of stall margin undershoots and the decrease of transient variations in component performance maps by over 29%.

EAP↗

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning↗

Liquid Nitrogen (Oxygen Simulent) Thermodynamic Venting System Test Data Analysis

In designing systems for the long-term storage of cryogens in low gravity space environments, one must consider the effects of thermal stratification on excessive tank pressure that will occur due to environmental heat leakage. During low gravity operations, a Thermodynamic Venting System (TVS) concept is expected to maintain tank pressure without propellant resettling. The TVS consists of a recirculation pump, Joule-Thomson (J-T) expansion valve, and a parallel flow concentric tube heat exchanger combined with a longitudinal spray bar. Using a small amount of liquid extracted by the pump and passing it though the J-T valve, then through the heat exchanger, the bulk liquid and ullage are cooled, resulting in lower tank pressure. A series of TVS tests were conducted at the Marshall Space Flight Center using liquid nitrogen as a liquid oxygen simulant. The tests were performed at fill levels of 90%, 50%, and 25% with gaseous nitrogen and helium pressurants, and with a tank pressure control band of 7 kPa. A transient one-dimensional model of the TVS is used to analyze the data. The code is comprised of four models for the heat exchanger, the spray manifold and injector tubes, the recirculation pump, and the tank. The TVS model predicted ullage pressure and temperature and bulk liquid saturation pressure and temperature are compared with data. Details of predictions and comparisons with test data regarding pressure rise and collapse rates will be presented in the final paper.

Hedayat, A.↗

Investigation and Diagnosis of Faulty Data Channels in CMS Outer Tracker Module Testing

The High-Luminosity Large Hadron Collider (HL-LHC) is currently undergoing upgrades to improve its luminosity. In parallel, this requires an upgrade to the Compact Muon Solenoid (CMS)’s Outer Tracker, consisting of Pixel-Strip (PS) and Strip-Strip (2S) modules that can accurately track the path of charged particles originating from the collisions. It follows that such complex modules call for extensive testing, requiring a sophisticated Data Acquisition (DAQ) system that can perform specific tests to assess their performance. In addition, errors caused by the hardware of a given testing station, and its associated data channel, need to be accurately identified to guarantee proper testing of modules. We have developed a software extension to the Phase-II Outer Tracker Analyzer of Test Outputs (POTATO), which is a specialized software designed to analyze and grade all of the module tests through a centralized database. This extension categorizes and analyzes module test results by its station and data channel. Its analysis can be used to identify trends in grading that indicate issues in these channels’ grading process rather than in the individual modules. This poster shows our methodology and results for identifying faulty data channels. Using this extension, we can quickly diagnose and address problems in our DAQ system, ensuring proper evaluation corrections for each module.

Chen, Angus [Fermilab]↗

Investigation and Diagnosis of Faulty Data Channels in CMS Outer Tracker Module Testing

The High-Luminosity Large Hadron Collider (HL-LHC) is currently undergoing upgrades to improve its luminosity. In parallel, this requires an upgrade to the Compact Muon Solenoid (CMS)’s Outer Tracker, consisting of Pixel-Strip (PS) and Strip-Strip (2S) modules that can accurately track the path of charged particles originating from the collisions. It follows that such complex modules call for extensive testing, requiring a sophisticated Data Acquisition (DAQ) system that can perform specific tests to assess their performance. In addition, errors caused by the hardware of a given testing station, and its associated data channel, need to be accurately identified to guarantee proper testing of modules. We have developed a software extension to the Phase-II Outer Tracker Analyzer of Test Outputs (POTATO), which is a specialized software designed to analyze and grade all of the module tests through a centralized database. This extension categorizes and analyzes module test results by its station and data channel. Its analysis can be used to identify trends in grading that indicate issues in these channels’ grading process rather than in the individual modules. This poster shows our methodology and results for identifying faulty data channels. Using this extension, we can quickly diagnose and address problems in our DAQ system, ensuring proper evaluation corrections for each module.

Chen, Angus [Fermilab]↗

Software Engineering Support of the Third Round of Scientific Grand Challenge Investigations: Earth System Modeling Software Framework Survey

One of the most significant challenges in large-scale climate modeling, as well as in high-performance computing in other scientific fields, is that of effectively integrating many software models from multiple contributors. A software framework facilitates the integration task, both in the development and runtime stages of the simulation. Effective software frameworks reduce the programming burden for the investigators, freeing them to focus more on the science and less on the parallel communication implementation. while maintaining high performance across numerous supercomputer and workstation architectures. This document surveys numerous software frameworks for potential use in Earth science modeling. Several frameworks are evaluated in depth, including Parallel Object-Oriented Methods and Applications (POOMA), Cactus (from (he relativistic physics community), Overture, Goddard Earth Modeling System (GEMS), the National Center for Atmospheric Research Flux Coupler, and UCLA/UCB Distributed Data Broker (DDB). Frameworks evaluated in less detail include ROOT, Parallel Application Workspace (PAWS), and Advanced Large-Scale Integrated Computational Environment (ALICE). A host of other frameworks and related tools are referenced in this context. The frameworks are evaluated individually and also compared with each other.

Talbot, Bryan↗

What Multilevel Parallel Programs do when you are not Watching: A Performance Analysis Case Study Comparing MPI/OpenMP, MLP, and Nested OpenMP

With the current trend in parallel computer architectures towards clusters of shared memory symmetric multi-processors, parallel programming techniques have evolved that support parallelism beyond a single level. When comparing the performance of applications based on different programming paradigms, it is important to differentiate between the influence of the programming model itself and other factors, such as implementation specific behavior of the operating system (OS) or architectural issues. Rewriting-a large scientific application in order to employ a new programming paradigms is usually a time consuming and error prone task. Before embarking on such an endeavor it is important to determine that there is really a gain that would not be possible with the current implementation. A detailed performance analysis is crucial to clarify these issues. The multilevel programming paradigms considered in this study are hybrid MPI/OpenMP, MLP, and nested OpenMP. The hybrid MPI/OpenMP approach is based on using MPI [7] for the coarse grained parallelization and OpenMP [9] for fine grained loop level parallelism. The MPI programming paradigm assumes a private address space for each process. Data is transferred by explicitly exchanging messages via calls to the MPI library. This model was originally designed for distributed memory architectures but is also suitable for shared memory systems. The second paradigm under consideration is MLP which was developed by Taft. The approach is similar to MPi/OpenMP, using a mix of coarse grain process level parallelization and loop level OpenMP parallelization. As it is the case with MPI, a private address space is assumed for each process. The MLP approach was developed for ccNUMA architectures and explicitly takes advantage of the availability of shared memory. A shared memory arena which is accessible by all processes is required. Communication is done by reading from and writing to the shared memory.

Jost, Gabriele↗

Assessing the Ability of Instantaneous Aircraft and Sonde Measurements to Characterize Climatological Means and Long-Term Trends in Tropospheric Composition

Over four decades of measurements exist that sample the 3-D composition of reactive trace gases in the troposphere from approximately weekly ozone sondes, instrumentation on civil aircraft, and individual comprehensive aircraft field campaigns. An obstacle to using these data to evaluate coupled chemistry-climate models (CCMs)the models used to project future changes in atmospheric composition and climateis that exact space-time matching between model fields and observations cannot be done, as CCMs generate their own meteorology. Evaluation typically involves averaging over large spatiotemporal regions, which may not reflect a true average due to limited or biased sampling. This averaging approach generally loses information regarding specific processes. Here we aim to identify where discrete sampling may be indicative of long-term mean conditions, using the GEOS-Chem global chemical-transport model (CTM) driven by the MERRA reanalysis to reflect historical meteorology from 2003 to 2012 at 2o by 2.5o resolution. The model has been sampled at the time and location of every ozone sonde profile available from the Would Ozone and Ultraviolet Radiation Data Centre (WOUDC), along the flight tracks of the IAGOSMOZAICCARABIC civil aircraft campaigns, as well as those from over 20 individual field campaigns performed by NASA, NOAA, DOE, NSF, NERC (UK), and DLR (Germany) during the simulation period. Focusing on ozone, carbon monoxide and reactive nitrogen species, we assess where aggregates of the in situ data are representative of the decadal mean vertical, spatial and temporal distributions that would be appropriate for evaluating CCMs. Next, we identically sample a series of parallel sensitivity simulations in which individual emission sources (e.g., lightning, biogenic VOCs, wildfires, US anthropogenic) have been removed one by one, to assess where and when the aggregated observations may offer constraints on these processes within CCMs. Lastly, we show results of an additional 31-year simulation from 1980-2010 of GEOS-Chem driven by the MACCity emissions inventory and MERRA reanalysis at 4o by 5o. We sample the model at every WOUDC sonde and flight track from MOZAIC and NASA field campaigns to evaluate which aggregate observations are statistically reflective of long-term trends over the period.

Atmospheric composition↗

Analysis of the Value Added When Deploying a Model-Based Approach for the Validation and Verification of the Medical Database Software

The Medical Database (MD) is a virtual repository consisting of two software components: Medical Item Database (MedID) and the Evidence Library (EL). MedID consists of engineering data and associated information for specific medical resource items (e.g., pharmaceutical, medical devices, and supporting components), while the EL is a tool which provides all of the medical evidence necessary. The MD will 1) serve as the single “source of truth” for the Informing Mission Planning via Analysis of Complex Tradespaces (IMPACT) tool suite for both medical evidence and medical resource engineering data and 2) will be used in conjunction with the IMPACT tool suite to inform research prioritizations and perform systematic trade study evaluations to aid stakeholders in making informed decisions regarding simulated human spaceflight missions. The MD project used a Model-Based Systems Engineering (MBSE) approach to support all life cycles of the software development, while in parallel the human factors engineering team used modeling to support Human Centered Design (HCD) strategies in an effort to improve software usability. HCD is a frequently used approach in design frameworks that develops resolutions to complexities and challenges by involving the human perspective in all steps of the problem-solving process. By integrating the model-based approaches used for systems engineering and human factors activities, the project is able to leverage the model-based artifacts originally created for HCD activities for system level and human factors validation. In this presentation, our team highlights the value added when leveraging these model-based artifacts to support the on-going verification and validation activities.

C. Laing↗

A real time neural net estimator of fatigue life

A neural net architecture is proposed to estimate, in real-time, the fatigue life of mechanical components, as part of the Intelligent Control System for Reusable Rocket Engines. Arbitrary component loading values were used as input to train a two hidden-layer feedforward neural net to estimate component fatigue damage. The ability of the net to learn, based on a local strain approach, the mapping between load sequence and fatigue damage has been demonstrated for a uniaxial specimen. Because of its demonstrated performance, the neural computation may be extended to complex cases where the loads are biaxial or triaxial, and the geometry of the component is complex (e.g., turbopump blades). The generality of the approach is such that load/damage mappings can be directly extracted from experimental data without requiring any knowledge of the stress/strain profile of the component. In addition, the parallel network architecture allows real-time life calculations even for high frequency vibrations. Owing to its distributed nature, the neural implementation will be robust and reliable, enabling its use in hostile environments such as rocket engines. This neural net estimator of fatigue life is seen as the enabling technology to achieve component life prognosis, and therefore would be an important part of life extending control for reusable rocket engines.

Troudet, T.↗

A real time neural net estimator of fatigue life

A neural network architecture is proposed to estimate, in real-time, the fatigue life of mechanical components, as part of the intelligent Control System for Reusable Rocket Engines. Arbitrary component loading values were used as input to train a two hidden-layer feedforward neural net to estimate component fatigue damage. The ability of the net to learn, based on a local strain approach, the mapping between load sequence and fatigue damage has been demonstrated for a uniaxial specimen. Because of its demonstrated performance, the neural computation may be extended to complex cases where the loads are biaxial or triaxial, and the geometry of the component is complex (e.g., turbopumps blades). The generality of the approach is such that load/damage mappings can be directly extracted from experimental data without requiring any knowledge of the stress/strain profile of the component. In addition, the parallel network architecture allows real-time life calculations even for high-frequency vibrations. Owing to its distributed nature, the neural implementation will be robust and reliable, enabling its use in hostile environments such as rocket engines.

Troudet, T.↗

Scalable Performance Environments for Parallel Systems

As parallel systems expand in size and complexity, the absence of performance tools for these parallel systems exacerbates the already difficult problems of application program and system software performance tuning. Moreover, given the pace of technological change, we can no longer afford to develop ad hoc, one-of-a-kind performance instrumentation software; we need scalable, portable performance analysis tools. We describe an environment prototype based on the lessons learned from two previous generations of performance data analysis software. Our environment prototype contains a set of performance data transformation modules that can be interconnected in user-specified ways. It is the responsibility of the environment infrastructure to hide details of module interconnection and data sharing. The environment is written in C++ with the graphical displays based on X windows and the Motif toolkit. It allows users to interconnect and configure modules graphically to form an acyclic, directed data analysis graph. Performance trace data are represented in a self-documenting stream format that includes internal definitions of data types, sizes, and names. The environment prototype supports the use of head-mounted displays and sonic data presentation in addition to the traditional use of visual techniques.

Reed, Daniel A.↗

Climate Ocean Modeling on a Beowulf Class System

With the growing power and shrinking cost of personal computers. the availability of fast ethernet interconnections, and public domain software packages, it is now possible to combine them to build desktop parallel computers (named Beowulf or PC clusters) at a fraction of what it would cost to buy systems of comparable power front supercomputer companies. This led as to build and assemble our own sys tem. specifically for climate ocean modeling. In this article, we present our experience with such a system, discuss its network performance, and provide some performance comparison data with both HP SPP2000 and Cray T3E for an ocean Model used in present-day oceanographic research.

Cheng, B. N.↗

Flexible All-Digital Receiver for Bandwidth Efficient Modulations

An all-digital high data rate parallel receiver architecture developed jointly by Goddard Space Flight Center and the Jet Propulsion Laboratory is presented. This receiver utilizes only a small number of high speed components along with a majority of lower speed components operating in a parallel frequency domain structure implementable in CMOS, and can currently process up to 600 Mbps with standard QPSK modulation. Performance results for this receiver for bandwidth efficient QPSK modulation schemes such as square-root raised cosine pulse shaped QPSK and Feher's patented QPSK are presented, demonstrating the flexibility of the receiver architecture.

Gray, Andrew↗

Real-Time Adaptive Lossless Hyperspectral Image Compression using CCSDS on Parallel GPGPU and Multicore Processor Systems

The proposed CCSDS (Consultative Committee for Space Data Systems) Lossless Hyperspectral Image Compression Algorithm was designed to facilitate a fast hardware implementation. This paper analyses that algorithm with regard to available parallelism and describes fast parallel implementations in software for GPGPU and Multicore CPU architectures. We show that careful software implementation, using hardware acceleration in the form of GPGPUs or even just multicore processors, can exceed the performance of existing hardware and software implementations by up to 11x and break the real-time barrier for the first time for a typical test application.

realtime↗

Automated and highly parallelized Bayesian optimization scheme for direct drive fusion experiments on OMEGA

Finding the optimal implosion design on existing experimental facilities for inertial confinement fusion requires an exhaustive search of the vast design parameter space. This is infeasible both with experiments and with simulations. Consequently, a large fraction of the experimentally realizable design space remains unexplored, and new design schemes are challenging to optimize in a reasonable time frame. On the OMEGA laser facility, predictive machine learning models have been developed to accurately forecast the result of an experiment using only inexpensive simulations and the large dataset of prior experimental data. However, the full design space remains vast enough to be unassailable with simple optimization techniques. Here we develop an automated and optimally parallel Bayesian optimization algorithm that can entirely optimize the target and pulse shape of a direct-drive ICF implosion under a given design paradigm. We use this algorithm to find a markedly improved design for the performance implosions on OMEGA that is predicted to hydroequivalently scale to ignition at 2.15 MJ.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

SPARTAN (Scalable Probabilistic Application Reconfigurable Tensor Autonomous Network)

The technical founder of Ludwig Computing Inc has been competitively selected for support by Cyclotron Road, a U.S. Department of Energy (DOE) Advanced Manufacturing Office (AMO) Lab-Embedded Entrepreneurship Program (LEEP) through an approved merit review process. Ludwig Computing Inc, supported by the U.S. Department of Energy's Advanced Manufacturing Office through the Cyclotron Road program, has investigated the advantages of probabilistic computing for real-world compute-intensive applications. This research adds to the understanding of alternative computing paradigms by exploring a unique hardware-software co-design that integrates quantum computing methods with nature-inspired problem-solving techniques. The project's focus on areas such as combinatorial optimization, graph analytics, and machine learning demonstrates the potential for significant advancements in computational efficiency and performance. By harnessing natural randomness to streamline large circuits into fewer devices, Ludwig's approach enables massive parallelism, potentially offering higher throughput, speed, and energy efficiency compared to conventional hardware solutions. This work benefits the public by paving the way for more efficient computing solutions that could address complex real-world problems while potentially reducing energy consumption in data-intensive industries.

97 MATHEMATICS AND COMPUTING↗