Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel Performance Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50

Flexible User-Defined Domain Decomposition in Kilometer-Scale E3SM Land Model Simulation

The Energy Exascale Earth System Model (E3SM) Land Model (ELM) has been extended to kilometer-scale (km-ELM) resolutions, enabling high-fidelity simulations of terrestrial processes at 1 km x 1 km grid spacing. In ELM, domain decomposition partitions the computational domain across processors, ensuring efficient parallel execution. Currently, round-robin decomposition is applied, providing a straightforward way to distribute computational workload. As ELM continues evolving at the kilometer-scale (km-scale), particularly with integrating lateral flow modeling, decomposition strategies must also account for the increased workload and data movement. This paper introduces a flexible user-defined domain decomposition framework, allowing users to customize domain partitioning based on application requirements. The impact of different decomposition strategies is evaluated across various applications concerning computation, communication, and I/O. Results demonstrate that while 1D partitioning yields superior I/O performance, k-nearest neighbors (KNN) clustering effectively reduces inter-process communication overhead. This study lays the groundwork for scalable partitioning in large-scale land surface simulations, enhancing next-generation Earth system modeling.

Wang, Dali [ORNL] (ORCID:0000000168065108)↗

Planning of reach-and-grasp movements: effects of validity and type of object information

Individuals are assumed to plan reach-and-grasp movements by using two separate processes. In 1 of the processes, extrinsic (direction, distance) object information is used in planning the movement of the arm that transports the hand to the target location (transport planning); whereas in the other, intrinsic (shape) object information is used in planning the preshaping of the hand and the grasping of the target object (manipulation planning). In 2 experiments, the authors used primes to provide information to participants (N = 5, Experiment 1; N = 6, Experiment 2) about extrinsic and intrinsic object properties. The validity of the prime information was systematically varied. The primes were succeeded by a cue, which always correctly identified the location and shape of the target object. Reaction times were recorded. Four models of transport and manipulation planning were tested. The only model that was consistent with the data was 1 in which arm transport and object manipulation planning were postulated to be independent processes that operate partially in parallel. The authors suggest that the processes involved in motor planning before execution are primarily concerned with the geometric aspects of the upcoming movement but not with the temporal details of its execution.

Cues↗

Simulation Schiaparelli's Entry and Comparison to Aerothermal Flight Data

The European Space Agency recently flew an entry, descent, and landing demonstrator module called Schiaparelli that entered the atmosphere of Mars on the 19th of October, 2016. The instrumentation suite included heatshield and backshell pressure transducers and thermocouples (known as AMELIA - Atmospheric Mars Entry and Landing Investigations and Analysis) and backshell radiation and direct heat flux-sensing sensors (known as COMARS (Combined Aerothermal and Radiometer Sensors Instrument Package) and ICOTOM (narrow band radiometers)). Due to the failed landing of Schiaparelli, only a subset of the flight data was transmitted before and after plasma black-out. The goal of this paper is to present comparisons of the flight data with calculations from NASA simulation tools, DPLR (Data Parallel Line Relaxation) / NEQAIR (NonEQuilibrium AIr Radiation) and LAURA (Langley Aerothermodynamic Upwind Relaxation Algorithm) / HARA (High-temperature Aerothermodynamic RAdiation ). DPLR and LAURA are used to calculate the flowfield around the vehicle and surface properties, such as pressure and convective heating. The flowfield data are passed to NEQAIR and HARA to calculate the radiative heat flux. Comparisons will be made to the COMARS total heat flux, radiative heat flux and pressure measurements. Results will also be shown against the reconstructed heat flux which was calculated from an inverse analysis of the AMELIA thermocouple data performed by Astrium. Preliminary calculations are presented in this abstract.

Brandis, A. M.↗

Sparse Linear Algebra Toolkit for Computational Aerodynamics

Finding solutions to sparse linear systems of equations is an essential step in Computational Engineering applications of interest to NASA. Linear systems of equations are composed and solved in almost every computational engineering application. The characteristics of linear systems vary greatly from one application to another. Accordingly, there are a wide variety of methods for the solution of linear systems of equations. The operations and methods prepared by the authors are focused on linear systems of interest to NASA, primarily those associated with Computational Fluid Dynamics (CFD), Aeroelasticity, and Aeroacoustics. The Sparse Linear Algebra Toolkit (SLAT) is a coordinated collection of software featuring operations, methods, and data structures that are useful when solving sparse linear systems of equations on modern computer architectures. The implemented operations and methods are designed and tuned for parallelism in shared memory, in distributed memory, and across the hybrid combination of distributed-shared memory. The toolkit includes novel methods and implementations for modern architectures and facilitates development of new approaches for meeting NASA’s evolving computational engineering challenges using evolving computer architectures that are not available in vendor libraries. In this paper, significant features and interfaces within SLAT are presented and verified for simulations performed with NASA’s CFD solver, FUN3D. The runtime and scaling performance of the Generalized Minimum Residual (GMRES) method implemented in SLAT is analyzed for the linear subproblems within the solution of turbulent Navier-Stokes equations employed in the simulation of high-lift configurations. Prior to this work, the SPARSKIT GMRES implementation was the only Krylov subspace method available within FUN3D. A strong scaling study shows the SLAT GMRES implementation facilitates accurate Reynolds-averaged Navier-Stokes CFD solutions between 15% and 56% faster than the SPARSKIT GMRES implementation.

Stephen L Wood↗

Evaluating the Sensitivity of Agricultural Model Performance to Different Climate Inputs: Supplemental Material

Projections of future food production necessarily rely on models, which must themselves be validated through historical assessments comparing modeled and observed yields. Reliable historical validation requires both accurate agricultural models and accurate climate inputs. Problems with either may compromise the validation exercise. Previous studies have compared the effects of different climate inputs on agricultural projections but either incompletely or without a ground truth of observed yields that would allow distinguishing errors due to climate inputs from those intrinsic to the crop model. This study is a systematic evaluation of the reliability of a widely used crop model for simulating U.S. maize yields when driven by multiple observational data products. The parallelized Decision Support System for Agrotechnology Transfer (pDSSAT) is driven with climate inputs from multiple sources reanalysis, reanalysis that is bias corrected with observed climate, and a control dataset and compared with observed historical yields. The simulations show that model output is more accurate when driven by any observation-based precipitation product than when driven by non-bias-corrected reanalysis. The simulations also suggest, in contrast to previous studies, that biased precipitation distribution is significant for yields only in arid regions. Some issues persist for all choices of climate inputs: crop yields appear to be oversensitive to precipitation fluctuations but under sensitive to floods and heat waves. These results suggest that the most important issue for agricultural projections may be not climate inputs but structural limitations in the crop models themselves.

simulation↗

Blade row dynamic digital compression program. Volume 2: J85 circumferential distortion redistribution model, effect of Stator characteristics, and stage characteristics sensitivity study

The results of dynamic digital blade row compressor model studies of a J85-13 engine are reported. The initial portion of the study was concerned with the calculation of the circumferential redistribution effects in the blade-free volumes forward and aft of the compression component. Although blade-free redistribution effects were estimated, no significant improvement over the parallel-compressor type solution in the prediction of total-pressure inlet distortion stability limit was obtained for the J85-13 engine. Further analysis was directed to identifying the rotor dynamic response to spatial circumferential distortions. Inclusion of the rotor dynamic response led to a considerable gain in the ability of the model to match the test data. The impact of variable stator loss on the prediction of the stability limit was evaluated. An assessment of measurement error on the derivation of the stage characteristics and predicted stability limit of the compressor was also performed.

Tesch, W. A.↗

Overview and extensions of a system for routing directed graphs on SIMD architectures

Many problems can be described in terms of directed graphs that contain a large number of vertices where simple computations occur using data from adjacent vertices. A method is given for parallelizing such problems on an SIMD machine model that uses only nearest neighbor connections for communication, and has no facility for local indirect addressing. Each vertex of the graph will be assigned to a processor in the machine. Rules for a labeling are introduced that support the use of a simple algorithm for movement of data along the edges of the graph. Additional algorithms are defined for addition and deletion of edges. Modifying or adding a new edge takes the same time as parallel traversal. This combination of architecture and algorithms defines a system that is relatively simple to build and can do fast graph processing. All edges can be traversed in parallel in time O(T), where T is empirically proportional to the average path length in the embedding times the average degree of the graph. Additionally, researchers present an extension to the above method which allows for enhanced performance by allowing some broadcasting capabilities.

Tomboulian, Sherryl↗

Analysis of the Capacity Potential of Current Day and Novel Configurations for New York's John F. Kennedy Airport

In 2015, a series of systems analysis studies were conducted on John F. Kennedy Airport in New York (NY) in a collaborative effort between NASA and the Port Authority of New York and New Jersey (PANYNJ). This work was performed to build a deeper understanding of NY airspace and operations to determine the improvements possible through operational changes with tools currently available, and where new technology is required for additional improvement. The analysis was conducted using tool-based mathematical analyses, video inspection and evaluation using recorded arrival/departure/surface traffic captured by the Aerobahn tool (used by Kennedy Airport for surface metering), and aural data archives available publically through the web to inform the video segments. A discussion of impacts of trajectory and operational choices on capacity is presented, including runway configuration and usage (parallel, converging, crossing, shared, independent, staggered), arrival and departure route characteristics (fix sharing, merges, splits), and how compression of traffic is staged. The authorization in March of 2015 for New York to use reduced spacing under the Federal Aviation Administration (FAA) Wake Turbulence Recategorization (RECAT) also offers significant capacity benefit for New York airports when fully transitioned to the new spacing requirements, and the impact of these changes for New York is discussed. Arrival and departure capacity results are presented for each of the current day Kennedy Airport configurations. While the tools allow many variations of user-selected conditions, the analysis for these studies used arrival-priority, no-winds, additional safety buffer of 5% to the required minimum spacing, and a mix of traffic typical for Kennedy. Two additional "novel" configurations were evaluated. These configurations are of interest to Port Authority and to their airline customers, and are believed to offer near-term capacity benefit with minimal operational and equipage changes. One of these is the addition of an Optimized Profile Descent (OPD) route to runways 22L and 22R, and the other is the simultaneous use of 4 runways, which is not currently done at Kennedy. The background and configuration for each of these is described, and the capacity results are presented along with a discussion of drawbacks and enablers for each.

Glaab, Patricia↗

Rapid code acquisition algorithms employing PN matched filters

The performance of four algorithms using pseudonoise matched filters (PNMFs), for direct-sequence spread-spectrum systems, is analyzed. They are: parallel search with fix dwell detector (PL-FDD), parallel search with sequential detector (PL-SD), parallel-serial search with fix dwell detector (PS-FDD), and parallel-serial search with sequential detector (PS-SD). The operation characteristic for each detector and the mean acquisition time for each algorithm are derived. All the algorithms are studied in conjunction with the noncoherent integration technique, which enables the system to operate in the presence of data modulation. Several previous proposals using PNMF are seen as special cases of the present algorithms.

Su, Yu T.↗

Parallel asynchronous hardware implementation of image processing algorithms

Research is being carried out on hardware for a new approach to focal plane processing. The hardware involves silicon injection mode devices. These devices provide a natural basis for parallel asynchronous focal plane image preprocessing. The simplicity and novel properties of the devices would permit an independent analog processing channel to be dedicated to every pixel. A laminar architecture built from arrays of the devices would form a two-dimensional (2-D) array processor with a 2-D array of inputs located directly behind a focal plane detector array. A 2-D image data stream would propagate in neuron-like asynchronous pulse-coded form through the laminar processor. No multiplexing, digitization, or serial processing would occur in the preprocessing state. High performance is expected, based on pulse coding of input currents down to one picoampere with noise referred to input of about 10 femtoamperes. Linear pulse coding has been observed for input currents ranging up to seven orders of magnitude. Low power requirements suggest utility in space and in conjunction with very large arrays. Very low dark current and multispectral capability are possible because of hardware compatibility with the cryogenic environment of high performance detector arrays. The aforementioned hardware development effort is aimed at systems which would integrate image acquisition and image processing.

Coon, Darryl D.↗

Remote Objects Message Exchange (ROME)

The performance of a single program running on a single processor is limited by the character of the processor. Moreover, the cost and difficulty of developing and sustaining programs tend to increase as their size and complexity increase. Clearly there ought to be some advantage in partitioning powerful application software in relatively small and simple components that can run in parallel on multiple processors; the software should run faster and it should be cheaper and easier to deploy. Remote Objects Message Exchange (ROME) is an attempt to provide a single relatively simple, universally available abstraction for data communication among C++ objects. It aims to enable the C++ application developer to specify objects' interactions with other objects wholly in terms of the application domain, without concern for details of interprocess communication. Every ROME-compliant object is conceptually a network peer of every other, as if each one were (for example) a separate UNIX process.

processors application software data communication↗

Validating Drag and Heating Coefficients for Hollow Reentry Objects in Continuum Flow Using a Mach 7 Ludwieg Tube

Drag and heating coefficient databases and models are crucial to destructive reentry simulation. The NASA Orbital Debris Program Office (ODPO) develops, maintains, and performs analysis with the Object Reentry Survival Analysis Tool (ORSAT), which comprises drag and heating models for free molecular, transitional, and continuum flow regimes. These models have, in the past, only included solid, convex, blunt shapes (such as boxes, spheres, and cylinders). Previous work led by ODPO includes the extension of these models to hollow cylinders and square boxes in free molecular and transitional flow using the Direct Simulation Monte Carlo (DSMC) method. Since 2019, the ODPO has continued its program of DSMC simulations and extended the project to include analyses with the NASA Data Parallel Line Relaxation (DPLR) program on hollow cylinders and boxes (with varying wall thickness-diameter ratio). In fall 2022, the ODPO began a collaboration with the University of Texas San Antonio (UTSA) to use the Mach 7 Ludwieg Tube facility to validate the model built using numerical simulations. This facility can replicate (at a scale of approximately 100:1) the conditions seen by reentering objects near typical demise altitudes. We present here the drag and heating coefficients derived from the continued DSMC simulations, the new DPLR simulations, and the 26-test series at UTSA.

Chris Ostrom↗

Artificial Neural Network Modeling for Airline Disruption Management

Since the 1970s, most airlines have incorporated computerized support for managing disruptions during flight schedule execution. However, existing platforms for airline disruption management (ADM) employ monolithic system design methods that rely on the creation of specific rules and requirements through explicit optimization routines, before a system that meets the specifications is designed. Thus, current platforms for ADM are unable to readily accommodate additional system complexities resulting from the introduction of new capabilities, such as the introduction of unmanned aerial systems (UAS), operations and infrastructure, to the system. To this end, we use historical data on airline scheduling and operations recovery to develop a system of artificial neural networks (ANNs), which describe a predictive transfer function model (PTFM) for promptly estimating the recovery impact of disruption resolutions at separate phases of flight schedule execution during ADM. Furthermore, we provide a modular approach for assessing and executing the PTFM by employing a parallel ensemble method to develop generative routines that amalgamate the system of ANNs. Our modular approach ensures that current industry standards for tardiness in flight schedule execution during ADM are satisfied, while accurately estimating appropriate time-based performance metrics for the separate phases of flight schedule execution.

Kolawole Ogunsina↗

Validating Drag and Heating Coefficients for Hollow Reentry Objects in Continuum Flow Using a Mach 7 Ludwieg Tube

Drag and heating coefficient databases and models are crucial to destructive reentry simulation. The NASA Orbital Debris Program Office (ODPO) develops, maintains, and performs analysis with the Object Reentry Survival Analysis Tool (ORSAT), which comprises drag and heating models for free molecular, transitional, and continuum flow regimes. These models have, in the past, only included solid, convex, blunt shapes (such as boxes, spheres, and cylinders). Previous work led by ODPO includes the extension of these models to hollow cylinders and square boxes in free molecular and transitional flow using the Direct Simulation Monte Carlo (DSMC) method. Since 2019, the ODPO has continued its program of DSMC simulations and extended the project to include analyses with the NASA Data Parallel Line Relaxation (DPLR) program on hollow cylinders and boxes (with varying wall thickness-diameter ratio). In fall 2022, the ODPO began a collaboration with the University of Texas San Antonio (UTSA) to use the Mach 7 Ludwieg Tube facility to validate the model built using numerical simulations. This facility can replicate (at a scale of approximately 100:1) the conditions seen by reentering objects near typical demise altitudes. We present here the drag and heating coefficients derived from the continued DSMC simulations, the new DPLR simulations, and the 26-test series at UTSA.

Chris Ostrom↗

Field-Programmable Gate Array Implementation of a Single Photon-Counting Receive Modem

We present a field-programmable gate array (FPGA) implementation of a single photon-counting receive modem for a pulse position modulated signal. The modem is compliant with the Consultative Committee for Space Data Systems (CCSDS) High Photon Efficiency (HPE) Optical Communications Coding and Synchronization standard and is capable of a maximum data rate of 267 Mbps. The system is designed on a commercial off-the-shelf FPGA platform and utilizes superconducting nanowire single photon counting detectors, analog to digital converters (ADC s) to sample the detectors, and two FPGAs. Symbol timing recovery, photon counting, convolutional deinterleaving, and codeword synchronization areis performed in the first FPGA. The second FPGA performs iterative decoding on each codeword of the serially concatenated pulse position modulated (SCPPM) signal. A digital filter is included to compensate for timing jitter of the detector, and the decoder throughput can be modified adjusted through reconfigurable parallelization. The decoder also implements a resource-efficient, algorithmic polynomial interleaver and deinterleaver. Both FPGAs can be reconfigured to switch between pulse position modulation (PPM)-16 and PPM-32 with code rates 1/3, 1/2, and 2/3. In this paper, we describe the receiver architecture and FPGA implementation of the timing recovery loop and SCPPM decoder, FPGA utilization for the different modes, and receive modem characterization test results.

Field-programmable Gate Array↗

DyG-DPCD: A Distributed Parallel Community Detection Algorithm for Large-Scale Dynamic Graphs

Dynamic (Temporal) graphs capture the valuable evolution of real-world systems, from the continuously evolving patterns of social interactions and genetic pathways to the dynamic fluctuations of economic forces. Detecting communities for such evolving networks poses unique challenges. Detecting and analyzing the evolution of communities within dynamic graphs unlocks valuable insights into the underlying structural and temporal patterns of real-world systems. However, the sheer volume of modern graph data and the inherent complexity of the temporal dimension pose significant challenges to scalable community detection algorithms. Addressing this gap, our work explores the limited landscape of scalable distributed-memory parallel methods specifically designed for dynamic network community detection. We propose a novel parallel algorithm, DyG-DPCD (Dynamic Graph Distributed Parallel Community Detection), to detect communities in dynamic networks using the Message Passing Interface (MPI) framework. We present a vertex-centric approach, allowing us to detect communities through local optimization. Furthermore, we enhance our baseline algorithm by incorporating three heuristics, which improve the algorithm’s performance significantly while maintaining the quality of the solutions. We demonstrate the efficiency of our algorithm by experimenting on several real-world large-scale networks with hundreds of millions of edges spanning diverse domains. Notably, DyG-DPCD achieves speedups between 25× and 30× for large networks that we experimented on using NERSC compute nodes. In conclusion, our algorithm outperforms the STINGER parallel re-agglomeration algorithm by 30×.

97 MATHEMATICS AND COMPUTING↗

America's Next Great Ship: Space Launch System Core Stage Transitioning from Design to Manufacturing

The Space Launch System (SLS) Program is essential to achieving the Nation's and NASA's goal of human exploration and scientific investigation of the solar system. As a multi-element program with emphasis on safety, affordability, and sustainability, SLS is becoming America's next great ship of exploration. The SLS Core Stage includes avionics, main propulsion system, pressure vessels, thrust vector control, and structures. Boeing manufactures and assembles the SLS core stage at the Michoud Assembly Facility (MAF) in New Orleans, LA, a historical production center for Saturn V and Space Shuttle programs. As the transition from design to manufacturing progresses, the importance of a well-executed manufacturing, assembly, and operation (MA&O) plan is crucial to meeting performance objectives. Boeing employs classic techniques such as critical path analysis and facility requirements definition as well as innovative approaches such as Constraint Based Scheduling (CBS) and Cirtical Chain Project Management (CCPM) theory to provide a comprehensive suite of project management tools to manage the health of the baseline plan on both a macro (overall project) and micro level (factory areas). These tools coordinate data from multiple business systems and provide a robust network to support Material & Capacity Requirements Planning (MRP/CRP) and priorities. Coupled with these tools and a highly skilled workforce, Boeing is orchestrating the parallel buildup of five major sub assemblies throughout the factory. Boeing and NASA are transforming MAF to host state of the art processes, equipment and tooling, the most prominent of which is the Vertical Assembly Center (VAC), the largest weld tool in the world. In concert, a global supply chain is delivering a range of structural elements and component parts necessary to enable an on-time delivery of the integrated Core Stage. SLS is on plan to launch humanity into the next phase of space exploration.

Birkenstock, Benjamin↗

Utilizing Gaps and Key Performance Parameters to Inform NASA Environmental Control and Life Support and Human Health and Performance Capability Technology Decisions

Human spaceflight is a complex endeavor requiring a multitude of capabilities for transportation, crew health, scientific goals, and safe return to Earth. The difference between spaceflight proven capabilities and those needed for a particular mission is defined as a capability gap. Capability gaps are not technology specific. Each capability gap is approachable with a wide array of technologies that have unique benefits and challenges. Determining what a capability’s relevant and distinguishing key performance parameters (KPPs) are for a mission is critical. Mass, power, and volume are always constrained and important, but defining these in a way normalized by performance is challenging. Additionally, KPP definition for reliability, dormancy, and integration needs are very important and still evolving. This paper provides the approach of the Environmental Control and Life Support – Crew Health and Performance (ECLSS-CHP) System Capability Leadership Team (SCLT) to defining gaps and KPPs in support of the NASA’s Capabilities Integration Team data call objectives. The nine ECLSS-CHP capability areas are decomposed to capabilities, gaps, and KPPs. Rather than defining very detailed gaps, ECLSS-CHP defines high-level gaps to be technology agnostic. Within a gap, detailed KPPs are defined to both compare technologies and measure progress within a technology over time. Ideally, KPPs are clearly defined, widely communicated both internally and externally, and provide a common nomenclature to describe the state of the art and the degree of improvement required for exploration missions. KPPs help define when the gap is closed and the core mission objectives can be accomplished. Further technology improvements to enhance the capability, as measured by improved KPPs, must then be weighed against investments in open capability gaps that prevent NASA from achieving its exploration missions. It is uncommon that a technology maturation to improve all the relevant KPPs simultaneously but using KPPs is a critical technology investment decision making component. In addition to traditional technology selections, KPPs are informing how investments in ground testing prior to and in parallel with ISS technology demonstrations are required to improve reliability KPPs. The collection of all major technology activities within a capability area are captured on technology roadmaps to communicate how diverse program activities are coordinated to close gaps and infuse into exploration mission needs. A selection of ECLSS-CHP gaps and KPPs and their formulation, current state, and how they inform capability roadmap planning are discussed. The paper will contain a summary of the approximately 60 gaps. Gaps are classified as to their type (architecture, knowledge, technology, developmental, or engineering) depending on the magnitude of the gap. The paper will provide brief overviews of a few major technology challenges and the technologies being considered, but will reference detailed papers for a more thorough treatment of the challenges and state of the art. Data analysis of the gaps is in work and results are not currently available for this abstract. It is anticipated the paper will include examples of select KPPs with descriptions as to why these are the relevant measures. Additionally some KPPs will be graphically presented over time to show progress to date and when performance targets need to be achieved to support exploration missions. Graphical summaries of how gaps closures with near term mission elements support follow-on mission elements will be provided.

Life Support↗