Search NASA⌕ Search

SEARCH · Search NASA

Results for “computational performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

HPC and Cloud Convergence Beyond Technical Boundaries: Strategies for Economic Sustainability, Standardization, and Data Accessibility

At the IEEE/ACM International Conference for High-Performance Computing, Networking, Storage, and Analysis (SC23), held in Denver, experts discussed the convergence of high-performance computing and cloud computing. Experts explored how this integration could address current scientific computing limitations, enhance computational capabilities, and foster global collaboration while focusing on economic, security, technical, and community challenges and opportunities.

97 MATHEMATICS AND COMPUTING↗

Computer controlled performance mapping of thermionic converters: effect of collector, guard-ring potential imbalances on the observed collector current-density, voltage characteristics and limited range performance map of an etched-rhenium, niobium planar converter

The effect of collector, guard-ring potential imbalance on the observed collector-current-density J, collector-to-emitter voltage V characteristic was evaluated in a planar, fixed-space, guard-ringed thermionic converter. The J,V characteristic was swept in a period of 15 msec by a variable load. A computerized data acquisition system recorded test parameters. The results indicate minimal distortion of the J,V curve in the power output quadrant for the nominal guard-ring circuit configuration. Considerable distortion, along with a lowering of the ignited-mode striking voltage, was observed for the configuration with the emitter shorted to the guard ring. A limited-range performance map of an etched-rhenium, niobium, planar converter was obtained by using an improved computer program for the data acquisition system.

Manista, E. J.↗

Job Superscheduler Architecture and Performance in Computational Grid Environments

Computational grids hold great promise in utilizing geographically separated heterogeneous resources to solve large-scale complex scientific problems. However, a number of major technical hurdles, including distributed resource management and effective job scheduling, stand in the way of realizing these gains. In this paper, we propose a novel grid superscheduler architecture and three distributed job migration algorithms. We also model the critical interaction between the superscheduler and autonomous local schedulers. Extensive performance comparisons with ideal, central, and local schemes using real workloads from leading computational centers are conducted in a simulation environment. Additionally, synthetic workloads are used to perform a detailed sensitivity analysis of our superscheduler. Several key metrics demonstrate that substantial performance gains can be achieved via smart superscheduling in distributed computational grids.

Shan, Hongzhang↗

Workload Characterization of CFD Applications Using Partial Differential Equation Solvers

Workload characterization is used for modeling and evaluating of computing systems at different levels of detail. We present workload characterization for a class of Computational Fluid Dynamics (CFD) applications that solve Partial Differential Equations (PDEs). This workload characterization focuses on three high performance computing platforms: SGI Origin2000, EBM SP-2, a cluster of Intel Pentium Pro bases PCs. We execute extensive measurement-based experiments on these platforms to gather statistics of system resource usage, which results in workload characterization. Our workload characterization approach yields a coarse-grain resource utilization behavior that is being applied for performance modeling and evaluation of distributed high performance metacomputing systems. In addition, this study enhances our understanding of interactions between PDE solver workloads and high performance computing platforms and is useful for tuning these applications.

Waheed, Abdul↗

Requirements for Next Generation Comprehensive Analysis of Rotorcraft

The unique demands of rotorcraft aeromechanics analysis have led to the development of software tools that are described as comprehensive analyses. The next generation of rotorcraft comprehensive analyses will be driven and enabled by the tremendous capabilities of high performance computing, particularly modular and scaleable software executed on multiple cores. Development of a comprehensive analysis based on high performance computing both demands and permits a new analysis architecture. This paper describes a vision of the requirements for this next generation of comprehensive analyses of rotorcraft. The requirements are described and substantiated for what must be included and justification provided for what should be excluded. With this guide, a path to the next generation code can be found.

Johnson, Wayne↗

Rasterization with Data-Parallel Primitives

Parallel rasterization can suffer from race conditions during fragment generation, which is traditionally addressed by using specialized hardware accessible via vendor graphics APIs. Unfortunately, graphics APIs are increasingly problematic on high-performance computers, either because they are not provided or because of concerns about dependencies with in situ visualization. In response, we present a hardware-agnostic rasterization algorithm that handles race conditions using only data-parallel primitives (DPPs), enabling efficient rendering on HPC systems without graphics API dependencies and aligning with recent efforts to deliver visualization software with DPPs. Our evaluation consists of three phases: (1) evaluating portability across different CPU and GPU architectures, (2) evaluating competitiveness with a community standard, and (3) evaluating performance across varying workloads and available parallelism. The supporting experiments run on both AMD and NVIDIA GPUs, considering data sets as large as 460 million triangles and 160 million pixels. While performance generally falls short of graphics API baselines, it achieves interactive frame rates on most workloads. As a result, we conclude our approach is a viable solution for rasterization on high-performance computers since our approach is portably performant across different architectures without the need for specialized vendor support.

Buckley, Makani [University of Oregon] (ORCID:0009↗

Development and Validation of a Fast, Accurate and Cost-Effective Aeroservoelastic Method on Advanced Parallel Computing Systems

Progress to date towards the development and validation of a fast, accurate and cost-effective aeroelastic method for advanced parallel computing platforms such as the IBM SP2 and the SGI Origin 2000 is presented in this paper. The ENSAERO code, developed at the NASA-Ames Research Center has been selected for this effort. The code allows for the computation of aeroelastic responses by simultaneously integrating the Euler or Navier-Stokes equations and the modal structural equations of motion. To assess the computational performance and accuracy of the ENSAERO code, this paper reports the results of the Navier-Stokes simulations of the transonic flow over a flexible aeroelastic wing body configuration. In addition, a forced harmonic oscillation analysis in the frequency domain and an analysis in the time domain are done on a wing undergoing a rigid pitch and plunge motion. Finally, to demonstrate the ENSAERO flutter-analysis capability, aeroelastic Euler and Navier-Stokes computations on an L-1011 wind tunnel model including pylon, nacelle and empennage are underway. All computational solutions are compared with experimental data to assess the level of accuracy of ENSAERO. As the computations described above are performed, a meticulous log of computational performance in terms of wall clock time, execution speed, memory and disk storage is kept. Code scalability is also demonstrated by studying the impact of varying the number of processors on computational performance on the IBM SP2 and the Origin 2000 systems.

Goodwin, Sabine A.↗

Advanced high area ratio nozzles

The objective is to develop computational techniques for the design of high-area-ratio nozzles and to validate these models by comparison with experiments and computations using other codes. Performance computations were added to the PARC2D code and the performance of the space shuttle main engine (SSME) nozzle was computed for inviscid, laminar and turbulent flow assuming a perfect gas with gamma = 1.2. The PARC2D code was modified in a non-CASP (Center for Advanced Space Propulsion) project to compute equilibrium flow about hypersonic blunt bodies. Progress has been made toward modifying this code to compute equilibrium H2/O2 flow through the SSME and related nozzles.

Raiszadeh, Farhad↗

LATTE: open-source, high-performance traveltime computation, tomography and source location in acoustic and elastic media

Traveltime-based tomography and source location are fundamental approaches for imaging subsurface structures and understanding the spatiotemporal distribution of seismicity from local to global scales. We present an open-source, high-performance framework integrating eikonal equation solvers and adjoint-state theory for traveltime computation, velocity tomography, source location and joint tomography-location in 2-D/3-D acoustic and elastic media. We introduce novel regularization schemes based on total generalized p-variation, structural similarity and multitask machine learning to enhance the fidelity and interpretability of inverted models and source locations. Key features of our implementation also include the ability to leverage both absolute-difference and double-difference traveltime misfits for high-fidelity velocity tomography and source parameter estimation; support for traveltime computation and inversion in diverse 2-D/3-D scenarios with arbitrary source and receiver distributions; and a perturbation-based optimal step-size estimation method to reduce computational costs. In addition, our implementation employs shared-memory and distributed-memory parallelization to provide an efficient solution for traveltime computation, tomography, and source location. In conclusion, we validate the efficacy and accuracy of our approach through multiple synthetic data examples.

58 GEOSCIENCES↗

Accuracy Enhancement of Nuclear Power Plant Simulators Utilizing High Accuracy Simulation Predictions

More recently, reactor core simulators for core designs associated with commercial nuclear power plants that utilize what is believed to be higher fidelity models have been developed. Features such as neutronics models that utilize transport equation solvers with fine spatial meshes and many energy-groups, thermal-hydraulic models that utilize sub-channel solvers with fine spatial mesh and capable of treating a wide range of fluid conditions, and fuel-coolant chemistry interaction models capable of treating CRUD deposition are to be found in these higher fidelity core simulators. These reactor core simulators require access to higher performance computers, characterized by many processors, cores and large memory. So associated with utilization of these simulators is access to high performance computers and ability to accommodate in one’s workflow longer execution times. By contrast, currently used core simulators by the nuclear industry can execute on engineering workstations and have execution times of seconds to minutes. The desirability for having short execution times is not only desired for support of time critical tasks but supports the mental process of decision making by engineers. The goal of the work reported upon here has the objective of retaining the fidelity of higher fidelity models while retaining the ability to utilize engineering workstations. Beyond the core simulator goal, additional goals of this work include incorporating the just described core simulator capability into a Nuclear Steam Supply System (NSSS) simulator, and to incorporate the resulting capability into an environment supportive of design and operational decision making associated with nuclear power stations. The model selected for the core neutronics model is the NESTLE code, for the core thermal-hydraulic model is the CTF code utilizing coarse mesh, and for the NSSS model is the RELAP5-3D code. WSC’s proprietary 3KEYMASTERTM platform is being used to provide software coupling, user interface, visualization, and reporting. The NESTLE core neutronics simulator was first integrated with the CTF core thermal-hydraulic simulator using CTF developed communication commands which are also used for CTF to communicate with RELAP5-3D under WSC’s proprietary 3KEYMASTERTM platform. To assure NESTLE prediction consistency with higher fidelity core neutronic simulators, buffer codes have been created to automatically generate from output files written by the VERA core simulator the NESTLE nodal neutronic parameter’ library, geometry, and pin-power reconstruction input files, thereby avoiding a number of challenges associated with utilizing lattice physics codes and providing consistency with VERA predictions. To treat absorber rod effects a multi-set library is utilized, where a set refers to a specific absorber rod fully inserted pattern. A coarse spatial mesh CTF model was developed with features added that support using CTF as envisioned in the engineering quality simulator. A hybrid meshing approach was implemented to allow for automated construction of models with mixed levels of refinement. Specifically, a core model could resolve some assemblies at a nodal level (4 subchannels per assembly) and others at a pin-resolution (one subchannel per coolant subchannel in the assembly). The intention is that this will allow for better resolution of limiting conditions such as DNBR and PCT, which are based on local rod and subchannel conditions. Further development was done of features that enhance the capabilities for the envisioned engineering quality simulator that has been developed, but now for RELAP-3D. The RELAP5-3D code development includes ability to model more than 999 components and the addition of the cross-channels turbulence mixing model and the void drift model that are implemented in CTF, aiming to achieve closer prediction agreement of the two codes for transient simulations, specifically, more accurate matches of the overall mass, momentum, and energy exchanges of both the liquid and gas phases between the neighboring core assemblies. Graphics were also developed for the Instructor Station for this project under WSC’s proprietary 3KEYMASTERTM platform to facilitate design and operational decision making.

42 ENGINEERING↗

Design analysis and computer-aided performance evaluation of shuttle orbiter electrical power system. Volume 1: Summary

Studies were conducted to develop appropriate space shuttle electrical power distribution and control (EPDC) subsystem simulation models and to apply the computer simulations to systems analysis of the EPDC. A previously developed software program (SYSTID) was adapted for this purpose. The following objectives were attained: (1) significant enhancement of the SYSTID time domain simulation software, (2) generation of functionally useful shuttle EPDC element models, and (3) illustrative simulation results in the analysis of EPDC performance, under the conditions of fault, current pulse injection due to lightning, and circuit protection sizing and reaction times.

Source record↗

High performance flight computer developed for deep space applications

The development of an advanced space flight computer for real time embedded deep space applications which embodies the lessons learned on Galileo and modern computer technology is described. The requirements are listed and the design implementation that meets those requirements is described. The development of SPACE-16 (Spaceborne Advanced Computing Engine) (where 16 designates the databus width) was initiated to support the MM2 (Marine Mark 2) project. The computer is based on a radiation hardened emulation of a modern 32 bit microprocessor and its family of support devices including a high performance floating point accelerator. Additional custom devices which include a coprocessor to improve input/output capabilities, a memory interface chip, and an additional support chip that provide management of all fault tolerant features, are described. Detailed supporting analyses and rationale which justifies specific design and architectural decisions are provided. The six chip types were designed and fabricated. Testing and evaluation of a brass/board was initiated.

Bunker, Robert L.↗

A fast and robust computational modeling approach for density and shape predictions in powder metallurgy hot isostatic pressing

Powder metallurgy hot isostatic pressing (PM-HIP) is an advanced manufacturing process that produces near-net-shape parts with high material utilization and uniform microstructures. PM-HIP is frequently used for producing small-scale parts with complicated geometries and is potentially economical for producing large-scale parts. However, excessive post-HIP shape distortions can reduce its effectiveness and economic advantage, especially for larger parts. A PM-HIP computational model can predict and help mitigate these distortions. However, due to complex deformation mechanisms and thermo-mechanical coupling present in PM-HIP processes, these non-linear computational models sometimes become numerically unstable. The numerical instabilities in these models can lead to very slow convergence or no convergence at all, which often translates to slow and unreliable models. These limitations are more pronounced in large models with complicated geometries. Hence, in this work, an alternative modeling approach is presented that improves numerical stability and computational performance. The presented approach achieves these improvements through approximating the fully coupled thermo-mechanical PM-HIP model as a decoupled model and adding inertial damping to the model’s mechanical part. In conclusion, a comparison with the fully coupled model indicated a slight dip in prediction accuracy (<5% error) but significant improvements in numerical stability (>20 times larger time step size) and computational performance (5-10 times speed-up with less computational resource usage) when using the presented approach.

Hot isostatic pressing↗

ChemComp: A Compilation Framework for Computing with Chemical Reaction Networks

The acceleration of scientific computation, data analytics, and artificial intelligence is driving a surge in computational requirements. Yet, state-of-the-art high-performance computing systems are approaching physical limitations that impede further significant improvements in energy efficiency. As we move towards post-exascale computing systems, innovative approaches are necessary to overcome this barrier in power consumption. Novel analog and hybrid digital-analog architectures hold promise for enhancing energy efficiency by several orders of magnitude. Biochemical computation stands out among the various solutions being explored due to its potential to enable new classes of devices with immense computational capabilities. These devices can capitalize on the inherent efficacy of biological cells in solving optimization problems and are scalable through increasing reaction system size or vessel capacity, potentially satisfying scientific computing's high-performance requirements. Nonetheless, several theoretical and practical limitations persist, including problem formulation and mapping to chemical reaction networks (CRNs) and implementation of actual CRN devices. In this paper, we propose a framework for biochemical computation using systems chemistry. We present the initial components of our approach: an abstract chemical reaction dialect implemented as a multi-level intermediate representation (MLIR) compiler extension and a pathway to represent mathematical problems with CRNs. To showcase the potential of this approach, we emulate a simplified chemical reservoir device. This work lays the groundwork for leveraging chemistry's computing potential in creating energy-efficient, high-performance computing systems tailored to contemporary computational needs.

artificial intelligence↗

An HPC benchmark survey and taxonomy for characterization

The field of High-Performance Computing (HPC) is defined by providing computing devices with highest performance for a variety of demanding scientific users. The tight co-design relationship between HPC providers and users propels the field forward, paired with technological improvements, achieving continuously higher performance and resource utilization. A key device for system architects, architecture researchers, and scientific users are benchmarks, allowing for well-defined assessment of hardware, software, and algorithms. Many benchmarks exist in the community, from individual niche benchmarks testing specific features, to large-scale benchmark suites for whole procurements. We survey the available HPC benchmarks, summarizing them in table form with key details and concise categorization, also through an interactive website. For categorization, we present a benchmark taxonomy for well-defined characterization of benchmarks.

Benchmarking↗

The Influence of Computer Architecture on Performance and Scaling for Hypersonic Flow Simulations

It is critical to understand how hypersonic simulation tools perform on a range of computational platforms. This information will aid in the acquisition of appropriate hardware and the potential refactoring of hypersonic codes to run on different systems. In this paper, we consider two representative high-speed reacting flow cases: a model Mach 8 hypersonic waverider glide vehicle and a model hydrocarbon-fueled hypersonic ramjet propulsion system. In both scenarios, the flow fields are in chemical non-equilibrium and are modeled by the multi-species reacting Navier-Stokes equations. For these simulations we use several hypersonic simulation tools, including US3D, Kestrel, FUN3D, and the JENRER© flow solver. We explore several high performance computing systems containing IntelR© XeonR© Platinum processors, AMD EPYCTM7702 processors, and NVIDIAR© Tesla V100 devices. We compare performance and strong scaling between the different systems.

CPU↗

NLR HPC Eagle Jobs Data and Additional Energy Metrics

Overview: Anonymized job-level records from the Eagle high-performance computing (HPC) system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, resource utilization, CPU/GPU energy consumption, and efficiency metrics. Sensitive fields (user, account, job name) are replaced with cryptographic hashes. System & Timeframe: Eagle was a 2,000-node, 8-petaflop system operated at NLR from 2019–2024. Data covers the full operational lifetime of the system. Slurm data was processed nightly; timestamps are in Mountain Time. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.eagle.job-anon.zip — Core anonymized job records (Hive-partitioned Parquet) esif.hpc.eagle.job-anon-energy-metrics.zip — Same records with additional iLO and Ganglia energy metrics datacard.md — Full dataset documentation ~13.8 million rows, 62 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct through a pipeline: Eagle Jobs API → Redpanda → StreamSets → HPCMON API → PostgreSQL. Node-level power from iLO (HP Integrated Lights-Out); GPU power from Ganglia monitoring, joined to jobs via node lists and time ranges. Preprocessing: Anonymization of name, user, and account fields via cryptographic hashing Derived columns: queue_wait, cpu_eff, max_mem_eff Simplified job state mapping (e.g., "CANCELLED BY 12345" → "CANCELLED") QoS accounting rules (buy-in, standby, or Slurm QoS value) CPU energy estimated from TDP (200W, Intel Xeon Gold 6154, 18 cores) Timezone-aware columns (_tz) sourced from LEX accounting database to correctly handle DST transitions Key Variables: Scheduling: job_id, partition, state_simple, submit_time_tz, start_time_tz, end_time_tz, queue_waitResources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, node_energy_total_watt_hours (iLO), gpu0/1_energy_total_watt_hours (Ganglia) Partitions: bigmem, bigmem-8600, bigscratch, csc, dav, ddn, debug, gpu, haswell, long, mono, short, standard Job States: CANCELLED, COMPLETED, FAILED, NODE_FAIL, OUT_OF_MEMORY, PENDING, RUNNING, TIMEOUT QoS Levels: Unknown, normal, buy-in, debug, penalty, high, standby Important Notes: Non-_tz timestamp columns may be off by one hour across DST boundaries; use _tz columns for time difference calculations Energy fields are null for jobs without monitoring coverage Job step records and raw Slurm JSONB fields are excluded from this extract Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING↗

NASA Center for Climate Simulation (NCCS) Advanced Technology AT5 Virtualized Infiniband Report

The NCCS is part of the Computational and Information Sciences and Technology Office (CISTO) of Goddard Space Flight Center's (GSFC) Sciences and Exploration Directorate. The NCCS's mission is to enable scientists to increase their understanding of the Earth, the solar system, and the universe by supplying state-of-the-art high performance computing (HPC) solutions. To accomplish this mission, the NCCS (https://www.nccs.nasa.gov) provides high performance compute engines, mass storage, and network solutions to meet the specialized needs of the Earth and space science user communities

cloud↗