Search NASA⌕ Search

SEARCH · Search NASA

Results for “Computer graphics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Speed Optimizations for Physics Ray Trace Algorithms

Ray tracing is a process used commonly in computer graphics and in physics to track light photons and particles, respectively. Much research was found on improving execution times for the computer graphics applications; however, in the short time frame of this literary review, almost no research was found on improving the execution times for the physics applications that were relevant to this problem. Two ray trace algorithms, a STL raytrace and a conebeam raytrace, were optimized using OpenMP and CUDA.

97 MATHEMATICS AND COMPUTING↗

Mechanical forces orchestrate the metabolism of the developing oilseed rape embryo

The initial free expansion of the embryo within a seed is at some point inhibited by its contact with the testa, resulting in its formation of folds and borders. Although less obvious, mechanical forces appear to trigger and accelerate seed maturation. However, the mechanistic basis for this effect remains unclear. Manipulation of the mechanical constraints affecting either the in vivo or in vitro growth of oilseed rape embryos was combined with analytical approaches, including magnetic resonance imaging and computer graphic reconstruction, immunolabelling, flow cytometry, transcriptomic, proteomic, lipidomic and metabolomic profiling. Our data implied that, in vivo, the imposition of mechanical restraints impeded the expansion of testa and endosperm, resulting in the embryo's deformation. An acceleration in embryonic development was implied by the cessation of cell proliferation and the stimulation of lipid and protein storage, characteristic of embryo maturation. The underlying molecular signature included elements of cell cycle control, reactive oxygen species metabolism and transcriptional reprogramming, along with allosteric control of glycolytic flux. Constricting the space allowed for the expansion of in vitro grown embryos induced a similar response. The conclusion is that the imposition of mechanical constraints over the growth of the developing oilseed rape embryo provides an important trigger for its maturation.

59 BASIC BIOLOGICAL SCIENCES↗

Metamodels for Rapid Analysis of Large Sets of Building Designs for Robotic Constructability: Technology Demonstration Using the NASA 3D Printed Mars Habitat Challenge

Disruptive robotic construction technologies such as additive deposition of cementitious materials like concrete (or "3D concrete printing") require the synchronous operation of multiple pieces of equipment in the production setup. In such an environment, it is crucial to simulate the robotic motions (for toolpath clashes) and the cementitious material behavior (for toolpath failures) to ensure fail-proof constructability of the envisioned building geometry. However, toolpath clash detection requires 4D simulations of the production setup, which are computationally graphics intensive, whereas toolpath failure detection requires actual 3D printing of test parts from the geometry to identify areas prone to failure while 3D printing, which is physically tedious. Both these processes, being computationally and physically intensive, have largely curtailed designers from simulating and exploring large sets of design options with varying geometries and toolpath configurations. To overcome this and allow designers to explore large sets of design possibilities, this paper proposes two novel computational metamodels capable of performing robotic toolpath clash detection and failure detection with significantly reduced times than the earlier approaches. The developed metamodels were used to rapidly simulate large sets of building design options for robotic constructability in the NASA 3D-Printed Mars Habitat Challenge.

clash detection↗

A MultiGPU Performance-Portable Solution for Array Programming Based on Kokkos

Today, multiGPU nodes are widely used in high-performance computing and data centers. However, current programming models do not provide simple, transparent, and portable support for automatically targeting multiple GPUs within a node on application areas of array programming. In this paper, we describe a new application programming interface based on the Kokkos programming model to enable array computation on multiple GPUs in a transparent and portable way across both NVIDIA and AMD GPUs. We implement different variations of this technique to accommodate the exchange of stencils (array boundaries) among different GPU memory spaces, and we provide autotuning to select the proper number of GPUs, depending on the computational cost of the operations to be computed on arrays, that is completely transparent to the programmer. We evaluate our multiGPU extension on Summit (#5 TOP500), with six NVIDIA V100 Volta GPUs per node, and Crusher that contains identical hardware/software as Frontier (#1 TOP500), with four AMD MI250X GPUs, each with 2 Graphics Compute Dies (GCDs)for a total of 8 GCDs per node. We also compare the performance of this solution against the use of MPI + Kokkos, which is the cur-rent de facto solution for multiple GPUs in Kokkos. Our evaluation shows that the new Kokkos solution provides good scalability for many GPUs and a faster and simpler solution (from a programming productivity perspective) than MPI + Kokkos.

Valero Lara, Pedro↗

Towards exascale for wind energy simulations

We examine large-eddy-simulation modeling approaches and computational performance of two open-source computational fluid dynamics codes for the simulation of atmospheric boundary layer flows that are of direct relevance to wind energy production. The first code, NekRS, is a high-order, unstructured-grid, spectral element code. The second code, AMR-Wind, is a second-order, block-structured, finite-volume code with adaptive mesh refinement capabilities. The objective of this study is to co-develop these codes in order to improve model fidelity and performance for each. These features will be critical for running ABL-based applications such as wind farm analysis on advanced computing architectures. To this end, we investigate the performance of NekRS and AMR-Wind on the Oak Ridge Leadership Facility supercomputers Summit, using 4 to 800 nodes (24 to 4,800 NVIDIA V100 GPUs), and Crusher, the testbed for the Frontier exascale system, using 18 to 384 Graphics Compute Dies on AMD MI250X GPUs. We compare strong- and weak-scaling capabilities, linear solver performance, and time to solution. We also identify leading inhibitors to parallel scaling.

17 WIND ENERGY↗

Nonlinear elasticity with the Shifted Boundary Method

Here, we propose a new unfitted/immersed computational framework for nonlinear solid mechanics, which bypasses the complexities associated with the generation of CAD representations and subsequent body-fitted meshing. This approach allows to speed up the cycle of design and analysis in complex geometry and requires relatively simple computer graphics representations of the surface geometries to be simulated, such as the Standard Tessellation Language (STL format). Complex data structures and integration on cut elements are avoided by means of an approximate boundary representation and a modification (shifting) of the boundary conditions to maintain optimal accuracy. An extensive set of computational experiments in two and three dimensions is included.

97 MATHEMATICS AND COMPUTING↗

A conservative Galerkin solver for the quasilinear diffusion model in magnetized plasmas

We propose a conservative Galerkin scheme for the quasilinear model in three-dimensional momentum space and three-dimensional spectral space, with cylindrical symmetry. We construct an unconditionally conservative weak form and use a discretization that preserves conservation properties independent of the wave emission probability. The discrete operators, combined with a consistent quadrature rule, preserve all the conservation laws rigorously. The proposed scheme is quite general: it works for both relativistic and non-relativistic systems, for both magnetized and unmagnetized plasmas, and even for problems with time-dependent dispersion relations. We represent the particle distribution by continuous basis functions and use discontinuous basis functions for the wave spectral energy density, which enables the application of a positivity-preserving technique. We adopt the marching simplex algorithm, designed initially for computer graphics, for numerical integration on the resonance manifold. Furthermore, the numerical examples with a bump-on-tail initial configuration show how the unstable waves produce strong momentum space diffusion.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

DGaaS: GPU as a Service on Distributed Computing System

In the rapidly evolving landscape of scientific computing, Graphics Processing Units (GPUs) have become indispensable for their unparalleled ability to handle parallel tasks in complex calculations, simulations, and data analysis. Their utility is further magnified in machine learning and AI applications, where they significantly accelerate model training and predictive analytics. Within this context, the Triton Inference Server emerges as a pivotal open-source tool, specializing in AI inferencing and optimizing GPU utilization across various platforms and frameworks. This paper presents an in-depth study on distributed High Throughput Computing (HTC), specifically focusing on the HTCondor framework and its resource provisioning tools, GlideinWMS and HEPCloud. These systems enable large-scale scientific experiments like CMS and DUNE to efficiently access and utilize vast computational resources. The paper explores the core architectural components of GlideinWMS, including jobs, user pools, and worker nodes, and discusses their integration with GPUs and the Triton server. The primary aim of this research is to develop a solution that optimizes GPU utilization by leveraging Glideins and containers. This approach allows computational jobs, particularly those involving AI models, to use GPUs only when essential, thereby facilitating efficient sharing of limited GPU resources. To validate this architecture, the study conducted three key tests involving custom scripts, container-based servers, and Triton server deployments. However, the study faces challenges, notably in locating the Triton server and ensuring secure remote access. To address these issues, future work will focus on developing a proxy mechanism and enhancing security protocols. In conclusion, this study offers a comprehensive roadmap for effective and efficient GPU utilization in distributed High Throughput Computing. It aims to contribute significantly to the scientific community by solving pressing problems and implementing robust solutions in collaboration with the GlideinWMS and HEPCloud teams. The research sets the stage for a more efficient, scalable, and cost-effective paradigm in scientific computing.

97 MATHEMATICS AND COMPUTING↗

Enabling Scientific Applications with Performance-Portability and High-Productivity for Multi-GPU Programming with JACC.Multi

This work bridges the gap between multi-GPU computing and high-productivity, performance-portable programming solutions. Our goal is to enhance scientific applications with a productive and portable solution—program once, deploy everywhere—for multi-GPU programming with no cost to programmability. To accomplish this, we implemented JACC.Multi, which is part of the Julia for ACCelerators (JACC) performance-portable framework. JACC. Multi is the only high-level, portable metaprogramming solution that targets multi-GPU environments and is integrated in a readily accessible programming language (e.g., Julia language). With transparent GPU-to-GPU communication, JACC. Multi is optimized for scientific application workloads and is portable for NVIDIA and AMD accelerators. For the evaluation, we use two modern multi-GPU systems: Hudson, which features two NVIDIA H100 Hopper GPUs per node, and Frontier, which features four AMD MI250X GPUs per node, each with two Graphics Compute Dies (GCDs) for a total of eight GCDs per node. Additionally, as part of the evaluation, we use JACC (one GPU), MPI+JACC, and JACC. Multi codes that implement well-known and widely used scientific algorithms/kernels such as the conjugate gradient algorithm and an explicit forward Euler solver that requires GPU-to-GPU communication. Overall, JACC. Multi codes achieve better performance than MPI+JACC codes and significant speedups over JACC (one GPU), with up to 1.9× on Hudson and 6× on Frontier.

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

Generalized fiducial inference on differentiable manifolds

We introduce a novel approach to inference on parameters that take values in a Riemannian manifold embedded in a Euclidean space. Parameter spaces of this form are ubiquitous across many fields, including chemistry, physics, computer graphics, and geology. Here, this new approach uses generalized fiducial inference (GFI) to obtain a posterior-like distribution on the manifold, without needing to know local parameterizations that map to the constrained space from an unconstrained Euclidean space. Using mathematical tools from Riemannian geometry, we construct a constrained generalized fiducial distribution (CGFD). A Bernstein-von Mises-type result for the CGFD, which provides intuition for how the desirable asymptotic qualities of the unconstrained generalized fiducial distribution are inherited by the CGFD, is provided. To illustrate the practical use of the CGFD, we provide a proof-of-concept example in the context of a linear logspline density estimation problem, and demonstrate that CGFD-based confidence sets exhibit desirable coverage properties via simulation. As an application, we fit a CGFD to COVID-19 case count data from North Carolina, USA.

97 MATHEMATICS AND COMPUTING↗

Updimensioning strategy derived from synthetic equiaxed grain structures for approximating 3D grain size distributions from 2D visualizations with 1D parameters

We generated synthetic equiaxed grain structures using computer graphics software to explore the relationship between various grain size determination methods and true three-dimensional (3D) grain diameters. Mirroring grain measurement techniques, the synthetic 3D grain structures are imaged as 2D micrographs which are measured to yield 1D grain size parameters. Synthetic grain structures provide data at a mass scale and permit exploration of both polished and fractured surface micrographs, revealing one-to-one correspondence between exposed 2D grain cross-sections and individual 3D grains. Analysis of this correspondence yielded a procedure to approximate 3D equiaxed grain size and volume distributions based on the mode of the 2D fractograph grain size distribution. The 3D approximation procedure is shown to be less susceptible to different imaging conditions that affect small, undiscernible grains compared to the standard planimetric and linear intercept methods, which by design also tend to underestimate the 3D grain diameter. The procedure requires larger sample sizes to lower variance and a deeper analysis which could become more practical with machine learning (ML) models for grain boundary segmentation, which synthetic grain structures can help train. This work lays the foundation for analyzing other grain distributions such as columnar and composite grains in similar depth.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Tusqh

SAND2025-00675O Tusqh is a software tool that generates cubical meshes in 2D and 3D and computes the homology of these meshes using persistent homology. It includes a grid cell in the output if its volume-fraction is above a selectable threshold, estimated by sampling points within the cell. Tusqh incorporates anti-aliasing algorithms to mitigate grid orientation and scale effects. It is designed for creating finite element meshes for simulations and can be used in various applications such as heat diffusion, mechanical simulations, and computer graphics rendering. The software outputs meshes in an open format compatible with downstream software. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

SciDAC↗

Interactive Quantum Chemistry Enabled by Machine Learning, Graphical Processing Units, and Cloud Computing

Modern quantum chemistry algorithms are increasingly able to accurately predict molecular properties that are useful for chemists in research and education. Despite this progress, performing such calculations is currently unattainable to the wider chemistry community, as they often require domain expertise, computer programming skills, and powerful computer hardware. In this review, we outline methods to eliminate these barriers using cutting-edge technologies. We discuss the ingredients needed to create accessible platforms that can compute quantum chemistry properties in real time, including graphical processing units–accelerated quantum chemistry in the cloud, artificial intelligence–driven natural molecule input methods, and extended reality visualization. We end by highlighting a series of exciting applications that assemble these components to create uniquely interactive platforms for computing and visualizing spectra, 3D structures, molecular orbitals, and many other chemical properties.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A 5G Enabled Adaptive Computing Workflow for Greener Power Grid

5G wireless technology can deliver higher data speeds, ultra low latency, more reliability, massive network capacity, increased availability, and a more uniform user experience to users. It brings additional power to help address the challenges brought by renewable integration and decarbonization. In this paper, a 5G enabled adaptive computing workflow tool has been presented that consists of various computing resources, such as 5G equipment, edge computing, cluster, Graphics processing unit (GPU) and cloud computing, with two examples showing technical feasibility for edge-grid-cloud interaction for real-time monitoring, security assessment, and forecasting. Benefiting from the high data transmission speed and massive connection capability of 5G, the workflow shows its potential to seamlessly integrate various applications at distributed and/or centralized locations to build more complex and powerful functions, with better flexibility.

5G technology, computational workflow, edge comput↗

Automated Hybrid Variance Reduction on Advanced Architectures in the Shift Monte Carlo Code

Monte Carlo transport methods are the most accurate schemes for solving problems with complex energy and spatial features, but they come with a high computational cost. Although hybrid methods have enabled the use of Monte Carlo transport for a large class of problems, they still require significant computing resources. Modern multicore CPUs with large numbers of compute cores and graphical processing units (GPUs) provide opportunities to optimize the memory and run-time costs of hybrid Monte Carlo methods. This paper documents the development and analysis of three Monte Carlo transport algorithms that support hybrid transport using the consistent adjoint-driven importance sampling (CADIS) and forward-weighted CADIS methods in the Shift Monte Carlo code: history-based transport using static and dynamic threading on multicore CPUs and event-based transport enabling weight window tracking on GPUs. The results are shown for two challenging hybrid problems on the Frontier supercomputer at the Oak Ridge Leadership Computing Facility. The results show that all three methods yield good performance and enable solutions of difficult fixed-source transport problems in less than 2 min on 20 nodes of Frontier. Dynamic threading was observed to give up to 20% better scaling behavior than static threading. Moreover, the AMD Instinct 250X GPU was found to give 9 to 11 times greater throughput per graphics compute die than the best CPU performance. In conclusion, additional opportunities for optimization of hybrid transport on GPUs are discussed.

Denovo↗

DG2DAG: Learning Directed Acyclic Graphs from Functional Priors

Physics-based systems-of-systems models are computationally expensive. Reduced graphical models can decrease computational complexity, but may not proffer an end-to-end model from upstream inputs to downstream outputs. We consequently are interested in reducing models on directed graphs to models on a directed acyclic subgraph such that preserves accurate reconstruction of nodes. The consequence is a model with a topological ordering, providing a one-way flow of computation, and a causal interpr

Voronin, Alexey [Sandia National Laboratories (SNL↗

ExaWind: Then and Now

The scientific goal of the ExaWind project is to advance our fundamental understanding of the flow physics governing whole wind plant performance, including wake formation, complex terrain impacts, and turbine-turbine-interaction effects. The primary application codes in the ExaWind environment are Nalu-Wind, an unstructured-grid computational fluid dynamics (CFD) code, AMR-Wind, a structured-grid CFD code, and OpenFAST, a whole-turbine simulation code. In this poster we present the current status of the ExaWind software stack in the context of the modeling and simulation capabilities when the project started in 2016.

computational fluid dynamics↗

Asynchronous GPU-based DEM solver embedded in commercial CFD software with polyhedral mesh support

A novel graphical processing unit-based discrete element method solver is introduced to improve stability, performance, and provide seamless integration into commercial or open-source computational fluid dynamics software. A key innovation is eliminating a need for network communication between solvers, which was previously required for cross-platform coupling. This is accomplished by a direct coupling method that employs dynamic-linked libraries. Furthermore, the solver optimizes memory usage by streamlining the particle-cell search algorithm by eliminating the cells' searching grid. This ensures the solver is compatible with a wide range of mesh types, providing high geometric flexibility. The approach simplifies the simulation process by directly incorporating computational fluid dynamics mesh information into the discrete element method solver. The performance analysis indicates about sixteen times boost in computational speed compared to benchmark central processing unit-based solvers. Finally, the solver's compatibility with polyhedral meshes, a vital advantage for complex geometries, is tested against a referenced study regarding the simulation of an immersed-tube fluidized bed.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗