Search NASA⌕ Search

SEARCH · Search NASA

Results for “distributed computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39

A Shell/3D Modeling Technique for the Analyses of Delaminated Composite Laminates

A shell/3D modeling technique was developed for which a local three-dimensional solid finite element model is used only in the immediate vicinity of the delamination front. The goal was to combine the accuracy of the full three-dimensional solution with the computational efficiency of a plate or shell finite element model. Multi-point constraints provided a kinematically compatible interface between the local three-dimensional model and the global structural model which has been meshed with plate or shell finite elements. Double Cantilever Beam (DCB), End Notched Flexure (ENF), and Single Leg Bending (SLB) specimens were modeled using the shell/3D technique to study the feasibility for pure mode I (DCB), mode II (ENF) and mixed mode I/II (SLB) cases. Mixed mode strain energy release rate distributions were computed across the width of the specimens using the virtual crack closure technique. Specimens with a unidirectional layup and with a multidirectional layup where the delamination is located between two non-zero degree plies were simulated. For a local three-dimensional model, extending to a minimum of about three specimen thicknesses on either side of the delamination front, the results were in good agreement with mixed mode strain energy release rates obtained from computations where the entire specimen had been modeled with solid elements. For large built-up composite structures modeled with plate elements, the shell/3D modeling technique offers a great potential for reducing the model size, since only a relatively small section in the vicinity of the delamination front needs to be modeled with solid elements.

Krueger, Ronald↗

Portable Software Environment for Ultrahigh-Resolution ELM Development on GPUs

This paper presents our endeavors in developing the large-scale, ultra-high-resolution E3SM Land Model (uELM), specifically designed for exascale computers furnished with accelerators such as Nvidia GPUs. The uELM is a sophisticated code that substantially relies on High-Performance Computing (HPC) environments, necessitating particular machine and software configurations. To facilitate community-based uELM developments employing GPUs, we have created a portable, standalone software environment preconfigured with uELM input datasets, simulation cases, and source code. This environment, utilizing Docker, encompasses all essential code, libraries, and system software for uELM development on GPUs. It also features a functional unit test framework and an offline model testbed for comprehensive numerical experiments. From a technical perspective, the paper discusses GPU-ready container generations, uELM code management, and input data distribution across computational platforms. Lastly, the paper demonstrates the use of environment for functional unit testing, end-to-end simulation on CPUs and GPUs, and collaborative code development.

E3SM Land Model↗

Optical analysis of parabolic dish concentrators for solar dynamic power systems in space

An optical analysis of a parabolic solar collection system operating in Earth orbit was performed using ray tracing techniques. The analysis included the effects of: (1) solar limb darkening, (2) parametric variation of mirror surface error, (3) parametric variation of mirror rim angle, and (4) parametric variation of alignment and pointing error. This ray tracing technique used numerical integration to combine the effects of rays emanating from different parts of the sun at different intensities with the effects of normally distributed mirror-surface errors to compute the angular intensity distribution of rays leaving the mirror surface. A second numerical integration was then performed over the surface of the parabolic mirror to compute the radial distribution of brightness at the mirror focus. Major results of the analysis included: (1) solar energy can be collected at high temperatures with high efficiency, (2) higher absorber temperatures can be achieved at lower efficiencies, or higher efficiencies can be achieved at lower temperatures, and (3) collection efficiency is near its maximum level across a broad plateau of rim angles from 40 deg to 70 deg.

Jefferies, K. S.↗

Compton scattering of microwave background radiation by gas in galaxy clusters

Based on data on the X-ray spectrum of the Coma cluster, interpreted as thermal bremsstrahlung, the expected brightness depletion from Compton scattering of the microwave background in the direction of the cluster is computed. The calculated depletion is about one-third that recently observed by Gull and Northover, and the discrepancy is discussed. In comparing the observed microwave depletion in the direction of other clusters which are X-ray sources it is found that there is no correlation with the cluster X-ray luminosity. Consequently, the microwave depletion observations cannot yet be taken as good evidence for a thermal bremsstrahlung origin for the X-ray emission. The perturbation from Compton scattering of photons on the high-frequency (Wien) tail of the blackbody distribution is computed and found to be much larger than predicted in previous calculations. In the Wien tail the effect is a relative increase in the blackbody intensity that is appreciably greater in magnitude than the depletion in the Rayleigh-Jeans domain.

Gould, R. J.↗

Simulation of radar reflectivity and surface measurements of rainfall

Raindrop size distributions (RSDs) are often estimated using surface raindrop sampling devices (e.g., disdrometers) or optical array (2D-PMS) probes. A number of authors have used these measured distributions to compute certain higher-order RSD moments that correspond to radar reflectivity, attenuation, optical extinction, etc. Scatter plots of these RSD moments versus disdrometer-measured rainrates are then used to deduce physical relationships between radar reflectivity, attenuation, etc., which are measured by independent instruments (e.g., radar), and rainrate. In this paper RSDs of the gamma form as well as radar reflectivity (via time series simulation) are simulated to study the correlation structure of radar estimates versus rainrate as opposed to RSD moment estimates versus rainrate. The parameters N0, D0 and m of a gamma distribution are varied over the range normally found in rainfall, as well as varying the device sampling volume. The simulations are used to explain some possible features related to discrepancies which can arise when radar rainfall measurements are compared with surface or aircraft-based sampling devices.

Chandrasekar, V.↗

Citizen Science

Scientists and engineers constantly face new challenges, despite myriad advances in computing. More sets of data are collected today from earth and sky than there is time or resources available to carefully analyze them. Some problems either don't have fast algorithms to solve them or have solutions that must be found among millions of options, a situation akin to finding a needle in a haystack. But all hope is not lost: advances in technology and the Internet have empowered the general public to participate in the scientific process via individual computational resources and brain cognition, which isn't matched by any machine. Citizen scientists are volunteers who perform scientific work by making observations, collecting and disseminating data, making measurements, and analyzing or interpreting data without necessarily having any scientific training. In so doing, individuals from all over the world can contribute to science in ways that wouldn't have been otherwise possible.

distributed computing↗

iDDS: intelligent distributed dispatch and scheduling for workflow orchestration

The intelligent distributed dispatch and scheduling (iDDS) service is a versatile workflow orchestration system designed for large-scale, distributed scientific computing. iDDS extends traditional workload and data management by integrating data-aware execution, conditional logic, and programmable workflows, enabling automation of complex and dynamic processing pipelines. Originally developed for the ATLAS experiment at the large hadron collider, iDDS has evolved into an experiment-agnostic platform that supports both template-driven workflows and a Function-as-a-Task model for Python-based orchestration. This paper presents the architecture and core components of iDDS, highlighting its scalability, modular message-driven design, and integration with systems such as PanDA and Rucio. We demonstrate its versatility through real-world use cases: fine-grained tape resource optimization for ATLAS, orchestration of large Directed Acyclic Graph (DAG) workflows for the Rubin Observatory, distributed hyperparameter optimization for machine learning applications, active learning for physics analyses, and AI-assisted detector design at the electron–ion collider. By unifying workload scheduling, data movement, and adaptive decision-making, iDDS reduces operational overhead and enables reproducible, high-throughput workflows across heterogeneous infrastructures. We conclude with current challenges and future directions, including interactive, cloud-native, and serverless workflow support.

97 MATHEMATICS AND COMPUTING↗

Optimizing High-Throughput Inference on Graph Neural Networks at Shared Computing Facilities with the NVIDIA Triton Inference Server

Abstract With machine learning applications now spanning a variety of computational tasks, multi-user shared computing facilities are devoting a rapidly increasing proportion of their resources to such algorithms. Graph neural networks (GNNs), for example, have provided astounding improvements in extracting complex signatures from data and are now widely used in a variety of applications, such as particle jet classification in high energy physics (HEP). However, GNNs also come with an enormous computational penalty that requires the use of GPUs to maintain reasonable throughput. At shared computing facilities, such as those used by physicists at Fermi National Accelerator Laboratory (Fermilab), methodical resource allocation and high throughput at the many-user scale are key to ensuring that resources are being used as efficiently as possible. These facilities, however, primarily provide CPU-only nodes, which proves detrimental to time-to-insight and computational throughput for workflows that include machine learning inference. In this work, we describe how a shared computing facility can use the NVIDIA Triton Inference Server to optimize its resource allocation and computing structure, recovering high throughput while scaling out to multiple users by massively parallelizing their machine learning inference. To demonstrate the effectiveness of this system in a realistic multi-user environment, we use the Fermilab Elastic Analysis Facility augmented with the Triton Inference Server to provide scalable and high-throughput access to a HEP-specific GNN and report on the outcome.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

T-FSM: A Scalable Distributed Task-Based System for Frequent Subgraph Pattern Mining from a Big Graph

Finding frequent subgraph patterns in a big graph is an important problem with many applications such as classifying chemical compounds and building indexes to speed up graph queries. Since this problem is NP-hard, some recent parallel and distributed systems have been developed to accelerate the mining. However, they often have a huge memory cost, very long running time, suboptimal load balancing, poor scale-out capability, and possibly inaccurate results. In this article, we propose an efficient system called T-FSM for parallel mining of frequent subgraph patterns in a big graph. T-FSM supports a new anti-monotonic frequentness measure called Fraction-Score, which is more accurate than the widely used MNI measure. The execution engine of T-FSM supports both intra-machine parallelism and inter-machine parallelism. For intra-machine parallelism, T-FSM adopts a novel task-based execution model to ensure high multithreading concurrency, bounded memory consumption, and effective load balancing. For inter-machine parallelism, T-FSM ensures good scale-out performance with a lightweight pattern rebalancing approach that reduces workload skewness of pattern evaluations among machines. To avoid recomputing the contexts for migrated patterns, we design a novel context cache table to support concurrent and asynchronous requesting and caching of remote context data, which can timely evict and garbage collect used pattern contexts that are no longer needed to keep memory consumption bounded. Extensive experiments show that T-FSM is orders of magnitude faster than existing state-of-the-art parallel systems (more than 10×, 51×, 131×, 55× speedup over ScaleMine, DistGraph, Pangolin and Peregrine, respectively) and distributed systems (more than 42× and 88× over ScaleMine and DistGraph, respectively) for frequent subgraph pattern mining, and it scales out satisfactorily to 512 CPU cores on the Polaris supercomputer at Argonne National Laboratory.

97 MATHEMATICS AND COMPUTING↗

Polyphony: A Workflow Orchestration Framework for Cloud Computing

Cloud Computing has delivered unprecedented compute capacity to NASA missions at affordable rates. Missions like the Mars Exploration Rovers (MER) and Mars Science Lab (MSL) are enjoying the elasticity that enables them to leverage hundreds, if not thousands, or machines for short durations without making any hardware procurements. In this paper, we describe Polyphony, a resilient, scalable, and modular framework that efficiently leverages a large set of computing resources to perform parallel computations. Polyphony can employ resources on the cloud, excess capacity on local machines, as well as spare resources on the supercomputing center, and it enables these resources to work in concert to accomplish a common goal. Polyphony is resilient to node failures, even if they occur in the middle of a transaction. We will conclude with an evaluation of a production-ready application built on top of Polyphony to perform image-processing operations of images from around the solar system, including Mars, Saturn, and Titan.

Space Exploration,↗

Optimal Communication Topology Determination and Sensor Selection for Independent Airspace Surveillance

The paper presents an approach to sensors selection and network topology determination for independent airspace surveillance with maximum outcome and minimum cost using ground based distributed sensing, computing and communication network infrastructure. The selection criteria includes minimum estimation error, maximum airspace coverage, minimum communication time and power consumption while guaranteeing the system observability and providing in-time quality information to a monitoring observer. The developed algorithm uses multi-objective optimization strategy taking into account trade-offs between conflicting objectives and relaxations for in time implementation. It is implemented utilizing graph theoretic tools. The approach is validated in a desktop simulation environment using synthetic sensors data generated for a simulated multi-vehicle flight scenario in the selected regional airspace.

Distributed sensing↗

Separation and Transition on the ROTEX-T Cone-Flare

As part of NATO STO AVT-346 “Predicting Hypersonic Boundary-Layer Transition on Complex Geometries,” coordinated experimental and computational studies were conducted on the ROTEX-T, a cone-flare geometry used in a successful flight-test experiment. At the as-flown conditions, a separation bubble existed at the compression corner. Separation, reattachment, and the multifaceted linear instability paths leading this bubble to transition to turbulence are challenging to predict, but have significant impact on surface pressure and heat flux. High-resolution background-oriented schlieren and infrared thermography measurements were made in the AFOSR–Notre Dame Large Mach-6 Quiet Tunnel at freestream unit Reynolds numbers from 5.8×10(exp 6) to 12.2×10(exp 6) m-1and nominally zero angle of attack. High-speed self-aligned focusing schlieren, infrared thermography, and focused laser differential interferometry measurements were made in the AFRL Mach-6 Ludwieg Tube from 2.2×10(exp 6) to 24.7×10(exp 6) m-1and nominally zero degrees angle of attack. The surface heat-flux and Stanton number distributions were computed. Separation and reattachment locations, as well as the flow state at each, were determined from the combination of surface and off-wall measurements. The convective and global boundary-layer instabilities of the axisymmetric laminar flow at the experimental conditions were investigated computationally. Amplification of Mack’s first and second modes were observed to have logarithmic amplification factors between 5 to 7.5 at the separation location, depending on conditions. The flow was found to be globally unstable to stationary three-dimensional disturbances concentrated in the reattachment region. Previous analysis of the ROTEX-T flight data had not assessed reattachment location or the flow state upon reattachment. Thanks to the insights gained from the coordinated, on- and off-wall ground-test measurements, these evaluations have now been made. The separation location indicated by laminar simulations is consistently numerically predicted to be upstream of the experimentally observed location for a transitional separation bubble. The cause of this difference is understood to lie in the steady-state and axisymmetric assumptions made by both solvers employed to compute the basic states analyzed, as flow topology considerations assert that unsteady two-dimensional or axisymmetric separation bubbles are structurally unstable and will become three-dimensional. Computed laminar heating rates prior to separation agreed well with experiment; transitional heating rates after reattachment were between laminar and turbulent computations.

CFD Validation↗

Separation and Transition on the ROTEX-T Cone-Flare

As part of NATO STO AVT-346 “Predicting Hypersonic Boundary-Layer Transition on Complex Geometries,” coordinated experimental and computational studies were conducted on the ROTEX-T, a cone-flare geometry used in a successful flight-test experiment. At the as-flown conditions, a separation bubble existed at the compression corner. Separation, reattachment, and the multifaceted linear instability paths leading this bubble to transition to turbulence are challenging to predict, but have significant impact on surface pressure and heat flux. High-resolution background-oriented schlieren and infrared thermography measurements were made in the AFOSR–Notre Dame Large Mach-6 Quiet Tunnel at freestream unit Reynolds numbers from 5.8×10(exp 6) to 12.2×10(exp 6) m-1and nominally zero angle of attack. High-speed self-aligned focusing schlieren, infrared thermography, and focused laser differential interferometry measurements were made in the AFRL Mach-6 Ludwieg Tube from 2.2×10(exp 6) to 24.7×10(exp 6) m-1and nominally zero degrees angle of attack. The surface heat-flux and Stanton number distributions were computed. Separation and reattachment locations, as well as the flow state at each, were determined from the combination of surface and off-wall measurements. The convective and global boundary-layer instabilities of the axisymmetric laminar flow at the experimental conditions were investigated computationally. Amplification of Mack’s first and second modes were observed to have logarithmic amplification factors between 5 to 7.5 at the separation location, depending on conditions. The flow was found to be globally unstable to stationary three-dimensional disturbances concentrated in the reattachment region. Previous analysis of the ROTEX-T flight data had not assessed reattachment location or the flow state upon reattachment. Thanks to the insights gained from the coordinated, on- and off-wall ground-test measurements, these evaluations have now been made. The separation location indicated by laminar simulations is consistently numerically predicted to be upstream of the experimentally observed location for a transitional separation bubble. The cause of this difference is understood to lie in the steady-state and axisymmetric assumptions made by both solvers employed to compute the basic states analyzed, as flow topology considerations assert that unsteady two-dimensional or axisymmetric separation bubbles are structurally unstable and will become three-dimensional. Computed laminar heating rates prior to separation agreed well with experiment; transitional heating rates after reattachment were between laminar and turbulent computations.

CFD Validation↗

Clustering at Massive Scale

ClaMS provides hierarchical clustering technology for use on massive, high-dimensional datasets that require distributed memory for processing. The algorithm employed is inspired by the popular HDBSCAN algorithm but makes use of computational kernels better suited for distributed computing. ClaMS is built on scalable nearest neighbor graph construction, metric forest completion, and approximate minimum spanning tree techniques.

Stanley, ThomasA [Lawrence Livermore National Labo↗

Cultivating an Emergent Earth Observation Analytics Ecosystem in the Cloud

A diverse set of data analytics systems for Earth Observations are sprouting up in the Earth Science community, with a wealth of processing algorithms and analysis methods. There is a similar wealth of data resources available via myriad data providers and clearinghouses, including large institutional systems like the Earth Observing System Data and Information System, Comprehensive Large Scale Array-data Stewardship System, and Federated Earth Observation Missions gateway. With Earth system science driving a need to work with more datasets together, and the community developing more analysis tools (some of them dataset-specific), how can we develop analysis workflows that incorporate far-flung datasets and leverage analysis resources from multiple organizations? Cloud computing points the way toward a solution in two different respects. Firstly, the access to and abstraction of virtually unlimited storage and computing power provides an environment that enables more straightforward means of pulling datasets and analysis resources together. Just as importantly, however, cloud computing serves as an example of an "ecosystem" of interoperating services, since the essence of cloud computing is the presentation of all resources as a service, from hardware to infrastructure to platform to software. This enables the combination of off-the-shelf, diverse services to construct entire systems that emerge out of an equally diverse community of architects and developers. This approach can be similarly applied to the data and analysis resources in the Earth Observation community. By exposing these resources via well understood services, and consuming resources in the same way, different organizations can construct bespoke analysis workflows and systems for their own purposes. The key leap the community needs to make is to develop analysis systems in components that interact with other components via services. The result would be a rich ecosystem of analytics components that can be combined to analyze datasets at scale and in conjunction with other datasets from other sources.

chaos↗

Fast Multilevel Implementation of Recursive Spectral Bisection for Partitioning Unstructured Problems

If problems involving unstructured meshes are to be solved efficiently on distributed-memory parallel computers, the meshes must be partitioned and distributed across processors in a way that balances tile computational load and minimizes communication. The recursive spectral bisection method (RSB) has been shown to be very effective for such partitioning problems compared to alternative methods, but RSB in its simplest form is expensive. Here a multilevel version of RSB is introduced that attains about an order-of-magnitude improvement in run time on typical examples.

Barnard, Stephen T.↗

An analytical transformation technique for generating uniformly spaced computational mesh

An analytical transformation method which can map arbitrary physical coordinate grid distribution into desired computational coordinate with uniform grid distribution is derived. The transformation function and its higher derivatives are differentiable. Salient features include; (1) precise control of grid sizes; (2) more than one location of clustered grids; (3) exact positioning of particular computational nodes in the physical plane; and (4) ensuring several patches of uniformly spaced grids in the physical plane for the higher accuracies (such as at the boundaries), etc., while keeping the variation of grid spacing continuous to avoid numerical instability.

Oh, Y. H.↗