Search NASA⌕ Search

SEARCH · Search NASA

Results for “supercomputing technologies”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

WorkflowHub: a registry for computational workflows

The rising popularity of computational workflows is driven by the need for repetitive and scalable data processing, sharing of processing know-how, and transparent methods. As both combined records of analysis and descriptions of processing steps, workflows should be reproducible, reusable, adaptable, and available. Workflow sharing presents opportunities to reduce unnecessary reinvention, promote reuse, increase access to best practice analyses for non-experts, and increase productivity. In reality, workflows are scattered and difficult to find, in part due to the diversity of available workflow engines and ecosystems, and because workflow sharing is not yet part of research practice. WorkflowHub provides a unified registry for all computational workflows that links to community repositories, and supports both the workflow lifecycle and making workflows findable, accessible, interoperable, and reusable (FAIR). By interoperating with diverse platforms, services, and external registries, WorkflowHub adds value by supporting workflow sharing, explicitly assigning credit, enhancing FAIRness, and promoting workflows as scholarly artefacts. The registry has a global reach, with hundreds of research organisations involved, and more than 800 workflows registered.

97 MATHEMATICS AND COMPUTING↗

Dark Energy Survey year 6 results: Magnification modeling and its impact on galaxy clustering and galaxy-galaxy lensing cosmology

Gravitational lensing magnification alters the observed spatial distribution of galaxies and must be accounted for to prevent biases in cosmological probes of the large-scale structure. We investigate its effects on the Dark Energy Survey Year 6 galaxy clustering and galaxy-galaxy lensing analyses using the fiducial lens (position tracer) sample M ag L im++. Magnification bias is parameterized by a coefficient that describes the response of the number of selected objects per unlensed area element to a change in the lensing convergence. We quantify this coefficient using the BALROG synthetic source injection catalog to account for the complexity of the selection function, and compare these results with simplified estimates. The resulting values of the magnification coefficients for each redshift bin are [3.16 ± 0.08, 2.76 ± 0.21, 4.09 ± 0.15, 4.42 ± 0.16, 4.90 ± 0.29, 4.83 ± 0.25]. Relative to Year 3, this analysis provides more precise and accurate magnification bias estimates through a larger BALROG area and reweighting to better match the data properties. Here, the cosmological results are robust when tested against various magnification parameter prior choices and also when adding cross-clustering between lens redshift bins. Neglecting magnification, however, introduces significant systematic shifts: relative to the fiducial analysis with Gaussian priors centered on the BALROG -derived estimates, we observe shifts of 1.37σ in S 8 and -0.84σ in Ω m (with cosmic shear included: -0.61σ in S 8 and -0.71σ in Ω m ), in agreement with findings from simulated data, demonstrating that magnification must be modeled to avoid biases. Freeing the magnification bias in lens bin 2 leads to unphysical negative values, further justifying its exclusion from the fiducial Year 6 analysis.

Cosmological parameters↗

The hierarchical growth of bright central galaxies and intracluster light as traced by the magnitude gap

Using a sample of 2800 galaxy clusters identified in the Dark Energy Survey across the redshift range 0.20 < z < 0.60, we characterize the hierarchical assembly of bright central galaxies (BCGs) and the surrounding intracluster light (ICL). To quantify hierarchical formation we use the stellar mass–halo mass (SMHM) relation, comparing the halo mass, estimated via the mass–richness relation, to the stellar mass within the BCG + ICL system. Moreover, we incorporate the magnitude gap (M14), the difference in brightness between the BCG (measured within 30 kpc) and fourth brightest cluster member galaxy within 0.5 $R_{200,c}$, as a third parameter in this linear relation. The inclusion of M14, which traces BCG hierarchical growth, increases the slope and decreases the intrinsic scatter, highlighting that it is a latent variable within the BCG + ICL SMHM relation. Moreover, the correlation with M14 decreases at large radii. However, the stellar light within the BCG + ICL transition region (30 –80 kpc) most strongly correlates with halo mass and has a statistically significant correlation with M14. Since the transition region and M14 are independent measurements, the transition region may grow due to the BCG’s hierarchical formation. Additionally, as M14 and ICL result from hierarchical growth, we use a stacked sample and find that clusters with large M14 values are characterized by larger ICL and BCG + ICL fractions, which illustrates that the merger processes that build the BCG stellar mass also grow the ICL. Furthermore, this may suggest that M14 combined with the ICL fraction can identify dynamically relaxed clusters.

79 ASTRONOMY AND ASTROPHYSICS↗

Prospects for gravitational wave and ultra-light dark matter detection with binary resonances beyond the secular approximation

Precision observations of orbital systems have recently emerged as a promising new means of detecting gravitational waves and ultra-light dark matter, offering sensitivity in new regimes with significant discovery potential. These searches rely critically on precise modeling of the dynamical effects of these signals on the observed system; however, previous analyses have mainly only relied on the secularly-averaged part of the response. We introduce here a fundamentally different approach that allows for a fully time-resolved description of the effects of oscillatory metric perturbations on orbital dynamics. We find that gravitational waves and ultra-light dark matter can induce large oscillations in the orbital parameters of realistic binaries, enhancing the sensitivity to such signals by orders of magnitude compared to previous estimates.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Cosmology with second- and third-order shear statistics for the Dark Energy Survey: Methods and simulated analysis

We present a new pipeline designed for the robust inference of cosmological parameters using both second- and third-order shear statistics. We build a theoretical model for rapid evaluation of three-point correlations using our fastnc code and integrate it into the cosmosis framework. We measure the two-point functions 𝜉 ± and the full configuration-dependent three-point shear correlation functions across all auto- and cross-redshift bins. We compress the three-point functions into the mass aperture statistic ⟨ℳ$^{3}_{ap}$⟩ for a set of 796 simulated shear maps designed to model the Dark Energy Survey Year 3 data. We estimate from it the full covariance matrix and model the effects of intrinsic alignments, shear calibration biases and photometric redshift uncertainties. We apply scale cuts to minimize the contamination from the baryonic signal as modeled through hydrodynamical simulations. We find a significant improvement of 83% on the figure of merit in the Ω m − 𝑆 8 plane when we add the ⟨ℳ$^{3}_{ap}$⟩ data to 𝜉 ± . Here, we present our findings for all relevant cosmological and systematic uncertainty parameters and discuss the complementarity of third-order and second-order statistics.

79 ASTRONOMY AND ASTROPHYSICS↗

Biasing from galaxy trough and peak profiles with the DES Y3 redMaGiC galaxies and the weak lensing mass map

We measure the correspondence between the distribution of galaxies and matter around troughs and peaks in the projected galaxy density, by comparing redMaGiC galaxies (0.15 < z < 0.65) to weak lensing mass maps from the Dark Energy Survey (DES) Y3 data release. We obtain stacked profiles, as a function of angle θ, of the galaxy density contrast δ g and the weak lensing convergence κ, in the vicinity of these identified troughs and peaks, referred to as ‘void’ and ‘cluster’ superstructures. The ratio of the profiles depend mildly on θ, indicating good consistency between the profile shapes. We model the amplitude of this ratio using a function $F(\boldsymbol{\eta }, \theta )$ that depends on cosmological parameters $\boldsymbol{\eta }$, scaled by the galaxy bias. We construct templates of $F(\boldsymbol{\eta }, \theta )$ using a suite of N-body (‘Gower Street’) simulations forward-modelled with DES Y3-like noise and systematics. We discuss and quantify the caveats of using a linear bias model to create galaxy maps from the simulation dark matter shells. We measure the galaxy bias in three lens tomographic bins (near to far): $2.32^{+0.86}_{-0.27}, 2.18^{+0.86}_{-0.23}, 1.86^{+0.82}_{-0.23}$ for voids, and $2.46^{+0.73}_{-0.27}, 3.55^{+0.96}_{-0.55}, 4.27^{+0.36}_{-1.14}$ for clusters, assuming the best-fit Planck cosmology. Similar values with ∼0.1σ shifts are obtained assuming the mean DES Y3 cosmology. The biases from troughs and peaks are broadly consistent, although a larger bias is derived for peaks, which is also larger than those measured from the DES Y3 3 × 2-point analysis. This method shows an interesting avenue for measuring field-level bias that can be applied to future lensing surveys.

cosmology: observations↗

Accelerating multiscale electronic stopping power predictions with time-dependent density functional theory and machine learning

Knowing the rate at which particle radiation releases energy in a material, the “stopping power,” is key to designing nuclear reactors, medical treatments, semiconductor and quantum materials, and many other technologies. While the nuclear contribution to stopping power, i.e., elastic scattering between atoms, is well understood in the literature, the route for gathering data on the electronic contribution has for decades remained costly and reliant on many simplifying assumptions, including that materials are isotropic. We establish a method that combines time-dependent density functional theory (TDDFT) and machine learning to reduce the time to assess new materials to hours on a supercomputer and provide valuable data on how atomic details influence electronic stopping. Our approach uses TDDFT to compute the electronic stopping from first principles in several directions and then machine learning to interpolate to other directions at a cost of 10 million times fewer core-hours. We demonstrate the combined approach in a study of proton irradiation in aluminum and employ it to predict how the depth of maximum energy deposition, the “Bragg Peak,” varies depending on the incident angle—a quantity otherwise inaccessible to modelers and far outside the scales of quantum mechanical simulations. The lack of any experimental information requirement makes our method applicable to most materials, and its speed makes it a prime candidate for enabling quantum-to-continuum models of radiation damage. The prospect of reusing valuable TDDFT data for training the model makes our approach appealing for applications in the age of materials data science.

36 MATERIALS SCIENCE↗

Brochure on the 2024 ASCR Workshop on Energy-Efficient Computing for Science

Large-scale computing has enabled numerous scientific discoveries, including ground-breaking achievements facilitated by the US Department of Energy (DOE) supercomputers and advances in applied mathematics and computer science. While important advances were made in energy efficiency to enable exascale computing, continued efforts are needed to dramatically improve the energy efficiency of the next generation of high-performance computing (HPC) systems and, more broadly, AI data centers. Without substantial improvements in energy efficiency, the energy consumption associated with computing could become a limiting factor for future scientific discovery, national security, and technological advancement.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A Brief Survey of Data Streaming Technologies

Streaming data is data that is emitted at variable volumes in a continuous, incremental manner with the goal of low-latency processing often at a different physical location. Network infrastructure is used to facilitate the connection between data sources and sinks, and must be robust to handle the requirements of the workflow. The U.S. Department of Energy Office of Science (DOE SC) a federal agency supporting fundamental scientific research for energy and the Nation’s largest supporter of basic research in the physical sciences. DOE SC has the responsibility for operating $\mathbf{1 0}$ National Laboratories, and 28 scientific user facilities supporting advanced supercomputers, particle accelerators, large x-ray light sources, neutron scattering sources, and other specialized facilities for nanoscience and genomics. This paper investigates the state of streaming data workfows, and details some of the approaches to this challenging problem.

Kissel, Ezra↗

Improving galaxy cluster selection with the outskirt stellar mass of galaxies

The number density and redshift evolution of optically selected galaxy clusters offer an independent measurement of the amplitude of matter fluctuations, 𝑆 8 . However, recent results have shown that clusters chosen by the redMaPPer algorithm show richness-dependent biases that affect the weak lensing signals and number densities of clusters, increasing uncertainty in the cluster mass calibration and reducing their constraining power. Here, in this work, we evaluate an alternative cluster proxy, outskirt stellar mass, 𝑀 out , defined as the total stellar mass within a [50, 100] kpc envelope centered on a massive galaxy. This proxy exhibits scatter comparable to redMaPPer richness, 𝜆, but is less likely to be subject to projection effects. We compare the Dark Energy Survey Year 3 redMaPPer cluster catalog with a 𝑀 out selected cluster sample from the Hyper-Suprime Camera survey. We use weak lensing measurements to quantify and compare the scatter of 𝑀 out and 𝜆 with halo mass. Our results show 𝑀 out has a scatter consistent with 𝜆, with a similar halo mass dependence, and that both proxies contain unique information about the underlying halo mass. We find 𝜆-selected samples introduce features into the measured Δ⁢Σ signal that are not well fit by a log-normal scatter only model, absent in 𝑀 out selected samples. Our findings suggest that 𝑀 out offers an alternative for cluster selection with more easily calibrated selection biases, at least at the generally lower richnesses probed here. Combining both proxies may yield a mass proxy with a lower scatter and more tractable selection biases, enabling the use of lower mass clusters in cosmology. Finally, we find the scatter and slope in the 𝜆 −𝑀 out scaling relation to be 0.49 ±0.02 and 0.38 ±0.09.

79 ASTRONOMY AND ASTROPHYSICS↗

Massively parallel phase-field simulations targeting exascale

The interface thickness in the phase-field (PF) method limits its simulation scales. Consequently, large-scale PF simulations become prohibitively expensive for resolving the extremely fine microstructures that typically form during rapid solidification processing. This challenge is significant in predicting microstructure evolution in metal additive manufacturing and has been identified by the United States Department of Energy’s Exascale Computing Project. Here, to address this, we develop a multi-GPU and MPI-based massively parallel simulation code, utilizing state-of-the-art algorithms, software, and libraries, for large-scale three-dimensional (3D) PF simulations. We report the first GPU-parallel PF simulations on Frontier (currently the second TOP500 exascale cluster) and Summit machines, taking dendritic growth as an example problem. We evaluate the parallel performance of our implementation using scaling studies with more than 24 000 GPUs (among the largest known computations to date) and the acceleration performance using large-scale simulations of dendritic growth in 3D. Finally, massively parallel GPUs in these supercomputers enabled the first coupled multiscale simulations of laser melting and subsequent dendritic solidification on the scale of a full melt-pool, demonstrating the feasibility of performing PF simulations with a point total over 2 billion grid points within an acceptable time.

Exascale↗

T-FSM: A Scalable Distributed Task-Based System for Frequent Subgraph Pattern Mining from a Big Graph

Finding frequent subgraph patterns in a big graph is an important problem with many applications such as classifying chemical compounds and building indexes to speed up graph queries. Since this problem is NP-hard, some recent parallel and distributed systems have been developed to accelerate the mining. However, they often have a huge memory cost, very long running time, suboptimal load balancing, poor scale-out capability, and possibly inaccurate results. In this article, we propose an efficient system called T-FSM for parallel mining of frequent subgraph patterns in a big graph. T-FSM supports a new anti-monotonic frequentness measure called Fraction-Score, which is more accurate than the widely used MNI measure. The execution engine of T-FSM supports both intra-machine parallelism and inter-machine parallelism. For intra-machine parallelism, T-FSM adopts a novel task-based execution model to ensure high multithreading concurrency, bounded memory consumption, and effective load balancing. For inter-machine parallelism, T-FSM ensures good scale-out performance with a lightweight pattern rebalancing approach that reduces workload skewness of pattern evaluations among machines. To avoid recomputing the contexts for migrated patterns, we design a novel context cache table to support concurrent and asynchronous requesting and caching of remote context data, which can timely evict and garbage collect used pattern contexts that are no longer needed to keep memory consumption bounded. Extensive experiments show that T-FSM is orders of magnitude faster than existing state-of-the-art parallel systems (more than 10×, 51×, 131×, 55× speedup over ScaleMine, DistGraph, Pangolin and Peregrine, respectively) and distributed systems (more than 42× and 88× over ScaleMine and DistGraph, respectively) for frequent subgraph pattern mining, and it scales out satisfactorily to 512 CPU cores on the Polaris supercomputer at Argonne National Laboratory.

97 MATHEMATICS AND COMPUTING↗

Overview of the ASDEX Upgrade results

After a 26-month vent ASDEX Upgrade (AUG) went back in operation with a newly designed upper W-divertor suitable for alternative divertor configurations (featuring in-vessel coils and cryo-pump). Parameter scans and an extensive set of measurements were obtained and their interpretation is ongoing. Prompted by the ITER wall change, dedicated experiments on non-boronized plasma startup were contrasted to that employing asymmetric and more symmetric boronizations. The asymmetric boronization proved to be as beneficial as the more symmetric one, which is in contrast to previous model calculations assuming perfect sticking of boron (measurements suggest sticking ≈ 0.3). In the startup phase also the impurity influxes at the outboard limiters were investigated contrasting the unboronized case featuring cold edges (low- Z radiation) to the boronized case, in which the lifetime of the boron layers could be estimated. Pedestal stability investigations revealed that the quasi-continuous exhaust (QCE) regime is obtained when ballooning modes are active in the vicinity of the separatrix and the global peeling-ballooning stability is high enough. Thus, at high enough shaping and high gas flux both can be achieved and QCE is a consequence. The closely related enhanced D-alpha (EDA) mode is not clearly distinguishable from QCE, e.g. the quasi-coherent mode characteristic for EDA also shows up in QCE. In QCE the impurity transport is behaving benign as could be measured for Ne with a novel analysis method making use of a comprehensive set of CXRS measurements. For high radiative fractions the regime of the X-point radiator (XPR) is accessible at AUG and the understanding of its access conditions and behaviour is further developed. Due to the localized radiative cooling at the X-point the XPR can be well diagnosed and thus controlled. For negative triangularity shapes, further experiments at increased shaping resulted in strongly heated L-mode plasmas avoiding ELMs. Two integrated modelling approaches towards ITER suggested that core W-accumulation will be no issue for ITER and that the fusion yield in ITER may be Q = 12 (i.e. ITPA20-IL scaling is too pessimistic). Further, investigations of the ITER ramp-down in AUG provide insights into maintaining position control. Various aspects of shattered pellet injection were investigated in AUG and one of the results show that with increasing Ne fraction the radiation during the current quench increases and the current decay becomes faster.

ASDEX Upgrade↗

Reducing Long‐Standing Surface Ozone Overestimation in Earth System Modeling by High‐Resolution Simulation and Dry Deposition Improvement

The overestimation of surface ozone concentration in low‐resolution global atmospheric chemistry and climate models has been a long‐standing issue. We first update the ozone dry deposition scheme in both high‐ (0.25°) and low‐resolution (1°) Community Earth System Model (CESM) version 1.3 runs, by adding the effects of leaf area index and correcting the sunlit and shaded fractions of stomatal resistances. With this update, 5‐year‐long summer simulations (2015–2019) using the low‐resolution CESM still exhibit substantial ozone overestimation (by 6.0–16.2 ppbv) over the U.S., Europe, eastern China, and ozone pollution hotspots. The ozone dry deposition scheme is further improved by adjusting the leaf cuticle conductance, reducing the mean ozone bias by 19%, and increasing the model resolution further reduces the ozone overestimation by 43%. We elucidate the mechanism by which model grid spacing influences simulated ozone, revealing distinctive pathways in urban versus rural areas. In rural areas, grid spacing mainly affects daytime ozone levels, where additional NO x emissions from nearby urban areas result in an ozone boost and overestimation in low‐resolution simulations. In contrast, over urban areas, daytime ozone overestimation follows a similar mechanism due to the influence of volatile organic compounds from surrounding rural areas. However, nighttime ozone overestimation is closely linked to weakened NO titration owing to the redistribution of urban NO x to rural areas. Additionally, stratosphere‐troposphere exchange may also contribute to reducing ozone bias in high‐resolution simulations, warranting further investigation. This optimized high‐resolution CESM may enhance understanding of ozone formation mechanisms, sources, and changes in a warming climate.

54 ENVIRONMENTAL SCIENCES↗

Dust Direct Radiative Effect Including Large Particles and Component Minerals

The direct radiative effect (DRE) of dust aerosols in Earth system models (ESMs) remains highly uncertain, largely due to inadequate representations of particle size distribution (PSD) and mineral composition. Using NASA's Earth Surface Mineral Dust Source Investigation (EMIT) soil mineralogy data and observed PSD in an ESM that resolves dust mineral composition and emitted diameters from 0.1 to 70 μm, we find a near-neutral global dust net DRE (−0.057 W m −2 ), weaker than most previous estimates. Large dust (diameter >10 μm) contributes 30% of the global longwave dust optical depth, providing observational constraints on large-particle abundance, and offsets 20% of the dust shortwave cooling over major source regions. Incorporation of EMIT mineralogy reduces shortwave uncertainty by more than 50%. The remaining uncertainty mainly exists in processes controlling dust abundance, particularly the poorly understood transport of large dust, and longwave optical properties, which require additional observational constraints to more accurately quantify the dust DRE.

Li, Longlei [Cornell Univ., Ithaca, NY (United Sta↗

Energy dataset of Frontier supercomputer for waste heat recovery

The Hewlett Packard Enterprise–Cray EX Frontier is the world’s first and fastest exascale supercomputer, hosted at the Oak Ridge Leadership Computing Facility in Tennessee, United States. Frontier is a significant electricity consumer, drawing 8–30 MW; this massive energy demand produces significant waste heat, requiring extensive cooling measures. Although harnessing this waste heat for campus heating is a sustainability goal at Oak Ridge National Laboratory (ORNL), the 30 °C–38 °C waste heat temperature poses compatibility issues with standard HVAC systems. Heat pump systems, prevalent in residential settings and some industries, can efficiently upgrade low-quality heat to usable energy for buildings. Thus, heat pump technology powered by renewable electricity offers an efficient, cost-effective solution for substantial waste heat recovery. However, a major challenge is the absence of benchmark data on high-performance computing (HPC) heat generation and waste heat profiles. This paper reports power demand and waste heat measurements from an ORNL HPC data centre, aiming to guide future research on optimizing waste heat recovery in large-scale data centres, especially those of HPC calibre.

97 MATHEMATICS AND COMPUTING↗

PSTN-019: The LSST Science Pipelines Software: Optical Survey Pipeline Reduction and Analysis Environment

The NSF-DOE Vera C. Rubin Observatory is executing the Legacy Survey of Space and Time (LSST) as its prime mission, producing a series of data releases over the ten-year survey. The LSST Science Pipelines Software will be used to create these data releases and to perform the nightly prompt processing and alert production. This paper provides an overview of the LSST Science Pipelines Software, describing the components and their integration into pipelines that generate science-ready data products.

79 ASTRONOMY AND ASTROPHYSICS↗

HPDR: High-Performance Portable Scientific Data Reduction Framework

The rapid growth in scientific data generation is outpacing advancements in computing systems necessary for efficient storage, transfer, and analysis, particularly in the context of exascale computing. With the deployment of first-generation exascale computing systems and next-generation experimental facilities, this gap is widening and necessitates effective data reduction techniques to manage enormous data volumes. Over the past decade, various data reduction methods, including lossless compression, error-controlled lossy compression, and data refactoring, have been developed to accelerate I/O in scientific workflows. Despite significant reductions in data volume, these methods introduce considerable computational overhead, which can become the new bottleneck in data processing. To mitigate this, GPU-accelerated data reduction algorithms have been introduced. However, challenges remain in their integration into exascale workflows, including limited portability across different GPU architectures, substantial memory transfer overhead, and reduced scalability on dense multi-GPU systems. To address these challenges, we propose HPDR, a high-performance and portable data reduction framework. HPDR is designed to enable the execution of state-of-the-art reduction algorithms across diverse processor architectures while reducing memory transfer overhead to 2.3 % of the original, resulting in up to 3.5× faster throughput compared to existing solutions. It also achieves up to 96% of the theoretical speedup in multi-GPU settings. In addition, evaluations on accelerating I/O operations at scale up to 1,024 nodes of the Frontier supercomputer demonstrate that HPDR can achieve up to 103 TB/s reduction throughput, providing up to 4× acceleration in parallel I/O performance compared to existing data reduction routines. This work highlights the potential of HPDR to significantly enhance data reduction efficiency in exascale computing environments.

Chen, Jieyang [University of Oregon]↗