Search NASASearch

SEARCH · Search NASA

Results for “GRACE”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Studying CPU and memory utilization of applications on Fujitsu A64FX and Nvidia Grace Superchip

ARM-based manycore CPU architectures are well-positioned to provide the rising memory throughput requirements of modern data intensive scientific applications in High Performance Computing (HPC). The Fujitsu A64FX CPU platform is based on the ARM v8.2A architecture, and is the processor of the flagship Japanese supercomputer - "Fugaku", which was previously ranked as the #1 supercomputer in the world according to the Top500 list. The Nvidia Grace superchip features 144 Neoverse V2 cores based on the ARMv9 architecture with 4x128b SVE2, providing exceptional computational power. The chip supports up to 480GB of memory, making it ideal for AI, machine learning, and scientific computing workloads. In this paper, we conduct a thorough performance exploration of a variety of parallel bandwidth-sensitive benchmarks and applications compiled with the native Fujitsu compiler on a Fugaku A64FX compute node and ARM (LLVM) Compiler on an NVIDIA Grace superchip compute node, engaging all the computational cores per cluster using OpenMP multithreading (assuming the cores can drive the available bandwidth). Our ultimate goals are to study the resource utilization of scientific applications and benchmarks on A64FX and Grace superchip, considering graph application scenarios ( GAP Benchmark suite) and eleven appli- cation proxies from the Rodinia heterogeneous benchmark suite (considering domains such as Data Mining, Bioinformatics, Fluid Dynamics, Pattern Recognition, etc.). Through exhaustive performance monitoring, we quantify the resource utilization of diverse OpenMP-based HPC applications on both the Fujitsu A64FX and the Nvidia Grace Superchip platforms.

benchmarking, Performance Analysis, High performan

GRACE Final Technical Report

The GRACE project (Grid that is Risk-Aware for Clean Electricity) was a five-year research initiative funded by the U.S. Department of Energy's Advanced Research Projects Agency- Energy (ARPA-E) under the PERFORM program. Led by Duke University's Nicholas School of the Environment, with contributions from The Ohio State University, North Carolina State University, Dartmouth College, and Pacific Northwest National Laboratory, the project addressed a fundamental challenge in grid management: conventional software plans for a single most-likely outcome and relies on reserves as a buffer, leaving utilities poorly equipped for the growing variability introduced by renewable energy. GRACE demonstrated a better approach: explicitly representing thousands of plausible future scenarios and choosing operating schedules that perform well across all of them.

14 SOLAR ENERGY

Preliminary Study on Fine-Grained Power and Energy Measurements on Grace Hopper GH200 with Open-Source Performance Tools

The increasing adoption of tightly integrated, heterogeneous architectures, combined with the slowdown of Moore’s law, has made application power and energy-driven optimizations critical to efficiently use high-performance computing systems. This paper introduces a newly developed open-source toolkit that seamlessly integrates the Linux real-time hardware monitoring program hwmon with the Performance Application Programming Interface and the Score-P performance measurement system, thereby enabling fine-grained power and energy measurements for high-performance computing applications. Our primary target platform is the Wombat test bed, which is a system based on the NVIDIA GH200 superchip. The toolkit can capture transient power peaks with high temporal resolution (50 ms) and, thanks to Score-P integration, can map power metrics to specific code regions, thereby providing actionable information on power-intensive operations and inefficiencies. The toolkit also provides a holistic view of both the power and the energy consumption of the entire GH200 superchip by covering all major components: the Grace CPU, the Hopper GPU, and the I/O subsystem. Experiments that use Locally Self-consistent Multiple Scattering, which is an application for first-principles calculations of materials developed at Oak Ridge National Laboratory, have demonstrated the tool’s ability to identify transient power spikes and uncover opportunities for energy-aware optimizations. Additionally, we introduce a Python-based utility for converting Open Trace Format 2 traces to Parquet format, thus enabling advanced data analysis for numerical integration methods applied to power data for accurate energy profiling.

Hernandez Mendoza, Oscar [ORNL] (ORCID:00000002538

Sum Reduction with OpenMP Offload on NVIDIA Grace-Hopper System

We evaluate the performance of the baseline and optimized reductions in OpenMP on an NVIDIA Grace-Hopper system. We explore the impacts of the number of teams, the number of elements to sum per loop iteration, and simultaneous execution on the central-processing unit (CPU) and the GPU in the unified memory (UM) mode upon the reduction performance. The experimental results show that the optimized reductions are 6.120X to 20.906X faster than the baselines on the GPU, and their efficiency ranges from 89% to 95% of the theoretical GPU memory bandwidth. Depending on where an input array is allocated in the program when co-running the reduction on the CPU and GPU in the UM mode, the average speedup over the GPU-only execution is approximately 2.484 or 1.067, and the speedup of the optimized reductions over the baseline reductions ranges from 0.996 to 10.654 or from 0.998 to 6.729.

Jin, Zheming

Deciphering the Role of Total Water Storage Anomalies in Mediating Regional Flooding

Regional floods result from various flood generation mechanisms. Traditional analyses mainly link flooding to extreme rainfall, with limited input from soil moisture. Total water storage (TWS) is a holistic measure of basin wetness, including additional storage components from surface water, snow, and groundwater. Utilizing a new 5-day Gravity Recovery and Climate Experiment and its Follow On (GRACE(-FO)) data set, we investigated the linkage between short-term TWS anomaly (TWSA) and regional flooding. The 5-day TWSA solutions revealed flood signals missed by monthly TWSA solutions. Global basins exhibit distinct storage-discharge co-evolution patterns, offering new insights into flood mechanisms and propensity. Our bivariate event analyses show the annual maximum river discharges co-occur more often with the TWSA maxima than with precipitation in many basins. Further analyses revealed TWSA's time-lagged effect on river discharge, particularly in basins susceptible to floods triggered by saturation-excess runoff. The 5-day TWSA provides a new source of information for enhancing global flood preparedness.

54 ENVIRONMENTAL SCIENCES

Climate-extreme modeling framework for sustainable flood management in the Arabian Peninsula

Evaluating extreme precipitation events (EPEs) is essential for building climate-resilient water management strategies, but it remains a major challenge in ungauged basins. Using the 26,070 km 2 Wadi al-Rummah basin in central Saudi Arabia as a case study, we developed an alternative, reliable, cost-effective satellite-based framework that combines empirically derived EPE thresholds, imagery-calibrated 2D hydrodynamic modeling, GRACE water-storage diagnostics, and bias-corrected CMIP6 projections to assess flood hazards and recharge potential under current and future climate scenarios in ungauged basins. The integrated approach and the resulting findings followed four key steps: (1) Identified a 22.5 mm EPE threshold, the 80th percentile of 3-day GPM/IMERG rainfall (2000–2024), aligned with flood-triggering events (Nov 2018: 23–28 mm; Apr, 2023: 42 mm); (2) Developed and calibrated a RiverFlow2D model using Sentinel-2 and PlanetScope imagery for the November 2018 flood, accurately reproducing flood depth and extent (RMSE ≤0.31 m; fuzzy-Dice ≥0.91), and estimating runoff (41 %), infiltration (25 %), and evaporation (34 %); (3) Independently validated the model with the April 2023 event (RMSE ≤0.35 m; fuzzy-Dice ≥0.86); (4) Conducted climate projections (2025–2100) from five bias-corrected NEX-GDDP CMIP6 models that revealed a 34 % increase in EPE intensity under SSP2-4.5 and 48 % under SSP5-8.5 scenarios, relative to 20th-century baselines. Our findings indicate that while intensifying extremes in the 21st century increase flood risk, the results highlight the potential for episodic recharge if effective retention strategies are employed, and offer a transferable model for climate-informed planning in data-scarce arid regions.

CMIP6

Distributed optimization for multi-commodity urban traffic control

A distributed method for concurrent traffic signal and routing control of traffic networks is proposed. The method is based on the multi-commodity store-and-forward model, in which the destinations are the commodities. The system benefits from the communication between vehicles and infrastructure, providing optimal signal timings to intersections and routes to vehicles on a link-by-link basis. Using the augmented Lagrangian to model the constraints into the objective, the baseline centralized problem is decomposed into a set of objective-coupled subproblems, one for each intersection, enabling the solution to be computed by a distributed- gradient projection algorithm. Further, the intersection agents only need to communicate and coordinate with neighboring intersections to ensure convergence to the optimal solution while tolerating suboptimal iterations that offer more flexibility, unlike other distributed approaches. Through microsimulation, we demonstrate the effectiveness of the proposed algorithm in traffic networks with time-varying demand. Computational analysis shows that the distributed problem is suitable for real-time applications. A robustness analysis show that the distributed formulation enables a graceful degradation of the system in case of failure.

Augmented Lagrangian

OpenUniverse2024: a shared, simulated view of the sky for the next generation of cosmological surveys

The OpenUniverse2024 simulation suite is a cross-collaboration effort to produce matched simulated imaging for multiple surveys as they would observe a common simulated sky. Both the simulated data and associated tools used to produce it are intended to uniquely enable a wide range of studies to maximize the science potential of the next generation of cosmological surveys. We have produced simulated imaging for approximately 70 deg 2 of the Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) Wide-Fast-Deep survey and the Nancy Grace Roman Space Telescope High-Latitude Wide-Area Survey, as well as overlapping versions of the ELAIS-S1 Deep-Drilling Field for LSST and the High-Latitude Time-Domain Survey for Roman. OpenUniverse2024 includes (i) an early version of the updated extragalactic model called Diffsky, which substantially improves the realism of optical and infrared photometry of objects, compared to previous versions of these models; (ii) updated transient models that extend through the wavelength range probed by Roman and Rubin; and (iii) improved survey, telescope, and instrument realism based on up-to-date survey plans and known properties of the instruments. It is built on a new and updated suite of simulation tools that improves the ease of consistently simulating multiple observatories viewing the same sky. The approximately 400 TB of synthetic survey imaging and simulated universe catalogs are publicly available, and we preview some scientific uses of the simulations.

large-scale structure of Universe

DAmodel: hierarchical Bayesian modelling of DA white dwarfs for spectrophotometric calibration

We use hierarchical Bayesian modelling to calibrate a network of 32 all-sky faint DA white dwarf (DA WD) spectrophotometric standards (⁠16.5 < V , 19.5⁠) alongside three CALSPEC standards, from 912 Å to 32 μm. The framework is the first of its kind to jointly infer photometric zero points and WD parameters (surface gravity log g⁠, effective temperature T eff ⁠, extinction A V ⁠, dust relation parameter R V ) by simultaneously modelling both photometric and spectroscopic data. We model panchromatic Hubble Space Telescope Wide Field Camera 3 (HST/WFC3) UVIS and IR photometry, HST/STIS UV spectroscopy, and ground-based optical spectroscopy to sub-per cent precision. Photometric residuals for the sample are the lowest yet yielding < 0.004 mag RMS on average from the UV to the NIR, achieved by jointly inferring time-dependent changes in system sensitivity and WFC3/IR count-rate nonlinearity. Our GPU-accelerated implementation enables efficient sampling via Hamiltonian Monte Carlo, critical for exploring the high-dimensional posterior space. The hierarchical nature of the model enables population analysis of intrinsic WD and dust parameters. Inferred spectral energy distributions from this model will be essential for calibrating the James Webb Space Telescope as well as next-generation surveys, including Vera Rubin Observatory’s Legacy Survey of Space and Time and the Nancy Grace Roman Space Telescope.

methods: statistical

Deep-field analytical calibration

The next generation of imaging surveys, including the Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST), Euclid, and the Nancy Grace Roman Space Telescope, will provide unprecedented constraints on cosmology using weak gravitational lensing. To fully exploit this statistical power, shear measurement methods must achieve sub- per cent accuracy while mitigating systematic biases from noise, the point-spread function (PSF), blending, and shear-dependent detection. The analytical calibration framework (AnaCal) has demonstrated such accuracy but requires adding noise to images, reducing effective depth. We introduce Deep-Field Analytical Calibration (DEEP-FIELD AnaCal), an extension of AnaCal that uses deep-field images to compute shear responses while preserving the statistical power of wide-field data. We validate DEEP-FIELD AnaCal on isolated and blended galaxy image simulations with LSST-like conditions, finding it meets the stringent requirement of multiplicative bias $|m| < 3\times 10^{-3}$ at 99.7 per cent confidence. Compared to standard AnaCal applied to wide-field images, DEEP-FIELD AnaCal increases the effective galaxy number density from 17 to 30 arcmin$^{-2}$ for simulated 10-yr LSST data. With deep fields $10\times$ longer than the wide field, we find pixel noise variance in shear estimation is reduced by 30 per cent and overall uncertainty by $\sim 25~{{\ \rm per\ cent}}$. Finally, using the LSST Deep Drilling Fields strategy, we assess sample variance and find an equivalent calibration uncertainty of $\lesssim 0.3~{{\ \rm per\ cent}}$. These results demonstrate that DEEP-FIELD AnaCal offers a promising path to achieve the required shear calibration for upcoming weak lensing surveys.

79 ASTRONOMY AND ASTROPHYSICS

Gravitational wave spectrum of chain inflation

Chain inflation is an alternative to slow-roll inflation in which the inflaton tunnels along a large number of consecutive minima in its potential. In this work we perform the first comprehensive calculation of the gravitational wave (GW) spectrum of chain inflation. In contrast to slow-roll inflation the latter does not stem from quantum fluctuations of the gravitational field during inflation, but rather from the bubble collisions during the first-order phase transitions associated with vacuum tunneling. Our calculation is performed within an effective theory of chain inflation which builds on an expansion of the tunneling rate capturing most of the available model space. The effective theory can be seen as chain inflation’s analog of the slow-roll expansion in rolling models of inflation. The near scale-invariance of the scalar power spectrum translates to a quasiperiodic shape of the inflaton potential in chain inflation, with the tunneling rate changing very slowly during the e-folds leading to cosmic microwave background observables. We show that chain inflation produces a very characteristic double-peak GW spectrum: a faint high-frequency peak associated with the gravitational radiation emitted during inflation, and a strong low-frequency peak associated with the graceful exit from chain inflation (marking the transition to the radiation-dominated epoch). There exist very exciting prospects to test the gravitational wave signal from chain inflation at the aLIGO-aVIRGO-KAGRA network, at LISA and /or at pulsar timing array experiments. A particularly intriguing possibility we point out is that chain inflation could be the source of the stochastic gravitational wave background recently detected by NANOGrav, PPTA, EPTA, and CPTA. We also show that the gravitational wave signal of chain inflation is often accompanied by running/ higher running of the scalar spectral index to be tested at future cosmic microwave background experiments. Published by the American Physical Society 2024

Freese, Katherine

Improving transition to IPv6-only via RFC8925 and IPv4 DNS Interventions

Nine years have passed since the American Registry for Internet Numbers exhausted its allocation of Internet Protocol version 4 (IPv4) addresses, and four years have passed since the United States Government mandated federal agencies to complete the transition to Internet Protocol version 6 (IPv6). Despite the IPv4 address shortage and IPv6 mandate, Federally Funded Research and Development Centers (FFRDCs) are still struggling to sunset IPv4. As demonstrated on SC23’s SC23v6 wireless network, newer tooling such as RFC8925 allows clients to disable their IPv4 protocol stack while retaining legacy IP connectivity via the RFC6145 translation algorithm. However, SC23v6 wireless clients without RFC8925 support or a disabled IPv6 stack would continue to receive internet access via legacy IPv4. This paper introduces a method of using poisoned IPv4 Domain Name System (DNS) records to gracefully inform IPv4-only clients at SC24’s SC24v6 wireless network about their inability to use the current version of internet protocol, with a goal of minimal impact to RFC8925 and dual-stack clients. When implemented as designed, this method may improve supportability and user experience of IPv6-only deployments at FFRDCs.

Costello, Thomas M

Reconstructing f ( T ) gravity and exploring the torsion driven warm inflationary cosmology

The current paper reports an investigation of a warm inflationary scenario in the context of f(T) gravity for a spatially flat FLRW universe. In our model, inflation is driven purely by the torsional sector of f(T) gravity, without introducing any additional scalar fields. We focus on the high dissipative regime (R >> 1), reconstruct the Hubble parameter as a function of the e-folding number N, and derive the slow-roll parameters ε 1 (N) and ε 2 (N). The study has encapsulated the dynamics of inflation and its duration under strong dissipation. The dissipative coefficient Γ is modeled with a temperature-dependent power-law form, linking the inflationary dynamics to thermal corrections and the particle content of the early universe. The analysis has affirmed that the torsion-induced energy density ρ T successfully transitions to radiation energy density ρ rad , facilitating a graceful exit from inflation. Finally, we have validated our model by comparing the scalar spectral index and tensor-to-scalar ratio with Planck 2018 results, demonstrating consistency within observational bounds. Additionally, it is verified that the thermal domination condition T * /H > 1 and the torsion dominance condition ρ T /ρ rad > 1 are satisfied.

Ghosh, Moli [Amity University, Kolkata (India); Mr

Enabling Command-and-Control in Advanced In Situ Workflows

Scientific discovery is progressing towards autonomous science with the combination of scientific instruments, high-performance computing, and artificial intelligence in complex workflows. This evolution introduces new requirements for managing scientific workflows, including feedback loops, near real-time constraints, and the ability to dynamically control workflow execution. In situ workflows that analyze and visualize data as it is generated are well-suited to satisfy stringent time constraints and their iterative nature offers greater opportunities for command-and-control. However, only a few of the many workflow management systems available have been specifically designed to manage in situ workflows and often lack support for automated feedback loops that allow analysis and visualization components to interact with the main scientific data producer. To address this need, we present in this paper how to add command-and-control capabilities to a workflow management system. We identify the functional design requirements of such a command-and-control system, detail its architecture, interface, and core mechanisms, and illustrate how advanced in situ workflows can leverage command-and-control in three use cases: graceful termination with checkpoint, dynamic and adaptive data reduction, and event-triggered analysis.

Mehta, Kshitij [ORNL] (ORCID:0000000297149981)

4th TDAMM Workshop White Paper

Time-Domain and Multi-Messenger Astrophysics (TDAMM) is entering a fundamentally new phase characterized by an unprecedented increase in the rate and diversity of astrophysical transient detections. The community is transitioning from a discovery-limited to a follow-up-limited era, driven by major investments across electromagnetic, gravitational-wave, and neutrino observatories. Upcoming facilities such as the Vera C. Rubin Observatory, the Nancy Grace Roman Space Telescope, and wide-field survey instruments will produce a deluge of time-domain alerts, reaching millions of events per night. Simultaneously, upgrades to the gravitational-wave network (LVK O5 and beyond) and neutrino observatories (IceCube Gen2) will significantly increase the detection rates of non-electromagnetic messengers. New high-energy missions and expansions of the InterPlanetary Network (IPN) will further enhance discovery capabilities across the gamma-ray and X-ray regimes. This convergence of capabilities represents a transformative opportunity: for the first time, the community will routinely detect rare and high-impact events across multiple messengers. However, the scientific return from these discoveries will depend critically on the ability to rapidly identify, prioritize, and coordinate follow-up observations across a heterogeneous and globally distributed set of facilities.

79 ASTRONOMY AND ASTROPHYSICS

Simulating Continuum-based Redshift Measurement in the Roman’s High Latitude Spectroscopic Survey

We investigate the capability of the Nancy Grace Roman Space Telescope’s (Roman) Wide-Field Instrument G150 slitless grism to detect red, quiescent galaxies based on the current reference survey. We simulate dispersed images for Roman reference High-Latitude Spectroscopic Survey (HLSS) and analyze two-dimensional spectroscopic data using the grism Redshift and Line Analysis (Grizli) software. This study focus on assessing Roman grism’s capability for continuum-level redshift measurement for a redshift range of 0.5 ≤ z ≤ 2.5. The redshift recovery is assessed by setting three requirements of: σ z = $\frac{|z–z_{true}|}{1+z}$ ≤ 0.01, signal-to-noise ratio≥ 5 and the presence of a single dominant peak in redshift likelihood function. We find that, for quiescent galaxies, the reference HLSS can reach a redshift recovery completeness of ≥50% for F158 magnitude brighter than 20.2 mag. We also explore how different survey parameters, such as exposure time and the number of exposures, influence the accuracy and completeness of redshift recovery, providing insights that could optimize future survey strategies and enhance the scientific yield of the Roman in cosmological research.

Astronomical simulations