Search NASA⌕ Search

SEARCH · Search NASA

Results for “Energy delay product”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Bridging the Gap: User-Centric Energy Monitoring for Policy-Driven Application Optimization in HPC Data Centers

Application energy optimization in HPC data centers face two critical gaps. Systematic methodologies that connect data center policies to application decisions and accessible monitoring tools that enable data-driven optimization. We address both gaps through two complementary pillars. First, we present a methodology based on extended weighted Energy Delay Product (EDP) to translate data center operational priorities and integrate energy considerations into the energy optimization workflow which starts from continuous monitoring through targeted optimization. Second, we present a user-space monitoring tool, Omnistat, that enables this methodology by providing developers with direct access to actionable energy telemetry. Through deployment on the Frontier supercomputer and case studies exploring performance-energy trade-offs, we show how these pillars help energy as an integral optimization target for developers as active participants in data center efficiency.

Shin, Woong [ORNL] (ORCID:0000000172077814)↗

ytopt: Autotuning Scientific Applications for Energy Efficiency at Large Scales

As we enter the exascale computing era, efficiently utilizing power and optimizing the performance of scientific applications under power and energy constraints has become critical and challenging. We propose a low-overhead autotuning framework to autotune performance and energy for various hybrid MPI/OpenMP scientific applications at large scales and to explore the tradeoffs between application runtime and power/energy for energy efficient application execution, then use this framework to autotune four ECP proxy applications—XSBench, AMG, SWFFT, and SW4lite. Our approach uses Bayesian optimization with a Random Forest surrogate model to effectively search parameter spaces with up to 6 million different configurations on two large-scale HPC production systems, Theta at Argonne National Laboratory and Summit at Oak Ridge National Laboratory. The experimental results show that our autotuning framework at large scales has low overhead and achieves good scalability. Using the proposed autotuning framework to identify the best configurations, we achieve up to 91.59% performance improvement, up to 21.2% energy savings, and up to 37.84% EDP (energy delay product) improvement on up to 4096 nodes.

Autotuning↗

Energy–Performance Trade-offs in Privacy-Preserving Federated Learning on SmartNIC-Enabled HPC Systems

Federated learning (FL) is increasingly deployed on accelerator-rich high-performance computing (HPC) systems, yet the system-level energy cost of privacy-aware FL remains poorly understood, particularly across heterogeneous networking and server-placement options. We present a measurement-driven study of energy–performance trade-offs for FL on GH200-class nodes across three deployment configurations: CPU-Ethernet, CPU-InfiniBand (RDMA-capable), and a DPU-hosted FL server over InfiniBand using a BlueField-3 SmartNIC/DPU. Using NVIDIA FLARE (NVFLARE), we align node-level power telemetry with per-round timing extracted from NVFLARE logs to quantify time-to-solution (TTS), energy-to-solution (ETS), energy-delay product (EDP), and synchronization behavior for three transformer models (ALBERT, DistilBERT, BERT), trained with and without differential privacy (DP). We find that interconnect choice is the dominant driver of runtime and energy: host-managed InfiniBand consistently reduces communication overhead versus Ethernet, yielding lower TTS/ETS/EDP. In contrast, in our NVFLARE deployment, placing the FL server on the DPU does not consistently match CPU-InfiniBand performance and can be slower—especially for larger models—highlighting that server placement alone is not sufficient to guarantee end-to-end gains. Finally, under our fixed-round protocol, DP increases per-round cost and runtime variance; ETS increases largely in proportion to TTS because average node power remains relatively stable across configurations.

Kotevska, Olivera [ORNL] (ORCID:0000000316772243)↗

Integrating ytopt and libEnsemble to autotune OpenMC

Ytopt is a Python machine-learning-based autotuning software package developed within the ECP PROTEAS-TUNE project. The ytopt software adopts an asynchronous search framework that consists of sampling a small number of input parameter configurations and progressively fitting a surrogate model over the input-output space until exhausting the user-defined maximum number of evaluations or the wall-clock time. libEnsemble is a Python toolkit for coordinating workflows of asynchronous and dynamic ensembles of calculations across massively parallel resources developed within the ECP PETSc/TAO project. libEnsemble helps users take advantage of massively parallel resources to solve design, decision, and inference problems and expands the class of problems that can benefit from increased parallelism. In this paper we present our methodology and framework to integrate ytopt and libEnsemble to take advantage of massively parallel resources to accelerate the autotuning process. Specifically, we focus on using the proposed framework to autotune the ECP ExaSMR application OpenMC, an open source Monte Carlo particle transport code. OpenMC has seven tunable parameters some of which have large ranges such as the number of particles in-flight, which is in the range of 100,000 to 8 million, with its default setting of 1 million. Setting the proper combination of these parameter values to achieve the best performance is extremely time-consuming. Therefore, we apply the proposed framework to autotune the MPI/OpenMP offload version of OpenMC based on a user-defined metric such as the figure of merit (FoM) (particles/s) or energy efficiency energy-delay product (EDP) on Crusher at Oak Ridge Leadership Computing Facility. In conclusion, the experimental results show that we achieve the improvement up to 29.49% in FoM and up to 30.44% in EDP.

Autotuning↗

An automated and portable method for selecting an optimal GPU frequency

Power consumption poses a significant challenge in current and emerging graphics processing unit (GPU) enabled high-performance computing systems. In modern GPUs, dynamic voltage frequency scaling (DVFS) appears to be a reliable control to regulate power consumption and performance. However, the DVFS design space is large - hence, brute-force approaches are infeasible to select the optimal frequency. Furthermore, no single frequency can be universally optimal for applications with varying computational intensities. Thus, the application's complexity and the availability of a wide range of frequency settings are a challenge in selecting the optimal frequency configuration for a given GPU workload. To that end, this paper proposes a systematic approach that consists of three steps. The feature characterization study identifies the fine-grain GPU utilization metrics that influence the power consumption and execution time of a given workload. To understand the performance, power, and energy consumption behaviors of a workload across GPU's DVFS design space, we derived analytical power and performance models using the identified fine-grain features. Here, it is shown that the same set of GPU utilization metrics can estimate both the power consumption and execution time while being agnostic of changes to frequency and input sizes. Applying a power control with the single objective of reducing power may cause performance degradation, leading to more energy consumption. A multi-objective approach is proposed to select the optimal GPU DVFS configuration for a workload that reduces power consumption with negligible degradation in performance. The evaluation was conducted using SPEC ACCEL benchmarks and three real applications - NAMD LAMMPS, and LSTM on NVIDIA GV100, GA100, and AMD MI210 GPUs. On average, real applications showed 29.6% energy savings with a performance loss of 5.2% on GA100 and 22.6% energy savings with a performance loss of 4.7% on GV100. Moreover, the proposed models are portable to real applications, GPU architectures, and vendors, and require metric collection at only the default frequency rather than all supported DVFS configurations. Additionally, we conducted a comparison between our models and the GPU assembly instructions (PTX)-based static models. The results revealed a significant reduction in the average error rates, with a decrease from 19.7% to 3.1% for power models and from 29.4% to 5.2% for performance models.

97 MATHEMATICS AND COMPUTING↗

A 3D Implementation of Convolutional Neural Network for Fast Inference

Low latency inference has many applications in edge machine learning. In this paper, we present a run-time configurable convolutional neural network (CNN) inference ASIC design for low-latency edge machine learning. By implementing a 5-stage pipelined CNN inference model in a 3D ASIC technology, we demonstrate that the model distributed on two dies utilizing face-to-face (F2F) 3D integration achieves superior performance. Our experimental results show that the design based on 3D integration achieves 43% better energy-delay product when compared to the traditional 2D technology.

Miniskar, Narasinga Rao↗

XploreNAS : Explore Adversarially Robust and Hardware-efficient Neural Architectures for Non-ideal Xbars

Compute In-Memory platforms such as memristive crossbars are gaining focus as they facilitate acceleration of Deep Neural Networks (DNNs) with high area and compute efficiencies. However, the intrinsic non-idealities associated with the analog nature of computing in crossbars limits the performance of the deployed DNNs. Furthermore, DNNs are shown to be vulnerable to adversarial attacks leading to severe security threats in their large-scale deployment. Thus, finding adversarially robust DNN architectures for non-ideal crossbars is critical to the safe and secure deployment of DNNs on the edge. This work proposes a two-phase algorithm-hardware co-optimization approach called XploreNAS that searches for hardware efficient and adversarially robust neural architectures for non-ideal crossbar platforms. We use the one-shot Neural Architecture Search approach to train a large Supernet with crossbar-awareness and sample adversarially robust Subnets therefrom, maintaining competitive hardware efficiency. Our experiments on crossbars with benchmark datasets (SVHN, CIFAR10, CIFAR100) show up to ~8–16% improvement in the adversarial robustness of the searched Subnets against a baseline ResNet-18 model subjected to crossbar-aware adversarial training. We benchmark our robust Subnets for Energy-Delay-Area-Products (EDAPs) using the Neurosim tool and find that with additional hardware efficiency–driven optimizations, the Subnets attain ~1.5–1.6× lower EDAPs than ResNet-18 baseline.

97 MATHEMATICS AND COMPUTING↗

First Observation of the β3αp Decay of 13 O via β-Delayed Charged-Particle Spectroscopy

The β-delayed proton decay of 13 O has previously been studied, but the direct observation of β-delayed 3⁢α⁢p decay has not been reported. Rare 3⁢α⁢p events from the decay of excited states in 13 N* provide a sensitive probe of cluster configurations in 13 N*. To measure the low-energy products following β-delayed 3⁢αp decay, the Texas Active Target (TexAT) time projection chamber was employed using the one-at-a-time β-delayed charged-particle spectroscopy technique at the Cyclotron Institute, Texas A&M University. A total of 1.9 × 10 5 13 O implantations were made inside the TexAT time projection chamber. Furthermore, a total of 149 3⁢αp events were observed, yielding a β-delayed 3⁢αp branching ratio of 0.078(6)%. Four previously unknown α-decaying excited states were observed in 13 N at 11.3, 12.4, 13.1, and 13.7 MeV decaying via the 3⁢α + p channel.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Real-time scattering in Ising field theory using matrix product states

We study scattering in Ising field theory (IFT) using matrix product states and the time-dependent variational principle. IFT is a one-parameter family of strongly coupled nonintegrable quantum field theories in 1+1 dimensions, interpolating between massive free fermion theory and Zamolodchikov's integrable massive 𝐸 8 theory. Particles in IFT may scatter either elastically or inelastically. In the postcollision wave function, particle tracks from all final-state channels occur in superposition; processes of interest can be isolated by projecting the wave function onto definite particle sectors, or by evaluating energy density correlation functions. Using numerical simulations we determine the time delay of elastic scattering and the probability of inelastic particle production as a function of collision energy. We also study the mass and width of the lightest resonance near the 𝐸 8 point in detail. Close to both the free fermion and 𝐸 8 theories, our results for both elastic and inelastic scattering are in good agreement with expectations from form-factor perturbation theory. Using numerical computations to go beyond the regime accessible by perturbation theory, we find that the high-energy behavior of the two-to-two particle scattering probability in IFT is consistent with a conjecture of Zamolodchikov. Our results demonstrate the efficacy of tensor-network methods for simulating the real-time dynamics of strongly coupled quantum field theories in 1+1 dimensions.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Production and suppression of delayed light in NaI(Tl) scintillators

Here we investigate a hypothesis that energy accumulation and the subsequent release in NaI(Tl) may lead to pulselike events in the few-keV energy regime, a phenomenon suggested by the crystal manufacturing company Saint-Gobain, who provided the crystals for DAMA/LIBRA. While we observed delayed long-lasting (days) light emission in a 3-inch NaI(Tl) crystal after exposing it to UV light, the delayed light consists primarily of single photons that are uncorrelated with each other. We also observe delayed light emission in NaI(Tl) following gamma radiation and large ionization events like cosmic-ray muons. We found that irradiating the crystal with red light after UV exposure significantly suppressed delayed photon emissions.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A flight research program to develop airborne systems for improved terminal area operations

The research program considered is concerned with the solution of operational problems for the approximate time period from 1980 to 2000. The problems are related to safety, weather effects, congestion, energy conservation, noise, atmospheric pollution, and the loss in productivity caused by delays, diversions, and schedule stretchouts. The terminal configured vehicle (TCV) program is to develop advanced flight-control capability. The various aspects of the TCV program are discussed, giving attention to avionics equipment, the piloted simulator, terminal-area environment simulation, the Wallops research facility, flight procedures, displays and human factors, flight activities, and questions of vortex-wake reduction and tracking.

Reeder, J. P.↗

Consideration of memory of spin and parity in the fissioning compound nucleus by applying the Hauser-Feshbach fission fragment decay model to photonuclear reactions

Prompt and β-delayed fission observables, such as the average number of prompt and delayed neutrons, the independent and cumulative fission product yields, and the prompt γ-ray energy spectra for the photonuclear reactions on 235,238 U and 239 Pu, are calculated with the Hauser-Feshbach fission fragment decay (HF 3 ⁢D) model and compared with available experimental data. Further, in the analysis of neutron-induced fission reactions to the case of photo-induced fission, an excellent reproduction of the delayed neutron yields supports a traditional assumption that the photo fission might be similar to the neutron-induced fission at the same excitation energies regardless of the spin and parity of the fissioning systems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

On the Nature of QPO Phase Lags in Black Hole Candidates

Observations of quasi-periodic oscillations (QPOs) in X-ray binaries hold a key to understanding many aspects of these enigmatic systems. Complex appearance of the Fourier phase lags related to QPOs is one of the most puzzling observational effects in accreting black holes. In this Letter we show that QPO properties, including phase lags, can be explained in a framework of a simple scenario, where the oscillating media provides a feedback on the emerging spectrum. We demonstrate that the QPO waveform is presented by the product of a perturbation and a time delayed response factors, where the response is energy dependent. The essential property of this effect is its non-linear and multiplicative nature. Our multiplicative reverberation model successfully describes the QPO components in energy dependent power spectra as well as the appearance of the phase lags between signal in different energy bands. We apply our model to QPOs observed by Rossi X-ray Timing Explorer in BH candidate XTE J1550-564. We briefly discuss the implications of the observed energy dependence of the QPO reverberation times and amplitudes to the nature of the power law spectral component and its variability.

Shaposhnikov, Nikolai↗

Fission Product Yield Modeling and Evaluation

Although independent and cumulative fission product yields have been a part of evaluated libraries for decades, there have been few updates over the years. The fission product yield sub-library in the ENDF/B-VIII.0 library is still largely based on the evaluation of England and Rider from the mid-90’s, with only more recent updates to the energy dependence of 239 Pu below 2 MeV and fixes to isomeric states and missing fission products. Over the past several years, there have been a wealth of new measurements of independent and cumulative fission product yields, particularly those with short half-lives, and there have been significant improvements in the modeling of prompt and delayed fission observables. Here, we describe recent progress in the improvement of fission product yield calculations, using the BeoH code and the underlying Hauser Fesh-bach Fission Fragment Decay (HF 3 D) model, developed at Los Alamos National Laboratory. We will describe our recent calculations for consistent prompt and delayed fission observables for major and minor actinides, including new work investigating isomeric ratios. We will detail the ongoing evaluation process for energy-dependent fission product yields from thermal up to 20 MeV incident neutron energy and some validation work that has been performed for these new fission product yield calculations. Additionally, we will discuss future perspectives of this work, highlighting the need for additional data.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

The time-delay spectrum of GX 5-1 in its horizontal branch

Using a cross-spectral technique we investigate time delays between intensity variations of GX 5-1 in 10 X-ray spectral channels. The data were taken during a 1989 Ginga observation during which the source was in its horizontal-branch spectral state. We develope a new method to measure 'time-delay spectra' in fixed Fourier frequency ranges and use it to determine the energy and intensity dependence of time delays in the low-frequency noise (nu less than 2 Hz), the horizontal branch quasi-periodic oscillations (QPO), and the QPO second harmonic. These are the first time-delay spectra of a Z-source in its horizontal branch, and the first detection of time delays in the second harmonic. We consider two mechanisms for the production of the time lags: Comptonization and evolving shots. We perform Monte Carlo simulations of Compton scattering in a homogeneous, isotropic, central corona and show that it qualitatively explain the observed energy and time-delay spectra, but that it cannot explain the differences in the QPO first and second harmomnic time-delay spectra, nor the observed dependence of the QPO fractional rms variability upon energy. We consider implications of our results for millisecond pulsar searches in low-mass X-ray binaries.

Vaughan, B.↗

Computational Analysis of a Chevron Nozzle Uniquely Tailored for Propulsion Airframe Aeroacoustics

A computational flow field and predicted jet noise source analysis is presented for asymmetrical fan chevrons on a modern separate flow nozzle at take off conditions. The propulsion airframe aeroacoustic asymmetric fan nozzle is designed with an azimuthally varying chevron pattern with longer chevrons close to the pylon. A baseline round nozzle without chevrons and a reference nozzle with azimuthally uniform chevrons are also studied. The intent of the asymmetric fan chevron nozzle was to improve the noise reduction potential by creating a favorable propulsion airframe aeroacoustic interaction effect between the pylon and chevron nozzle. This favorable interaction and improved noise reduction was observed in model scale tests and flight test data and has been reported in other studies. The goal of this study was to identify the fundamental flow and noise source mechanisms. The flow simulation uses the asymptotically steady, compressible Reynolds averaged Navier-Stokes equations on a structured grid. Flow computations are performed using the parallel, multi-block, structured grid code PAB3D. Local noise sources were mapped and integrated computationally using the Jet3D code based upon the Lighthill Acoustic Analogy with anisotropic Reynolds stress modeling. In this study, trends of noise reduction were correctly predicted. Jet3D was also utilized to produce noise source maps that were then correlated to local flow features. The flow studies show that asymmetry of the longer fan chevrons near the pylon work to reduce the strength of the secondary flow induced by the pylon itself, such that the asymmetric merging of the fan and core shear layers is significantly delayed. The effect is to reduce the peak turbulence kinetic energy and shift it downstream, reducing overall noise production. This combined flow and noise prediction approach has yielded considerable understanding of the physics of a fan chevron nozzle designed to include propulsion airframe aeroacoustic interaction effects.

Massey, Steven J.↗

Realistic application of short-lived fission product delayed neutron, gamma-ray analysis for simultaneous nondestructive trace quantification of U, Pu mixtures on cellulose swipes

Detection and characterization of fissile traces are of interest to the international nuclear nonproliferation community, including the International Atomic Energy Agency. Pre-inspection check samples are analyzed by neutron activation analysis at the High Flux Isotope Reactor operated by the Oak Ridge National Laboratory under the umbrella of the IAEA Network of Analytical Laboratories. The simultaneous quantification of U and Pu mixtures was accomplished using the combined delayed neutron (DN) delayed gamma-ray (DG) method to analyze cellulose swipes with actinide loading <1ng in a blind field trial. The total fissile quantity was measured by the DN counts and the relative proportions of U, Pu were determined by calibration of the 104 Tc / 141 Ba fission product count ratio using known mixtures. The DNDG method demonstrated high accuracy in flagging the presence of 239 Pu in uranium down to <100 pg mass loading. In conclusion, peak significance tests helped to control false positive Pu flagging and simultaneous quantification of U and Pu loading was accomplished on samples that passed the significance tests.

36 MATERIALS SCIENCE↗

Addition of HCl to the double-pulse copper chloride laser

Addition of small amounts of hydrogen chloride to the buffer gas of a double-pulse CuCl laser causes an increase in the production of copper atoms in the ground state. A maximum laser energy increase of 15% was observed and the span of delay times for which laser action occurred increased.

Vetter, A. A.↗