Search NASA⌕ Search

SEARCH · Search NASA

Results for “Energy delay product”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Bridging the Gap: User-Centric Energy Monitoring for Policy-Driven Application Optimization in HPC Data Centers

Application energy optimization in HPC data centers face two critical gaps. Systematic methodologies that connect data center policies to application decisions and accessible monitoring tools that enable data-driven optimization. We address both gaps through two complementary pillars. First, we present a methodology based on extended weighted Energy Delay Product (EDP) to translate data center operational priorities and integrate energy considerations into the energy optimization workflow which starts from continuous monitoring through targeted optimization. Second, we present a user-space monitoring tool, Omnistat, that enables this methodology by providing developers with direct access to actionable energy telemetry. Through deployment on the Frontier supercomputer and case studies exploring performance-energy trade-offs, we show how these pillars help energy as an integral optimization target for developers as active participants in data center efficiency.

Shin, Woong [ORNL] (ORCID:0000000172077814)↗

ytopt: Autotuning Scientific Applications for Energy Efficiency at Large Scales

As we enter the exascale computing era, efficiently utilizing power and optimizing the performance of scientific applications under power and energy constraints has become critical and challenging. We propose a low-overhead autotuning framework to autotune performance and energy for various hybrid MPI/OpenMP scientific applications at large scales and to explore the tradeoffs between application runtime and power/energy for energy efficient application execution, then use this framework to autotune four ECP proxy applications—XSBench, AMG, SWFFT, and SW4lite. Our approach uses Bayesian optimization with a Random Forest surrogate model to effectively search parameter spaces with up to 6 million different configurations on two large-scale HPC production systems, Theta at Argonne National Laboratory and Summit at Oak Ridge National Laboratory. The experimental results show that our autotuning framework at large scales has low overhead and achieves good scalability. Using the proposed autotuning framework to identify the best configurations, we achieve up to 91.59% performance improvement, up to 21.2% energy savings, and up to 37.84% EDP (energy delay product) improvement on up to 4096 nodes.

Autotuning↗

Energy–Performance Trade-offs in Privacy-Preserving Federated Learning on SmartNIC-Enabled HPC Systems

Federated learning (FL) is increasingly deployed on accelerator-rich high-performance computing (HPC) systems, yet the system-level energy cost of privacy-aware FL remains poorly understood, particularly across heterogeneous networking and server-placement options. We present a measurement-driven study of energy–performance trade-offs for FL on GH200-class nodes across three deployment configurations: CPU-Ethernet, CPU-InfiniBand (RDMA-capable), and a DPU-hosted FL server over InfiniBand using a BlueField-3 SmartNIC/DPU. Using NVIDIA FLARE (NVFLARE), we align node-level power telemetry with per-round timing extracted from NVFLARE logs to quantify time-to-solution (TTS), energy-to-solution (ETS), energy-delay product (EDP), and synchronization behavior for three transformer models (ALBERT, DistilBERT, BERT), trained with and without differential privacy (DP). We find that interconnect choice is the dominant driver of runtime and energy: host-managed InfiniBand consistently reduces communication overhead versus Ethernet, yielding lower TTS/ETS/EDP. In contrast, in our NVFLARE deployment, placing the FL server on the DPU does not consistently match CPU-InfiniBand performance and can be slower—especially for larger models—highlighting that server placement alone is not sufficient to guarantee end-to-end gains. Finally, under our fixed-round protocol, DP increases per-round cost and runtime variance; ETS increases largely in proportion to TTS because average node power remains relatively stable across configurations.

Kotevska, Olivera [ORNL] (ORCID:0000000316772243)↗

Integrating ytopt and libEnsemble to autotune OpenMC

Ytopt is a Python machine-learning-based autotuning software package developed within the ECP PROTEAS-TUNE project. The ytopt software adopts an asynchronous search framework that consists of sampling a small number of input parameter configurations and progressively fitting a surrogate model over the input-output space until exhausting the user-defined maximum number of evaluations or the wall-clock time. libEnsemble is a Python toolkit for coordinating workflows of asynchronous and dynamic ensembles of calculations across massively parallel resources developed within the ECP PETSc/TAO project. libEnsemble helps users take advantage of massively parallel resources to solve design, decision, and inference problems and expands the class of problems that can benefit from increased parallelism. In this paper we present our methodology and framework to integrate ytopt and libEnsemble to take advantage of massively parallel resources to accelerate the autotuning process. Specifically, we focus on using the proposed framework to autotune the ECP ExaSMR application OpenMC, an open source Monte Carlo particle transport code. OpenMC has seven tunable parameters some of which have large ranges such as the number of particles in-flight, which is in the range of 100,000 to 8 million, with its default setting of 1 million. Setting the proper combination of these parameter values to achieve the best performance is extremely time-consuming. Therefore, we apply the proposed framework to autotune the MPI/OpenMP offload version of OpenMC based on a user-defined metric such as the figure of merit (FoM) (particles/s) or energy efficiency energy-delay product (EDP) on Crusher at Oak Ridge Leadership Computing Facility. In conclusion, the experimental results show that we achieve the improvement up to 29.49% in FoM and up to 30.44% in EDP.

Autotuning↗

An automated and portable method for selecting an optimal GPU frequency

Power consumption poses a significant challenge in current and emerging graphics processing unit (GPU) enabled high-performance computing systems. In modern GPUs, dynamic voltage frequency scaling (DVFS) appears to be a reliable control to regulate power consumption and performance. However, the DVFS design space is large - hence, brute-force approaches are infeasible to select the optimal frequency. Furthermore, no single frequency can be universally optimal for applications with varying computational intensities. Thus, the application's complexity and the availability of a wide range of frequency settings are a challenge in selecting the optimal frequency configuration for a given GPU workload. To that end, this paper proposes a systematic approach that consists of three steps. The feature characterization study identifies the fine-grain GPU utilization metrics that influence the power consumption and execution time of a given workload. To understand the performance, power, and energy consumption behaviors of a workload across GPU's DVFS design space, we derived analytical power and performance models using the identified fine-grain features. Here, it is shown that the same set of GPU utilization metrics can estimate both the power consumption and execution time while being agnostic of changes to frequency and input sizes. Applying a power control with the single objective of reducing power may cause performance degradation, leading to more energy consumption. A multi-objective approach is proposed to select the optimal GPU DVFS configuration for a workload that reduces power consumption with negligible degradation in performance. The evaluation was conducted using SPEC ACCEL benchmarks and three real applications - NAMD LAMMPS, and LSTM on NVIDIA GV100, GA100, and AMD MI210 GPUs. On average, real applications showed 29.6% energy savings with a performance loss of 5.2% on GA100 and 22.6% energy savings with a performance loss of 4.7% on GV100. Moreover, the proposed models are portable to real applications, GPU architectures, and vendors, and require metric collection at only the default frequency rather than all supported DVFS configurations. Additionally, we conducted a comparison between our models and the GPU assembly instructions (PTX)-based static models. The results revealed a significant reduction in the average error rates, with a decrease from 19.7% to 3.1% for power models and from 29.4% to 5.2% for performance models.

97 MATHEMATICS AND COMPUTING↗

A 3D Implementation of Convolutional Neural Network for Fast Inference

Low latency inference has many applications in edge machine learning. In this paper, we present a run-time configurable convolutional neural network (CNN) inference ASIC design for low-latency edge machine learning. By implementing a 5-stage pipelined CNN inference model in a 3D ASIC technology, we demonstrate that the model distributed on two dies utilizing face-to-face (F2F) 3D integration achieves superior performance. Our experimental results show that the design based on 3D integration achieves 43% better energy-delay product when compared to the traditional 2D technology.

Miniskar, Narasinga Rao↗

XploreNAS : Explore Adversarially Robust and Hardware-efficient Neural Architectures for Non-ideal Xbars

Compute In-Memory platforms such as memristive crossbars are gaining focus as they facilitate acceleration of Deep Neural Networks (DNNs) with high area and compute efficiencies. However, the intrinsic non-idealities associated with the analog nature of computing in crossbars limits the performance of the deployed DNNs. Furthermore, DNNs are shown to be vulnerable to adversarial attacks leading to severe security threats in their large-scale deployment. Thus, finding adversarially robust DNN architectures for non-ideal crossbars is critical to the safe and secure deployment of DNNs on the edge. This work proposes a two-phase algorithm-hardware co-optimization approach called XploreNAS that searches for hardware efficient and adversarially robust neural architectures for non-ideal crossbar platforms. We use the one-shot Neural Architecture Search approach to train a large Supernet with crossbar-awareness and sample adversarially robust Subnets therefrom, maintaining competitive hardware efficiency. Our experiments on crossbars with benchmark datasets (SVHN, CIFAR10, CIFAR100) show up to ~8–16% improvement in the adversarial robustness of the searched Subnets against a baseline ResNet-18 model subjected to crossbar-aware adversarial training. We benchmark our robust Subnets for Energy-Delay-Area-Products (EDAPs) using the Neurosim tool and find that with additional hardware efficiency–driven optimizations, the Subnets attain ~1.5–1.6× lower EDAPs than ResNet-18 baseline.

97 MATHEMATICS AND COMPUTING↗

First Observation of the β3αp Decay of 13 O via β-Delayed Charged-Particle Spectroscopy

The β-delayed proton decay of 13 O has previously been studied, but the direct observation of β-delayed 3⁢α⁢p decay has not been reported. Rare 3⁢α⁢p events from the decay of excited states in 13 N* provide a sensitive probe of cluster configurations in 13 N*. To measure the low-energy products following β-delayed 3⁢αp decay, the Texas Active Target (TexAT) time projection chamber was employed using the one-at-a-time β-delayed charged-particle spectroscopy technique at the Cyclotron Institute, Texas A&M University. A total of 1.9 × 10 5 13 O implantations were made inside the TexAT time projection chamber. Furthermore, a total of 149 3⁢αp events were observed, yielding a β-delayed 3⁢αp branching ratio of 0.078(6)%. Four previously unknown α-decaying excited states were observed in 13 N at 11.3, 12.4, 13.1, and 13.7 MeV decaying via the 3⁢α + p channel.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Real-time scattering in Ising field theory using matrix product states

We study scattering in Ising field theory (IFT) using matrix product states and the time-dependent variational principle. IFT is a one-parameter family of strongly coupled nonintegrable quantum field theories in 1+1 dimensions, interpolating between massive free fermion theory and Zamolodchikov's integrable massive 𝐸 8 theory. Particles in IFT may scatter either elastically or inelastically. In the postcollision wave function, particle tracks from all final-state channels occur in superposition; processes of interest can be isolated by projecting the wave function onto definite particle sectors, or by evaluating energy density correlation functions. Using numerical simulations we determine the time delay of elastic scattering and the probability of inelastic particle production as a function of collision energy. We also study the mass and width of the lightest resonance near the 𝐸 8 point in detail. Close to both the free fermion and 𝐸 8 theories, our results for both elastic and inelastic scattering are in good agreement with expectations from form-factor perturbation theory. Using numerical computations to go beyond the regime accessible by perturbation theory, we find that the high-energy behavior of the two-to-two particle scattering probability in IFT is consistent with a conjecture of Zamolodchikov. Our results demonstrate the efficacy of tensor-network methods for simulating the real-time dynamics of strongly coupled quantum field theories in 1+1 dimensions.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Production and suppression of delayed light in NaI(Tl) scintillators

Here we investigate a hypothesis that energy accumulation and the subsequent release in NaI(Tl) may lead to pulselike events in the few-keV energy regime, a phenomenon suggested by the crystal manufacturing company Saint-Gobain, who provided the crystals for DAMA/LIBRA. While we observed delayed long-lasting (days) light emission in a 3-inch NaI(Tl) crystal after exposing it to UV light, the delayed light consists primarily of single photons that are uncorrelated with each other. We also observe delayed light emission in NaI(Tl) following gamma radiation and large ionization events like cosmic-ray muons. We found that irradiating the crystal with red light after UV exposure significantly suppressed delayed photon emissions.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Consideration of memory of spin and parity in the fissioning compound nucleus by applying the Hauser-Feshbach fission fragment decay model to photonuclear reactions

Prompt and β-delayed fission observables, such as the average number of prompt and delayed neutrons, the independent and cumulative fission product yields, and the prompt γ-ray energy spectra for the photonuclear reactions on 235,238 U and 239 Pu, are calculated with the Hauser-Feshbach fission fragment decay (HF 3 ⁢D) model and compared with available experimental data. Further, in the analysis of neutron-induced fission reactions to the case of photo-induced fission, an excellent reproduction of the delayed neutron yields supports a traditional assumption that the photo fission might be similar to the neutron-induced fission at the same excitation energies regardless of the spin and parity of the fissioning systems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Fission Product Yield Modeling and Evaluation

Although independent and cumulative fission product yields have been a part of evaluated libraries for decades, there have been few updates over the years. The fission product yield sub-library in the ENDF/B-VIII.0 library is still largely based on the evaluation of England and Rider from the mid-90’s, with only more recent updates to the energy dependence of 239 Pu below 2 MeV and fixes to isomeric states and missing fission products. Over the past several years, there have been a wealth of new measurements of independent and cumulative fission product yields, particularly those with short half-lives, and there have been significant improvements in the modeling of prompt and delayed fission observables. Here, we describe recent progress in the improvement of fission product yield calculations, using the BeoH code and the underlying Hauser Fesh-bach Fission Fragment Decay (HF 3 D) model, developed at Los Alamos National Laboratory. We will describe our recent calculations for consistent prompt and delayed fission observables for major and minor actinides, including new work investigating isomeric ratios. We will detail the ongoing evaluation process for energy-dependent fission product yields from thermal up to 20 MeV incident neutron energy and some validation work that has been performed for these new fission product yield calculations. Additionally, we will discuss future perspectives of this work, highlighting the need for additional data.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Accelerating technology development to monitor and minimize effects from land‐based wind energy on birds and bats

While wind energy is a key sector of domestic energy production for the United States, operation of wind turbines directly and indirectly adversely affects certain species of birds and bats. The cumulative effect of wind turbine strikes can have both biological and regulatory consequences, and, in some cases, delay permitting and construction or affect ongoing operations. Technology can help quantify and minimize these effects, but the pace of development, acceptance, and adoption of technological solutions is slow. Although adopting cost‐effective technologies may reduce negative effects on wildlife and help achieve both energy production and conservation goals, consensus is lacking among developers, regulators, and the conservation community regarding how to define technology effectiveness and acceptance and how to develop a standardized process for doing so. Removing barriers to technology advancement requires deviating from the status quo. Changes include 1) creating incentives to mitigate impacts, 2) establishing options for research as mitigation, 3) rethinking how research is funded, 4) increasing stakeholder coordination, and 5) increasing the efficiency of research and development. We recommend the creation of a national framework to establish clear criteria and protocols for technology evaluation and adoption.

17 WIND ENERGY↗

Realistic application of short-lived fission product delayed neutron, gamma-ray analysis for simultaneous nondestructive trace quantification of U, Pu mixtures on cellulose swipes

Detection and characterization of fissile traces are of interest to the international nuclear nonproliferation community, including the International Atomic Energy Agency. Pre-inspection check samples are analyzed by neutron activation analysis at the High Flux Isotope Reactor operated by the Oak Ridge National Laboratory under the umbrella of the IAEA Network of Analytical Laboratories. The simultaneous quantification of U and Pu mixtures was accomplished using the combined delayed neutron (DN) delayed gamma-ray (DG) method to analyze cellulose swipes with actinide loading <1ng in a blind field trial. The total fissile quantity was measured by the DN counts and the relative proportions of U, Pu were determined by calibration of the 104 Tc / 141 Ba fission product count ratio using known mixtures. The DNDG method demonstrated high accuracy in flagging the presence of 239 Pu in uranium down to <100 pg mass loading. In conclusion, peak significance tests helped to control false positive Pu flagging and simultaneous quantification of U and Pu loading was accomplished on samples that passed the significance tests.

36 MATERIALS SCIENCE↗

A control-inspired approach for energy transition planning under uncertainty

As the global carbon footprint continues to grow, many countries are implementing carbon emission reduction policies which have incentivized the expansion of low-carbon and renewable technologies. However, the speed and scale of deployment falls short of that needed to meet climate goals. Energy system models serve as key tools for guiding investment decisions and helping policymakers evaluate the effects of various policies on the development of an energy system. This study focuses on the energy system of the United States and builds upon prior work by incorporating more geographic granularity to account for the trade of commodities and addresses transmission congestion through electricity price adjustments. Furthermore, real-world characteristics, such as delays in constructing new liquid fuel production and electricity generation facilities, are integrated using a sequential decision-making approach that better reflects how decisions can be updated as uncertainties unfold. Results demonstrate that stochastic programming combined with sequential decision-making produces energy transition pathways that are robust to multiple uncertain futures. Additionally, considering real-world characteristics significantly impacts the deployment of renewable technologies and the ability to meet carbon emission reduction goals while also reliably meeting demand. These findings highlight the importance of accounting for uncertainty and real-world characteristics to avoid overly optimistic projections in energy system planning.

energy systems↗

Modeling Systems’ Disruption and Social Acceptance—A Proof-of-Concept Leveraging Reinforcement Learning

As the need for a just and equitable energy transition accelerates, disruptive clean energy technologies are becoming more visible to the public. Clean energy technologies, such as solar photovoltaics and wind power, can substantially contribute to a more sustainable world and have been around for decades. However, the fast pace at which they are projected to be deployed in the United States (US) and the world poses numerous technical and nontechnical challenges, such as in terms of their integration into the electricity grid, public opposition and competition for land use. For instance, as more land-based wind turbines are built across the US, contention risks may become more acute. This article presents a methodology based on reinforcement learning (RL) that minimizes contention risks and maximizes renewable energy production during siting decisions. As a proof-of-concept, the methodology is tested on a case study of wind turbine siting in Illinois during the 2022–2035 period. Results show that using RL halves potential delays due to contention compared to a random decision process. This approach could be further developed to study the acceptance of offshore wind projects or other clean energy technologies.

17 WIND ENERGY↗

Continuing Analysis of Charge Current Interactions in ANNIE

The Accelerator Neutrino Neutron Interaction Experiment (ANNIE) is a gadolinium-loaded water Cherenkov detector on the Fermilab Booster Neutrino Beam (BNB).Using νμ in the energy range of 500 to 1000 MeV, ANNIE is designed to measure final-state neutrons.In this poster we will cover the ongoing analysis work exploring Charged-Current neutrino interactions. Charged-Current (CC) νμ interactions in this regime have uncertainties in the relative contributions of quasielastic and res- onance production. Together with intranuclear final-state interactions (FSI) and missing hadronic energy, these interactions drive important biases in neutrino energy reconstruction. ANNIE mit- igates these effects by combining muon kinematics from the downstream Muon Range Detector (MRD) with neutron identification via delayed gamma cascades from thermal captures on gadolin- ium. We characterize CC0π and Δ samples by measuring neutron multiplicity versus event topology and reconstructed kinematics, comparing neutron-tagged data to interaction-model predictions to probe resonance production and pion FSI/absorption

Fleming, Dylon [UC, Davis; Fermilab] (ORCID:000000↗

Particle production by time-varying dark energy and the end of cosmic expansion

We consider various possible consequences of time-varying dark energy due to a quintessence scalar field whose energy density is partially converted to particles as the field evolves down its potential. This particle production acts as a source of thermal friction on the field that can make it difficult to distinguish whether dark energy is due to a radiating field rolling down a steep potential, a purely self-interacting field moving down a flatter potential, or a cosmological constant. By reducing the acceleration of the scalar field, thermal friction increases the amount of accelerated expansion and can cause a sizable bump in the quintessence equation of state. We take special interest in the case where a steep potential rapidly changes from positive to negative as the field evolves, resulting in the end of cosmic expansion and the beginning of contraction. Even in this case, we find that thermal friction lengthens the period of accelerated expansion and consequently delays the end of cosmic expansion, making it challenging to detect the impending transition to contraction using conventional cosmological tests. However, particle production can also provide alternative avenues for detection by generating a background of thermal dark radiation, partly comprised of neutrinos or other particles, whose energy density exceeds the remnant photon energy density.

cosmological neutrinos↗