Search NASA⌕ Search

SEARCH · Search NASA

Results for “FLOPS”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

34 records · Page 2

Hidden features in the OH-stretching spectra of amino acid decorated air–water interfaces

Chemical reactivity at the air–water interface is governed by the interfacial solvation of reactive species. For instance, during aqueous amino acid-based CO 2 absorption, water reorganizes around the reactive sites and couples dynamically with reaction pathways, facilitating the reaction. In this context, surface-sensitive vibrational sum-frequency generation (vSFG) spectroscopy can probe the OH stretch vibrations of interfacial water and determine the solvation structures around reactants and products, thereby furthering our understanding of the role of interfacial solvation. However, vSFG spectra of the air–water interface in the presence of charged species can be remarkably complex; key species-bound local water structures with distinct orientations may be hidden beneath prominent vSFG peaks arising from water–water hydrogen bonds and remain difficult to resolve. Here, we measure and compute vSFG spectra of the water OH stretch at air–water interfaces decorated with amino acids in their zwitterionic and anionic forms, as well as equimolar mixtures of these forms with bicarbonate. The latter represents post-CO 2 -absorption conditions. We find that computing depth- and frequency-dependent spectral densities—decomposed into contributions from water molecules hydrogen-bonded exclusively to other water molecules, exclusively to amines, exclusively to carboxylates, or shared between these polar/charged groups—is indispensable for accurate interpretation of the vSFG spectra. Key findings include orientational flip-flop in water sub-layers, strong carboxylate-water H-bonding, and water orientational ordering extending into the bulk aqueous phase induced by anionic amino acids. Here, this study provides a computational spectroscopic platform for improved understanding of interfacial solvation relevant to interfacial reactivity.

Air-water interface↗

Spin Decoherence Dynamics of Er 3+ in CeO 2 Films

Developing telecom-compatible spin-photon interfaces is essential towards scalable quantum networks. Erbium ions (Er 3+ ) exhibit a unique combination of a telecom (1.5 mu m) optical transition and an effective spin-1=2 ground state, but identifying a host that enables heterogeneous device integration while preserving long optical and spin coherence remains an open challenge. In this work, we explore the potential of Er 3+ :CeO 2 films on silicon and study the Er 3+ spin coherence, offering low nuclear spin density and the potential for on-chip integration. We demonstrate a 38.8 mu s spin coherence, which can be extended to 176.4 mu s with dynamical decoupling. Pairing experiments with cluster correlation expansion calculations, we identify spectral diffusion induced by bath Er 3+ spin flip-flops as the dominant decoherence mechanism and provide pathways to millisecond-scale coherence.

36 MATERIALS SCIENCE↗

Magnetism and magnetoelastic effect in the two-dimensional van der Waals multiferroic CuCrP 2 ⁢S 6

Here, we report a magnetic and neutron diffraction study on the ground state magnetism and field evolution of single-crystal van der Waals multiferroic CuCrP 2 ⁢S 6 . The ordered moments align along the 𝑏 axis in the "𝐴-type" antiferromagnetic configuration with a spin-flop transition along the same direction. Field application along 𝑎 introduces a smooth transition to a fully polarized ferromagnetic state via in-plane spin rotation. These findings resolve the ambiguity of the ground state magnetization direction in CuCrP 2 ⁢S 6 and uncover its field responses, providing a firm basis for future magnetoelectric study. A magnetoelastic coupling effect connecting the interlayer spacing and the magnetic order was further revealed, highlighting the out-of-plane strain as an effective control knob for tuning magnetism both in this system and in related van der Waals magnets.

Guo, Jiasen [Oak Ridge National Laboratory (ORNL),↗

Formation of a simple cubic antiferromagnet through charge ordering in a double Dirac material

Following the topological classification of electronic phases, interest has grown in materials with unique electronic or magnetic properties driven by topology and interactions. Here, in this study, we report that the topologically nontrivial mixed valent intermetalllic EuPd 3 S 4 undergoes long-range charge ordering at 𝑇 𝐶⁢𝑂 = 340 K wherein 𝐽 = 7/2 Eu 2+ and Van Vleck 𝐽 = 0 Eu 3+ ions on a body-centered-cubic lattice separate into two interpenetrating simple cubic sublattices. The reduced symmetry transmutes 8-fold double Dirac states into 4-fold Dirac states and leads to a simple cubic Heisenberg antiferromagnet with G-type antiferromagnetic order for 𝑇 <⁢ 𝑇 𝑁 = 2.85⁢(6) K. While time reversal symmetry is broken, its combination with nearest neighbor lattice translation can form a nonsymmorphic symmetry preserving the 4-fold Dirac point. The application of a magnetic field yields a spin flop transition at the lowest temperatures but, as a consequence of the extreme isotropy of Eu, that phase transition turns into a cross-over at higher temperatures. Our work demonstrates how charge order modifies topology in EuPd 3 S 4 and exposes an archetypal simple cubic Heisenberg antiferromagnet.

Berry, Tanya [Johns Hopkins Univ., Baltimore, MD (↗

Noncollinear spin order, field-induced transitions, and short-range correlations in Cu 4 ⁢SO 4 ⁢(OH) 6

We report a comprehensive study of spin-$\frac{1}{2}$ quantum magnet brochantite, Cu 4 ⁢SO 4 ⁢(OH) 6 , combining high-field thermodynamic measurements, polarized neutron diffraction, inelastic neutron scattering, and nuclear magnetic resonance spectroscopy. Using bulk magnetization and specific heat measurements we construct the magnetic 𝐻−𝑇 phase diagram for magnetic fields applied along main crystallographic directions up to 24.1 T. Polarized neutron diffraction reveals a noncollinear magnetic ground state confined to the 𝑎⁢𝑏 plane. A field-induced transition is observed for magnetic fields in the 𝑎⁢𝑏 plane whose critical field and character evolve continuously with field direction in the 𝑎⁢𝑏 plane. While the transition for 𝐻 ∥ 𝑏 might be a spin-flop-like transition, the one for 𝐻 ∥ 𝑎 resembles the short-range correlated state above 𝑇 N , suggesting a magnetic configuration related to the quasi-two-dimensional correlations in the 𝑏⁢𝑐 planes. The experimental results demonstrate that the ground-state and field-induced phases cannot be explained by a simple XXZ model with a single dominant exchange interaction and that weaker interchain and anisotropic interactions have to be taken into account. Our work establishes brochantite as a low-symmetry Cu-based quantum magnet with noncollinear order and complex field-induced behavior.

Prokhnenko, Oleksandr [Helmholtz-Zentrum Berlin (H↗

A Parallel Alternative for Energy-Efficient Neural Network Training and Inferencing

Energy efficiency of training and inferencing with large neural network models is a critical challenge facing the future of sustainable large-scale machine learning workloads. This paper introduces an alternative strategy, called phantom parallelism, to minimize the net energy consumption of traditional tensor (model) parallelism, the most energy-inefficient component of large neural network training. The approach is presented in the context of feed-forward network architectures as a preliminary, but comprehensive, proof-of-principle study of the proposed methodology. We derive new forward and backward propagation operators for phantom parallelism, implement them as custom autograd operations within an end-to-end phantom parallel training pipeline and compare its parallel performance and energy-efficiency against those of conventional tensor parallel training pipelines. Formal analyses that predict lower bandwidth and FLOP counts are presented with supporting empirical results on up to 256 GPUs that corroborate these gains. Experiments are shown to deliver ∼50% reduction in the energy consumed to train FFNs using the proposed phantom parallel approach when compared with conventional tensor parallel methods. Additionally, the proposed approach is shown to train smaller phantom models to the same model loss on smaller GPU counts as larger tensor parallel models on larger GPU counts offering the possibility for even greater energy savings.

Seal, Sudip [ORNL] (ORCID:0000000332330656)↗

Toward Energy-Efficient HPC: Insights from Power Profiling a Cloud-Resolving Earth System Model

Power is a fundamental constraint as supercomputing advances to exascale. Efficient operation within strict power budgets requires application-aware power management based on a detailed understanding of application-level power behavior. This work analyzes the Energy Exascale Earth System Model (E3SM) atmosphere component, SCREAM, on Perlmutter (NERSC) and Frontier (OLCF). We characterize power variation across inputs, concurrency levels, and power caps, evaluate the energy impact of code optimizations, and attribute energy within the code using a newly developed GPU energy model. Results show that SCREAM’s peak power remains stable during its core execution phase and decreases gradually as concurrency increases. Power capping experiments reveal a performance–energy "sweet spot". On Perlmutter, limiting GPU power to 50% of thermal design power (TDP) achieves up to 15% energy savings with a 7% performance penalty. On Frontier, a 40% TDP cap yields up to 10% energy savings with less than 10% performance loss. Code optimizations reduce SCREAM energy by shortening run time without increasing power. Modeling reveals a critical insight: data movement accounts for approximately 70% of SCREAM’s GPU energy. This fundamentally shifts the optimization focus from FLOPS to data transfer reduction for this class of applications, offering the most impactful strategy for improving energy efficiency. This work establishes a foundation for practical, application-aware power management at exascale.

Zhao, Zhengji [Lawrence Berkeley National Laborato↗

Laboratory tests of laser control of electron beams for future colliders

Laser-driven Compton backscattering (CBS) has been proposed as method for controlling the intensity of colliding bunches in the FCC-ee so as to avoid the flip-flop instability caused by intensity asymmetry in colliding bunches. Laser-based collimation has also been proposed as an indestructible collimator for high-intensity electron beams. We have initiated a laboratory-based test program of these concepts with the E344 experiment at FACET-II. In this paper, we describe simulations of laser–beam interactions at FACET-II and the relevant scaling for FCC-ee. We also describe the experimental setup and diagnostics that will be used to make the measurements at FACET-II.

accelerator physics↗

RingX: Scalable Parallel Attention for Long-Context Learning on HPC

The attention mechanism has become foundational for remarkable AI breakthroughs since the introduction of the Transformer, driving the demand for increasingly longer context to power frontier models such as large-scale reasoning language models and high-resolution image/video generators. However, its quadratic computational and memory complexities present substantial challenges. Current state-of-the-art parallel attention methods, such as ring attention, are widely adopted for long-context training but utilize a point-to-point communication strategy that fails to fully exploit the capabilities of modern HPC network architectures. In this work, we propose ringX, a scalable family of parallel attention methods optimized explicitly for HPC systems. By enhancing workload partitioning, refining communication patterns, and improving load balancing, ringX achieves up to 3.4 × speedup compared to conventional ring attention on the Frontier supercomputer. Optimized for both bi-directional and causal attention mechanisms, ringX demonstrates its effectiveness through training benchmarks of a Vision Transformer (ViT) on a climate dataset and a Generative Pre-Trained Transformer (GPT) model, Llama3 8B. Our method attains an end-to-end training speedup of approximately 1.5 × in both scenarios. To our knowledge, the achieved 38% model FLOPs utilization (MFU) for training Llama3 8B with a 1M-token sequence length on 4,096 GPUs represents one of the highest training efficiencies reported for long-context learning on HPC systems. Our code implementation is available at https://github.com/jqyin/ringX-attention.

Yin, Junqi [ORNL] (ORCID:0000000338435520)↗

ZFP: A compressed array representation for numerical computations

HPC trends favor algorithms and implementations that reduce data motion relative to FLOPS. We investigate the use of lossy compressed data arrays in place of traditional IEEE floating point arrays to store the primary data of calculations. Simulation is fundamentally an exercise in controlled approximation, and error introduced by finite-precision arithmetic (or lossy compression) is just one of several sources of error that need to be managed to ensure sufficient accuracy in a computed result. We describe ZFP, a compressed numerical format designed for in-memory storage of multidimensional arrays, and summarize theoretical results that demonstrate that the error of repeated lossy compression can be bounded and controlled. Furthermore, we establish a relationship between grid resolution and compression-induced errors and show that, contrary to conventional floating point, ZFP reduces finite-difference errors with finer grids. We present example calculations that demonstrate data reduction by 4x or more with negligible impact on solution accuracy. Our results further demonstrate several orders-of-magnitude increase in accuracy using ZFP over IEEE floating point and Posits for the same storage budget.

Lindstrom, Peter↗

Measuring the Energy Consumption and Efficiency of Deep Neural Networks: An Empirical Analysis and Design Recommendations

Addressing the "Red-AI" trend of rising energy consumption by large-scale neural networks, this study investigates the measured energy consumption of training various fully connected neural network architectures. We introduce the BUTTER-E dataset, an augmentation to the BUTTER Empirical Deep Learning dataset, containing energy consumption and performance data from 41,129 individual experimental runs spanning 30,582 distinct configurations: 13 datasets, 20 sizes (trainable parameters), 8 "shapes", and 14 depths on both CPUs and GPUs using node-level watt-meters. This dataset reveals the complex relationship between dataset size, network structure, and energy use. Our analysis uncovers a surprising, hardware-mediated non-linear relationship between energy efficiency and network design, challenging the assumption that reducing the number of parameters or FLOPs is the best way to achieve greater energy efficiency. We propose a straightforward and effective energy model that accounts for network size, computing, and memory hierarchy. Highlighting the need for cache-considerate algorithm development, we suggest a codesign approach to energy efficient network, algorithm, and hardware design. This work contributes to the fields of sustainable computing and Green AI, offering practical guidance for creating more energy-efficient neural networks and promoting sustainable AI.

97 MATHEMATICS AND COMPUTING↗

ExtremeMETA: High-speed Lightweight Image Segmentation Model by Remodeling Multi-channel Metamaterial Imagers

Deep neural networks (DNNs) have heavily relied on traditional computational units, such as CPUs and GPUs. However, this conventional approach brings significant computational burden, latency issues, and high power consumption, limiting their effectiveness. This has sparked the need for lightweight networks such as ExtremeC3Net. Meanwhile, there have been notable advancements in optical computational units, particularly with metamaterials, offering the exciting prospect of energy-efficient neural networks operating at the speed of light. Yet, the digital design of metamaterial neural networks (MNNs) faces precision, noise, and bandwidth challenges, limiting their application to intuitive tasks and low-resolution images. In this study, we proposed a large kernel lightweight segmentation model, ExtremeMETA. Based on ExtremeC3Net, our proposed model, ExtremeMETA maximized the ability of the first convolution layer by exploring a larger convolution kernel and multiple processing paths. With the large kernel convolution model, we extended the optic neural network application boundary to the segmentation task. To further lighten the computation burden of the digital processing part, a set of model compression methods was applied to improve model efficiency in the inference stage. The experimental results on three publicly available datasets demonstrated that the optimized efficient design improved segmentation performance from 92.45 to 95.97 on mIoU while reducing computational FLOPs from 461.07 MMacs to 166.03 MMacs. The large kernel lightweight model ExtremeMETA showcased the hybrid design’s ability on complex tasks.

large convolution kernel↗

Entropy Analysis of FPGA Interconnect and Switch Matrices for Physical Unclonable Functions

Random variations in microelectronic circuit structures represent the source of entropy for physical unclonable functions (PUFs). In this paper, we investigate delay variations that occur through the routing network and switch matrices of a field-programmable gate array (FPGA). The delay variations are isolated from other components of the programmable logic, e.g., look-up tables (LUTs), flip-flops (FFs), etc., using a feature of Xilinx FPGAs called dynamic partial reconfiguration (DPR). A set of partial designs is created to fix the placement of a time-to-digital converter (TDC) and supporting infrastructure to enable the path delays through the target interconnect and switch matrices to be extracted by subtracting out common-mode delay components. Delay variations are analyzed in the different levels of routing resources available within FPGAs, i.e., local routing and across-chip routing. Data are collected from a set of Xilinx Zynq 7010 devices, and a statistical analysis of within-die variations in delay through a set of the randomly-generated and hand-crafted interconnects is presented.

97 MATHEMATICS AND COMPUTING↗

Comparative Uptake Patterns of Radioactive Iodine and [18F]-Fluorodeoxyglucose (FDG) in Metastatic Differentiated Thyroid Cancers

Background: Metastatic differentiated thyroid cancer (DTC) represents a molecularly heterogeneous group of cancers with varying radioactive iodine (RAI) and [ 18 F]-fluorodeoxyglucose (FDG) uptake patterns potentially correlated with the degree of de-differentiation through the so-called “flip-flop” phenomenon. However, it is unknown if RAI and FDG uptake patterns correlate with molecular status or metastatic site. Materials and Methods: A retrospective analysis of metastatic DTC patients (n = 46) with radioactive 131-iodine whole body scan (WBS) and FDG-PET imaging between 2008 and 2022 was performed. The inclusion criteria included accessible FDG-PET and WBS studies within 1 year of each other. Studies were interpreted by two blinded radiologists for iodine or FDG uptake in extrathyroidal sites including lungs, lymph nodes, and bone. Cases were stratified by BRAF V600E mutation status, histology, and a combination of tumor genotype and histology. The data were analyzed by McNemar’s Chi-square test. Results: Lung metastasis FDG uptake was significantly more common than iodine uptake (WBS: 52%, FDG: 84%, p = 0.04), but no significant differences were found for lymph or bone metastases. Lung metastasis FDG uptake was significantly more prevalent in the papillary pattern sub-cohort (WBS: 37%, FDG: 89%, p = 0.02) than the follicular pattern sub-cohort (WBS: 75%, FDG: 75%, p = 1.00). Similarly, BRAF V600E+ tumors with lung metastases also demonstrated a preponderance of FDG uptake (WBS: 29%, FDG: 93%, p = 0.02) than BRAF V600E- tumors (WBS: 83%, FDG: 83%, p = 1.00) with lung metastases. Papillary histology featured higher FDG uptake in lung metastasis (WBS: 39%, FDG: 89%, p = 0.03) compared with follicular histology (WBS: 69%, FDG: 77%, p = 1.00). Patients with papillary pattern disease, BRAF V600E+ mutation, or papillary histology had reduced agreement between both modalities in uptake at all metastatic sites compared with those with follicular pattern disease, BRAF V600E- mutation, or follicular histology. Low agreement in lymph node uptake was observed in all patients irrespective of molecular status or histology. Conclusions: The pattern of FDG-PET and radioiodine uptake is dependent on molecular status and metastatic site, with those with papillary histology or BRAF V600E+ mutation featuring increased FDG uptake in distant metastasis. Further study with an expanded cohort may identify which patients may benefit from specific imaging modalities to recognize and surveil metastases.

60 APPLIED LIFE SCIENCES↗

Laboratory Tests of Laser Control of Electron Beams for Future Colliders

Laser-driven Compton backscattering (CBS) has been proposed as method for controlling the intensity of colliding bunches in the FCC-ee so as to avoid the flip-flop instability caused by intensity asymmetry in colliding bunches. Laser-based collimation has also been proposed as an indestructible collimator for high-intensity electron beams. We have initiated a laboratory-based test program of these concepts with the E344 experiment at FACET-II. In this paper, we describe simulations of laser-beam interactions at FACET-II and the relevant scaling for FCC-ee. We also describe the experimental setup and diagnostics that will be used to make the measurements at FACET-II.

Accelerator Physics (physics.acc-ph)↗

Surrogate Neural Architecture Codesign Package (SNAC-Pack)

Neural architecture search (NAS) is a powerful approach for automating model design, but existing methods often optimize for accuracy alone or rely on proxy metrics such as bit operations (BOPs) that correlate poorly with hardware cost. This gap is particularly large for FPGA deployment, where cost is dominated by a multi-dimensional budget of lookup tables, DSPs, flip-flops, BRAM, and latency. We present the Surrogate Neural Architecture Codesign Package (SNAC-Pack), an open-source AutoML framework for hardware-aware neural architecture codesign and end-to-end FPGA deployment. SNAC-Pack runs a multi-objective global search with Optuna and NSGA-II, loading trials to a shared SQLite store that enables parallel workers across compute nodes. A hardware surrogate model outputs per-trial resource and latency estimates, avoiding the synthesis cost that would otherwise dominate the search loop. A local search stage then applies quantization-aware training (QAT) together with iterative magnitude pruning in a combined compression loop, after which the final model is synthesized to FPGA firmware via the hls4ml Python library. A YAML configuration and an optional agentic frontend let users run the pipeline on new datasets without modifying the framework. We demonstrate SNAC-Pack on jet classification at the Large Hadron Collider and superconducting qubit readout, discovering compact architectures that match or exceed strong baselines on the task metric while reducing FPGA resource utilization and, in the qubit readout case, reducing the design space exploration process from months of manual fine-tuning to hours of automated search.

Weitz, Jason [UC, San Diego]↗