Search NASASearch

SEARCH · Search NASA

Results for “performance optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Advanced flip-coil system for magnetic field integral measurements of insertion devices

A novel flip-coil measurement system has been developed for the National Synchrotron Light Source II (NSLS-II) at Brookhaven National Laboratory. This paper describes the design, implementation, and commissioning of the new measurement bench, highlighting its key features, including improved mechanical stability, advanced data acquisition, and enhanced reproducibility. The system enables precise characterization of field integrals and multipole components, ensuring the optimal performance of Insertion Devices (IDs) before installation in the NSLS-II storage ring. The flip-coil system incorporates an innovative approach to minimize mechanical and electrical errors, which significantly improves the reproducibility of measurements. In addition, the system features a state-of-the-art data acquisition system that enables real-time monitoring and analysis, further enhancing the efficiency and accuracy of the measurement process. Furthermore, preliminary tests have demonstrated that the new system meets the stringent requirements for magnetic field characterization of advanced insertion devices, making it an essential tool for future ID commissioning and quality assurance at NSLS-II.

36 MATERIALS SCIENCE

Wavelet flow for extragalactic foreground simulations

Extragalactic foregrounds in cosmic microwave background (CMB) observations are both a source of cosmological and astrophysical information and a nuisance to the CMB. Effective field-level modeling that captures their non-Gaussian statistical distributions is increasingly important for optimal information extraction, particularly given the low-noise observations from current and upcoming experiments. Here, we explore the use of Wavelet Flow (WF) models to tackle the novel task of modeling the field-level probability distributions of multi-component CMB secondaries and foregrounds. Specifically, we jointly train correlated CMB lensing convergence (κ) and cosmic infrared background (CIB) maps with a WF model and obtain a network that statistically recovers the input to high accuracy — the trained network generates samples of κ and CIB fields whose average power spectra are within a few percent of the inputs across all scales, and whose Minkowski functionals are similarly accurate compared to the inputs. Leveraging the multiscale architecture of these models, we fine-tune both the model parameters and the priors at each scale independently, optimizing performance across different resolutions. These results demonstrate that WF models can accurately simulate correlated components of CMB secondaries, supporting improved analysis of cosmological data. Our code and trained models can be found on this GitHub repo.

cosmological simulations

Integrated modeling of RF-induced tungsten erosion at ICRH antenna structures in the WEST tokamak *

This paper introduces STRIPE (Simulated Transport of RF Impurity Production and Emission), an advanced modeling framework developed to analyze material erosion and the global transport of eroded impurities originating from radio-frequency (RF) antenna structures in magnetic confinement fusion devices. STRIPE integrates multiple physics modules: SolEdge3x for scrape-off-layer plasma profiles, COMSOL for 3D RF rectified sheath potentials, RustBCA for erosion yields and surface interactions, and global impurity transport for 3D ion energy-angle distributions and impurity transport. The framework is applied to an ion cyclotron RF-heated L-mode discharge (#57877) in the WEST tokamak, where it predicts a thirty-fold increase in gross tungsten erosion at antenna limiters during the transition from ohmic to ICRH operation. Additionally, under ICRH conditions, a tenfold enhancement in erosion is observed when comparing RF sheath effects to purely thermal sheath conditions. High-charge-state oxygen ions ($\mathrm{O}$ 6+ and above) are identified as the dominant contributors to tungsten sputtering. To validate the model, a synthetic diagnostic tool based on inverse photon efficiency (S/XB coefficients) from the ColRadPy collisional-radiative model enables direct comparison with spectroscopic measurements. Model predictions using a plasma composition of 1% oxygen and 99% deuterium show good agreement with observed W − I (400.9 nm) emission for discharge #57877, supporting the accuracy of the STRIPE framework. This study focuses specifically on gross erosion calculations to demonstrate STRIPE’s capabilities. Future extensions of this work will incorporate net erosion, re-deposition, self-sputtering effects, and whole-device modeling of sputtered tungsten impurity transport. STRIPE is also being applied to other RF-heated linear and toroidal devices, offering valuable insights for antenna design, impurity control, and performance optimization in next-generation fusion reactors.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Iso-cost-performance of thermal energy storage

As thermal energy storage (TES) systems gain increasing recognition as next-generation energy storage solutions, evaluating their techno-economic performance is crucial. This perspective analyzes the cost-performance of latent heat-based TES systems and introduces the concept of iso-cost-performance—maintaining a constant cost per unit energy ($\$$/kWh) despite material degradation. Using an empirical degradation model, we show that after 1200 thermal cycles, the thermal conductivity and volumetric energy density of a paraffin-based phase change material (PCM) decreased by 36.0% and 26.1%, respectively, resulting in a 31.2% reduction in the figure of merit (FOM) and, consequently, the system cost-performance. However, by introducing thermally conductive additives to enhance effective thermal conductivity, the TES system can recover its initial FOM, achieving iso-cost-performance operation. This framework quantitatively demonstrates how degradation-mitigation strategies—such as improving thermal conductivity—can offset material degradations and maintain long-term cost-effectiveness. Beyond PCM-based TES, the proposed FOM-based approach provides a generalized pathway for cost-performance optimization across various TES technologies.

25 ENERGY STORAGE

Overcoming sparse datasets with multi-task learning as applied to high entropy alloys

Abstract The design of novel High Entropy Alloys for use in high-temperature applications is an area of active interest due to their potential to provide exceptional properties compared to conventional alloys. Since the increased popularity of machine learning, an important cog in the design process has been training surrogate models on alloy properties. However, these Single-Task models are trained on individual mechanical properties and do not take advantage of the relatedness between properties. Multi-Task models can capture the interdependencies between tasks, leading to potentially more accurate predictions for all tasks. In this paper, we investigate if Multi-Task models can show improvement over Single-Task models when used for predicting the mechanical properties of these alloys. To ensure fair evaluation between the models, we apply L 0 regularization and skip connections to the models, which allows them to adjust the number of model parameters and depth for optimal performance. We find that the Multi-Task models can leverage task relationships to perform better than Single-Task models, especially for high amounts of missing data in the tasks. Furthermore, adding simple auxiliary targets can boost Multi-Task performance even further despite not being effective as input descriptors to single-task models themselves. We anticipate that the proposed strategies can achieve more accurate predictions and consequently enable better design capabilities for such data-constrained domains without incurring much additional computational cost.

Debnath, Arindam (ORCID:0000000194274499)

Towards Mitigating Electron Beam-Induced Damage in MOFs Using Low Dose Electron Microscopy for In-Situ Room-Temperature Measurements

Metal-Organic Frameworks (MOFs) have emerged as a versatile class of materials with applications in gas storage and separation, catalysis, and drug delivery [1]. Understanding their damage mechanisms in under various environmental conditions is crucial for optimizing performance and stability [2]. This study employs a combination of Length Structural Coherence (LSC) analysis in electron scattering and Electron Energy Loss Spectroscopy (EELS) to elucidate the damage mechanisms in two representative MOFs. MIL101(Fe), and MIL101(Cr). Here, the primary objective is to reveal structural and chemical changes that occur under electron beam irradiation in Scanning Transmission Electron Microscopy (STEM), both in the pristine state and during in-situ CO and CO 2 adsorption. By combining LSC in Radial Distribution Function (RDF) analysis with EELS, comprehensive insights into the nanoscale alterations are obtained and correlated to the electron dose rate.

dos Santos, Gabriel T. [Northwestern University, E

Interplay Between Time and Energy in Bosonic Noisy Quantum Metrology

Quantum entanglement and coherence often allow for protocols that outperform classical ones in estimating a system’s parameter. When using infinite-dimensional probes (such as a bosonic mode), one could, in principle, obtain infinite precision in a finite time for both classical and quantum protocols, which makes it hard to quantify potential quantum advantage. However, such a situation is unphysical, as it would require infinite resources, so one needs to impose some additional constraint: typically the average energy employed by the probe is finite. Here we treat both energy and time as a resource, showing that, in the presence of noise, there is a nontrivial interplay between the average energy and the time devoted to the estimation. Our results are valid for the most general metrological schemes (e.g., adaptive schemes, which may involve entanglement with external ancillae or any kind of continuous measurement). We apply recently derived precision bounds for all parameters characterizing the paradigmatic case of a bosonic mode, subject to Lindbladian noise. We show how the time employed in the estimation should be partitioned in order to achieve the best possible precision. In most cases, the optimal performance may be obtained without the necessity of adaptivity or entanglement with ancilla. We compare results with classical strategies. Interestingly, for temperature estimation, applying a fast-prepare-and-measure protocol with Fock states provides better scaling with the number of photons than any classical strategy.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

A Machine Learning based Approach of Estimating Equivalent Circuit Model Parameters at Different SoCs of Li-ion Batteries from Voltage Relaxation

Abstract: In this study, an approach of estimating the equivalent circuit model (ECM) parameters for Li-ion batteries (LIBs) is proposed based on the voltage value at different intervals while relaxing the LIB after discharge. The typical approach for estimating ECM parameters of a LIB is to conduct electrochemical impedance spectroscopy (EIS) measurements at different frequencies and fit them to a predefined circuit model, which requires additional measuring arrangements and specialized devices. The proposed methodology utilizes four different voltages at 0s, 60s, 360s, and 1800s alongside the specific state of charge (SoC) value for a specific constant discharge current value of ~1C until the relaxation stage to train and evaluate three regression-based machine learning models— Support Vector Regression (SVR), Extreme Gradient Boosting (XGBoost), and Gaussian Process Regression (GPR)—for estimating the ECM parameters of the selected model. Bayesian optimization is employed for hyperparameter tuning to achieve optimal performance for all the regressor models, among which, the GPR provided the best performance with the root-mean-squared error (RMSE) of less than 4x10-4 on average for the resistive components and less than 0.27 for capacitive components with excellent R2 scores. The simplicity of the approach enables it to eliminate the need for sophisticated measuring equipment and computation power.

Sagar, Md. Samiul [The University of Alabama (UA)]

Real-Time Lifetime Prediction of Semiconductor Devices Using Hardware-in-the-Loop

This paper presents a unique approach to enable real-time lifespan prediction of semiconductor power modules using a Hardware-in-the-Loop (HIL) system. By integrating the module's overall loss characteristics-specifically switching and conduction losses-with a thermoelectric model of the thermal management system, this research demonstrates that the model can dynamically estimates the junction temperature profile of the semiconductor devices in response to a changing torque demand profile for the motor drive system. This capability enables continuous monitoring of the module's operational time and cumulative stress induced on the devices to compute accumulated remaining lifetime or time-to-failure (TTF). This study provides an architectural framework for the HIL system with high-fidelity component models of multiple physical domains, allowing simulation of dynamic behaviors of a closely-coupled motor drive system. The advanced real-time computation and measurement functionalities of the HIL system allow for both dynamic lifetime calculations based on simulated data and aggregate lifetime predictions utilizing historical data. Moreover, this paper details an algorithm that not only computes cumulative damage but also synthesizes these data into a comprehensive aggregated lifetime metric. This methodology can enhance the maintenance scheduling strategies and operational reliability of semiconductor devices in critical applications, ultimately extending their service life while optimizing performance.

hardware-in-the-loop (HIL)

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (distributed parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

graph algorithms, high performance comptuing

Automatic Extraction of Network Configurations for Realistic Simulation and Validation

Popular HPC network interconnection simulators such as SST Macro provide a variety of configurable parameters to explore the design space of hardware components such as network links and switches. While such knobs provide flexibility to explore design trade-offs for novel hardware, manually configuring simulations for existing hardware to focus on topology exploration can be cumbersome and error-prone, leading to widely inaccurate simulations. This challenge is compounded when specifications of various (proprietary) technologies are not readily available or are intentionally omitted. In this work, we provide a methodology to automatically tune the simulation configuration of the multiple network models running within SST Macro using Bayesian optimization. We perform this optimization in the context of multiple messaging regimes (i.e., small to large and latency to bandwidth-bound messages) and provide a detailed analysis of the simulation error for four systems. With our automated framework, we achieve a 5x improvement in accuracy over best-effort configurations based on available hardware specifications.

Suetterlein, Joshua D.

MDLoader: A Hybrid Model-Driven Data Loader for Distributed Graph Neural Network Training

Scalable data management is essential for processing large scientific dataset on HPC platforms for distributed deep learning. In-memory distributed storage is preferred for its speed, enabling rapid, random, and frequent data access required by stochastic optimizers. Processes use one-sided or collective communication to fetch remote data, with optimal performance depending on (i) dataset characteristics, (ii) training scale, and (iii) interconnection network. Empirical analysis shows collective communication excels with larger mini-batch sizes and/or fewer processes, whereas one-sided communication outperforms at larger scales. We propose MDLoader, a hybrid in-memory data loader for distributed graph neural network training. MDLoader features a model-driven performance estimator that dynamically selects between one-sided and collective communication at the beginning of training using Tree of Parzen Estimators (TPE). Evaluations on NERSC Perlmutter and OLCF Summit show MDLoader outperforms single-backend loaders by up to 2.83 × and predicts the suitable communication method with 96.3% (Perlmutter) and 94.3% (Summit) success rate.

Bae, Jonghyun

An Experimental Setup for Mechanical Vibration Analysis Using VLC

This study explores the potential applications across various domains, including earthquake detection and warning systems, where the system’s sensitivity to ground vibrations can contribute to early seismic event detection. Additionally, the study paves the way of developing applications of VLC/T in mechanical vibration and stability analysis of engines and platforms, offering insights into structural integrity and performance optimization. These multifaceted applications underscore the adaptability and potential of VLC/T systems in diverse fields, heralding advancements in sensing, communication, and security technologies. To achieve this, in this study, Peak to Average Power Ratio (PAPR) is proposed to represent the impact of mechanical shocks and vibrations generated by several weights dropped onto the platform with which the receiver is fixed. Even though non-contact measurement methodology is preferred for various reasons, the proposed measurement campaign obtains the data in contact form; however, the system and signal model proposed in this study could easily be extended into non-contact form. Considering the fact that the proposed measurement campaign employs off-the-shelf products, it is cost-effective and very scalable.

Yilmaz, Ahmet Mucahit

Memory-Aware External Facelist Calculation: A Data-Parallel Atomic Hash Counting Approach

Unstructured volumetric meshes serve as fundamental data representations in various scientific simulations and analyses. They play a crucial role in representing complex computational domains and are essential for important numerical techniques, such as finite element analysis. Whenever such a mesh is read from a file, streamed in-situ, or generated by algorithms, scientific visualization libraries rely on calculating the external surface of a geometry, named “external facelist”, to produce a polygonal mesh for rendering. Consequently, external facelist calculation has become one of the most widely used algorithms in the scientific visualization domain, necessitating optimal performance. In this paper, we explore relevant work on external facelist calculation algorithms in two common visualization libraries, VTK and Viskores, assess their performance and memory constraints, and introduce a novel memory-aware external facelist calculation algorithm employing an atomic hash counting approach. This algorithm fully leverages Viskores' data-parallel primitive operations, facilitating its execution across diverse many-core architectures. Our algorithm features the lowest memory footprint on the GPU and the second-lowest on the CPU among all evaluated methods, and it also delivers the fastest performance on both CPU and GPU. It has been made available under an open-source license in the VTK and Viskores visualization systems.

Tsalikis, Spiros [Kitware] (ORCID:0000000151137195

Strangers in a foreign land: ‘Yeastizing’ plant enzymes

Abstract Expressing plant metabolic pathways in microbial platforms is an efficient, cost‐effective solution for producing many desired plant compounds. As eukaryotic organisms, yeasts are often the preferred platform. However, expression of plant enzymes in a yeast frequently leads to failure because the enzymes are poorly adapted to the foreign yeast cellular environment. Here, we first summarize the current engineering approaches for optimizing performance of plant enzymes in yeast. A critical limitation of these approaches is that they are labour‐intensive and must be customized for each individual enzyme, which significantly hinders the establishment of plant pathways in cellular factories. In response to this challenge, we propose the development of a cost‐effective computational pipeline to redesign plant enzymes for better adaptation to the yeast cellular milieu. This proposition is underpinned by compelling evidence that plant and yeast enzymes exhibit distinct sequence features that are generalizable across enzyme families. Consequently, we introduce a data‐driven machine learning framework designed to extract ‘yeastizing’ rules from natural protein sequence variations, which can be broadly applied to all enzymes. Additionally, we discuss the potential to integrate the machine learning model into a full design‐build‐test cycle.

59 BASIC BIOLOGICAL SCIENCES

Sixteen multiple-amplifier sensing charge-coupled devices and characterization techniques targeting the next generation of astronomical instruments

We present a candidate sensor for future spectroscopic applications, such as a Stage-5 Spectroscopic Survey Experiment or the Habitable Worlds Observatory. This type of charge-coupled device (CCD) sensor features multiple in-line amplifiers at its output stage allowing multiple measurements of the same charge packet, either in each amplifier or in the different amplifiers. Recently, the operation of an eight-amplifier sensor has been experimentally demonstrated, and we present the operation of a 16-amplifier sensor. This new sensor enables a noise level of ∼1 erms− with a single sample per amplifier. In addition, it is shown that sub-electron noise can be achieved using multiple samples per amplifier. In addition to demonstrating the performance of the 16-amplifier sensor, we aim to create a framework for future analysis and performance optimization of this type of detectors. New models and techniques are presented to characterize specific parameters, which are absent in conventional CCDs and Skipper CCDs: charge transfer between amplifiers and independent and common noise in the amplifiers and their processing.

16 multiple-amplifer sensing CCD (MAS-CCD)

OpenARC

OpenARC is an open-sourced, very High-Level Intermediate Representation (HLIR)-based, extensible compiler framework, where various performance optimizations, traceability mechanisms, fault tolerance techniques, etc., can be built for better debuggability/performance/resilience on the complex accelerator computing. OpenARC is the first OpenACC compiler supporting Altera FPGAs, in addition to NVIDIA GPUs, AMD GPUs, and Intel Xeon Phis.

Lee, Seyong [Oak Ridge National Laboratory (ORNL),

Vedizar Fingerprinter

SAND2025-03289O Vedizar Fingerprinter simplifies the process of identifying devices on a network by analyzing traffic data. It uses a unique library to recognize different devices, making it easier for users to understand what is happening on their networks. This software is ideal for IT and operational technology environments, helping organizations monitor their networks effectively. By saving results in a database, it allows for easy access and review of device information. Users can enhance their network security and optimize performance without needing specialized hardware or technical expertise. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Jacobellis, John [Sandia National Lab. (SNL-CA), L