Search NASASearch

SEARCH · Search NASA

Results for “performance optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Development and Evaluation of a Novel Fuel Injector Design Method using Hybrid-Additive Manufacturing (Final Report)

The widespread application of metal additive manufacturing (AM) technologies has enabled exploration of complex design spaces to achieve optimally performing components. Current optimization techniques make use of several advanced methods to provide designs that are superior to existing versions. However, they seldom discuss the manufacturability of the optimal designs. The objective of this project was to develop a design optimization tool that simultaneously optimizes fuel injector hardware and the combustor flow field with optimization functions and constraints that consider both combustor performance and manufacturability using advanced AM methods and post-processing. In this way, the resultant hardware design is inherently imbued with our most advanced knowledge of combustion physics and AM methods from its conception.

36 MATERIALS SCIENCE

A Pulsar-Inspired Timing Framework for Power System: Optimization and Performance Evaluation

Due to their excellent stability, neutron pulsar stars are considered promising candidate timing sources for power system applications. However, the complexity of pulsar signals necessitates advanced processing algorithms to provide accurate timing references. This paper presents the foundational framework for pulsar signal processing, serving as the basis for further optimization. To enhance the timing accuracy and computation efficiency in pulsar period searches, three algorithms are proposed as the initial optimization step: wavelet de-noising, fast folding, and cross-correlation for profile evaluation. Wavelet de-noising improves signal-to-noise ratio (SNR) by 36%–70%. Fast folding reduces computation time from hundreds of seconds to mere milliseconds. Cross-correlation works better than traditional SNR-based methods by effectively identifying the optimal period. The performance of the proposed algorithms is evaluated using observation data from telescopes. Together, these algorithms significantly improve pulsar timing performance, reducing the error of the Pulse Per Second (PPS) signal from hundreds to tens of microseconds.

Wu, Ori [ORNL] (ORCID:0000000326723410)

Spike-Free Adaptive Sliding Mode Control: Application to Permanent Magnet Synchronous Motors

A new methodology for adaptive sliding mode control (ASMC) has been widely used to improve the control performance in various systems. This method exhibits several advantages, including low sliding mode control (SMC) chattering, no knowledge of the system disturbance bound, and no overestimation of the control gain. Despite its advantages, this method can be hampered by the spike phenomenon, slow control gain convergence, and difficulty in achieving optimal performance under varying disturbances. Consequently, this article proposes a spike-free ASMC method with a disturbance observer (DOB) to address these problems. Previous ASMC methods have been analyzed via simulations to verify the aforementioned problems. Here, this analysis highlights the need for disturbance compensation and improvements in the SMC gain adaptation law. Therefore, a DOB is designed to mitigate the spike phenomenon by compensating for disturbances. Subsequently, an SMC gain adaptation law based on disturbance error estimation is designed to eliminate the spike phenomenon completely. The proposed adaptation law makes the SMC gain to converge to a slightly higher value than the disturbance estimation error. Consequently, the proposed method not only eliminates the spike phenomenon, but also ensures optimal performance under varying disturbances. The performance of the proposed method is experimen tally validated through a comparative study.

42 ENGINEERING

The Simons Observatory: Design, Optimization, and Performance of Low-Frequency Detectors

The Simons Observatory (SO) is a cosmic microwave background (CMB) experiment located in the Atacama Desert in Chile that will make precise temperature and polarization measurements over six spectral bands ranging from 27 to 285 GHz. Three small aperture telescopes (SATs) and one large aperture telescope (LAT) will house ~60,000 detectors and cover angular scales between one arcminute and tens of degrees. We present the performance of the dichroic, low-frequency (LF) lenslet-coupled sinuous antenna transition-edge sensor (TES) bolometer arrays with bands centered at 27 and 39 GHz. The LF focal plane will primarily characterize Galactic synchrotron emission as a critical part of foreground subtraction from CMB data. We will discuss the design, optimization, and current testing status of these pixels.

79 ASTRONOMY AND ASTROPHYSICS

pnnl/EZBattery

Cell performance optimization is important for improving the system efficiency of a redox flow battery. To gain better insights into key controlling factors Of system efficiency, this work first proposed a theoretical model for a unit cell by extending a two-dimensional analytic model to a full cell. The model is then used for cell performance optimization after validating it with experimental and numerical modeling data.

Bao, Jie

Etching-Chemistry-Driven Ruthenium Doping on Ti 3 C 2 T x MXene for Optimizing Electrochemical Performance

We demonstrate that the etching chemistry used during MXene synthesis from Ti 3 AlC 2 MAX phase significantly influences surface functionalization and structural vacancies, which in turn affect ruthenium (Ru) ion interactions. Using hydrofluoric acid (HF) and ammonium bifluoride (NH 4 HF 2 ) as etchants, we obtained MXene surfaces with distinct functional groups and Ti vacancies that impact Ru ion interactions and electrochemical performance. Both MXene variants (labeled MX(H) and MX(N), respectively) exhibited negative zeta potentials in their pristine state, but upon the addition of Ru the zeta potential for MX(H) reached 12.9 mV while that for MX(N) remained negative at −6.4 mV. This adsorption resulted in a 14.4-fold increase in the specific capacitance of MX(H)/Ru compared to pristine MX(H), whereas MX(N)/Ru exhibited only a 4.4-fold increase over its pristine counterpart. X-ray diffraction analysis identified the formation of ammonium titanium oxide fluoride, (NH 4 ) 3 TiOF 5 , on MX(N), which likely contributed to its reduced Ru adsorption. X-ray photoelectron spectroscopy suggested the presence of Ti vacancies in both MXene variants; however, their behavior toward Ru accommodation differed markedly, with MX(H) showing the most obvious shift in the Ti 2p peak in the XPS survey spectrum, while MX(N) showed the most obvious shift in the C 1s peak. Electron paramagnetic resonance spectroscopy further demonstrated a distinct alteration in the spectral signatures of MX(H) upon Ru addition, in contrast to the negligible changes in MX(N), indicating effective passivation of the Ti defect sites in MX(H) via vacancy-assisted Ru doping. Cyclic voltammetry showed that Ru-incorporated MX(H) nanocomposites exhibit more efficient redox-active sites, as reflected in their higher capacitance values. These findings highlight the pivotal role of MXene surface chemistry in controlling cation adsorption, providing valuable insights for the rational design of high-performance electrodes.

2D surface engineering

Criticality analysis of nuclear binding energy neural networks

Machine learning methods, in particular deep learning methods such as artificial neural networks (ANNs) with many layers, have become widespread and useful tools in nuclear physics. However, these ANNs are typically treated as ‘black boxes’, with their architecture (width, depth, and weight/bias initialization) and the training algorithm and parameters chosen empirically by optimizing learning based on limited exploration. We test a non-empirical approach to understanding and optimizing nuclear physics ANNs by adapting a criticality analysis based on renormalization group flows in terms of the hyperparameters for weight/bias initialization, training rates, and the ratio of depth to width. This treatment utilizes the statistical properties of neural network initialization to find a generating functional for network outputs at any layer, allowing for a path integral formulation of the ANN outputs as a Euclidean statistical field theory. We use a prototypical example to test the applicability of this approach: a simple ANN for nuclear binding energies. We find that with training using a stochastic gradient descent optimizer, the predicted criticality behavior is realized, and optimal performance is found with critical tuning. However, the use of an adaptive learning algorithm leads to somewhat superior results without concern for tuning and thus obscures the analysis. Nevertheless, the criticality analysis offers a way to look within the black box of ANNs, which is a first step towards potential improvements in network performance beyond using adaptive optimizers.

artificial neural network

Insights from Optimizing HPL Performance on Exascale Systems: A Comparative Analysis of Panel Factorization

High performance LINPACK (HPL) remains the primary benchmark for evaluating supercomputing performance. It includes many parts with substantial internal complexity, and its performance is affected by a large number of parameters that interact in ways that are difficult to predict on large-scale heterogeneous supercomputer systems. We present a comprehensive performance analysis of HPL on Frontier, the world’s first exascale supercomputer, which achieved HPL performance of 1.35 exaflops. Through empirical parameter tuning, detailed modeling, and comparative evaluation, we uncover critical performance insights, share lessons learned, and outline best practices for effective parameter tuning on exascale systems. We introduce and evaluate two novel PDFACT strategies: a dedicated-thread (DT) variant and a GPU-based variant (GPUPDFACT) implementation using HIP cooperative groups, demonstrating that GPU-based factorization outperforms conventional CPU-based PDFACT on Frontier’s architecture. Our findings establish key performance factors for HPL on exascale systems and offer valuable guidance for future high-performance computing and benchmarking efforts.

Lu, Hao [ORNL] (ORCID:000000018941870X)

Distributed Wind-Energy-Based Hybrids

Presentation defining distributed wind-based hybrids and introducing the Hybrid Optimization Performance Platform (HOPP) an open-source tool that helps design and optimize buildable hybrid power plants.

distributed wind-based hybrids

AutoFocus: AI/ML-driven real-time wavefront diagnostics to autonomously align and optimize X-ray optics

We present an integrated system that combines advanced wavefront diagnostics with artificial intelligence (AI) to automate and optimize X-ray optics at synchrotron beamlines. This system couples real-time wavefront sensing with AI-driven control algorithms to achieve precise beam alignment, stabilization, and performance optimization. A key feature is the use of multi-fidelity transfer learning, which enables knowledge gained from both real-world beamline optimizations and ultra-realistic digital twin simulations to be effectively applied to in situ optimization. By leveraging multi-objective bayesian optimization, the system continuously refines its performance, reducing optimization time and minimizing the need for manual adjustments. Designed for seamless deployment, it operates with existing beamline hardware and provides an intuitive graphical interface. Initial deployments at the advanced photon source beamlines have demonstrated its ability to enhance beam stability, improve reproducibility, and significantly streamline alignment procedures. This AI-enhanced control framework represents a significant step toward fully autonomous beamline operation in next-generation synchrotron facilities.

Rebuffi, Luca [Argonne National Laboratory (ANL),

Data-Driven Optimization of Pixelated CdZnTe Spectrometers for Uranium Enrichment Assay

Here, in recent work [Vavrek et al. (2025)], we developed the performance optimization framework spectre-ml for gamma spectrometers with variable performance across many readout channels. The framework uses non-negative matrix factorization (NMF) and clustering to learn groups of similarly-performing channels and sweep through various learned channel combinations to optimize the performance tradeoff of including worse-performing channels for better total efficiency. In this work, we integrate the pyGEM uranium enrichment assay code with our spectre-ml framework, and show that the U-235 enrichment relative uncertainty can be directly used as an optimization target. We find that this optimization reduces relative uncertainties after a 30 -minute measurement by an average of 20%, as tested on six different H3D M400 CdZnTe spectrometers, which can significantly improve uranium non-destructive assay measurement times in nuclear safeguards contexts. Additionally, this work demonstrates that the spect re-ml optimization framework can accommodate arbitrary end-user spectroscopic analysis code and performance metrics, enabling future optimizations for complex Pu spectra.

Gamma-ray detection

ROSE

Developed at Lawrence Livermore National Laboratory (LLNL), ROSE is an open source compiler infrastructure to build source-to-source program transformation and analysis tools for large-scale C (C89 to C23), C++ (C++98 to C++23), UPC, Fortran (Fortran4, 66, 77, 95, 2003), OpenMP, Java, Python, and Binary applications. ROSE users range from experienced compiler researchers to library and tool developers who may have minimal compiler experience. ROSE is particularly well suited for building custom tools for static analysis, program optimization, arbitrary program transformation, domain-specific optimizations, complex loop optimizations, performance analysis, and cyber-security. ROSE is: A library (and set of associated tools) to quickly and easily apply compiler techniques to one's code in order to improve application performance and developer productivity. A research and development compiler infrastructure for for writing custom source-to-source translators to perform source code transformations, analysis, and optimizations. Is

Pinnow, NathanT [Lawrence Livermore National Labor

Dynamical Sketching for Enhanced Communication Efficiency in Federated Learning

Federated learning (FL) has revolutionized distributed machine learning by enabling collaborative model training without sharing local data. However, communication efficiency and privacy guarantees remain significant challenges. This paper introduces a dynamic sketching mechanism in FL, optimizing the trade-off between communication efficiency and model accuracy. By dynamically selecting the sketch matrix size, our approach adapts to the evolving characteristics of the data and the model, ensuring optimal performance across diverse scenarios. We leverage Bayesian optimization to systematically tune the sketch parameters, achieving an effective balance between resource efficiency and model performance. Experimental results on the MNIST dataset using a convolutional neural network (CNN) architecture validate the proposed method's efficiency and scalability. Our dynamic sketching approach significantly outperforms fixed-size sketching techniques, achieving higher compression ratios (up to 62x) and providing better privacy guarantees while maintaining high model accuracy. These findings highlight the robustness and versatility of our approach and make it a valuable solution for privacy-preserving, communication-efficient federated learning.

Afrose, Sharmin [ORNL]

A High-Quality Workflow for Multi-Resolution Scientific Data Reduction and Visualization

Multi-resolution methods such as Adaptive Mesh Refinement (AMR) can enhance storage efficiency for HPC applications generating vast volumes of data. However, their applicability is limited and cannot be universally deployed across all applications. Furthermore, integrating lossy compression with multi-resolution techniques to further boost storage efficiency encounters significant barriers. To this end, we introduce an innovative workflow that facilitates high-quality multi-resolution data compression for both uniform and AMR simulations. Initially, to extend the usability of multi-resolution techniques, our workflow employs a compression-oriented Region of Interest (ROI) extraction method, transforming uniform data into a multi-resolution format. Subsequently, to bridge the gap between multi-resolution techniques and lossy compressors, we optimize three distinct compressors, ensuring their optimal performance on multi-resolution data. These optimizations can improve the compression ratio of SOTA approaches by up to 3.3× under the same data quality loss. Lastly, we incorporate an advanced uncertainty visualization method into our workflow to understand the potential impacts of lossy compression. Experimental evaluation demonstrates that our workflow achieves significant compression quality improvements.

Wang, Daoce

Workflow for evaluating enzyme immobilization and performance for continuous flow manufacturing

Enzymes have shown promise in various industries due to their functional specificity, catalytic efficiency, and environmental sustainability. These biological catalysts can be a pivotal component of manufacturing pipelines like continuous flow chemistry. For this, there exists a need to robustly immobilize enzymes on solid supports and assess the effects of the solid supports on catalytic performance and stability. Here, we use an industrially relevant model enzyme, C. ensiformis (Jack bean) urease, to demonstrate immobilization and assess performance in the context of continuous flow manufacturing. Various immobilization strategies were screened focusing on immobilization efficiency, protocol simplicity, and urease biocatalyst kinetics. Based on this, CDI-agarose and NHS-agarose resins were identified as the best-performing immobilization strategies for urease. CDI-agarose-urease and NHS-agarose-urease were then scaled up and applied to a large-scale continuous flow reactor to evaluate product yields, operational stability, and long-term stability. These experiments identified differences in stability and performance depending on the immobilization method tested. This highlights the importance of screening immobilization methods and subsequent enzyme performance for each candidate biocatalyst used in manufacturing to promote optimal performance and stability. As such, this work provides a framework for evaluating enzyme biocatalyst immobilization approaches to improve performance and enable transition into industrial processes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Humins-Derived Hard Carbon as a Low-Cost Material for Sodium-Ion Battery Anodes

The growing demand for sodium-ion batteries (SIBs) in grid storage underscores the need for electrode materials that balance performance and cost, including sustainable and robust carbon sources. The equitable and abundant distribution of materials for SIBs, along with their superior low-temperature performance, safety, and fast-charging capability, further distinguishes them from lithium-ion batteries (LIBs). Hard carbon (HC) is the state-of-the-art anode for SIBs, but current commercial HC production is localized mostly to one region of the world, raising concerns about supply chain vulnerability, critical material dependency, and environmental aspects. Here, the first demonstration of humins, an abundant biorefinery byproduct, as a precursor for HC anodes for sodium-ion storage is reported. Humins were carbonized at 1100 degrees C-1300 degrees C, and the resulting materials were subjected to comprehensive materials and electrochemical characterization. Among the temperatures studied, humins-derived hard carbon synthesized at 1200 degrees C delivers an initial reversible capacity 270mA h g-1 , with stable cycling performance up to 500 cycles and excellent rate capability, representing the optimal performance. This study establishes humins as a promising and low-cost carbon source that provides a route to mitigate supply chain risks and valorizes an underutilized biorefinery waste stream for high-performance SIB anodes.

25 ENERGY STORAGE

Exploring code portability solutions for HEP with a particle tracking test code

Traditionally, high energy physics (HEP) experiments have relied on x86 CPUs for the majority of their significant computing needs. As the field looks ahead to the next generation of experiments such as DUNE and the High-Luminosity LHC, the computing demands are expected to increase dramatically. To cope with this increase, it will be necessary to take advantage of all available computing resources, including GPUs from different vendors. A broad landscape of code portability tools—including compiler pragma-based approaches, abstraction libraries, and other tools—allow the same source code to run efficiently on multiple architectures. In this paper, we use a test code taken from a HEP tracking algorithm to compare the performance and experience of implementing different portability solutions. While in several cases portable implementations perform close to the reference code version, we find that the performance varies significantly depending on the details of the implementation. Achieving optimal performance is not easy, even for relatively simple applications such as the test codes considered in this work. Several factors can affect the performance, such as the choice of the memory layout, the memory pinning strategy, and the compiler used. The compilers and tools are being actively developed, so future developments may be critical for their deployment in HEP experiments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training

Large Language Models (LLMs) have demonstrated remarkable performance in various natural language processing tasks. However, the training of these models is computationally intensive and susceptible to faults, particularly in the attention mechanism, which is a critical component of transformer-based LLMs. In this paper, we investigate the impact of faults on LLM training, focusing on INF, NaN, and near-INF values in the computation results with systematic fault injection experiments. We observe the propagation patterns of these errors, which can trigger non-trainable states in the model and disrupt training, forcing the procedure to load from checkpoints. To mitigate the impact of these faults, we propose ATTNChecker, the first Algorithm-Based Fault Tolerance (ABFT) technique tailored for the attention mechanism in LLMs. ATTNChecker is designed based on fault propagation patterns of LLM and incorporates performance optimization to adapt to both system reliability and model vulnerability while providing lightweight protection for fast LLM training. Evaluations on four LLMs show that ATTNChecker on average incurs on average 7% overhead on training while detecting and correcting all extreme errors. Compared with the state-of-the-art checkpoint/restore approach, ATTNChecker reduces recovery overhead by up to 49×.

Liang, Yuhang [University of Alabama - Birmingham]