Search NASA⌕ Search

SEARCH · Search NASA

Results for “performance optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41

Data imbalance in drug response prediction: multi-objective optimization approach in deep learning setting

Abstract Drug response prediction (DRP) methods tackle the complex task of associating the effectiveness of small molecules with the specific genetic makeup of the patient. Anti-cancer DRP is a particularly challenging task requiring costly experiments as underlying pathogenic mechanisms are broad and associated with multiple genomic pathways. The scientific community has exerted significant efforts to generate public drug screening datasets, giving a path to various machine learning models that attempt to reason over complex data space of small compounds and biological characteristics of tumors. However, the data depth is still lacking compared to application domains like computer vision or natural language processing domains, limiting current learning capabilities. To combat this issue and improves the generalizability of the DRP models, we are exploring strategies that explicitly address the imbalance in the DRP datasets. We reframe the problem as a multi-objective optimization across multiple drugs to maximize deep learning model performance. We implement this approach by constructing Multi-Objective Optimization Regularized by Loss Entropy loss function and plugging it into a Deep Learning model. We demonstrate the utility of proposed drug discovery methods and make suggestions for further potential application of the work to achieve desirable outcomes in the healthcare field.

Biochemistry & Molecular Biology↗

A Benchmark Suite for Evaluating Scientific AI Workloads on GPUs

AI applications have been steadily increasing in the allocation portfolio among leadership computing facilities. These applications depend on deep learning frameworks with hardware acceleration and underlying software systems. With the rapid development of applications, software stacks, and hardware devices, it is essential to evaluate the performance of core operations in AI workloads for direction of optimizations and procurement of next-generation high-performance computing (HPC) infrastructures. Currently, most benchmarks lack scientific AI workloads. So, we present DeepKernelBench and the experimental results of evaluating the benchmark suite for early observations and performance comparisons on datacenter GPUs using representative workloads for scientific AI, including Attentions, General matrix multiplications, Geometrics and Fourier neural operations.

Jin, Zheming [Advanced Micro Devices (AMD)]↗

Generic Discretization Library

The GenDiL library is a collection of C++ software abstractions designed to discretize and solve partial differential equations (PDEs) for high-performance computing (HPC) applications. Its primary focus is on modern C++ generic programming, which helps ensure portability across various hardware architectures. The central idea behind the library is to provide building blocks for numerical algorithms-such as discretization methods and iteration patterns-so that domain experts can focus on the math, rather than the low-level details of hardware or implementation. By defining abstractions for data types, iteration over computational grids, and scheduling of operations, the library isolates the high-level PDE algorithms from the platform-specific optimizations needed to achieve efficient performance.

Dudouit, Yohann [Lawrence Livermore National Labor↗

Physics-informed machine learning for building performance simulation-A review of a nascent field

Building performance simulation (BPS) is critical for understanding building dynamics and behavior, analyzing the performance of the built environment, optimizing energy efficiency, improving demand flexibility, and enhancing building resilience. However, conducting BPS is not trivial. Traditional BPS relies on accurate building energy models, which are primarily physics-based and heavily dependent on detailed building information, expert knowledge, and case-by-case model calibrations, significantly limiting their scalability. With the development of sensing technology and the increased availability of data, there is growing attention and interest in data-driven BPS. However, purely data-driven models often suffer from limited generalization ability and a lack of physical consistency, resulting in poor performance in real-world applications. To address these limitations, recent studies have begun integrating physics priors into data-driven models, a methodology known as physics-informed machine learning (PIML). PIML is an emerging field where its definitions, methodologies, evaluation criteria, application scenarios, and future directions remain open. To bridge those gaps, this study systematically reviews the state-of-the-art PIML for BPS, offering a comprehensive definition of PIML and comparing it to traditional BPS approaches regarding data requirements, modeling effort, performance, and computational cost. We also summarize the commonly used methodologies, validation approaches, application domains, available data sources, open-source packages, and testbeds. In addition, this study provides a general guideline for selecting appropriate PIML models based on BPS applications. Finally, this study identifies key challenges and outlines future research directions, providing a solid foundation and valuable insights to advance R&D of PIML in BPS.

Jiang, Zixin↗

Dissimilar Material Joining via Interlocking Metasurfaces

Background The integration of dissimilar materials poses a significant challenge in engineering, necessitating innovative solutions for robust and reliable joining. Interlocking metasurfaces (ILMs) are a new joining technology comprising arrays of autogenous features patterned across two surfaces that interlock to form robust structural joints. Objective Here, this study elucidates the factors influencing the tensile performance of ILM joints formed between dissimilar materials. Methods We employed parametric optimization to identify optimal unit cell geometries for maximal yield strength based on the hypothesis that the elastic tensile properties of the materials are the primary determinants of tensile performance. Experimental validation was performed by mechanically testing the theorized optimal ILM geometry and a range of ILM geometries to capture the overall behavior trends of joints between two additively manufactured polymers, VeroPureWhite (VW) and RGDA8430-DM (8430). Results Experimental validation of optimized designs revealed that additional factors, e.g. flexural strength and localized plasticity, also strongly influenced the tensile performance of T-slot ILMs joining dissimilar materials. The proposed optimal design remained the best performer. Conclusions This study demonstrates the viability of ILMs as a joining method for dissimilar materials. ILMs can join dissimilar materials with no loss in joint yield strength compared to joints composed solely of the weaker of the two constitutive materials. ILMs demonstrated their potential as a versatile and effective joining technology in diverse engineering applications.

Elbrecht, Benjamin James [Sandia National Laborato↗

Catching Rays: How Bifacial_Radiance Sheds Light on the Future of Solar PV

The challenge of energy transition is immediate and immense, with current projections targeting 75 TW of photovoltaics (PV) capacity globally by 2050. Alongside the rapid deployment is the "solar-coaster" ride the PV industry experiences with evolving technologies and novel installation methods. In 2016, NREL developed bifacial_radiance, a python open-source modeling tool for bifacial PV. This tool is a wrapper of the raytracing engine Radiance, which you all know better than us at this workshop. Bifacial_radiance integrates the many characteristics of common PV systems to model irradiance on both the front and rear sides of bifacial PV technology - a technology that now represents 75% of utility-scale deployment in the US. Bifacial_radiance has been pivotal for understanding bifacial system performance, shading, and edge effects, and now agrivoltaics research. It has also helped develop simplified models used in PV due diligence tools for optimizing new deployments or evaluating the performance of existing projects. Now, it's the go-to comparison tool for many university, and industry-developed systems modeling tools, and a pivotal tool for further research in photovoltaics. This talk will cover the needs bifacial_radiance addresses as an open-source tool, its development path, and the opportunity for any raytracer to shine light on the solar industry through research and practical application of modeling in regular site installations and novel setups like agrivoltaics and vertical panels at high latitudes (and even the South Pole!).

agrivoltaics↗

Investigating the Effects of Individual Neutron-Induced Defects in Bipolar Junction Transistors

Here, this study investigates neutron-induced displacement damage in Bipolar Junction Transistors (BJTs) using TCAD models informed by Deep-Level-Transient-Spectroscopy (DLTS) data. These models are calibrated and validated against experimental measurements performed at various neutron fluences. Both npn and pnp transistor configurations are studied to analyze the effects of individual traps on carrier recombination and base leakage currents. In npn transistors, deep traps (0.42 eV from the conduction band) dominate at low voltages, while shallow traps (0.17 eV from the conduction band) become prominent at higher voltages. Conversely, pnp transistors have base leakage current predominantly due to deep-level traps. The study observes a notable trend in trap density versus fluence, characterized by a linear relationship on a log-log scale. These insights into defect evolution under radiation conditions are crucial for optimizing semiconductor device reliability and performance in radiation-prone environments.

42 ENGINEERING↗

Learning nuclear cross sections across the chart of nuclides with graph neural networks

We explore the use of deep learning techniques to learn how nuclear cross sections change as we add or remove protons and neutrons. As a proof of principle, we focus on the neutron-induced reactions in the fast energy regime. Our approach follows a two-stage learning framework. First, we apply representation learning to encode cross section data into a latent space using either variational autoencoders (VAEs) or implicit neural representations (INRs). Then, we train graph neural networks (GNNs) on the resulting embeddings to predict missing values across the nuclear chart by leveraging the topological structure of neighboring isotopes. We demonstrate accurate cross section predictions within a 9 × 9 block of missing nuclei. We also find that the optimal GNN training strategy depends on the type of latent representation used, with VAE embeddings performing best under end-to-end optimization in the original space, while INR embeddings achieve better results when the GNN is trained only in the latent space. Furthermore, using clustering algorithms, we map groups of latent vectors into regions of the nuclear chart and show that VAEs and INRs can discover some of the neutron magic numbers. These findings suggest that deep-learning models based on the representation encoding of cross sections combined with graph neural networks hold significant potential in augmenting nuclear theory models, e.g., by providing reliable estimates of covariances of cross sections, including cross-material covariances.

Machine learning↗

CGSim: A Simulation Framework for Large Scale Distributed Computing Environment

Large-scale distributed computing infrastructures such as the Worldwide LHC Computing Grid (WLCG) require comprehensive simulation tools for evaluating performance, testing new algorithms, and optimizing resource allocation strategies. However, existing simulators suffer from limited scalability, hardwired algorithms, lack of real-time monitoring, and inability to generate datasets suitable for modern machine learning approaches. We present CGSim, a simulation framework for large-scale distributed computing environments that addresses these limitations. Built upon the validated SimGrid simulation framework, CGSim provides high-level abstractions for modeling heterogeneous grid environments while maintaining accuracy and scalability. Key features include a modular plugin mechanism for testing custom workflow scheduling and data movement policies, interactive real-time visualization dashboards, and automatic generation of event-level datasets suitable for AI-assisted performance modeling. We demonstrate CGSim’s capabilities through a comprehensive evaluation using production ATLAS PanDA workloads, showing significant calibration accuracy improvements across WLCG computing sites. Scalability experiments show near-linear scaling for multi-site simulations, with distributed workloads achieving 6 × better performance compared to single-site execution. The framework enables researchers to simulate WLCG-scale infrastructures with hundreds of sites and thousands of concurrent jobs within practical time budget constraints on commodity hardware.

Vatsavai, Sairam Sri [Brookhaven National Laborato↗

A Full-scale Demonstration of Pressurized Water Reactor Core Design Optimization using Multi-Cycle Optimization Methodology

The U.S. nuclear sector encounters a difficulty in upholding essential safety standards while also securing economic viability for continued operation. Safety stands as a pivotal factor across all facets of operations within light-water reactor nuclear power plants. Achieving economic feasibility alongside safety can be facilitated through the utilization of a risk-informed framework, exemplified by the ongoing development within the Risk-Informed Systems Analysis Pathway under the auspices of the U.S. Department of Energy's LWRS Program. This initiative advocates for a diverse array of research and development endeavors aimed at optimizing both safety and economic efficacy within nuclear power plants, particularly pertinent as many plants contemplate second license renewals. The Risk-Informed Systems Analysis Pathway has two main goals: deploy methodologies and technologies that better represent safety margins and cost and safety factors and develop advanced applications that enable cost-effective plant operation. This report assesses the potential for resolving multi-cycle plant reload challenges through real-world scenarios utilizing the Plant ReLoad Optimization (PRLO) framework. This framework offers reactor core design developers analytic tools of reactor safety and fuel performance with the assistance of artificial intelligence (AI) to enhance core design solutions. Multi-objective genetic algorithm alongside acceleration techniques is explored as an enabling technology for improving fuel efficiency while upholding safety thresholds. The demonstration of multi-cycle core design optimization is performed. This report investigates the practical application of the PRLO platform in addressing real-world core design challenges, supporting AI efforts, and contrasting outcomes with those derived from heuristic or conventional algorithms.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Faster Randomized Dynamical Decoupling

We present a randomized dynamical decoupling (DD) protocol that can substantially improve the performance of any given deterministic DD scheme for suppressing coherent noise by using no more than two additional pulses. Our construction is implemented by probabilistically applying sequences of pulses, which, when combined, effectively eliminate the error terms that scale linearly with the system-environment coupling strength. As a result, we show that a randomized protocol using a few pulses can outperform deterministic DD protocols that require considerably more pulses. Furthermore, we prove that the randomized protocol provides an improvement compared to deterministic DD sequences that aim to reduce the error in the system’s Hilbert space, such as Uhrig DD, which had been previously regarded to be optimal. To rigorously evaluate the performance, we introduce new analytical methods suitable for analyzing higher-order DD protocols that might be of independent interest. Here, we also present numerical simulations confirming the significant advantage of using randomized protocols compared to widely used deterministic protocols.

Quantum algorithms & computation↗

Utilizing waste heat in wastewater treatment plants for water desalination: Modeling and Multi-Objective optimization of a Multi-Effect desalination system using Decision Tree Regression and Pelican optimization algorithm

This paper examines the feasibility of using waste heat from wastewater treatment plants (WWTPs) for water desalination. A model was developed to utilize waste heat from the gensets at As Samra WWTP in Jordan, using real data and TRNSYS® software to calculate available waste heat. The desalination process was then modeled with ASPEN PLUS® software, focusing on multi-effect desalination (MED). Both series and parallel configurations for the MED system were compared. The study investigated the effects of system feeding flow rate, feeding pressure, and heat input on productivity, performance ratio, and recovery ratio. The study also introduces a novel optimization technique combining machine learning and modern optimization algorithms to maximize system productivity and performance. Initially, a decision tree regression (DTR) model is developed to establish relationships between key independent variables (flow rate, feed pressure, and heat input) and dependent variables (productivity, performance ratio, and recovery ratio). The Pelican Optimization Algorithm (POA) is then used to identify the optimal values of the independent variables for maximum productivity and performance. The results show that using a series configuration yields a system productivity of 3984.2 kg/hr, a performance ratio of 3.78, and a recovery ratio of 0.991 at a feed flow rate of 4000 kg/hr, feed pressure of 3 bars, and heat input of 719 kW. Optimal productivity (4421 kg/hr), performance ratio (3.81), and recovery ratio (0.851) are achieved at a feed flow rate of 5166 kg/hr, feed pressure of 3.2 bars, and heat input of 794 kW. In conclusion, the techno-economic assessment indicates a levelized cost of water of 1.63 USD/m 3 for parallel configurations and 1.65 USD/m 3 for series configurations, with a payback period of less than two years.

42 ENGINEERING↗

SYCL for Performance Portability: Application Experience with Coupled Cluster Formalism in Quantum Chemistry on Exascale Systems

The exascale computing has brought unprecedented heterogeneity in node architectures, with systems such as Frontier and Aurora featuring diverse GPU accelerators, network connectivity among others. Ensuring performance portability across these platforms is a key challenge. To address this, we employ the SYCL programming model to develop portable, high-performance quantum chemistry workloads. As a representative application, we focus on the non-iterative Triples component of the coupled-cluster CCSD(T) method, a key driver in quantum chemistry. In this work, we report on our experience deploying SYCL-based implementations using both DPC++ and AdaptiveCPP across two flagship exascale platforms: OLCF Frontier with AMD MI250X GPUs and ALCF Aurora with Intel GPUs. Our results demonstrate that SYCL enables efficient, single-source implementations that scale to thousands of nodes, delivering performance on par with vendor-optimized HIP solutions. We highlight key insights into runtime behavior, kernel portability, and scaling characteristics, showing that SYCL offers a viable path for performance-portable computing.

Bagusetty, Abhishek [Argonne National Laboratory (↗

Solid State Power Substation DC Node Optimization and Controller Hardware-In-The-Loop Demonstration

A solid state power substation (SSPS) node is a microgrid that integrates distributed energy resources and loads and injects/absorbs power to/from the SSPS distribution network. It is an essential building block of a futuristic distribution grid network. This paper presents the development and demonstration of optimization use cases of a SSPS DC node. By adopting multi-layer hierarchical control architecture and developing automatic device identification and dynamic optimization formulation algorithms, the SSPS DC node can perform plug-and-play resource integration and seamless transition of the optimized node operation under on and off grid condition without sophisticated algorithms, control mode changes, and user interactions. Four optimization use cases including economic dispatches with price signal changes, a sudden PV power drop, and a single directional meter and its associated costs with sending power back to the grid, and resiliency under a grid inverter trip condition were demonstrated through the real-time controller hardware-in-the-loop simulation.

Kim, Namwon↗

Roll-to-Roll Manufacturing of Solid Oxide Fuel Cells

The overall goal of this project is to develop a high-volume electrode electrolyte assembly (EEA) production capability to significantly increase throughput of solid oxide fuel cell (SOFC) manufacturing and reduce the cost while maintaining the same level of performance. Specifically, four approaches will be adopted: 1) optimization of the lamination process and correlation of the EEA properties and performance with the lamination conditions; 2) scale up of the lamination process and demonstration of >10 ft of EEA; 3) further increase of the EEA throughput via slot-die coating and demonstration of > 5 m/min in coating the thick anode layer; and 4) minimization of the anode thickness to reduce material cost.

30 DIRECT ENERGY CONVERSION↗

Optimal load attachment of a deeply embedded ring anchor in clay

A Deeply Embedded Ring Anchor (DERA) system has been developed as a cost-effective solution for mooring arrays of floating offshore wind turbines (FOWTs) to the seabed. The DERA boasts several key features, including its versatility in various soil types, compact size, compatibility with diverse mooring systems, multi-line potential, and robust performance even under unintentional loading conditions. While prior preliminary studies have provided valuable insights into how the DERA can enhance cost-effectiveness by offering a high load capacity, these studies have predominantly focused on optimizing anchor performance under translational horizontal and vertical loading. However, to design the DERA optimally, we must also consider its ability to handle inclined loading conditions in addition to lateral and axial loadings. Due to its shorter length compared to a conventional caisson, the DERA has less resistance to moments, making it more sensitive to horizontal load capacity and the optimal load attachment depth concerning load angle. For this reason, our study introduces an analytical approach to evaluate the effects of inclined loading on anchor performance, utilizing the previously validated upper bound plastic limit analysis (PLA) method. In investigating the optimal load attachment of the DERA, this paper conducts a parametric study to analyze how factors such as load attachment depth, anchor aspect ratio, and load inclination affect the DERA’s load capacity. Our findings indicate that PLA can serve as a valuable analytical tool for assessing the ultimate load capacity of the DERA, particularly under inclined loading conditions.

42 ENGINEERING↗

Hydrogen Recombiner Catalyst Evaluations for Waste Storage

Radiolysis of water in nuclear waste storage generates hydrogen gas that can accumulate within sludge style waste and be rapidly released during agitation events, creating a significant flammability hazard. Engineering controls are therefore required to limit hydrogen concentrations during both quiescent storage and transient disturbances. Catalytic recombination of hydrogen in waste storage offgas is a proven mitigation strategy, maintaining hydrogen levels below flammability limits and managing sudden concentration spikes. Conventional recombiners rely on platinum and/or palladium catalysts, with development efforts focused on extending service life, increasing active surface area, and ensuring safe deployment in radioactive environments. Savannah River National Laboratory (SRNL) is evaluating a newly developed hydrogen recombiner catalyst from Canadian Nuclear Laboratories as a cost-effective and durable alternative for nuclear waste applications. Testing was conducted in SRNL’s Shielded Cells facility, which enables reduced-scale experimental modeling under radiation fields and near-use-case conditions relevant to radioactive waste storage. Catalyst performance was evaluated using a custom offgas characterization system designed for near-zero flow conditions. The experimental apparatus consisted of a gas-tight 2.7 L PTFE vessel equipped with temperature monitoring, gas flow controls, and a variable-speed mixer to simulate sludge agitation. Offgas composition was monitored using a dedicated gas chromatograph with argon carrier gas and a krypton internal standard. Measurements were obtained for an empty vessel, the vessel containing a well characterized radioactive tank waste sample, and the same configuration with the candidate catalyst installed. Results demonstrate that the new catalyst effectively reduced hydrogen concentrations in the offgas within the constraints of the experimental design. In addition to confirming catalytic activity, the testing provided valuable insights into experimental optimization and considerations for future performance evaluations. These findings support the potential scalability of the technology and highlight its applicability to broader nuclear waste management operations, offering improved safety and reduced operational costs through enhanced catalyst durability and lower replacement frequency.

Tener, Zachary P. [Savannah River National Laborat↗

Expediting field-effect transistor chemical sensor design with neuromorphic spiking graph neural networks

Improving the sensitive and selective detection of analytes in a variety of applications requires accelerating the rational design of field-effect transistor (FET) chemical sensors. Achieving high-performance detection relies on identifying optimal probe materials that can effectively interact with target analytes, a process traditionally driven by chemical intuition and time-consuming trial-and-error methods. To address the difficulties in probe screening for FET sensor development, this work presents a methodology that combines neuromorphic machine learning (ML) architectures, specifically a hybrid spiking graph neural network (SGNN), with an enriched dataset of physicochemical properties through semi-automated data extraction using large language models. Achieving a classification accuracy of 0.89 in predicting sensor sensitivity categories, the SGNN model outperformed traditional ML techniques by leveraging its ability to capture both global physicochemical properties and sparse topological features through a hybrid modeling framework. Next-generation sensor design was informed by the actionable insights into the connections between material properties and sensing performance offered by the SGNN framework. Through virtual screening for the detection of per- and polyfluoroalkyl substances (PFAS) as a use case, the effectiveness of the SGNN model was further validated. Density functional theory simulations confirmed graphene as a promising active material for PFAS detection as suggested by the SGNN framework. By bridging gaps in predictive modeling and data availability, this integrated approach provides a strong foundation for accelerating advancements in FET sensor design and innovation.

Ferreira, Rodrigo Pires [Univ. of Chicago, IL (Uni↗