Search NASA⌕ Search

SEARCH · Search NASA

Results for “Design Space Exploration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

LENS: Learning Enabled Network Synthesis

RTRC and UMD have developed novel machine learning based methods under the ARPA-E DIFFERENTIATE program for rapid acceleration of hypothesis generation in complex architecture design spaces involving both discrete choices of component inclusion and interconnection and continuous parametric decisions. The project named Learning Enabled Network Synthesis (LENS) further demonstrated the developed methods on challenging electrical power converter design problems by identifying the most suitable circuit topologies and simultaneously selecting the most appropriate components to achieve optimized design of power converter with improved performances. We demonstrated that LENS could enable exploration of very large design space of circuit topologies and components by addressing the limitations of conventional design process in non-linear, high switching speed, multi-dimensional power converter design and optimization. The key innovation developed in LENS is the seamless integration of statistical learning and logical reasoning techniques and building on the individual strengths of these techniques for rapid hypothesis discovery. The main component of LENS comprises of: 1) Graph Reasoning Engine (GRE) to enforce composition rules that rapidly reject all discrete architectures that are composed incorrectly and generates an adaptive database of feasible designs which can be used by ML modules, 2) Graph Generative Learning module which is a deep neural network based generative model for graph architectures which can enable design space exploration beyond the dataset generated by the GRE, 3) Graph Reduced Order Model (ROM) for graph domains for accelerating computation of output metrics, and 4) Active learning and Rule Discovery module for sample efficient learning and extracting logical rules from the learned ML models which will be integrated in the GRE to enhance the filtering effectiveness. LENS approach can be applied to any design domains where designs can be represented as multi-attribute graphs. The LENS team integrated the various technical innovations listed above into an optimization pipeline and exercised the optimization pipeline on the converter design problem. The LENS project demonstrated that the developed AI/ML technologies can be used to generate novel converter circuits >45x faster than experts on chosen use-cases. This can enable faster design space exploration and identification of new designs which are not considered by experts due to the increasing design space complexity. This has significant potential impact on the public and energy needs of the country. It is currently estimated that 30% of all electrical powers generated passes through power converters. The future estimate is that 80% of all power generated would be passing through converters. LENS fills a critical gap in this space since by accelerating the design process the designers would be able to generate more efficient converters which can lead to significant energy savings for the country.

42 ENGINEERING↗

COSMIC DAWN: Distributed Analysis of Wireless at Nextscale

Distributed Analysis of Wireless at Nextscale (DAWN) is a novel simulation framework for large-scale design-space exploration (DSE) of unmodified software-defined radio (SDR) applications interacting in a scalable, high-fidelity, virtual physics environment. The software-defined nature of the coupled software-physics simulation leverages hardware emulation to permit in-depth examination and modification of not only the electromagnetic environment, including each signal in flight, but also the precise state of system software and components. DAWN supports modular, customizable physics environments allowing realistic propagation effects so that computationally efficient empirical models, reduced order/surrogate models, or large-scale, high-fidelity, site-specific simulations can be used as a propagation medium based on scenario requirements. This paper introduces DAWN’s design and initial implementation, detailing key architectural components, including the Physics Realization Engine (PhyRE), Runtime Infrastructure for Simulation Environments (RISE), and the design space exploration (DSE) suite. It concludes with demonstrations using unmodified 4G/LTE software available from srsRAN on computing resources ranging from a small cluster to ORNL’s Frontier Exascale system.

Wise, Mike [ORNL] (ORCID:0000000266120641)↗

Subsampling of Parametric Models with Bifidelity Boosting

Least squares regression is a ubiquitous tool for building emulators (a.k.a. surrogate models) of problems across science and engineering for purposes such as design space exploration and uncertainty quantification. When the regression data are generated using an experimental design process (e.g., a quadrature grid) involving computationally expensive models, or when the data size is large, sketching techniques have shown promise at reducing the cost of the construction of the regression model while ensuring accuracy comparable to that of the full data. However, random sketching strategies, such as those based on leverage scores, lead to regression errors that are random and may exhibit large variability. To mitigate this issue, we present a novel boosting approach that leverages cheaper, lower-fidelity data of the problem at hand to identify the best sketch among a set of candidate sketches. This in turn specifies the sketch of the intended high-fidelity model and the associated data. We provide theoretical analyses of this bifidelity boosting (BFB) approach and discuss the conditions the low- and high-fidelity data must satisfy for a successful boosting. In doing so, we derive a bound on the residual norm of the BFB sketched solution relating it to its ideal, but computationally expensive, high-fidelity boosted counterpart. Finally, empirical results on both manufactured and PDE data corroborate the theoretical analyses and illustrate the efficacy of the BFB solution in reducing the regression error, as compared to the nonboosted solution.

97 MATHEMATICS AND COMPUTING↗

Inverse design of hypoeutectoid pearlite steel microstructures using a deep learning and genetic algorithm optimization framework

Goal-oriented microstructure design in metallic materials is a challenging task due to complex structure-property relationships. Traditional experimental and computational approaches are time-intensive and economically inefficient, limiting their applicability for large-scale design space exploration. Here, in this work, we propose an end-to-end framework that integrates deep learning models with genetic optimization to design microstructures with targeted mechanical properties. Deep learning models enable accurate forward design, while their integration with genetic optimization enables efficient inverse design within a few hours, compared to days or weeks using conventional finite element simulations. The framework combines experimental characterization and finite element modeling to analyze the influence of microstructural features on the mechanical behavior of hypoeutectoid steels. Data from both experiments and simulations are used to train the deep learning models. To demonstrate its effectiveness, we apply the framework to 0.63% carbon steel with proeutectoid ferrite and pearlite phases, commonly used in industrial applications. In this study, 2D microstructures were used for modeling, selected primarily for computational efficiency and to establish proof of concept. The framework successfully optimizes microstructures for targeted yield strength, ultimate strength, and stress concentration factors while significantly reducing computational time. Beyond hypoeutectoid steels, this scalable framework can be extended to other material systems and integrated with additive manufacturing, offering an efficient approach for accelerating microstructure design for specific engineering applications.

ConvLSTM↗

A Synthesis Methodology for Intelligent Memory Interfaces in Accelerator Systems

Domain-specific systems improve the performance of a specific set of applications compared to general-purpose processing systems by deploying custom hardware accelerators. These hardware accelerators are generated using high-level synthesis (HLS) tools. The HLS tools enable a comprehensive design space exploration to optimize the compute performance of the generated accelerators. However, they often ignore the challenges of implementing the accelerators in a system-on-chip, particularly how the accelerators access memory. Our work introduces a buffering system design that improves accelerators' memory accesses by intelligently employing burst transactions to prefetch useful data from external memory to on-chip local buffers. Our design is dynamic, parametric, and transparent to the accelerators generated by HLS tools. We derive the buffering system parameters using appropriate compiler-based analysis passes and memory channel latency constraints. The proposed buffering system design results in, on average, 8.8x performance improvements while lowering memory channel utilization on average by 53.2% for a set of PolyBench kernels.

Limaye, Ankur M. (ORCID:0000000194062584)↗

Multiphysics Design Optimization and Additive Manufacturing of Nuclear Components (Final CRADA Report - Executive Summary)

Westinghouse Electric Company (WEC) actively participated in the advancement of the nuclear fuel and reactor design space and requested the help of Oak Ridge National Laboratory (ORNL) in the creation of a new design tool set. This report details the creation of a collection of software tool sets that are linked together to collectively assist WEC design engineers in developing novel ideas outside the normal scope of traditional nuclear fuel and reactor design formulas. Specifically, Siemens HEEDS, a design space exploration and parametric optimization software, monitored and changed parameters in a collection of softwares to meet the team’s objective. The HEEDS parametric optimization method, SHERPA, was developed to control the Siemens NX CAD platform to adjust the native CAD of a hexahedral spacer grid. This new geometry can be used to execute a topological design optimization by the NX Topology software add-in. The resulting geometry is additively manufacturable. This topological optimization occurred twice—once on the spacer grid’s spring, and once on the dimple geometry. These new geometries were imported by Siemens’ STAR-CCM+, a multiphysics structural and fluid dynamic computational solver in which the spring geometry is deflected to match the rod insertion configuration. Along with the dimple geometry, this new deflected spring was used to complete a hydraulic assessment of a single-unit cell comprising one rod, one spring, and two dimples. The HEEDS SHERPA algorithm ranks the design based on the final mass of the unit cell and the hydraulic pressure drop performance. The ORNL team demonstrated the ability to use this software and provided engineering judgement to apply modern aerospace aerodynamic design. The effort has been focused on thinking outside the conventional design space and redesigning a spacer grid to perform beyond the WEC set objectives. Furthermore, the ORNL team also demonstrated that the HEEDS optimization routine can independently develop a design that meets the WEC design goals. Although these designs were at a low technology readiness level, their demonstration confirmed the team’s capability to create novel advanced nuclear concepts.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Theoretical modeling of a bottom-raised oscillating surge wave energy converter structural loadings and power performances

Here, this study presents theoretical formulations to evaluate the fundamental parameters and performance characteristics of a bottom-raised oscillating surge wave energy converter (OSWEC) device. Employing a flat plate assumption and potential flow formulation in elliptical coordinates, closed-form equations for the added mass, radiation damping, and excitation forces/torques in the relevant pitch-pitch and surge-pitch directions of motion are developed and used to calculate the system's response amplitude operator and the forces and moments acting on the foundation. The model is benchmarked against numerical simulations using WAMIT and WEC-Sim, showcasing excellent agreement. The sensitivity of plate thickness on the analytical hydrodynamic solutions is investigated over several thickness-to-width ratios ranging from 1:80 to 1:10. The results show that as the thickness of the benchmark OSWEC increases, the deviation of the analytical hydrodynamic coefficients from the numerical solutions grows from 3% to 25%. Differences in the excitation forces and torques, however, are contained within 12%. While the flat plate assumption is a limitation of the proposed analytical model, the error is within a reasonable margin for use in the design space exploration phase before a higher-fidelity (and thus more computationally expensive) model is employed. A parametric study demonstrates the ability of the analytical model to quickly sweep over a domain of OSWEC dimensions, illustrating the analytical model's utility in the early phases of design.

13 HYDRO ENERGY↗

Extending High-Level Synthesis with AI/ML Methods

Artificial Intelligence (AI) and Machine Learning (ML) methods provide significant opportunities of improving quality of results when performing high-level synthesis (HLS). For example, they can be used to model and predict metrics of the final design (e.g., area, considering aspects such as interconnect overhead for different device technologies), facilitating exploration when searching for the best design trade-offs. They can also enable identifying hidden correlations across the various phases of the synthesis and the various optimizations performed, identifying the most effective pipelines. Finally, in more general terms, bio-inspired heuristic algorithms can improve the design space exploration for the synthesis process in terms of time and quality of the result. This paper discusses opportunities and challenges to augment HLS with AI/ML using as example flow the SODA Synthesizer, an open-source hardware generation toolchain which includes SODA-OPT, a hardware/software partitioning and pre-optimization tool developed with the MLIR framework, and PandA-Bambu, a state-of-the art HLS tool. SODA interfaces with OpenROAD to provide a complete end-to-end toolchain.

artificial intelligence↗

A novel design optimization framework to sustain remanufacturability

The ever-increasing global carbon emissions have urged the need for environmentally conscious/sustainable product design, for which the design for remanufacturing (DfRem) is one potential approach. DfRem targets at designing products that have multiple life cycles, thus significantly reducing raw material usage, energy consumption, and carbon emissions. In this paper, we develop a three-stage framework that consists of (1) systematic design space exploration and a multi-objective optimization formulation to minimize the likelihood of failure causes (such as fatigue and wear) and environmental footprint, (2) topology optimization to further reduce material usage without significantly affecting the load-carrying capability of the product, and (3) post-topology optimization design verification to ensure the proposed design satisfies all design constraints. The environmental impact can be assessed at varying comprehensiveness levels (e.g., design and manufacturing phase, use phase) and in terms of carbon or GHG emission, energy use, and waste generation. Because the novel design framework predominantly adjusted the geometry, we focused on mass-based change and energy savings due to sustained remanufacturability. The multi-objective optimization formulation in the first step results in a Pareto optimal set of possible design solutions that the designer can use for the second step. Finally, we demonstrate the utility of this framework through a case study of an engine cylinder head subjected to thermo-mechanical loads, where we find that about 5% of the product mass can be conserved with only about a 3% increase in surface area that has a fatigue life less than 10,000 cycles.

42 ENGINEERING↗

An Integrated Framework for Memory-Centric Analysis: From Trace Collection to Co-Design

The memory wall phenomenon—where advances in processor performance significantly outpace those in memory subsystems—poses a fundamental challenge for contemporary computing systems. In memory-bound applications, memory subsystem behavior dominates performance, yet existing analysis approaches present significant limitations: detailed microarchitectural simulators require days to weeks to simulate modest workloads; hardware performance counters provide only aggregate statistics that obscure temporal and spatial access patterns; and scaled simulation approaches face challenges in capturing certain behaviors that emerge at larger scales. These limitations reflect a processor-centric design philosophy increasingly misaligned with memory-bound workloads where detailed understanding of memory access patterns, cache hierarchy interactions, and contention is critical for effective optimization. This paper presents an integrated framework for memory-centric analysis that enables effective hardware-software co-design. We describe practical trace collection techniques, including hardware-assisted processor tracing with minimal overhead and portable software-based instrumentation with statistical sampling. We present multi-perspective analysis methods that examine memory behavior from temporal, sequential, spatial, and relational viewpoints, revealing distinct optimization opportunities invisible in aggregate metrics. We detail an architectural modeling framework that uses sampled traces with temporal interpolation and confidence-based filtering to evaluate cache and memory configurations. Evaluation on representative benchmarks demonstrates that this framework achieves practical accuracy (L2 cache errors of 2.64\%, confidence-filtered L3 errors of 9.92\%, bandwidth errors of 7.33\%) while providing substantial speedup (26.8×) over cycle-accurate simulation, enabling rapid design space exploration. We demonstrate how this integrated framework enables systematic identification of both hardware optimizations (memory controller tuning, bank partitioning, NUMA configuration) and software optimizations (data layout restructuring, prefetching strategies, memory-aware scheduling). Through this comprehensive treatment of the memory-centric analysis pipeline—from trace collection through architectural modeling to co-design application—we provide researchers and practitioners with practical techniques for addressing memory bottlenecks in contemporary computing systems.

Gajaria, Dhruv Mayur↗

Performance Impact and Trade-Offs for Tuning Key Architectural Parameters on CPU+GPU Systems

In this work, we performed an initial design space exploration of an accelerated processing unit (APU)—a hybrid CPU+GPU architecture that integrates both compute units (CUs) and memory into a unified system. This integration aims to reduce data movement, enhance memory locality, and improve energy efficiency by enabling the CPU and GPU to share memory directly. This effort focused on the interplay of key design components—cache line size, the number of CUs, and main memory technology—and the trade-offs of each configuration were analyzed. This paper highlights the various configurations’ impact on memory accesses, data reuse, and power utilization. The results provide valuable insights that can be leveraged to optimize APU architectures for high-performance and energy-efficient computing and thus create a balanced architecture. This optimization can be achieved by adopting dynamic cache management, runtime CU scaling, and advanced memory integration, highlighting the potential of APUs to address critical challenges in compute, data movement, and memory power consumption.

Asifuzzaman, Kazi [ORNL] (ORCID:0000000240044791)↗

From Cell to System: Accelerated hpc Simulations of BESS Aging under Frequency Regulation and Arbitrage use cases

Lithium-ion battery energy storage systems (BESS) packs have emerged as a leading solution for grid-scale energy storage, enhancing resiliency and balancing load fluctuations. Yet, experimental characterization of large-format LIB packs-particularly to assess performance and degradation over hundreds of cycles - demands substantial hardware investment and multi-year testing campaigns. In this work, we couple a hierarchical, physics-based modeling framework agnostic to electrode chemistries with high-performance computing to accelerate systems level evaluation by upto two orders of magnitude. Building on the open-source liionpack platform, we implement cell, module, and pack-scale electrochemical models enriched with mechanistic aging mechanisms and deploy them on an HPC cluster to simulate 150−200kWh systems over 500 - 1,000 cycles with in days. We subject these virtual B ESS to both constant-current cycling and realistic grid service profiles spanning frequency regulation, ramp-rate support, and energy arbitrage-and quantify the resulting degradation patterns. Our results reveal that localized cell aging can induce substantial nonuniformity at module and pack levels, with service-specific cycling protocols driving distinct aging modes. This rapid, multiscale modeling approach provides a powerful design-space exploration tool for optimizing electrical architecture, control strategies, and operational schedules to prolong pack lifetime and lower total cost of ownership.

Ayalasomayajula, Surya [ORNL] (ORCID:0009000860788↗

Accelerating Traction Motor Optimization Design with AI Surrogate Models

The advancement of artificial intelligence systems enables the use of data-driven physics-based surrogate models to explore design spaces rapidly and deeply for engineering projects. This work presents a surrogate model workflow that accelerates electric traction motor design optimization by replacing finite element analysis (FEA) with an artificial neural network (ANN) and using this model in a genetic algorithm for design optimization. A baseline interior permanent-magnet motor is parameterized and sampled to generate FEA-labeled training data, after which a feed-forward ANN predicts key outputs (e.g., loss components and weight). The validated surrogate enables genetic-algorithm optimization and deep search over the design space without new FEA runs, producing Pareto-optimal trade-offs between weight and losses and set of optimized designs for rapid downselection of manufacturable motor designs.

Ribeiro, Pedro [ORNL] (ORCID:0009000921026641)↗

Throughput Estimation of Data Transport Networks From Digital Twin Measurements

Digital twins of networked infrastructures, known as Virtual Infrastructure Twins (VITs), are increasingly used for software development, pre-deployment testing, and design space exploration. While VITs avoid the costs and potential disruptions associated with experiments on operational networks, their throughput measurements are typically not sufficiently accurate for performance profiling of wide-area networks that they emulate. Here, machine learning (ML) methods are developed to transform these inaccurate VIT network throughput measurements to closely match in peak and overall profile of those from a physical testbed or production network. First, a micro kernel network reflecting a physical network is utilized to collect one-time measurements on a host to support this ML transformation. Then, a generic multi-modal ML method is developed to learn a map that transforms measurements from subsequent VITs on the same host to match past, current and follow-on testbed and cloud networks. ML generalization equations are derived to establish its correctness and probabilistically guarantee its generalization accuracy. Experimental results are presented for a variety of VIT hosts with target testbed and cloud networks; they include a case study of a four-site science ecosystem wherein inaccurate convex VIT measurement profiles are transformed into accurate concave profiles of target networks.

97 MATHEMATICS AND COMPUTING↗

Surrogate Model Integration with MOOSE XFEM for Creep Crack Growth

Ferritic-martensitic steels are key structural materials for advanced reactors but experience time-dependent deformation and damage under prolonged high temperature and irradiation, leading to creep-driven crack initiation and growth. High-fidelity models—crystal plasticity with irradiation mechanisms, phase-field for microstructural evolution, and continuum-damage viscoplasticity—capture the underlying physics but are too computationally intensive for broad design-space exploration and uncertainty quantification. This milestone advances a scalable alternative by integrating a microstructure-sensitive surrogate creep model into the Multiphysics Object-Oriented Simulation Environment (MOOSE) finite element framework and extending it to fracture via the extended finite element method (XFEM). The surrogate model, developed with collaborators at Sandia and Los Alamos National Laboratories, maps relevant microstructural descriptors to the viscoplastic response of HT9. We embed this surrogate within a coupled deformation-damage workflow in MOOSE/XFEM to simulate creep-driven crack initiation and propagation. Implementation enhancements include updates to the material interface, a plastic correction phase involving microstructure evolution, and fracture criteria to ensure numerical robustness and compatibility with the surrogate structure. Demonstrations on canonical creep benchmarks spanning uniaxial and multiaxial states show that the surrogate reproduces key trends of high-fidelity models while substantially reducing computational cost. The resulting capability bridges physics fidelity and performance, providing a practical path to a predictive, microstructure-aware assessment of creep and fracture in reactor materials.

36 - MATERIALS SCIENCE↗

Fast Machine Learning for Quantum Control of Microwave Qudits on Edge Hardware

Quantum optimal control is a promising approach to improve the accuracy of quantum gates, but it relies on complex algorithms to determine the best control settings. CPU or GPU-based approaches often have delays that are too long to be applied in practice. It is paramount to have systems with extremely low delays to quickly and with high fidelity adjust quantum hardware settings, where fidelity is defined as overlap with a target quantum state. Here, we utilize machine learning (ML) models to determine control-pulse parameters for preparing Selective Number-dependent Arbitrary Phase (SNAP) gates in microwave cavity qudits, which are multi-level quantum systems that serve as elementary computation units for quantum computing. The methodology involves data generation using classical optimization techniques, ML model development, design space exploration, and quantization for hardware implementation. Our results demonstrate the efficacy of the proposed approach, with optimized models achieving low gate trace infidelity near $10^{-3}$ and efficient utilization of programmable logic resources.

Sanders, Flor [Columbia U.]↗

Towards Automated Generation of Chiplet-Based Systems

The Software Defined Architectures (SODA) Synthesizer is an open-source compiler-based tool able to automatically generate domain-specialized systems targeting Application- Specific Integrated Circuits (ASICs) or Field Programmable Gate Arrays (FPGAs) starting from high-level programming. SODA is composed of a high-level frontend, SODA-OPT, which leverages the multilevel intermediate representation (MLIR) framework to interface with productive programming tools (e.g., machine learning frameworks), identify kernels suitable for acceleration, and perform high-level optimizations, and of a state-of-the-art high-level synthesis backend, Bambu from the PandA framework, to generate custom accelerators. One specific application of the SODA Synthesizer is the generation of accelerators to enable ultra-low latency inference and control on autonomous systems for scientific discovery (e.g., electron microscopes, sensors in particle accelerators, etc.). This talk will discuss ongoing work on the SODA synthesizer to enable no-human-in-the-loop generation and design space exploration of the chiplets for highly specialized artificial intelligence accelerators. Connecting these highly specialized chiplets to general-purpose cores or programmable accelerators will allow to quickly deploy autonomous systems for scientific discovery.

Limaye, Ankur M.↗