Search NASA⌕ Search

SEARCH · Search NASA

Results for “hardware complexity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Preparing angular momentum eigenstates using engineered quantum walks

Coupled angular-momentum eigenstates are widely used in atomic and nuclear physics calculations and are building blocks for spin networks and the Schur transform. To combine two angular momenta J 1 and J 2 , forming eigenstates of their total angular momentum J=J 1 +J 2 , we develop a quantum-walk scheme that does not require inputting O(j 3 ) nonzero Clebsch–Gordan (CG) coefficients classically. In fact, our scheme may be regarded as a unitary method for computing CG coefficients on quantum computers with a typical complexity of O⁡(j) and a worst-case complexity of O⁡(j 3 ). Equivalently, our scheme provides decompositions of the dense CG unitary into sparser unitary operations. Our scheme prepares angular-momentum eigenstates using a sequence of Hamiltonians to move an initial state deterministically to desired final states, which are usually highly entangled states in the computational basis. In contrast with usual quantum walks, whose Hamiltonians are prescribed, we engineer the Hamiltonians in su⁡(2)×su⁡(2), which are inspired by, but different from, Hamiltonians that govern magnetic resonances and dipole interactions. To achieve a deterministic preparation of both ket and bra states, we use projection and destructive interference to double pinch the quantum walks, such that each step is a unit-probability population transfer within a two-level system. We test our state preparation scheme on classical computers, reproducing tables of CG coefficients. Finally, we also implement small test problems on current quantum hardware.

97 MATHEMATICS AND COMPUTING↗

Flexible AI Models for Grid Resilience

The rapid growth in size and complexity of artificial intelligence (AI) and machine learning (ML) models has led to increased energy demands, posing a threat to the reliability of the existing power grid. This project addresses the challenge of highly intermittent and energy-intensive inference workloads by (1) developing fidelity-adaptive neural networks capable of dynamic response to grid conditions and (2) integrating these networks with power flow simulations to assess their impact on power grid reliability. We will explore both top-down and bottom-up approaches to create hierarchies of submodels that provide a controlled trade-off between power draw and prediction accuracy. The top-down method utilizes NN pruning to reduce a flagship model into progressively smaller, energy-efficient variants. The bottom-up approach employs geometrically principled weight setting strategies to construct depth-efficient models from the ground up. A real-time hardware-in-the-loop (HIL) platform will be developed to simulate a scaled AC power grid, integrating live AI workload power draw and enabling dynamic model switching in response to grid feedback. This work will provide a novel framework for evaluating the impact of flexible AI/ML workloads on grid performance and establish new methodologies for energy-aware computing in data centers. The outcomes will demonstrate that adaptive AI/ML can play a critical role in improving grid stability while advancing NREL's leadership in energy-efficient computing research.

24 POWER TRANSMISSION AND DISTRIBUTION↗

ExaCA v2.0: A versatile, scalable, and performance portable cellular automata application for additive manufacturing solidification

The previously established ExaCA software for performance portable alloy grain structure simulation has been updated to better represent the solidification behavior during complex alloy processing conditions, such as those encountered during metal additive manufacturing (AM), and for improved performance and scalability. Here, an extension to the time–temperature history input data format and the core ExaCA algorithm to include an arbitrary number of melting and solidification events yielded improved prediction of texture for various melt pool geometries, expanding the range of AM-relevant conditions that can be accurately simulated. Improved heat transport process simulation coupling, including the creation of large raster datasets from single track time–temperature history data and in-memory coupling with the new, performance portable finite difference code Finch, were also demonstrated in example studies on the effect of multilayer AM microstructure predictions on hatch spacing and cell size, respectively. Additional new features are detailed and demonstrated, including the ability to perform simulations using various interfacial response function forms, execute simulations on state-of-the-art hardware, improved usability through post-processing versatility, and improved strong and weak scaling performance. The performance, physics, and versatility improvements demonstrated here will further enable large-scale studies on AM process–microstructure relationships that were not previously possible. Furthermore, the usability improvements and ability to run coupled AM process–microstructure simulations using the Finch-ExaCA workflow will facilitate broader use of this open-source software by the computational materials community.

36 MATERIALS SCIENCE↗

SmartFuse: Reconfigurable Smart Switches to Accelerate Fused Collectives in HPC Applications

Communication switches have sometimes been augmented to process collectives (e.g., the IBM BlueGene project and the Mellanox SHArP switch). In this work, we find that there is a great acceleration opportunity through the further augmentation of switches to accelerate more complex functions that combine communication with computation. We consider three types of such functions. The first is fully-fused collectives built by fusing multiple existing collectives like Allreduce with Alltoall. The second is semi-fused collectives built by combining a collective with another computation. The third we refer to as higher-order collectives built by combining multiple computations and communications, such as to perform matrix-matrix multiply (PGEMM). In this work, we propose a framework called SmartFuse to accelerate fused collective functions. The core of SmartFuse is a reconfigurable smart switch to support these operations. The semi/fully fused collectives are implemented with a CGRAlike architecture, while higher-order collectives are implemented with a more specialized computational unit that can also schedule communication. Supporting our framework is software to evaluate and translate relevant parts of the input program, compile them into a control data flow graph, and then map this graph to the switch hardware. The proposed framework, once deployed, has the strong potential to accelerate existing HPC applications transparently by encapsulation within an MPI implementation. Experimental results show that this approach improves the performance of the PGEMM kernel, MINIFE, and AMG by, on average, 94%, 15%, and 13%, respectively.

Haghi, Pouya↗

Classical combinatorial optimization scaling for random Ising models on 2D heavy-hex graphs

Motivated by near term quantum computing hardware limitations, combinatorial optimization problems that can be addressed by current quantum algorithms and noisy hardware with little or no overhead are used to probe capabilities of quantum algorithms such as the quantum approximate optimization algorithm. In this study, a specific class of near term quantum computing hardware defined combinatorial optimization problems, Ising models on heavy-hex graphs both with and without geometrically local cubic terms, are examined for their classical computational hardness via empirical computation time scaling quantification. Specifically the time-to-solution (TTS) metric using the classical heuristic simulated annealing is measured for finding optimal variable assignments (ground states), as well as the time required for the optimization software Gurobi to find an optimal variable assignment. Because of the sparsity of these Ising models, the classical algorithms are able to find optimal solutions efficiently even for large instances (i.e. 100 000 spin variables). The Ising models both with and without geometrically local cubic terms exhibit average-case linear-time or weakly quadratic scaling when solved exactly using Gurobi, and the Ising models with no cubic terms show evidence of exponential-time TTS scaling when sampled using simulated annealing. These findings point to the necessity of developing and testing more complex, namely more densely connected, optimization problems in order for quantum computing to ever have a practical advantage over classical computing. Our results are another illustration that different classical algorithms can indeed have exponentially different running times, thus making the identification of the best practical classical technique important in any quantum computing vs. classical computing comparison.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Advances on CHP District Energy and Microgrids Deployment: Simplified Tool for Rapidly Deploying Feasibility Analytics for the Non-Technical User (Final Technical Report)

Community energy systems have proven to have the potential to improve cost efficiency, resilience, and decarbonize. However, investing in community energy systems such as community microgrids or district energy systems is a complex decision due to the high initial investment and the uncertainties associated with the long development time and lifecycle of the project. Tools that make feasibility assessments accessible to non-technical users like investors, policymakers, and other stakeholders will result in more feasibility analyses completed, more candidate projects identified, and more community energy systems deployed. The pilot tool developed under this award is named Energy Fellow. Energy Fellow allows technical and non-technical users to complete feasibility analyses for district energy systems and community microgrids. This is the first software tool of its kind designed for non-technical users and available at no cost. Its scope was adjusted to a 25x25-mile region within the Houston area in Texas to make its development compatible with the funding available. However, the findings and models developed make this pilot tool easily scalable to the US. The lessons learned during the design, implementation, and testing stages have helped find trade-off solutions to software and hardware challenges related to implementing 3D models in online tools. Green software strategies has been successfully applied to the design and operations of the tool, and the team has researched the aspects of the (non-technical) user experience that will make commercial developments of this tool even more impactful.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

High Temperature Additive Architectures for 65% Efficiency (Final Technical Report)

This project aimed to develop advanced high-temperature additive components that contribute towards the DOE’s goal for advanced gas turbines that are capable of at least 65% efficiency in combined cycle application. The objective was to leverage state-of-the-art additive manufacturing to develop an innovative stage 1 turbine nozzle (S1N) that can provide cooling flow savings while maintaining the component durability expected in today’s gas turbines. The program had two phases. Phase I was a conceptual phase for novel advanced cooling designs enabled by additive manufacturing, as well as proposals for validation. Phase II included execution of the Phase I conceptual design, including manufacturing of prototype hardware and validation within an environment that is similar to engine operation. During Phase I of the program, the team devised a concept to reduce cooling air usage. The cooling air used in the side walls is filmed out along the side walls, while the cooling air used in the airfoil is eventually directed to near-wall channels and exits holes along the airfoil trailing end. During this program, the team performed additive trials to analyze the geometric limitations of additive manufacturing. This helped the team understand minimum wall thicknesses, hole sizes, and cooling channel dimensions among other limits. Phase II of this program pushed GE Vernova beyond its previous experience of designing and manufacturing an additively manufactured hot gas path component. Modern hot gas path components utilize material chemistries that are traditionally hard to weld, such as cast Renè 108, and exhibit solidification cracking when additively manufactured using Direct Metal Laser Melting (DMLM). Note that AM108 is a powder form of Renè 108. A S1N with advanced cooling is larger and more complex than parts previously built by additive manufacturing and required new learnings to resolve risks around solidification cracking. Finally, the team validated the design in a combustion rig that replicated operation in a gas turbine. In order to properly quantify the benefits of the new additive design, a baseline was also tested in the rig and operated under the same conditions. In addition, an uncertainty analysis was done to quantify any sources of error that could impact the results. At the end of the validation effort, it was determined that the additive nozzle exceeded the 15% reduction in cooling flow goal even with the worst-case assumptions for uncertainty.

03 NATURAL GAS↗

Dynamic, symmetry-preserving, and hardware-adaptable circuits for quantum computing many-body states and correlators of the Anderson impurity model

We present a hardware-reconfigurable ansatz on N q -qubits for the variational preparation of many-body states of the Anderson impurity model (AIM) with N imp + N bath = N q /2 sites, which conserves total charge and spin z component within each variational search subspace. The many-body ground state of the AIM is determined as the minimum over all minima of O(N$^2_ q$) distinct charge-spin sectors. Hamiltonian expectation values are shown to require ω(N q ) < N meas. $\leqslant$ O(N imp N bath ) symmetry-preserving, parallelizable measurement circuits, each amenable to postselection. To obtain the one-particle impurity Green’s function we show how initial Krylov vectors can be computed via midcircuit measurement and how Lanczos iterations can be computed using the symmetry-preserving ansatz. For a single-impurity Anderson model with a number of bath sites increasing from one to seven, we show using numerical emulation that the ease of variational ground-state preparation is suggestive of linear scaling in circuit depth and subquartic scaling in optimizer complexity. We therefore expect that, combined with time-dependent methods for Green’s function computation, our ansatz provides a useful tool to account for electronic correlations on early fault-tolerant processors. Finally, with a view towards computing real materials properties of interest like magnetic susceptibilities and electron-hole propagators, we provide a straightforward method to compute many-body, time-dependent correlation functions using a combination of time evolution, midcircuit measurement-conditioned operations, and the Hadamard test.

36 MATERIALS SCIENCE↗

Fully implicit crystal plasticity models representing orientations with modified Rodrigues parameters

Here, this work describes a crystal plasticity formulation combining several mathematical, numerical, and implementation choices to produce a highly efficient model. Specifically, the key choices in the implementation are (1) representing orientations with modified Rodrigues parameters, (2) implementing a fully coupled implicit time integration for the elastic stretch, the crystal orientations, and the model internal variables, (3) implementing the model in the NEML2 constitutive modeling framework, based on PyTorch, to vectorize the calculations and port the computation to GPUs and other hardware accelerators, and (4) an exact implementation of the consistent tangent matrix, even for arbitrary coupling to other field variables beyond the displacements, like temperature, neutron fluence, etc. The first two features of the model are, to our knowledge, novel. The paper considers each of these choices individually as well as the final model as a whole. This includes a full description of modified Rodrigues parameters, their advantages over other representations of orientations, the mathematical formulae and tools required to implement a model with modified Rodrigues parameters, and a detailed description of the geometry of the space of modified Rodrigues parameters (in an appendix). It also includes a description of a fully implicit time integration scheme for the orientations and the advantages in representing orientations with modified Rodrigues parameters in implementing such a model. The work then assess, via numerical examples, the advantages of fully coupled implicit time integration versus more common decoupled and explicit time integration schemes. These studies demonstrate the computational advantages of fully coupled integration versus other time integration algorithms, though the performance of the competing models depends on the complexity of the underlying single crystal model. The study concludes by demonstrating that the choice of time integration method affects the sharpness of the predicted texture, with explicit methods for integrating the orientations overestimating texture sharpness and implicit methods underestimating texture sharpness.

Crystal plasticity↗

IRIS: Exploring Performance Scaling of the Intelligent Runtime System and its Dynamic Scheduling Policies

High-Performance Computing is becoming increasingly heterogeneous, relying on a diverse mix of hardware to achieve good performance. Paradoxically, current drivers and frameworks for these devices typically require separate languages and implementations for each vendor. Furthermore, there are few tools and little support to schedule codes between these devices in a truly heterogeneous manner-partly because of this fragmentation between vendors and the languages each supports. To overcome both limitations, the Intelligent Runtime System (IRIS) was developed. It allows a common task abstraction to automatically be shared among contemporary vendors and is run from a single host-side API. At runtime, IRIS queries the host system and registers which frameworks and drivers are available, these determine which kernels can be used by the scheduler-CPUs via OpenMP, Nvidia GPUs (CUDA), AMD GPUs (HIP), and Intel and Xilinx FPGAs with OpenCL. IRIS enables tasks to be scheduled to any heterogeneous device and resolves to the appropriate kernel binary at runtimeit only uses the devices supported by the system on which it is run. IRIS supports single-task and graph-based expressions of dependencies of tasks. Additionally, IRIS features a range of dynamic scheduling policies, allowing complex chains of tasks and interactions to be executed, relieving the programmer/user from considering the system to assign tasks to devices optimally. This paper presents the peak performance attainable by IRIS over a range of systems-each with different numbers and types of accelerator devices, it highlights the flexibility of IRIS since these devices are truly heterogeneous, relying on different backends (drivers, frameworks, and languages) which historically required unique implementations to utilize them. We then use this peak performance as a baseline to compare increasingly complex chains of tasks (with increasingly complex task dependencies) and evaluate how IRIS copes. Finally, we consider the performance of different IRIS scheduling policies on this range of task graphs.

Johnston, Beau↗

Development of message passing-based graph convolutional networks for classifying cancer pathology reports

Abstract Background Applying graph convolutional networks (GCN) to the classification of free-form natural language texts leveraged by graph-of-words features (TextGCN) was studied and confirmed to be an effective means of describing complex natural language texts. However, the text classification models based on the TextGCN possess weaknesses in terms of memory consumption and model dissemination and distribution. In this paper, we present a fast message passing network (FastMPN), implementing a GCN with message passing architecture that provides versatility and flexibility by allowing trainable node embedding and edge weights, helping the GCN model find the better solution. We applied the FastMPN model to the task of clinical information extraction from cancer pathology reports, extracting the following six properties: main site, subsite, laterality, histology, behavior, and grade. Results We evaluated the clinical task performance of the FastMPN models in terms of micro- and macro-averaged F1 scores. A comparison was performed with the multi-task convolutional neural network (MT-CNN) model. Results show that the FastMPN model is equivalent to or better than the MT-CNN. Conclusions Our implementation revealed that our FastMPN model, which is based on the PyTorch platform, can train a large corpus (667,290 training samples) with 202,373 unique words in less than 3 minutes per epoch using one NVIDIA V100 hardware accelerator. Our experiments demonstrated that using this implementation, the clinical task performance scores of information extraction related to tumors from cancer pathology reports were highly competitive.

59 BASIC BIOLOGICAL SCIENCES↗

Distributed Quantum-Enhanced Optimization: A Topographical Preconditioning Approach for High-Dimensional Search

Optimization problems become fundamentally challenging as the number of variables increases. Because the volume of the search space grows exponentially, classical algorithms frequently fail to locate the global minimum of non-convex functions. While quantum optimization offers a potential alternative, mapping continuous problems onto near-term quantum hardware introduces severe scaling limits and barren plateaus. To bridge this gap, we propose the Distributed Quantum-Enhanced Optimization (D-QEO) framework. Instead of forcing the quantum processor to find the exact minimum, we use it simply as a topographical preconditioner. The QPU maps the landscape to locate the most promising basin of attraction, generating high-quality seed points for a classical GPU-accelerated solver to refine. To make this approach viable for utility-scale problems, we exploit the mathematical structure of separable functions. This allows us to cut a 50-qubit (i.e., $2^{50}$) global search space into independent and manageable sub-spaces using 5-qubit subcircuits. By executing these fragments concurrently with CUDA-Q, we completely bypass the overhead of cross-register entanglement and classical tensor knitting for separable functions. Benchmarks on the 10-dimensional Rastrigin and Ackley functions show that D-QEO prevents the exponential failure rates observed in purely classical algorithms. Furthermore, this quantum warm-start significantly reduces the number of classical BFGS iterations required to converge, providing a highly practical blueprint for utilizing near-term quantum resources in complex global search.

Soos, Dominik [Old Dominion U.]↗

Improving Cyber Situational Understanding

Effective cybersecurity operations require the ability to analyze large amounts of information to assess security risks and formulate defensive strategies against adversaries. This has become more complex in recent years as the sprawl and interconnectivity of devices grows through implementation of virtualization, cloud computing, and Internet of Things (IoT). The amount of data and analysis required for effective cybersecurity command and control decisions far exceeds humans’ capacity to perform manually. We characterize the analysis problem as cyber situational understanding. The research presented to improve cyber situational understanding focuses on vulnerability analysis and threat intelligence. Regarding vulnerabilities, entities must analyze and plan work for between thousands and tens of thousands of software vulnerabilities annually. Entities heavily use network firewalls to limit vulnerability exposure. As a result, some of these vulnerabilities permit exposure to adversarial exploitation, whereas others are inaccessible and therefore present negligible risk of exploitation. Distinguishing between high and low risk software vulnerabilities requires a deep understanding of the vulnerability, network firewall protection, and characteristics of the targeted device. This problem is solved by extracting network service features from vulnerability data features using both machine-learning and natural language processing. Then, the network firewall topology is parsed to determine which vulnerabilities are reachable by adversaries. Ultimately, a state-based safety analysis ascertains which vulnerabilities are unsafe. A related vulnerability analysis problem occurs in cybersecurity operations when associating an entity’s hardware and software assets to public vulnerability databases. Assets often reveal hardware and software through installation artifacts and network service identification, and entities store these artifacts in inventory databases. However, software and hardware vendors apply a standard Common Platform Enumeration (CPE) naming convention when publicly reporting vulnerabilities. Associating these two datasets often requires many hours to days of manual inspection. The proposed solution automates the mapping approach of human analysts using fuzzy matching techniques, natural language processing, and, ultimately, machine learning to present a small set of recommendations for mapping the two datasets. The result significantly reduces human analysis time and reduces the occurrence of false positives in vulnerability notifications. Finally, cyber threat intelligence (CTI) requires associating cyber observable artifacts, such as IP addresses, URIs, and file hashes, with cyber threat tactics, techniques, and procedures. Unfortunately, most CTI data is compartmentalized across multiple organizations and cannot be shared due to the legal and reputational risk with cyber threat being associated with the entity. The approach to solving this problem inovlves using a distributed ledger with anonymous token spending and authentication. This allows a consortium of semi-trusted entities to share the workload of curating CTI for a threat sharing community’s cooperative benefit.

Huff, Philip↗

Holistic Small-Signal Stability Analysis for Large-Scale Inverter-Intensive Power Systems with Coupled and Full-Order Dynamics from Control Systems and Power Networks

The increasing penetration of inverter-based resources (IBRs) into the existing power systems introduces tremendous benefits for enhanced sustainability but also poses inevitable challenges in terms of insufficient inertia, potential instability, and complex network dynamics, among others. However, the additional coupling introduced by the interactions among gridfollowing (GFL) and grid-forming (GFM) IBRs and the other components (i.e., synchronous generators [SGs], loads, and network, etc.) has not been clearly explored. A holistic, scalable, and quantitative stability analysis framework with the control systems and power networks is still missing. Here, in this paper, to fill in the technical gaps, a holistic small-signal model of the entire system with both rotating generation units and IBRs is established. An extended power flow model with operation dynamics from both generator control schemes and power networks is proposed to provide the varying steady-state operating points for small-signal modeling. The proposed method is compared with MATLAB solvers, and the results show that the proposed approach has a minimum calculation time, which can be less than 12 seconds for a large-scale power system with up to 2,000 buses. Furthermore, a quantitative method is developed to identify the impacts of IBRs on system performance with emphases on the potential stability issues with GFL IBRs, additional benefits of employing GFM IBRs, the feasibility of replacing SGs with GFM IBRs, and the impact of penetration level of different kinds of generation units. Finally, a field island power system is used to verify the proposed approach, and hardware-in-the-loop (HIL) tests are provided to further demonstrate the effectiveness of the proposed analysis.

14 SOLAR ENERGY↗

SuperLab 2.0 Showcase: Connecting Five Labs to Tackle Grid Complexity and Unlock Unique Grid Asset Potential

SuperLab 2.0 (5-Lab Demo) is a collaborative, national-scale experiment showcasing the coordination of geographically distributed energy assets in real time. The demonstration integrates 25 physical and digital assets, spanning wind, PV, batteries, electrolyzers, DC fast chargers, microgrid controllers, building automation systems, small modular reactor (SMR), control centers, and gas turbines, across five DOE national laboratories-NLR, INL, NETL, LBNL, and SNL. These assets are unified using Energy Sciences Network (ESnet), a low-latency, high-performance U.S. Department of Energy's (DOE) network, and controlled via a centralized energy controller hosted at NLR's ARIES facility. The demonstration validates the ability to stress-test hybrid energy systems under dynamic scenarios to de-risk advanced control strategies for greater resilience and flexibility. SuperLab 2.0 (5-Lab Demo) showcased a major advancement in federated national laboratory collaboration, enabling real-time, cross-laboratory experimentation to coordinate geographically dispersed distributed energy resources (DERs) using various communication protocols and networks. SuperLab 2.0 (5-Lab Demo) built on previous demonstrations conducted between NLR-PNNL and NLR-INL connecting diverse assets including distant protection devices, a SMR simulator, and a high temperature electrolyzer (HTE). Previous demos were based on a single connection between two labs with minimal coordination challenges. The 5-Lab demo with a centralized controller, distributed testbeds across different geographical locations, and use of protocols-based communication represents a scenario closer to real-world grid operations that coordinate resources across a region to meet system needs. This experiment studied how local DER controllers interact with a centralized energy controller during normal and abnormal events to maintain reliability. The SuperLab team across the five labs implemented a notional power system model equivalent of transmission and distribution lines, represented by the data networks interconnecting the labs. Each lab continuously exchanged local parameters (such as P and Q) from its Hardware-In-Loop (CHIL) and Power Hardware-In-Loop (PHIL) assets through centralized energy controller at NLR, enabling real-time interaction and coordination across sites. By leveraging ESnet as the communication backbone, the team successfully operated the distributed assets as a unified power system, with each bus represented by a different laboratory. This setup mirrors how assets interact in real-world power systems across dispersed locations with various protocols and latencies. At each lab site, assets were operated using their own local controllers which were coordinated through an overarching operation and control layer of centralized energy controller, equivalent to how an energy management system (EMS) orchestrates assets across a regional or national grid. SuperLab's federated connectivity utilized a Digital Real-Time Simulators (DRTS)-type gateway to connect Controller Hardware-In-Loop (CHIL) and PHIL assets between labs. To enable this federated connection through ESnet, a deterministic network was established where latency variations were consistent. This consistency allowed the development of digital filters for the power system assets across CHIL and PHIL interfaces to avoid unstable and unreliable grid conditions. This report provides an overview of the cross-laboratory configuration and offers insights into interconnecting geographically distributed research assets to test them as if they were co-located. This experiment represents a step toward linking nine DOE national laboratories, enabling nation-wide simulations that can address utility-driven challenges with grid resilience, flexibility, and modernization.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Optimization and Evaluation of Energy Savings for Connected and Autonomous Off-Road Vehicles

Off-road vehicles, such as wheel loaders, excavators, and harvesters, are extensively utilized across a wide range of industries, including construction, agriculture, and mining. These machines have become indispensable in supporting the day-to-day operational needs of a nation, playing a critical role in various sectors' infrastructure and productivity. However, despite their utility, off-road vehicles are significant consumers of fossil fuels, resulting in substantial emissions that contribute to environmental degradation. This highlights the pressing need for research and technological advancements aimed at improving their energy efficiency and reducing their carbon footprint. There are, however, two primary challenges that must be addressed to achieve these goals. First, off-road vehicles typically perform both driving and working tasks simultaneously, which introduces a high level of complexity into their overall dynamic systems. Analysis the interactions between these functions is challenging. Second, research into off-road vehicles is inherently interdisciplinary, demanding expertise across several domains such as fluid power systems, vehicle dynamics, control theory, optimization techniques, and real-world implementation. Recognizing these challenges, we proposed the project titled "Optimization and Evaluation of Energy Savings for Connected and Autonomous Off-Road Vehicles" as a comprehensive solution to enhance fuel efficiency while simultaneously improving productivity. This project specifically focuses on autonomous off-road vehicles, with particular attention to wheel loaders, and seeks to develop novel methods to optimize energy consumption without sacrificing operational performance. The project integrates real-time control algorithms, vehicle dynamics modeling, and co-optimization of powertrain system and vehicle system to achieve these goals. Our optimization strategy dynamically co-optimizes critical parameters at both the powertrain and vehicle levels, including vehicle speed, working tool movements, powertrain dynamics, and engine operations in real-time. To streamline this optimization process, we developed a vehicle model that captures the key dynamics while significantly enhancing computational efficiency. This allows the system to intelligently minimize fuel consumption, all while maintaining or even improving productivity through real-time calculations during various off-road operations. To validate the effectiveness of this energy optimization method, we introduced a state-of-the-art Hardware-in-the-Loop (HIL) testbed. This reconfigurable testbed seamlessly integrates the actual engine with virtual models of the wheel loader's subsystems, allowing for accurate emulation of real-world operational loads and environments. By simulating these conditions, the HIL testbed enables us to evaluate the wheel loader’s performance under diverse working scenarios, ensuring the developed solution is applicable in real-world operations. This testbed proved to be instrumental in validating the optimization algorithms and demonstrating the system's practical effectiveness. During the evaluation and testing phase, we employed the HIL testbed to rigorously assess the energy savings and productivity improvements generated by the optimized system. The results were highly encouraging, revealing that the automated wheel loader achieved over 30% fuel savings compared to traditional, human-operated cycles, with comparable or even enhanced levels of productivity. The insights gained from this HIL-based testing provided critical validation of our approach and highlighted the potential for deploying these optimized autonomous technologies in real-world off-road vehicles.

33 ADVANCED PROPULSION SYSTEMS↗

FiberFlex: Real-time FPGA-based Intelligent and Distributed Fiber Sensor System for Pedestrian Recognition

In recent years, security monitoring of public places and critical infrastructure has heavily relied on the widespread use of cameras, raising concerns about personal privacy violations. To balance the need for effective security monitoring with the protection of personal privacy, we explore the potential of optical fiber sensors for this application. This article proposes FiberFlex, an intelligent and distributed fiber sensor system. Ultizing Field Programmable Gate Arrays (FPGA) high-level synthesis (HLS) acceleration, FiberFlex offers real-time pedestrian detection by co-designing the entire pipeline of optical signal acquisition, processing, and recognition networks based on the principles of optical fiber sensing. As a promising alternative to traditional camera-based monitoring systems, FiberFlex achieves pedestrian detection by analyzing the vibration patterns caused by pedestrian footsteps, enabling security monitoring while preserving individual privacy. FiberFlex comprises three modules: First , fiber-optic sensing system: A fiber-optic distributed acoustic sensing (DAS) system is built and used to measure the ground vibration waves generated by people walking. Second , algorithms: We first collect the training data by measuring the ground vibration waves, label the data, and use the data to train the neural network models to perform pedestrian recognition. Third , hardware accelerators: We use HLS tools to design hardware modules on FPGA for data collection and pre-processing and integrate them with the downstream neural network accelerators to perform in-line real-time pedestrian detection. The final detection results are sent back from FPGA to the host CPU. We implement our system FiberFlex with the in-house built DAS system and AMD/Xilinx Kintex7 FPGA KC705 board and verify the whole system using the real-world collected data. We conduct recognition tests on five test subjects of varying ages, heights, and weights in a fixed sensing area. Each subject experienced 20 real-time recognition tests using their daily walking habits, and the subjects were given adequate rest between tests. After 100 tests on five test subjects, the overall real-time recognition accuracy exceeded \(88.0\%\) . The whole system uses 55 W of power, 33 W in the optical DAS system and 22 W in the FPGA. Relying on its end-to-end interdisciplinary design, FiberFlex seamlessly combines fiber-optic sensors with FPGA accelerators to enable low-power real-time security monitoring without compromising privacy, making it a valuable addition to the existing security monitoring network. According to FiberFlex, more valuable research can be conducted in the future, such as fall monitoring for the elderly, migration of identification networks between different application scenarios, and improvement of anti-interference performance in more complex environments. In future perception networks, where the “eyes” are not feasible, let’s use fiber optic touch instead.

Distributed↗

Reduced-order modeling on a near-term quantum computer

Quantum computing is an advancing area of research in which computer hardware and algorithms are developed to take advantage of quantum mechanical phenomena. In recent studies, quantum algorithms have shown promise in solving linear systems of equations as well as systems of linear ordinary differential equations (ODEs) and partial differential equations (PDEs). Reducedorder modeling (ROM) algorithms for studying fluid dynamics have shown success in identifying linear operators that can describe flowfields, where dynamic mode decomposition (DMD) is a particularly useful method in which a linear operator is identified from data. In this work, DMD is reformulated as an optimization problem to propagate the state of the linearized dynamical system on a quantum computer. This reformulation was chosen as a means of facilitating implementation on a near-term quantum computer. Quadratic unconstrained binary optimization (QUBO), a technique for optimizing quadratic polynomials in binary variables, allows for quantum annealing algorithms to be applied. A quantum circuit model (quantum approximation optimization algorithm, QAOA) is utilized to obtain predictions of the state trajectories. Results are shown for the quantum-ROM predictions for flow over a 2D cylinder at Re = 220 and flow over a NACA0009 airfoil at Re = 500 and α = 15°. The quantum-ROM predictions are found to depend on the number of bits utilized for a fixed point representation and the truncation level of the DMD model. Comparisons with DMD predictions from a classical computer algorithm are made, as well as an analysis of the computational complexity and prospects for future, more fault-tolerant quantum computers.

97 MATHEMATICS AND COMPUTING↗