Search NASASearch

SEARCH · Search NASA

Results for “relevancy algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Tactical Analysis for Calculating Contextual Risk at Boundaries: Summary of Laboratory Directed Research & Development Effort

The Tactical Analysis for Calculating Contextual Risk at Boundaries (TACCRAB) tool is an innovative digital twin (DT) platform and automated risk algorithm designed to transform operational decision-making in structured screening environments, with an initial focus on Southern Border Land Ports of Entry (POEs). The invention provides integration points for advanced artificial intelligence, predictive modeling, and real-time data analysis to produce a comprehensive risk management tool that enables proactive, data-informed security strategies. The core inventive features of TACCRAB center on its unique risk algorithm, which dynamically calculates contextual risk by synthesizing historical data, near real-time streaming data from the checkpoints themselves, and AI-generated predictions. Unlike traditional risk assessment methods, TACCRAB utilizes a DT to provide comprehensive operational insights, allowing stakeholders to visualize, simulate, and optimize checkpoint configurations with unprecedented speed and contextual awareness. TACCRAB's key innovation lies in its ability to combine multiple complex inputs - including technology detection probabilities, resource availability, screening pathway characteristics, and threat actor behavioral patterns - into a unified risk calculation and update these inputs based on changing operational and environmental conditions. By leveraging a DT that continuously updates and learns from linked data, TACCRAB can suggest adaptive mitigation strategies that minimize risk while maintaining operational efficiency. Particularly novel is the platform's approach to decision support, which goes beyond static risk assessment. The DT provides dynamic metrics such as wait times, resource allocation effectiveness, and potential emerging threat scenarios, enabling users to view sophisticated, relevant what-if simulations and optimize checkpoint operations in near real-time. The system's architecture allows for generalized application across different screening environments, such as secure facilities, ports of entry, and soft targets, making it a versatile tool for security and operational management. The invention distinguishes itself through its comprehensive integration of predictive modeling, AI-driven pattern discovery, and user-friendly interface design. By combining these elements, TACCRAB transforms complex risk data into actionable insights, supporting decision-makers at various organizational levels - from booth agents making split-second screening decisions to checkpoint managers optimizing the day's resource allocation to strategic planners managing long-term investments.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF

New NDA Methods for Thorium Fuel Cycle Safeguards (Final Report)

This project developed portable Neutron Resonance Transmission Analysis (pNRTA) as a new non-destructive assay (NDA) method for thorium fuel cycles safeguards and other applications where multiple isotopes must be measured when present together. pNRTA leverages epithermal neutron resonances to assay multiple safeguards-relevant isotopes (e.g., 233 U and 235 U) when they are present together in a sample. Existing techniques are challenged by this task, driving the need for new active interrogation methods. With selected detectors, pNRTA works in high gamma-ray backgrounds from fission and activation products and 232 U progeny expected in thorium fuel cycle samples. This project leveraged a pNRTA system developed at Pacific Northwest National Laboratory (PNNL) and collaboration with the Massachusetts Institute of Technology (MIT). The system uses a commercially available deuterium-tritium (DT) neutron generator at short standoff (2 m). Key achievements in this project included: first-of-a-kind pNRTA quantitative measurements of 233 U oxide samples, an assessment of neutron detector technologies suitable for pNRTA in high gamma-ray background environments, experimentally demonstrating quantitative assay of samples containing 233 U and 235 U, and modeling studies showing the applicability of pNRTA to a wide range of material forms. Further, a custom algorithm was developed at MIT, which provided mean bias of 9% and relative standard deviation of 36% in assaying 233 U, 235 U, 238 U, and 232 Th content in eight measured samples. These outcomes form a solid technical basis for pNRTA as a new promising capability for international safeguards verification that is portable, non-destructive, quantitative, and isotopic specific.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Bridging Control and Deployment: A Cross-Layer Analysis of Scalable Building Cluster Control

Building cluster control has emerged as a promising approach for enabling flexible and coordinated operation of distributed building systems, yet its transition from pilot demonstrations to routine grid-interactive operation remains limited. This paper argues that this gap cannot be explained by control algorithms alone. Instead, it arises from interacting barriers in communication infrastructure, data and semantic interoperability, uncertainty management, stakeholder participation, market design, and policy support. Accordingly, the paper reviews both technical and non-technical barriers to building cluster control. Technical challenges include heterogeneous devices and protocols, communication latency and reliability, distributed decision-making, and uncertainty propagation across aggregated loads. Non-technical barriers include user participation, stakeholder coordination, incentive allocation, and data governance. Existing solution approaches are synthesized, including semantic interoperability frameworks, edge and hierarchical communication architectures, distributed and transactive control strategies, uncertainty-aware optimization, policy mechanisms, and market reforms. Based on this analysis, two research directions are identified: testing infrastructures that can evaluate control performance under realistic multi-building conditions, and abstraction methods that allow building clusters to interact with other energy sectors through standardized flexibility representations. Overall, the paper provides a structured review of how building cluster control can move from isolated demonstrations toward reproducible, market-compatible, and grid-relevant implementation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Static Subspace Approximation for Random Phase Approximation Correlation Energies: Implementation and Performance

Developing theoretical understanding of complex reactions and processes at interfaces requires using methods that go beyond semilocal density functional theory to accurately describe the interactions between solvent, reactants and substrates. Methods based on many-body perturbation theory, such as the random phase approximation (RPA), have previously been limited due to their computational complexity. However, this is now a surmountable barrier due to the advances in computational power available, in particular through modern GPU-based supercomputers. In this work, we describe the implementation of RPA calculations within BerkeleyGW and show its favorable computational performance on large complex systems relevant for catalysis and electrochemistry applications. Our implementation builds off of the static subspace approximation which, by employing a compressed representation of the frequency dependent polarizability, enables the evaluation of the RPA correlation energy with significant acceleration and systematically controllable accuracy. We find that the computational cost of calculating the RPA correlation energy scales only linearly with system size for systems containing up to 50 thousand bands, and is expected to scale quadratically thereafter. We also show excellent strong scaling results across several supercomputers, demonstrating the performance and portability of this implementation.

algorithmic development

Custom-trained Machine-learning Interatomic Potentials: ZnCl2 Aqueous Solution

This dataset was generated using an iterative active-learning strategy implemented in the ArcaNN software package (https://github.com/arcann-chem/arcann_training) to train machine-learning interatomic potentials for aqueous ZnCl2 solutions. Each active-learning cycle consisted of three stages: training, exploration, and labeling. The initial training set combined configurations generated in this work from enhanced-sampling ab initio molecular dynamics simulations with configurations from a previously reported neural-network-potential study of aqueous ZnCl2. The enhanced-sampling ab initio molecular dynamics simulations involved Zn–Cl separation and the chloride coordination number around Zn²? as collective variables. These configurations served as the seed dataset. Subsequent active-learning cycles expanded the training set by identifying and labeling configurations that were poorly represented by the current models, thereby improving coverage of ion-association states and changes in local coordination and charge-state environments relevant to the solution free-energy landscape. For all selected configurations, single-point calculations of the total energies and atomic forces were performed within density functional theory using the CP2K Quickstep module. Reference calculations employed the revPBE-D3 and r2SCAN exchange-correlation functionals. Motivated by recent work on aqueous Zn²?, the main revPBE calculations omitted D3 dispersion contributions involving Zn²?, while retaining the D3 correction for water and chloride. For comparison, fully dispersion-corrected revPBE-D3 reference calculations were also performed, with D3 applied to all species, including Zn²?. Valence electrons were treated explicitly, while core electrons were represented using norm-conserving Goedecker–Teter–Hutter pseudopotentials. The wave functions were expanded using the mixed Gaussian-and-plane-wave scheme with TZV2P-MOLOPT basis sets for all elements and a 600 Ry auxiliary plane-wave cutoff for the electron density. Self-consistent-field convergence was accelerated using the orbital-transformation and Direct Inversion in the Iterative Subspace algorithms, with a convergence threshold of 10?6. All single-point calculations were performed in periodic orthorhombic cells. The CELL_REF keyword in CP2K was used to define a fixed reference cell with a box length of 25 Å. This treatment ensured a consistent reference for configurations extracted from NpT trajectories with fluctuating cell dimensions. The resulting DFT energies and atomic forces constitute the ground-truth labels used to train the MLIPs. The resulting MLIP was trained for aqueous ZnCl2 solutions spanning concentrations from 0 to 30 molal and a broad pH range, from strongly acidic to strongly basic conditions. Representative examples of configurations included in the MLIP training dataset are provided below. These include 1) Representative configurations from the dataset labeled at the revPBE-D3 level, with D3 dispersion interactions involving Zn2+ excluded (revPBE-wo-D3). 2) Representative configurations from the dataset labeled at the fully dispersion-corrected revPBE-D3 level, with D3 interactions applied to all species, including Zn2+ (revPBE-D3). 3) Representative configurations from the dataset labeled at the r2SCAN level of theory (r2SCAN).

Dinpajooh, Mohammadhasan [Pacific Northwest Nation

Accelerating Bilevel Optimization With Hierarchical Many-Threaded Parallel Differential Evolution

Bilevel optimization is encountered in many relevant real-world applications. The main feature of this type of problem is that an upper-level optimization problem is constrained by a nested lower-level optimization problem. Because of this nested structure, bilevel problems (BLPs) are usually computationally expensive to solve. Differential evolution (DE) has demonstrated promising results in solving BLPs of relatively small scales. As the problem scale increases, the decision space becomes intrinsically larger, requiring a growing number of function evaluations for the method to work properly. In this context, heavy parallelization and high-performance computing techniques are indispensable to enable the resolution of more complex and challenging optimization problems. Hence, we propose a hierarchical many-threaded parallel DE approach for BLPs, where both levels are parallelized. The computational experiments demonstrate that the parallel implementation achieved runtime speeds ranging from 44 to 2559 times faster than the sequential version on a well-known scalable SMD benchmark test problem when executed on an NVIDIA A100 GPU. The findings indicate that the algorithm’s convergence is strongly influenced by the number of both upper- and lower-level generations. Moreover, the success of experiments with large-scale problems is closely linked to the choice of small population sizes.

Dufek, Amanda S

Evaluating Machine Learning-Based MRI Reconstruction Using Digital Image Quality Phantoms

Quantitative and objective evaluation tools are essential for assessing the performance of machine learning (ML)-based magnetic resonance imaging (MRI) reconstruction methods. However, the commonly used fidelity metrics, such as mean squared error (MSE), structural similarity (SSIM), and peak signal-to-noise ratio (PSNR), often fail to capture fundamental and clinically relevant MR image quality aspects. To address this, we propose evaluation of ML-based MRI reconstruction using digital image quality phantoms and automated evaluation methods. Our phantoms are based upon the American College of Radiology (ACR) large physical phantom but created in k-space to simulate their MR images, and they can vary in object size, signal-to-noise ratio, resolution, and image contrast. Our evaluation pipeline incorporates evaluation metrics of geometric accuracy, intensity uniformity, percentage ghosting, sharpness, signal-to-noise ratio, resolution, and low-contrast detectability. We demonstrate the utility of our proposed pipeline by assessing an example ML-based reconstruction model across various training and testing scenarios. The performance results indicate that training data acquired with a lower undersampling factor and coils of larger anatomical coverage yield a better performing model. The comprehensive and standardized pipeline introduced in this study can help to facilitate a better understanding of the performance and guide future development and advancement of ML-based reconstruction algorithms.

47 OTHER INSTRUMENTATION

TrustDER: Trusted, Private and Scalable Coordination of Distributed Energy Resources

In this project, the Stanford and SLAC Teams have developed a Trusted, Private and Scalable platform for coordinating Coordination of Distributed Energy Resources (TrustDER). This is a layered system that ensures private, trusted and scalable coordination and monitoring of DERs. It accommodates a variety of resources, such as solar generation, gensets and loads, with a particular focus on battery systems-based resources, as they are a transformational technology experiencing fast growth in adoption by large critical facilities. The platform can be used as standalone or added to existing aggregation systems to enable trust, privacy and resilience. TrustDER consists of layers that address each of the shortcomings of the existing state of the art. Each layer in the platform can operate independently but provides information to the layers above it to enable a novel form of overall coordination architecture. The project consists of several tasks, with each task dedicated to the design of each layer. Task 2 Resource Virtualization defined a software abstraction layer for distributed energy resources (DERs). The goal of this abstraction was to simplify the implementation of algorithms utilizing cooperation of DERs resources in a variety of use cases. Task 3 is on Secure ID for Asset Authentication. Identity Management Systems (IDMS) are a foundational infrastructure for interactions between entities (organizations, users, devices, and services). Secure ID is blockchain-based a distributed identity management system allowing (1) identity provisioning, (2) authentication, (3) authorization, and (4) identity data sharing for IoT-enabled assets on the electricity grid. In this project, the SLAC team focused on designing and testing Keymaker, a protocol for authenticating device identity managed by Secure ID. Task 5 Private and Safe Integration is focused on the design and evaluation of a DER cooperation scheme which allows for the aggregation of DERs without impacting network reliability. The approach is designed based on realistic assumptions regarding data availability, communication infrastructure limitations, and privacy. Task 6 Scalable Distributed Privacy for Information explored how virtualized batteries could be managed privately. Specifically, it examined the case in which a principal provides a partitioned battery to multiple clients. Task 7 Use Cases was to ensure that this technology was applied in relevant situations and scenarios. Primarily, this means that virtualization needed to be employed in a manner that either improved flexibility, bolstered security or privacy, or decreased costs.

25 ENERGY STORAGE

Capturing many-body correlation effects with quantum and classical computing

Theoretical descriptions of excited states of molecular systems in high-energy regimes are crucial for supporting and driving many experimental efforts at light source facilities. However, capturing their complicated correlation effects requires formalisms that provide a hierarchical infrastructure of approximations. These approximations lead to an increased overhead in classical computing methods and, therefore, decisions regarding the ranking of approximations and the quality of results must be made on purely numerical grounds. The emergence of quantum computing methods has the potential to change this situation. Here, in this study, we demonstrate the efficiency of the quantum phase estimator (QPE) in identifying core-level states relevant to x-ray photoelectron spectroscopy. We compare and validate the QPE predictions with exact diagonalization and real-time equation-of-motion coupled-cluster formulations, which are some of the most accurate methods for states dominated by collective correlation effects.

74 ATOMIC AND MOLECULAR PHYSICS

Co-Occurring Atmospheric Features and Their Contributions to Precipitation Extremes

Object-based identification algorithms for atmospheric features are commonly utilized to attribute global precipitation. This study employs a systematic approach to examine feature co-occurrences and their relationships to mean and extreme precipitation. Four features are identified using existing data sets for atmospheric rivers (ARs), mesoscale convective systems (MCSs), low-pressure systems (LPSs), and fronts (FTs). Often, a single atmospheric phenomenon satisfies the criteria set by multiple feature identification algorithms, yielding an association between precipitation and multiple features. Over the extra-tropics, the number of features attributed to a single event typically increases with precipitation intensity. Over two-thirds of the precipitation is from co-occurring features, with a considerable fraction related to AR-FT co-occurrences. Over the tropics, about one-quarter of precipitation is associated with co-occurring features, with LPS-MCS co-occurrences contributing substantially in monsoon regions. MCSs are the leading single-feature contributors over tropical land and oceans. In the extra-tropics, FTs, ARs, and their co-occurrences account for over half of the total precipitation over oceans. AR-FT-MCS and FT-MCS co-occurrences contribute to extremes (precipitation exceeding the 95th percentile) over both oceans (over 30%) and land (over 20%). Any combination of features involving MCSs shows a larger contribution to high percentiles of precipitation intensity. A case analysis indicates that AR-FT-MCS co-occurrences exhibit convective instability and deep vertical motion, suggesting that the feature trackers and reanalysis are capturing physics relevant to both convective and frontal systems. The results here emphasize the need for simultaneous identifications of multiple features when attributing precipitation to atmospheric phenomena.

54 ENVIRONMENTAL SCIENCES

Exploring Uncertainty in Moment Estimation for Small Earthquakes in Southern Nevada Using the Coda Envelope Method

Compiling source parameter estimates for small earthquakes is important both for our understanding of earthquake physics and for accurately assessing earthquake hazard. Reliable source parameter estimates are difficult to achieve for small earthquakes, in part due to our inability to accurately model the relevant physical processes at high frequencies. The coda envelope methodology developed by Mayeda and Walter (1996) and Mayeda et al. (2003) can mitigate this concern and estimate the moment of small earthquakes by determining the parameters that control the shape of the S-wave coda envelope while eliminating path effects by minimizing the scatter between seismic stations. Here, we use an open-source implementation of this technique called the Coda Calibration Tool (CCT; Barno, 2017) to calculate CCT-based moment magnitude estimates of small earthquakes (M L 0–3) in the Rock Valley, Nevada, region within the Nevada National Security Site. The Rock Valley data set is of particular interest because it allows us to explore the changes in uncertainties of the coda calibration method with earthquake size and depth. We found that a consistent linear relationship exists between the local magnitude M L and our coda-derived M w estimates for earthquakes as small as M L 0–3, but that current CCT workflows do not accurately characterize very shallow events. We also demonstrate that the epistemic uncertainty in the apparent stress value assumed by the CCT algorithm can influence magnitude estimates of small earthquakes. In conclusion, these results provide valuable insight into the seismicity of this region, and inform future analysis and modeling efforts for nuclear monitoring and seismic hazard.

58 GEOSCIENCES

Promise of Graph Sparsification and Decomposition for Noise Reduction in QAOA: Analysis for Trapped-Ion Compilations

We develop new approximate compilation schemes that significantly reduce the expense of compiling the Quantum Approximate Optimization Algorithm (QAOA) for solving the Max-Cut problem. Our main focus is on compilation with trapped-ion simulators using Pauli-X operations and all-to-all Ising Hamiltonian HIsing evolution generated by Molmer-Sorensen or optical dipole force interactions, though some of our results also apply to standard gate-based compilations. Our results are based on principles of graph sparsification and decomposition; the former reduces the number of edges in a graph while maintaining its cut structure, while the latter breaks a weighted graph into a small number of unweighted graphs. Though these techniques have been used as heuristics in various hybrid quantum algorithms, there have been no guarantees on their performance, to the best of our knowledge. This work provides the first provable guarantees using sparsification and decomposition to improve quantum noise resilience and reduce quantum circuit complexity. For quantum hardware that uses edge-by-edge QAOA compilations, sparsification leads to a direct reduction in circuit complexity. For trapped-ion quantum simulators implementing all-to-all HIsing pulses, we show that for a (1−ϵ) factor loss in the Max-Cut approximation (ϵ>0), our compilations improve the (worst-case) number of HIsing pulses from O(n2) to O(nlog(n/ϵ)) and the (worst-case) number of Pauli-X bit flips from O(n2) to O(nlog(n/ϵ)ϵ2) for n-node graphs. This is an asymptotic improvement for any constant ϵ>0. We demonstrate that significant improvements to the approximation ratio are obtained using decomposition in simulated trapped-ion experiments with dephasing noise. We further present a generic argument showing that sparsification results in an exponentially improved circuit fidelity lower bound in digital computing schemes based on one- and two-qubit gates, which are relevant to a wide variety of hardwares such as superconducting qubits and certain neutral atom or trapped ion setups, and more sophisticated noise models. We anticipate these approximate compilation techniques will be useful tools in a variety of future quantum computing experiments.

Moondra, Jai [Georgia Institute of Technology]

Qutrit and qubit circuits for three-flavor collective neutrino oscillations

We explore the utility of qutrits and qubits for simulating the flavor dynamics of dense neutrino systems. The evolution of such systems impacts some important astrophysical processes, such as core-collapse supernovae and the nucleosynthesis of heavy nuclei. Many-body simulations require classical resources beyond current computing capabilities for physically relevant system sizes. Quantum computers are therefore a promising candidate to efficiently simulate the many-body dynamics of collective neutrino oscillations. Previous quantum simulation efforts have primarily focused on properties of the two-flavor approximation due to their direct mapping to qubits. Furthermore, we present new quantum circuits for simulating three-flavor neutrino systems on qutrit- and qubit-based platforms, and demonstrate their feasibility by simulating systems of two, four, and eight neutrinos on IBM and Quantinuum quantum computers.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Mixed quantum-classical methods for polaron spectral functions

In this work, using two distinct semiclassical approaches—namely, the mean-field Ehrenfest method and the mapping approach to surface hopping—we investigate the spectral function of a single charge interacting with phonons on a lattice. This quantity is relevant for the description of angle-resolved photoemission experiments. Focusing on the one-dimensional Holstein model, we compare the performance of these approaches across a range of coupling strengths and lattice sizes, exposing the relative strengths and weaknesses of each. We demonstrate that these approaches can be efficiently applied with reasonable accuracy to ab initio polaron models. Furthermore, our work provides a route to the calculation of spectral properties in realistic electron–phonon-coupled systems in a computationally inexpensive manner with encouraging accuracy.

Atomic and molecular spectra

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY

Even Higher-Level Synthesis: An Exploration of AI Hardware Accelerators using HLS4ML

With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.

Di Guglielmo, Giuseppe [Fermilab]

Low-Cost Heliostat for High-Flux Small-Area Receivers (Final Technical Report)

This project analyzed a two-stage heliostat concept consisting of a tracking stage and a concentrating stage. The tracking stage uses mirrors mounted on a common drive that move to track the sun. The concentrating stage consists of stationary mirrors that each have a unique angle to direct rays towards a small-area, high-flux, point-focused receiver. By splitting the collection and concentrating process into two stages, multiple small, inexpensive mirrors can share a structure and be controlled by a single drive in the tracking stage. The project effort developed modeling techniques that were specifically relevant to this two-stage heliostat concept. Both field-level and unit-level models were developed. The field-level model does not explicitly consider unit-level losses which are predicted by the unit-level model and then integrated into the field-level model through a correlation referred to as an efficiency modifier. This approach is referred to as the two-model approach; the development and demonstration of this two-model approach for a multi-stage heliostat technology is a key outcome of this work. The field-level model is used to design a field that hits a specific design day power given a set of heliostat design parameters. An oversized field is simulated and then heliostat units are removed based on their annual energy production in order to generate the highest performing field. The field reduction procedure fits a smooth curve fit to annual energy production as a function of position in the field which has the effect of reducing the noise that is otherwise caused by the Monte Carlo ray tracing technique. This approach is referred to as the annual energy fit method and substantially reduces computational run time for a given field level modeling accuracy. The annual energy fit approach enables the selection of a properly sized, high-performing field using orders of magnitude fewer rays than would otherwise be possible and the development of this approach is a second key outcome of this work. These models are used within a genetic optimization algorithm in order to optimize the geometric parameters associated with a heliostat in order to achieve the lowest cost per unit of collected design day power. The cost modeling that underlies the optimization is a simple, scaling type analysis backed up by a much more detailed Design for Manufacture and Assembly (DFMA) analysis. Although the figure of merit used for optimization was not cost per mirror area, this metric is reasonable to use as a means of comparison. The optimally designed 500 kW design has a tracking mirror specific cost of $181.85/m 2 , which is significantly larger than the target value and also larger than the current state of the art. The cost of the torque-tube type linkages contributed substantially to the overall cost. Based on this observation, potentially attractive alternative design configuration utilizing a capstan type actuation system should be investigated. Finally, NREL compared the performance of the two-stage heliostat to the performance of a focused and different sized flat conventional heliostats and showed that, as expected, additional losses versus the convention heliostat caused by a worse cosine efficiency, two stages of reflection, and interstage interactions. The two-stage heliostat requires around 75% more reflective area than a flat 1x1 meter conventional heliostat (similar to a focused heliostat) and 40% more than a flat 2x2 meter conventional heliostat.

14 SOLAR ENERGY

Autonomous hybrid optimization of a SiO 2 plasma etching mechanism

Computational modeling of plasma etching processes at the feature scale relevant to the fabrication of nanometer semiconductor devices is critically dependent on the reaction mechanism representing the physical processes occurring between plasma produced reactant fluxes and the surface, reaction probabilities, yields, rate coefficients, and threshold energies that characterize these processes. The increasing complexity of the structures being fabricated, new materials, and novel gas mixtures increase the complexity of the reaction mechanism used in feature scale models and increase the difficulty in developing the fundamental data required for the mechanism. This challenge is further exacerbated by the fact that acquiring these fundamental data through more complex computational models or experiments is often limited by cost, technical complexity, or inadequate models. In this paper, we discuss a method to automate the selection of fundamental data in a reduced reaction mechanism for feature scale plasma etching of SiO 2 using a fluorocarbon gas mixture by matching predictions of etch profiles to experimental data using a gradient descent (GD)/Nelder–Mead (NM) method hybrid optimization scheme. These methods produce a reaction mechanism that replicates the experimental training data as well as experimental data using related but different etch processes.

36 MATERIALS SCIENCE