Search NASA⌕ Search

SEARCH · Search NASA

Results for “online optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

A hybrid neural architecture: Online attosecond x-ray characterization

The emergence of high-repetition-rate x-ray free-electron lasers (XFELs), such as SLAC’s LCLS-II, serves as our canonical example for autonomous controls that necessitate high-throughput diagnostics paired with streaming computational pipelines capable of single-shot analysis with extremely low latency. We present the deterministic characterization with an integrated parallelizable hybrid resolver architecture, a hybrid machine learning framework designed for fast, accurate analysis of XFEL diagnostics using angular streaking-based sinogram images. This architecture integrates convolutional neural networks and bidirectional long short-term memory models to denoise input, identify x-ray sub-spike features, and extract sub-spike relative delays with sub-30 attosecond temporal resolution. Deployed on low-latency hardware, it achieves over 10 kHz throughput with 168.3 μs inference latency, indicating scalability to 14 kHz with field-programmable gate array integration. By transforming regression tasks into classification problems and leveraging optimized error encoding, we achieve high precision with low-latency performance that is critical for real-time streaming event selection and experimental control feedback signals. This represents a key development in real-time control pipelines for next-generation autonomous science, generally, and high repetition-rate x-ray experiments in particular.

Accelerator Physics (physics.acc-ph)↗

Studying the Impact of Momentary Cessation of IBRs Compliant With IEEE Std. 2800-2022 on Transmission Line Protection Elements

This paper examines the impact of inverter-based resource (IBR) momentary cessation timing on transmission system protection relays. IEEE Std. 2800 mandates voltage ride-through, requiring IBRs to remain online during voltage disturbances. However, for severe voltage dips (|V| = 0.1 p.u.), the standard permits current blocking. In practice, vendors implement momentary cessation with varying delays, potentially affecting relay operation if IBRs cease too early. This article studies the impact of the time IBRs enter momentary cessation on transmission line protective relays. Using an electromagnetic transient (EMT) simulation, this study models IBR momentary cessation with tunable delays in a realistic transmission network. COMTRADE files generated from EMT simulations are analyzed in MATLAB using detailed protective relay models. Results indicate that IBRs must sustain operation for at least one cycle to ensure proper relay response; otherwise, relays may fail to oper- ate. These findings guide IBR vendors in optimizing momentary cessation settings to enhance transmission line protection.

24 POWER TRANSMISSION AND DISTRIBUTION↗

ELM 1016706669-AA B581 GDE01 Generator Summary pg 4-5 only

The table below shows the general load breakdown for 581GDE01. Although the total connected load exceeds the generator’s 150kW nameplate capacity, the normal configured load during standby power mode is less, which was measured at 81kW, or 54%, during preventive maintenance activities on 12/12/24. The generator’s available spare capacity must account for dynamically changing loads that can increase the total power demand at any time. If the two online VFD’s are operated at full speed the actual standby mode load is estimated to increase by 38kW, which would bring the total configured load during standby power mode to 119kW, or 79%, still within 581GDE01’s acceptable capacity. Per the LLNL Site 200 Generator Consolidation Final Study 2022, “Standby Emergency Generator nameplate ratings are based upon operation with varying load averaging 70% of the nameplate for 200 hours per year. Continuous loading between 70% and 100% will reduce a generator’s expected lifetime before a major overhaul. This is never a problem with Laboratory machines because of conservative application of generators and the reliability of the normal power system combines to keeps the load and hours down”. To achieve optimal performance and prolong generator life, the recommended generator loading is between 40% and 70%, optimally at 70%, which 581GDE01 appropriately falls within.

42 ENGINEERING↗

Towards a Robust Adaptive Digital Twin for Fusion Applications

The development of a digital twin system for fusion applications is essential for enhancing the prediction, analysis, and optimization of complex plasma processes. Machine learning (ML), particularly deep learning has demonstrated strong capabilities in modeling such highly nonlinear and intricate systems. However, two critical challenges limit the deployment of deep learning-based digital twins: Uncertainty Quantification (UQ) and data drift. UQ is vital for ensuring trustworthy predictions, especially in decision-support scenarios. Additionally, data-driven models are often sensitive to changes in the underlying data distribution, such as shot-to-shot variations in fusion experiments, which can lead to performance degradation over time. To address these challenges, we are developing an uncertainty-aware, adaptive digital twin framework. Our approach incorporates deep learning models enhanced with Gaussian Process approximations for predictive uncertainty estimation, coupled with an online learning mechanism that enables continuous model adaptation to new experimental data. This adaptive capability allows the data driven models to respond effectively to evolving plasma behaviors and equipment conditions. Specifically, to mitigate the effects of shot-to-shot drift, our system updates itself incrementally as new data becomes available, improving both robustness and fidelity. Our vision is to evolve this data driven model into a self-sustaining digital twin system that leverages UQ based feedback to continuously refine itself and potentially support real-time decision making. This presentation will cover a brief background on uncertainty quantification for ML, our ongoing effort on development of UQ capabilities for ML, our data science pipeline from data collection to model development and analysis and online learning framework for modeling coil deflection at DIII-D. I will also briefly touch upon opportunities and challenges in development of digital twin framework.

Sammuli, Brian [General Atomics]↗

From Data to Knowledge: A Graph-Based Reliability Approach to Assess System Health

With the goal of maximizing plant reliability and availability, complex systems such as nuclear power plants continuously monitor and record the performance and the health status of many components, assets, and systems. Such data may take the form of online monitoring data, condition reports, and maintenance reports and it carries the potential to provide system engineers with insights into anomalous behaviors or degradation trends as well as the possible causes behind them and to predict their direct consequences. The analysis of such data poses however few challenges. While some of these challenges are technical in nature (i.e., data are often distributed over several physical servers or databases), others are conceptual in nature (i.e., data elements come in different formats, numeric or textual), and measured values have different scales (e.g., vibration spectra and oil temperature). This paper directly tackles these challenges, and it focuses on the integration of all these data elements in order to assist plant system engineers in analyzing component, assets, and systems performances and optimize maintenance activities. This is performed by 1) extracting knowledge from textual data via technical language processing methods, and 2) quantifying system, asset, and component health from numeric condition-based data. We rely on model-based system engineering (MBSE) models of systems and assets to identify their architecture and functional (i.e., cause and effect) relations. Numeric and textual data elements are then associated with an MBSE graph element, based on their nature. This bonding of MBSE models and data elements constitutes a first-of-its-kind knowledge graph of a nuclear power plants system, with data elements being organized in a structured manner that enables system engineers to identify cause-effect trends in data elements and carry out appropriate actions in response.

97 MATHEMATICS AND COMPUTING↗

Advances on CHP District Energy and Microgrids Deployment: Simplified Tool for Rapidly Deploying Feasibility Analytics for the Non-Technical User (Final Technical Report)

Community energy systems have proven to have the potential to improve cost efficiency, resilience, and decarbonize. However, investing in community energy systems such as community microgrids or district energy systems is a complex decision due to the high initial investment and the uncertainties associated with the long development time and lifecycle of the project. Tools that make feasibility assessments accessible to non-technical users like investors, policymakers, and other stakeholders will result in more feasibility analyses completed, more candidate projects identified, and more community energy systems deployed. The pilot tool developed under this award is named Energy Fellow. Energy Fellow allows technical and non-technical users to complete feasibility analyses for district energy systems and community microgrids. This is the first software tool of its kind designed for non-technical users and available at no cost. Its scope was adjusted to a 25x25-mile region within the Houston area in Texas to make its development compatible with the funding available. However, the findings and models developed make this pilot tool easily scalable to the US. The lessons learned during the design, implementation, and testing stages have helped find trade-off solutions to software and hardware challenges related to implementing 3D models in online tools. Green software strategies has been successfully applied to the design and operations of the tool, and the team has researched the aspects of the (non-technical) user experience that will make commercial developments of this tool even more impactful.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Enhancements and Deployment of the TDAQ System for the Mu2e Experiment

The Real Time Processing Systems Division at Fermilab has deployed new features to the Off-The-Shelf Data Acquisition framework (otsdaq) for the Mu2e experiment. The Mu2e experiment will search for the coherent neutrino-less conversion of a muon into an electron in the field of an aluminum nucleus with a sensitivity improvement of 10,000 times over existing limits. Such a charged lepton flavor-violating reaction probes new physics at a scale unavailable at present or planned high-energy colliders. The Mu2e Trigger and Data Acquisition (TDAQ) system uses otsdaq as its online Data Acquisition System (DAQ) framework. otsdaq integrates the artdaq and art frameworks for event transfer, filtering, and processing. otsdaq is a web-based DAQ software suite focusing on flexibility and scalability and provides a multi-user interface accessible through a web browser. artdaq handles the entire data stream, which is read over the peripheral component interconnect express (PCIe) bus to a software filter algorithm that selects events combined with the data flux coming from a cosmic-ray veto (CRV) system. Detector front-ends are configured through the PCIe bus by customized otsdaq plugins. The otsdaq slow controls infrastructure has been further developed using the experimental physics and industrial control system (EPICS) open-source platform for monitoring, controlling, alarming, and archiving. The detector control system (DCS) for Mu2e has been integrated into otsdaq. The production TDAQ and DCS system has been deployed at the experimental hall and is being debugged and optimized for experiment operations. We report on the feature enhancements and deployment of otsdaq for Mu2e.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Unsupervised Process Anomaly Detection and Identification Using the Leave-One-Variable-Out Approach

Automated anomaly detection and identification can signal equipment issues and pinpoint causes in large-scale industrial systems. For systems with limited failure history, unsupervised machine learning methods can be utilized as they do not require past failures. This study introduces the leave-one-variable-out (LOVO) model, which masks one variable at a time to predict the others, learning underlying process correlations. Detection performance was assessed with synthetic and experimental data, while identification performance used only synthetic data due to its ability to generate labeled anomaly types. For detection using synthetic data, the LOVO model generally outperformed comparative models; while using experimental data, the comparative methods outperformed the LOVO model. However, the comparative methods required selecting a latent size, and these conclusions pertain to using the optimal size. In practice, it would not be feasible to always select the optimal value, and incorrect selections impacted performance. In contrast, the LOVO model does not require a latent space. For identification using synthetic data, the LOVO model was slightly outperformed in interpretability and repeatability but still demonstrated impressive results. These outcomes suggest that the LOVO model is an effective model and may be more easily implemented without the challenging tuning process of selecting a latent size.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Visual Analytics of Multivariate Networks With Representation Learning and Composite Variable Construction

Multivariate networks are commonly found in real-world data-driven applications. Uncovering and understanding the relations of interest in multivariate networks is not a trivial task. This article presents a visual analytics workflow for studying multivariate networks to extract associations between different structural and semantic characteristics of the networks (e.g., what are the combinations of attributes largely relating to the density of a social network?). The workflow consists of a neural-network-based learning phase to classify the data based on the chosen input and output attributes, a dimensionality reduction and optimization phase to produce a simplified set of results for examination, and finally an interpreting phase conducted by the user through an interactive visualization interface. A key part of our design is a composite variable construction step that remodels nonlinear features obtained by neural networks into linear features that are intuitive to interpret. We demonstrate the capabilities of this workflow with multiple case studies on networks derived from social media usage and also evaluate the workflow with qualitative feedback from experts.

97 MATHEMATICS AND COMPUTING↗

Regional surrogates for predictive control of digital twins

Digital twins of complex systems must involve a model that is fast, generalizable, and usable for real-time control. For example, high-fidelity nonlinear multiphysics simulations can capture laser-material interactions, but are too slow for optimization or model predictive control (MPC). Reduced-order models, used to accelerate such computation, frequently fail to generalize to unseen inputs or control states. We show theoretically that this failure is intrinsic, i.e., that a learned model is non-unique outside the sampled subspace when its low-rank structure arises from limited excitation and clustered eigenvalues, rather than from a user-imposed truncation alone. Motivated by this result, we propose a control-ready regional surrogate-construction framework for both autonomous and nonautonomous dynamics; it employs Koopman lifting to represent nonlinearities, while preserving spatial locality. We illustrate our approach by constructing a control-ready surrogate for the digital twin of a thermal component of additive-manufacturing process. Our surrogate, localized in space through a von Neumann stencil, is learned from noisy high-fidelity simulations that emulate thermal-camera images collected during the manufacturing. It is linear in thermo-physically augmented states so that MPC reduces to a convex quadratic program. The surrogate requires no online correction, generalizes to unseen scan paths and power profiles of the laser, and is more than three orders of magnitude faster than a finite-difference solver. Furthermore, when the MPC sequence computed on the digital twin is applied to this solver, closed-loop temperature regulation is recovered, showing that the surrogate preserves control-relevant input-output behavior.

Data-driven model↗

Developing Fluorescence-Based Sensors to Support Rare Earth Element Separation

Rare earth elements (REEs) are essential to most renewable energy technologies. Unfortunately, as we transition to sustainable energy production, the demand for REEs is rapidly growing well beyond current rates of production. As a result, novel means of efficient, scalable, and easily adaptable methods for processing primary and recycle feedstocks are needed. Development and integration of sensors for highly selective in-line monitoring can support more efficient design and testing of such novel separation processes, as well as more cost-effective deployment of those separation flowsheets. Work here will explore the application of fluorescence spectroscopy, a highly sensitive and selective technique, to quantify multiple lanthanides in complex mixtures including known interferents or quenching agents. Results include identification of the optimal excitation wavelength and the limit of detection of various rare earth elements as well as the performance of data-science-based quantification approaches in streams where “unknowns” are present. Overall, the data science tools in conjunction with optical sensor data were able to quantify analytes in the presence of other lanthanides which can be anticipated in the actual industrial stream. Here we include characterization of lanthanides in a microfluidic device similar to those used in new process development. This study demonstrates the capability of utilizing fluorescence spectroscopy to quantify analytes in a complicated solution matrix, suggesting this is a successful approach for in-line monitoring to optimize the separation efficiency in an industrial stream.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Quantifying Temperature Dependence of Pu(IV) Absorbance Spectra for Advanced Online Monitoring of Nuclear Processes

This article presents a systematic study of Pu(IV) absorbance spectral features as a function of temperature to develop an understanding of this parameter’s effect on chemometric models that can be used as online monitoring tools to support nuclear processing. The descriptive and predictive models that provide real-time feedback of these processes are usually constructed with data collected in conditions typical of a laboratory environment, which can differ drastically from a processing environment. To assess the impact of temperature on Pu(IV) absorbance spectra, 11 samples of Pu(IV) were synthesized with varying HNO 3 concentrations ranging from 0.6 to 9.5 M and heated between 15 and 45 °C. Ultraviolet (UV)–visible (vis)–near-infrared (NIR) absorption spectra collected at different HNO 3 concentrations and temperatures revealed that features associated with Pu(IV) are sensitive to temperature at all HNO 3 concentrations and that changes in features depend on HNO 3 concentration. The contributions of temperature and HNO 3 concentration to variation in Pu(IV) spectral features were evaluated using the principal component analysis of spectra that were baseline-corrected with an asymmetric least-squares method. Furthermore, predictive modeling for HNO 3 concentration with partial least-squares regression of UV–vis–NIR spectra highlighted the importance of accounting for temperature in the calibration set to optimize model performance. This methodology constitutes a new, systematic approach to account for the effect of temperature on the absorption spectra of metal ions and is useful for process monitoring applications in many industries.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

ReMU: regional minimal updating for model-based derivative-free optimization

Derivative-free optimization (DFO) problems are optimization problems where derivative information is unavailable or extremely difficult to obtain. Model-based DFO solvers have been applied extensively in scientific computing. Powell's NEWUOA (2004) [Powell, The NEWUOA software for unconstrained optimization without derivatives, in Large-Scale Nonlinear Optimization, Nonconvex Optimization and its Applications Vol. 83, G. Di Pillo and M. Roma, eds., Springer, 2006, pp. 255–297] and Wild's POUNDerS (2014) [Wild, Solving derivative-free nonlinear least squares problems with POUNDERS, in Advances and Trends in Optimization with Engineering Applications, T. Terlaky, M.F. Anjos, and S. Ahmed, eds., SIAM, 2017, pp. 529–540] explore the numerical power of the minimal norm Hessian (MNH) model for DFO and contributed to the open discussion on building better models with fewer data to achieve faster numerical convergence. Another decade later, we propose the regional minimal updating (ReMU) models, and extend the previous models into a broader class, including the H 2 norm models [Xie and Yuan, Least H 2 norm updating of quadratic interpolation models for derivative-free trust-region algorithms, IMA J. Numer. Anal. 46 (2025), pp. 21–50]. This paper shows motivation behind ReMU models, computational details, theoretical and numerical results on particular extreme points and the barycentre of ReMU's weight coefficient region, and the associated KKT matrix error and distance. Novel metrics, such as the truncated Newton step error, are proposed to numerically understand the new models' properties. A new algorithmic strategy, based on iteratively adjusting the ReMU model type, is also proposed, and shows numerical advantages by combining and switching between the barycentric model and the classic least Frobenius norm model in an online fashion.

derivative-free trust-region methods↗

Hydrogen, Methane, Brine Flow Behavior, and Saturation in Sandstone Cores During H 2 and CH 4 Injection and Displacement

Large-scale underground hydrogen storage (UHS) is a critical component in the emerging hydrogen economy. Knowledge of multiphase flow behavior involving hydrogen in storage reservoir formations is crucial to characterizing hydrogen transport properties and essential for the deliverability and storage operations of UHS. There are still many gaps in fully understanding hydrogen–methane–brine multiphase phase flow that require further investigation. In this work, H 2 and CH 4 were injected through brine-saturated sandstone cores using a tri-axial core holder system fitted with flow rate meters and pressure transducers, while the effluent gas concentrations were analyzed using an online micro gas chromatograph. Brine displacement, permeability, and gas breakthrough curves were measured. We studied the flow behavior of hydrogen and methane in sandstone cores through testing brine displacement by gas injection and comparing the hydrogen displacement of methane with the methane displacement of hydrogen. We also tested the differences between horizontal and vertical flow in brine displacement. The results showed that brine displacement was more efficient in a core with higher permeability and porosity, resulting in a higher initial gas saturation. A higher gas injection rate brought about faster gas breakthrough measured by pore volume and sharper concentration curves. Hydrogen did not exhibit abnormal flow in the sandstone when the flow was horizontal and downward vertical. Gas overriding was observed in brine displacements when the flow was horizontal, with hydrogen showing this behavior more profoundly compared to methane. Downward vertical gas injection induced higher efficiency brine displacement compared to horizontal displacement and resulted in a higher initial gas saturation in the sandstone cores. These findings address critical knowledge gaps regarding gas flow patterns and displacement behaviors during hydrogen injection and recovery phases in UHS facilities using methane as the cushion gas. The insights from this research offer valuable guidance for optimizing UHS systems, ensuring operational efficiency, and advancing sustainable energy solutions in alignment with decarbonization goals.

08 HYDROGEN↗

Pyrolysis of high-density polyethylene: Degradation behaviors, kinetics, and product characteristics

Pyrolysis is a promising technology for converting plastic waste into valuable raw materials while offering a potential solution to the global plastic pollution crisis. In this study, the thermal pyrolysis of high-density polyethylene (HDPE) is investigated in a drop tube reactor under nearly isothermal conditions. The impact of reaction temperature and gas/volatile residence time on carbon conversion and product distribution is examined across a range of 500–900°C and 3.6–32.2s, respectively. Non-condensable gas products detected by online mass spectrometry are H 2 , CH 4 , C 2 H 4 , C 2 H 6 , C 3 H 6 , and C 3 H 8 . At elevated temperatures and prolonged residence time, H 2 yield reaches as high as 8.6 wt% of the initial HDPE mass due to intensified cracking reactions of C 2 –C 3 hydrocarbons and long-chain aliphatic compounds. Consequently, pyrolysis tars consist mainly of polycyclic aromatic hydrocarbons (PAHs) with 5–7 rings, accompanied by visible coke deposition within the reactor. HDPE decomposition to volatiles is an endothermic process and it is complete at a temperature between 492°C and 525°C, depending on the heating rate employed, from non-isothermal thermogravimetric analysis and differential scanning calorimetry (TGA-DSC) measurements. The thermal degradation of HDPE pellets follows the two-dimensional nucleation growth model for conversion levels up to 0.8 with an apparent activation energy of 259–270 kJ/mol and a pre-exponential factor of 4.83 × 10 17 –1.37 × 10 19 min -1 , determined from various isoconversional methods such as Flynn-Wall-Ozawa (FWO), Kissinger-Akahira-Sunose (KAS), and Starink, along with Criado's master plots. Further, these findings provide valuable insights into optimizing process parameters and refining reactor design for pyrolysis, which can be integrated with gasification and reforming processes to enhance hydrogen production on a larger scale.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Scale-up Unlearnable Examples Learning with High-performance Computing

Recent advancements in AI models, like ChatGPT, are structured to retain user interactions, which could inadvertently include sensitive healthcare data. In the healthcare field, particularly when radiologists use AI-driven diagnostic tools hosted on online platforms, there is a risk that medical imaging data may be repurposed for future AI training without explicit consent, spotlighting critical privacy and intellectual property concerns around healthcare data usage. Addressing these privacy challenges, a novel approach known as Unlearnable Examples (UEs) has been introduced, aiming to make data unlearnable to deep learning models. A prominent method within this area, called Unlearnable Clustering (UC), has shown improved UE performance with larger batch sizes but was previously limited by computational resources (e.g., a single workstation). To push the boundaries of UE performance with theoretically unlimited resources, we scaled up UC learning across various datasets using Distributed Data Parallel (DDP) training on the Summit supercomputer. Our goal was to examine UE efficacy at high-performance computing (HPC) levels to prevent unauthorized learning and enhance data security, particularly exploring the impact of batch size on UE’s unlearnability. Utilizing the robust computational capabilities of the Summit, extensive experiments were conducted on diverse datasets such as Pets, MedMNist, Flowers, and Flowers102. Our findings reveal that both overly large and overly small batch sizes can lead to performance instability and affect accuracy. However, the relationship between batch size and unlearnability varied across datasets, highlighting the necessity for tailored batch size strategies to achieve optimal data protection. The use of Summit’s high-performance GPUs, along with the efficiency of the DDP framework, facilitated rapid updates of model parameters and consistent training across nodes. Our results underscore the critical role of selecting appropriate batch sizes based on the specific characteristics of each dataset to prevent learning and ensure data security in deep learning applications. The source code is publicly available at https: // github. com/ hrlblab/ UE_ HPC .

Zhu, Yanfan [Vanderbilt University, Nashville, TN,↗

High Temperature Copper Metallization: Demand, Hurdles and Reliability

As newer cells structures come online, the pressing need to replace silver in the metallization pastes has renewed interest in alternative technologies employing base metals. Copper typically leads the charge with its abundance and lower cost but has faced numerous obstacles from relatively higher oxidation and diffusion rates which can damage the lifetime of the devices. In this study, a low-cost alternative to silver metallization pastes has been shown on PERC cells. The screen printable copper paste can be fired in air at temperatures >500 degrees C, and the impact of processing conditions and equipment on the performance and reliability of 274 cm2 cells have been evaluated. Through damp heat testing of micro-modules using 16 cm2 cells, routes that can lead to both the failure and success of durable contacts have been demonstrated.

copper↗

FPGA-accelerated SpeckleNN with SNL for real-time X-ray single-particle imaging

We present the implementation of a specialized version of our previously published unified embedding model, SpeckleNN, for real-time speckle pattern classification in X-ray Single-Particle Imaging (SPI), using the SLAC Neural Network Library (SNL) on an FPGA platform. This hardware realization transitions SpeckleNN from a prototypic model into a practical edge solution, optimized for running inference near the detector in high-throughput X-ray free-electron laser (XFEL) facilities, such as those found at the Linac Coherent Light Source (LCLS). To address the resource constraints inherent in FPGAs, we developed a more specialized version of SpeckleNN. The original model, which was designed for broader classification across multiple biological samples, comprised ~5.6 million parameters. The new implementation, while reducing the parameter count to 64.6K (a 98.8% reduction), focuses on maintaining the model's essential functionality for real-time operation, achieving an accuracy of 90%. Furthermore, we compressed the latent space from 128 to 50 dimensions. This implementation was demonstrated on the KCU1500 FPGA board, utilizing 71% of available DSPs, 75% of LUTs, and 48% of FFs, with an average power consumption of 9.4W according to the Vivado post-implementation report. The FPGA performed inference on a single image with a latency of 45.015 microseconds at a 200 MHz clock rate. In comparison, running the same inference on an NVIDIA A100 GPU resulted in an average power consumption of ~73W and an image processing latency of around 400 microseconds. Our FPGA-accelerated version of SpeckleNN demonstrated significant improvements, achieving an 8.9 × speedup and a 7.8 × reduction in power consumption compared to the GPU implementation. Key advancements include model specialization and dynamic weight loading through SNL, which eliminates the need for time-consuming FPGA design re-synthesis, allowing fast and continuous deployment of models (re)trained online. These innovations enable real-time adaptive classification and efficient vetoing of speckle patterns, making SpeckleNN more suited for deployment in XFEL facilities. This implementation has the potential to significantly accelerate SPI experiments and enhance adaptability to evolving experimental conditions.

47 OTHER INSTRUMENTATION↗