Search NASA⌕ Search

SEARCH · Search NASA

Results for “Reliability and Mitigation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

From Obsolete to Optimal (ATR Demineralized Water System)

My project was migrating a Human-Machine Interface (HMI) and Programmable Logic Controller (PLC) program from legacy software to a modern platform, enhancing operational efficiency and eliminating unsupported, obsolete equipment. Initially, I used a migration tool to transfer the programs to the new software, then manually updated variables to align with the new PLC. I optimized the PLC code by leveraging new scaling features in the I/O cards, removing over 45 obsolete rungs, and enhancing clarity with functional tag aliases. Next, I modernized the HMI's visuals with an updated color scheme and 3D buttons. I developed multiple interface versions iteratively refining them based on operator feedback. I also designed and created a dedicated "Rounds" screen incorporating insights from a previous intern’s project to streamline daily rounds. To improve clarity, consistency, and efficiency I redesigned the entire interface. Adding visual indicators to enhance operational awareness: red is used for values outside normal specifications, green for normal operating modes, yellow for manual mode, and grey for offline units. Additionally, I added a dynamic visual representation of tank levels, creating PLC logic to show the level in inches as well as inches and feet. I then incorporated a navigation sidebar to facilitate efficient screen transitions. Risk of human error was mitigated by defining narrow input ranges for tank level shutoff and password protecting setpoint changes. I also fixed issues with datatypes that caused inaccuracies in the original code. Overall, this comprehensive upgrade significantly enhanced the HMI's functionality, user experience, and operational reliability.

47 - OTHER INSTRUMENTATION↗

Exploring Uncertainty in Moment Estimation for Small Earthquakes in Southern Nevada Using the Coda Envelope Method

Compiling source parameter estimates for small earthquakes is important both for our understanding of earthquake physics and for accurately assessing earthquake hazard. Reliable source parameter estimates are difficult to achieve for small earthquakes, in part due to our inability to accurately model the relevant physical processes at high frequencies. The coda envelope methodology developed by Mayeda and Walter (1996) and Mayeda et al. (2003) can mitigate this concern and estimate the moment of small earthquakes by determining the parameters that control the shape of the S-wave coda envelope while eliminating path effects by minimizing the scatter between seismic stations. Here, we use an open-source implementation of this technique called the Coda Calibration Tool (CCT; Barno, 2017) to calculate CCT-based moment magnitude estimates of small earthquakes (M L 0–3) in the Rock Valley, Nevada, region within the Nevada National Security Site. The Rock Valley data set is of particular interest because it allows us to explore the changes in uncertainties of the coda calibration method with earthquake size and depth. We found that a consistent linear relationship exists between the local magnitude M L and our coda-derived M w estimates for earthquakes as small as M L 0–3, but that current CCT workflows do not accurately characterize very shallow events. We also demonstrate that the epistemic uncertainty in the apparent stress value assumed by the CCT algorithm can influence magnitude estimates of small earthquakes. In conclusion, these results provide valuable insight into the seismicity of this region, and inform future analysis and modeling efforts for nuclear monitoring and seismic hazard.

58 GEOSCIENCES↗

ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training

Large Language Models (LLMs) have demonstrated remarkable performance in various natural language processing tasks. However, the training of these models is computationally intensive and susceptible to faults, particularly in the attention mechanism, which is a critical component of transformer-based LLMs. In this paper, we investigate the impact of faults on LLM training, focusing on INF, NaN, and near-INF values in the computation results with systematic fault injection experiments. We observe the propagation patterns of these errors, which can trigger non-trainable states in the model and disrupt training, forcing the procedure to load from checkpoints. To mitigate the impact of these faults, we propose ATTNChecker, the first Algorithm-Based Fault Tolerance (ABFT) technique tailored for the attention mechanism in LLMs. ATTNChecker is designed based on fault propagation patterns of LLM and incorporates performance optimization to adapt to both system reliability and model vulnerability while providing lightweight protection for fast LLM training. Evaluations on four LLMs show that ATTNChecker on average incurs on average 7% overhead on training while detecting and correcting all extreme errors. Compared with the state-of-the-art checkpoint/restore approach, ATTNChecker reduces recovery overhead by up to 49×.

Liang, Yuhang [University of Alabama - Birmingham]↗

A Framework for Assessing Economic and Environmental Trade-offs of Internalized Emission Costs in ERCOT Grid Planning

The power grid is on the cusp of a massive transition driven by three major areas: 1) the growth in demand for electricity, 2) efforts to decarbonize the United States economy, and 3) a desire to mitigate social disparities from the impact of electricity generation on local populations. However, most studies of the electricity sector do not include equity impacts in their models. This study seeks to do so by developing a comprehensive and generalizable model tailored to the Electric Reliability Council of Texas (ERCOT) grid, designed to incorporate the equity impacts of electricity generation in a decarbonized and resilient framework. To integrate equity into our research, we incorporate environmental externalities into our capacity expansion model of ERCOT. Specifically, we factor in intermediate-level local marginal damages of precursor pollutants (NH3, NOx, primary PM2.5, SO2, and VOC) and global pollutant CO2 into the cost of generating electricity. We do this by taking into account county population, county ambient pollution concentration, and generator emission rates. Leveraging open-source modeling tools, such as PowerGenome, pyGRETA, and GenX we construct a county-level model to account for these costs. We integrate these marginal damages into the variable operations and maintenance costs of generators, for both existing and potential future builds. This study’s findings suggest that the value of a dynamic social cost of carbon (SSC) will cover criteria pollutant marginal damages within the ERCOT grid and solar and wind is expected to increase out to 2035. Key metrics evaluated within this research include fuel mix distribution across technologies, transmission and grid infrastructure costs, CO2 emissions, local pollutants marginal damages, and the variation in generation capacity built by the model. These results and framework can be used to support grid decisions that explicitly include distributional and procedural equity within a decarbonized and sustainable grid framework.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Stability Analysis of Parallel Connected Bidirectional WPT System

This paper presents a stability analysis of parallel-connected bi-directional series-series resonant network wireless power transfer (WPT), optimized for Electric Vehicle (EV) charging and vehicle-to-grid (V2G) applications. The study addresses critical stability challenges in systems integrated with diverse distributed energy resources (DERs), including photovoltaics, fuel cells, wind turbines, energy storage systems, and the AC grid. The stability of such integrated DC grid systems is paramount for ensuring reliable operation, particularly under varying power flow conditions and dynamic interactions between parallel WPT systems. The analysis included system impedance characterization, state-space modeling, and open and closed-loop stability evaluations. The results demonstrated that the integration of a robust control architecture effectively mitigates instability risks and supports scalable, efficient operation. This work underscores the converter's adaptability and its potential for large-scale deployment in wireless EV charging infrastructures and integrated DC grid systems.

Asa, Erdem [ORNL] (ORCID:0000000190884812)↗

Integrating AI Data Centers with the Power Grid

The rapid expansion of artificial intelligence (AI) has triggered an unprecedented surge in electricity demand, with US data center energy use projected to double or triple 2023 levels by 2028. This exponential growth places strain on grid infrastructure, which can hinder timely construction of desired computing capacity. To bridge this supply-demand gap, utilities and AI developers are increasingly turning to demand flexibility, a strategy that incentivizes shifting or reducing power use during peak periods of grid stress. Data centers are uniquely equipped for flexible operations due to their digital workloads, built-in redundancy, and onsite energy assets. This article outlines four primary mechanisms to enable data center flexibility: computational load flexibility (shifting tasks temporally or geographically), flexible use of core facility infrastructure adjustments, energy storage utilization, and onsite electricity generation. To encourage adoption, utilities are deploying new tariff designs, including voluntary interruptible service riders, mandated flexibility requirements, and streamlined interconnection processes for flexible loads. For the highly capitalized and rapidly growing AI industry, the primary motivators for embracing these strategies are expediting facility interconnection, satisfying emerging regulatory mandates, and mitigating community resistance. While demand flexibility cannot substitute the long-term need for new bulk power generation, it serves as an essential, immediate solution for enabling near-term deployment. By transforming data centers from grid stressors into stabilizing assets, flexible operations can ensure reliable grid integration, ease market pressures, and support a resilient power system.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Noncontact excitation of multi-GHz lithium niobate electromechanical resonators

Abstract The demand for high-performance electromechanical resonators is ever-growing across diverse applications, ranging from sensing and time-keeping to advanced communication devices. Among the electromechanical materials being explored, thin-film lithium niobate stands out due to its strong piezoelectric properties and low acoustic loss. However, in nearly all existing lithium niobate electromechanical devices, the configuration is such that the electrodes are in direct contact with the mechanical resonator. This configuration introduces an undesirable mass-loading effect, producing spurious modes and additional damping. Here, we present an electromechanical platform that mitigates this challenge by leveraging a flip-chip bonding technique to separate the electrodes from the mechanical resonator. By offloading the electrodes from the resonator, our approach yields a substantial increase in the quality factor of these resonators, paving the way for enhanced performance and reliability for their device applications.

Instruments & Instrumentation↗

ICALEPCS 2025: Managing Technical Debt Across Large-Scale Control Systems

This presentation provides an overview of technical debt in the context of control systems for large-scale physics facilities. We explore various forms, common causes, and potential consequences on system reliability, maintainability, and extensibility. Drawing on experiences from multiple projects, including the ACORN control system modernization at Fermilab, we present a range of strategies for proactively managing technical debt, including best practices in design, development, testing, and documentation, as well as reactive approaches for identifying and mitigating existing issues.

Watts, Adam [Fermilab]↗

Interior soft x-ray tomography with sparse global sampling

To investigate the feasibility of interior imaging reconstruction in soft X-ray tomography for higher-resolution cellular imaging, including whole-cell imaging, we develop an alignment and reconstruction algorithm that combines a small number of sparse whole-cell images with a high-resolution local interior scan. Based on numerical simulations, we demonstrate that combined reconstructions mitigate the depth-of-field limitation in high-resolution scans, enable radiation dose optimization, and yield quantitative X-ray absorption values with sparse sampling. We further validate our numerical approach using experimental data from two different cell types and show that the combined reconstruction reliably provides high spatial resolution within an interior region of interest of a whole cell. The resulting sparse reconstruction framework offers robust, faithful visualization of cellular organelles in soft X-ray tomography. This mesoscale imaging strategy allows one to ‘scout’ and zoom into selected subcellular volumes of interest, enabling increased spatial resolution without sacrificing larger-volume imaging and providing information on the relative positions of all organelles within a cell.

3D imaging↗

PSA 2025 DPRA for Cyber Optimization

Cyberattacks can have many different attack paths, durations, and goals. There are also many different mitigation options involving hardware, software, and/or humans. Evaluating defense options should include quantitative evaluation of overall effectiveness to make cost and risk-informed decisions. Typical cyberattack modeling methods only provide a qualitative evaluation and have difficulty with time dependent scenarios. The main areas of cybersecurity are confidentiality, integrity, and availability. For companies with cyber-physical systems such as advanced nuclear reactors, cyber-related integrity is a requirement set by the U.S. Nuclear Regulatory Commission. But companies are also concerned about availability or reliability as a business case. As cyber threats are evolving to a business-for-hire structure, more attacks focus on disrupting business success and reliability, causing financial and economic stability risk. Companies want reliability analysis while optimizing cost, which requires more than safety modeling methods. Dynamic-state-based and Markov-based modeling provides a method for better cyber scenario modeling with timing and conditional features not found in other numerical evaluation methods. EMRALD (Event Modeling Risk Assessment using Lined Diagrams) is a dynamic risk analysis modeling and simulation tool and has features that reduce modeling issues such as state-base explosion found in Markov-based tools. It has been used to model different time-dependent events including plant behavior and operator procedures. As a general modeling tool, EMRALD can also be used to model cyberattack scenarios with varying mitigation options and quantify effectiveness, producing numerical data for risk-informed decisions. This paper uses EMRALD to demonstrate that dynamic risk analysis can be used for cyber threat modeling to provide insights for design decision-making and optimize defense strategies.

97 - MATHEMATICS AND COMPUTING↗

Reduced‐Order Probabilistic Emulation of Physics‐Based Ring Current Models: Application to RAM‐SCB Particle Flux

Abstract In this work, we address the computational challenge of large‐scale physics‐based simulation models for the ring current. Reduced computational cost allows for significantly faster than real‐time forecasting, enhancing our ability to predict and respond to dynamic changes in the ring current, valuable for space weather monitoring and mitigation efforts. Additionally, it can also be used for a comprehensive investigation of the system. Thus, we aim to create an emulator for the Ring current‐Atmosphere interactions Model with Self‐Consistent magnetic field (RAM‐SCB) particle flux that not only improves efficiency but also facilitates forecasting with reliable estimates of prediction uncertainties. The probabilistic emulator is built upon the methodology developed by Licata and Mehta (2023), https://doi.org/10.1029/2022sw003345 . A novel discrete sampling is used to identify 30 simulation periods over 20 years of solar and geomagnetic activity. Focusing on a subset of particle flux, we use Principal Component Analysis for dimensionality reduction and Long Short‐Term Memory (LSTM) neural networks to perform dynamic modeling. Hyperparameter space was explored extensively resulting in about 5% median symmetric accuracy across all data sets for one‐step dynamic prediction. Using a hierarchical ensemble of LSTMs, we have developed a reduced‐order probabilistic emulator (ROPE) tailored for time‐series forecasting of particle flux in the ring current. This ROPE offers accurate predictions of omnidirectional flux at a single energy with no pitch angle information, providing robust predictions on the test set with an error score below 11% and calibration scores under 8% with bias under 2% providing a significant speed up as compared to the full RAM‐SCB run.

79 ASTRONOMY AND ASTROPHYSICS↗

Machine Learning for Anomaly Detection in Neural Network Security and SRF Cavities

This dissertation explores the development and deployment of machine learning approaches to address critical challenges in anomaly detection across two distinct domains: neural network security in federated learning settings and cavity behavior analysis in particle accelerator operations at Jefferson Lab in Newport News, Virginia. Anomaly detection identifies deviations from expected patterns, safeguarding systems in cybersecurity, industry, and research against malicious activities and failures. This dissertation demonstrates how our machine learning approaches enhance detection accuracy and efficiency in both neural network security and industrial applications. First, we investigate vulnerabilities in deep neural networks deployed in federated learning. Although federated learning preserves user privacy by training models locally, it remains vulnerable to backdoor attacks, in which malicious participants embed hidden triggers that induce targeted misbehavior. We propose a self-supervised contrastive learning framework to detect and mitigate such backdoor attacks. In our experiments, this method achieves higher detection accuracy and lower false positive rates than existing defenses, while operating without access to local model updates or original training data and thus preserving the privacy guarantees of the federated setting. Second, we address the operational reliability of superconducting radio-frequency (SRF) cavities at the Continuous Electron Beam Accelerator Facility (CEBAF). Our research leverages an unsupervised learning approach, combined with Principal Component Analysis (PCA) and k-means clustering, to identify anomalous behaviors in SRF cavities. Our method detects subtle anomalous behavior by analyzing SRF signal data. This knowledge allows for the early detection and resolution of potential faults, significantly improving the efficiency and reliability of operations. Third, we extend these insights to time-series anomaly detection more broadly. We design a contrastive-learning based model tailored to increasingly dynamic environments and academic research. This model improves detection accuracy in settings that require real-time monitoring and predictive maintenance. Our research underscores the broader applicability and impact of advanced machine learning techniques in anomaly detection. By extracting meaningful patterns from complex data, machine learning can significantly enhance security in distributed neural networks and improve the efficiency of particle accelerator operations. This dissertation serves as a stepping stone for future investigations into the vast possibilities of anomaly detection, inspiring further exploration and development of machine learning techniques in this field.

Ferguson, Hal [Old Dominion University]↗

Performing Numerical Analysis of Cybersecurity Options Using Dynamic Risk Analysis Tool EMRALD

Cyberattacks can have many different attack paths, durations, and goals. There are also many different mitigation options involving hardware, software, and/or humans. Considering a cyber threat should involve defense-in-depth methods and a quantitative or numerical evaluation of overall effectiveness against dynamic, time-dependent attacks to make cost and risk-informed decisions. Typical cyberattack modeling methods only provide a qualitative evaluation. The main areas of cybersecurity are confidentiality, integrity, and availability. For companies with cyber-physical systems such as advanced nuclear reactors, cyber-related safety is a requirement set by North American Electric Reliability and the U.S. Nuclear Regulatory Commission. They are also concerned about availability or reliability as a business case. As cyber threats are evolving to a business-for-hire structure, more attacks may focus on disrupting business success and reliability, causing financial and economic stability risk. Companies want to know business reliability and recovery from those threats, and that requires modeling physical behavior of the targets. Dynamic-state-based and Markov-based modeling provides a method for better cyber scenario modeling with different tools having issues such as state-base explosion. Dynamic modeling enables time and conditional features not found in other numerical evaluation methods. EMRALD (Event Modeling Risk Assessment using Lined Diagrams) is a dynamic risk analysis modeling and simulation tool and has features that reduce modeling issues. It has been used to model different time-dependent events including plant behavior and operator procedures. As a general modeling tool, EMRALD can also be used to model cyberattack scenarios with varying mitigation options and quantify effectiveness, producing numerical data for risk-informed decisions. This paper uses EMRALD to demonstrate that dynamic numerical risk analysis can be used for cyber threat modeling to provide insights for design decision-making and optimize defense strategies. Keywords: cyber modeling; cyber-physical systems; numerical cyber modeling

97 - MATHEMATICS AND COMPUTING↗

Diaspora: Resilience-Enabling Services for Real-Time Distributed Workflows

The need for real-time processing to enable automated decision making and experimental steering has driven a shift from high-performance computing workflows on a centralized system to a distributed approach that integrates remote data sources, edge devices, and diverse compute facilities. Under this paradigm, data can be processed close to the source where it is generated, thus reducing latency and bandwidth usage. System resilience is thus a key challenge, requiring distributed workflows to survive component failures and to meet stringent quality-of-service requirements, which results in the need to mitigate anomalies such as congestion and low availability of resources. To address these challenges, we propose Diaspora, a unified resilience framework that is inspired by event-driven communication patterns used in public clouds. Specifically, we propose an event fabric that extends across sites, facilities, and computations to provide timely, reliable, and accurate information about data, application, and resource status. On top of the event fabric, we build resilience-enabling services that combine QoS-aware data streaming, resilient data views, resilient compute and data resources, and anomaly detection and prediction, all of which collectively enhance workflow resilience for these scientific cases.

Rao, Nageswara↗

REFSafE: A RAG-Enabled Framework for Predictive Risk Analysis and Automated Safety Report Generation in Mission-Critical Environments

Operational safety in mission-critical environments requires AI systems that are accurate, interpretable, and resistant to hallucination. We present an agentic Retrieval-Augmented Generation (RAG) framework, REFSafe, for grounded hazard analysis and automated safety report generation. The system integrates Large Language Models (LLMs) with structured operational data, historical incident repositories, policy documents, and external authoritative sources. Through iterative agentic reasoning, the framework retrieves, verifies, and synthesizes evidence prior to generation, enforcing citation-backed outputs with explicit source attribution (documents, links, and prior events) to ensure traceability and trust. To mitigate hallucinations and unsupported claims, all risk assessments and forecasts are constrained to retrieved evidence, with confidence signals derived from retrieval relevance and source consistency. A transparent pipeline enables subject matter experts (SMEs) to validate predictions, and provide structured feedback, forming a continuous performance calibration loop. Preliminary deployment demonstrates improved reliability in hazard detection and safety/vulnerability report generation. This work advances trustworthy, evidence-grounded AI for predictive safety intelligence in mission-critical operations.

Das, Sanjay [ORNL] (ORCID:0009000542591915)↗

Uncertainty-Aware and Explainable Human Error Detection in the Operation of Nuclear Power Plants

The timely and accurate identification of incidents, such as human factor error, is important to restore nuclear power plants (NPPs) to a stable state. However, the identification of abnormal operating conditions is difficult because of the existence of multiple scenarios. In addition, to implement mitigation actions rapidly after an incident occurs, operators must accurately identify an incident by monitoring the trends of many variables. The mental burden posed by this can increase human error and cause failure in identifying incidents. Failure to identify incidents directly results in erroneous mitigation measures, which are detrimental to NPPs. In this study, we leverage uncertainty-aware models to identify such errors and thereby increase the chances of mitigating them. We use the data collected from a physical test bed. The goal is to identify both certain and accurate models. For this, the two main aspects of focus in this study are explainable artificial intelligence (XAI) and uncertainty quantification (UQ). While XAI elucidates the decision pathway, UQ evaluates decision reliability. Their integration paints a comprehensive picture, signifying that understanding decisions and their confidence should be interlinked. Thus, in this study we leverage UQ measures (e.g. entropy and mutual information) along with Shapley additive explanations to gain insights into the features contributing to both accuracy and uncertainty in error identification. Furthermore, our results show that uncertainty-aware models combined with XAI tools can explain the artificial intelligence–prescribed decisions, with the potential of better explaining errors for the operators.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

The Design and Evaluation of Zero Trust Architecture for Electric Vehicle Charging Infrastructure: EVs @ Scale Series on EV Charging Station Cybersecurity

Implementing a zero trust architecture can significantly bolster the security of electric vehicle (EV) charging infrastructure. EV charging infrastructure includes numerous networked interfaces, each of which can present potential vulnerabilities. When these vulnerabilities are exploited, they can compromise the entire system, leading to severe operational and security risks. Zero trust is a security model that operates on the principle of "never trust, always verify," which helps manage the attack surface and limit the scope of any potential compromises. Fundamentally, this model ensures that no entity, whether inside or outside the network, is trusted by default. The design principles of zero trust include continuous verification, strict deny-by-default access controls, and micro-segmentation. Continuous verification ensures that every request is thoroughly checked, regardless of its origin. Strict access controls enforce the principle of least privilege, allowing users and devices only the minimum necessary access to perform their functions. Micro-segmentation involves dividing the network into smaller, isolated segments to prevent lateral movement in case of a breach. In the context of EV charging infrastructure, zero trust can be implemented through various strategies. For example, multi-factor authentication (MFA) can be required for engineers to access the management interfaces and control systems of charging stations. Real-time monitoring and analysis of network traffic can help detect and respond to anomalies. Systems that do not need to communicate with each other can be micro-segmented to enhance security. All communications should adhere to predefined policies to be permitted. Additionally, encrypting communications can protect sensitive information exchanged between chargers and management systems. This paper presents a zero trust architecture specifically designed for EV charging infrastructure. Implementing zero trust not only mitigates risks but also builds a resilient infrastructure capable of withstanding and quickly recovering from cyber threats. The architecture addresses six defined security objectives. A comprehensive test plan is developed to assess the architecture against these objectives, and the results of the evaluation are reported. This approach is essential for maintaining the reliability and integrity of EV charging services in an increasingly interconnected and vulnerable digital landscape. This is the first in a planned series of papers exploring the implementation of zero trust in EV charging infrastructure. Each paper will delve into different aspects and applications of zero trust, highlighting how various work processes and requirements can lead to distinct architectural designs. These architectures will be tailored to address specific security challenges and operational needs within the EV charging ecosystem, ensuring a robust and adaptable security framework.

33 ADVANCED PROPULSION SYSTEMS↗

Continual Learning for Production-Level Machine Learning in Particle Accelerators

Particle accelerators operate in complex environments where data distribution can change dynamically, leading to data drifts that significantly challenge Machine Learning (ML) models. These non-stationary conditions often cause ML models to deteriorate in performance, making it difficult to maintain reliable predictions in operation. The primary sources of data drifts are changes in accelerator settings and changes in equipment performance which cannot be measured directly. To bridge this gap between ML development and long-term deployment in operational settings, we identify key areas within particle accelerators where continual learning can help mitigate drift-induced performance degradation. We will provide a practical guide on selecting the appropriate method given resource constraints and desired stability plasticity trade offs. As a concrete example, we will present a real-world use case for anomaly detection to predict errant beams at the Spallation Neutron Source accelerator, where continual learning has been employed to demonstrate stable performance on drifting data streams. We will present practical challenges, lessons learned, and the results from the deployed ML model.

Rajput, Kishansingh [Thomas Jefferson National Acc↗