Search NASASearch

SEARCH · Search NASA

Results for “reliability modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Improving Reliability of Large Language Models for Nuclear Power Plant Diagnostics [Poster]

Large Language Models (LLMs) struggle out of the box when answering factually about detailed questions, especially in domains that are sparsely represented in their training data. This causes hallucinations and reduces reliability making it difficult for them to be used in practice. This work shows that using RAG techniques can improve factual accuracy and reliability, allowing for the application of LLMs in specialized areas, even when those areas that aren’t extensively covered in their initial training.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Improving Reliability of Large Language Models for Nuclear Power Plant Diagnostics Technical Presentation

Large Language Models (LLMs) struggle out of the box when answering factually about detailed questions, especially in domains that are sparsely represented in their training data. This causes hallucinations and reduces reliability making it difficult for them to be used in practice. This work shows that using RAG techniques can improve factual accuracy and reliability, allowing for the application of LLMs in specialized areas, even when those areas that aren’t extensively covered in their initial training.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Improving neutrino-nuclei interaction models: Recommendations and case studies on Peelle’s Pertinent Puzzle

Improving the modeling of neutrino-nuclei interactions using data-driven methods is crucial for high-precision neutrino oscillation experiments. This paper investigates Peelle’s Pertinent Puzzle (PPP) in the context of neutrino measurements, a longstanding challenge to fitting theoretical models to experimental data. Inconsistencies in data-model comparisons hinder efforts to enhance the accuracy and reliability of model predictions. We analyze various sources contributing to these inconsistencies and propose strategies to address them, supported by practical case studies. We advocate for incorporating model fitting exercises as a standard practice in cross section publications to enhance the robustness of results. We use a common analysis framework to explore PPP-related challenges with MicroBooNE and T2K data in an unified manner. Our findings offer valuable insights for improving the accuracy and reliability of neutrino-nuclei interaction models, particularly by systematically tuning models using data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Hierarchical Bayesian Modeling for Cosmology: Can NPE reliably replace MCMC?

Hierarchical neural posterior estimation has its place Hierarchical Bayesian Modeling (HBM) combined with MCMC algorithms has been shown to provide more robust and accurate inference for real-world phenomena in which nature takes a nested form. However, MCMC-based inference can be computationally expensive, and its performance often suffers for complex posterior geometries. These costs are especially pertinent for HBM. Studies have recently demonstrated the potential for a flexible, expressive, and amortized hierarchical neural posterior estimator (HNPE) built on Normalizing Flows. These studies have mostly been performed on simple datasets, or they focus on a single parameter from each level of the hierarchy. A systematic study analyzing how both hierarchical methods compare for more complex and realistic datasets is necessary before applying HNPE for scientific measurements. Here, we re-explore the theory behind HNPE and conduct comparative numerical experiments of HNPE and MCMC-based HBM methods on real and synthetic data, including strong gravitational lensing simulations. In particular, we use a suite of diagnostics to show trade-offs in terms of accuracy, precision, time to train or sample, reproducibility, and the need for expert domain knowledge. Especially for higher dimensional and complex posteriors, HNPE is expected to drastically improve on time for inference, accuracy, and precision with an upfront training time cost.

Hur, Rachel [Chicago U.] (ORCID:000900089890445X)

A System-Level Cost Modeling Framework for Design for Remanufacturing: A Case Study of an Agricultural Machine Transmission

Remanufacturing offers significant environmental and economic benefits by restoring end-of-life products to as-new conditions. Although extensive research has been conducted on the topic, the adoption of remanufacturing practices remains limited across various industries. A primary barrier to broader implementation is the substantial upfront investment required, which necessitates reliable cost modeling to justify potential future savings. Most existing models treat components independently and ignore inter-component dependencies. We develop a probabilistic, system-level cost modeling framework that integrates reliability, reusability, and a dependency matrix to capture cascading effects across components over multiple life cycles. Our model identifies those critical components that maximize remanufacturing benefits across a product's many lives. A toy example and an industry case study (John Deere PowrQuad transmission subassembly) illustrate how design alternatives affect cumulative cost. Using a Monte Carlo simulation (MCS) to perform life cycle cost analysis on the system with different design changes, we show the normalized average cost savings after three remanufacturing cycles. Accounting for dependencies meaningfully alters cost projections and ignoring them underestimates accumulated cost by up to 20% in our examples. Furthermore, the results of our study confirm that accounting for component interdependencies is necessary to produce cost estimates that meet industry standards.

Life Cycle Analysis and Design

Reliability Analysis of Power Grids Considering Component Failures of Variable Energy Resources

This paper proposes an improved model for the reliability assessment of power systems considering component failures of variable energy resources (VER). The inherent intermittency of VER such as solar photovoltaic (PV) and wind farms, along with their susceptibility to component failures, present significant challenges to reliable system operation. These issues, combined with power grid operation and network constraints, complicate the reliable operation of VER-integrated power systems. Here, to address these concerns, this paper introduces a reliability assessment framework that considers VER input variability, its impact on component availability, and their resulting impact on overall system reliability. Stochastic models based on discrete Markov processes are developed to incorporate variable irradiance, wind speeds, and their effects on PV and wind component failure rates. A next-event and state transition-based approach is then developed to integrate the stochastic models into a mixed-timing sequential Monte Carlo simulation framework for composite reliability assessment. Case studies on the RTS-GMLC system demonstrate the effectiveness of the proposed model in evaluating the reliability of VER-integrated systems.

Pandit, Dilip [Sandia National Laboratories (SNL-N

Active operator learning with predictive uncertainty quantification for partial differential equations

With the increased prevalence of neural operators being used to provide rapid solutions to partial differential equations (PDEs), understanding the accuracy of model predictions and the associated error levels is necessary for deploying reliable surrogate models in scientific applications. Existing uncertainty quantification (UQ) frameworks employ ensembles or Bayesian methods, which can incur substantial computational costs during both training and inference. Here, we propose a lightweight predictive UQ method tailored for Deep operator networks (DeepONets) that also generalizes to other operator networks. Numerical experiments on linear and nonlinear PDEs demonstrate that the framework’s uncertainty estimates are unbiased and provide accurate out-of-distribution uncertainty predictions with a sufficiently large training dataset. Our framework provides fast inference and uncertainty estimates that can efficiently drive outer-loop analyses that would be prohibitively expensive with conventional solvers. We demonstrate how predictive uncertainties can be used in the context of Bayesian optimization and active learning problems to yield improvements in accuracy and data-efficiency for outer-loop optimization procedures. In the active learning setup, we extend the framework to Fourier Neural Operators (FNO) and describe a generalized method for other operator networks. To enable real-time deployment, we introduce an inference strategy based on precomputed trunk outputs and a sparse placement matrix, reducing evaluation time by more than a factor of five. Our method provides a practical route to uncertainty-aware operator learning in time-sensitive settings.

97 MATHEMATICS AND COMPUTING

Assessing the reliability of medical resource demand models in the context of COVID-19

Abstract Background Numerous medical resource demand models have been created as tools for governments or hospitals, aiming to predict the need for crucial resources like ventilators, hospital beds, personal protective equipment (PPE), and diagnostic kits during crises such as the COVID-19 pandemic. However, the reliability of these demand models remains uncertain. Methods Demand models typically consist of two main components: hospital use epidemiological models that predict hospitalizations or daily admissions, and a demand calculator that translates the outputs of the epidemiological model into predictions for resource usage. We conducted separate analyses to evaluate each of these components. In the first analysis, we validated various hospital use epidemiological models using a recent validation framework designed for epidemiological models. This allowed us to quantify the accuracy of the models in predicting critical aspects such as the date and magnitude of local COVID-19 peaks, among other factors. In the second analysis, we evaluated a range of demand calculators for ventilators, medical gowns, and COVID-19 test kits. To achieve this, we decoupled these demand calculators from the underlying epidemiological models and provided ground truth data for their inputs. This approach enabled a direct comparison of the demand calculators, comparing them against each other and actual usage data when available. The code is available athttps://doi.org/10.5281/zenodo.13712387. Results Performance varied greatly across the epidemiological models, with greater variability in COVID-19 hospital use predictions than for COVID-19 deaths as analyzed previously. Some models did not have any peaks. Among those that did, the models under-estimated date of peak approximately as often as they over-estimated, but were more likely to under-estimate magnitude of peak, with typical relative errors around 50%. Regarding demand calculator predictions, there was significant variability, including five-fold differences in predictions for gown models. Validation against actual or surrogate usage data illustrated the potential value of demand models while demonstrating their limitations. Conclusions The emerging field of demand modeling holds promise in averting medical resource shortages during future public health emergencies. However, achieving this potential necessitates focused efforts on standardization, transparency, and rigorous model validation before placing reliance on demand models in critical public health decision-making.

Medical Informatics

ACCELERATED DEPLOYMENT OF NOVEL MATERIALS BASED ON RELIABILITY INTEGRITY MANAGEMENT USING CUMULATIVE DAMAGE MODELING

There is currently no widely agreed, detailed general method for licensing a novel plant incorporating novel materials (or materials being deployed in novel environments); in many such situations, there are no directly applicable engineering code cases for decision-makers (including regulators) to rely on. This paper discusses a framework for solving this problem that is based on the Reliability and Integrity Management (RIM) approach delineated in ASME BPVC Section XI Division 2. NRC Regulatory Guide 1.246, Rev. 0, endorses, with conditions, the subject portion of the 2019 ASME Code. The proposed framework is meant to support development of a licensing case by addressing certain remaining technical challenges. The framework discussed here is compatible with the Licensing Modernization Project, but applying it in a specific case will call for advances in the state of practice, if not the state of the art. The RIM approach calls for applicants to (a) allocate reliability targets to plant structures, systems, and components (SSCs), (b) show that they are able to relate the currently observed physical condition of each SSC in the program to its failure probability well enough to determine whether the target reliability allocations are being satisfied, allowing for uncertainty related to the novelty of the materials/designs/operating environments, and (c) be able to demonstrate that the proposed program of surveillances will reliably detect unacceptable degradation of an SSC before SSC failure occurs. A modeling approach potentially applicable to item (b), based on cumulative damage modeling rather than failure rates, is briefly illustrated.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Assessing and Enabling Trustworthy Predictions for High-Consequence Decisions

Predictions from physics-based computational models provide critical information to inform high consequence decisions, e.g., engineering design decisions. The ability to assess the reliability of such predictions is therefore critical. However, to date, reliability assessment rely heavily on expert judgment and qualitative arguments. This report details the efforts of LDRD 233072 to develop quantitative methods to assess reliability of model predictions, especially in the context of simplifying assumptions that can impact their reliability.

42 ENGINEERING

Reliability Assessment of Cooling Fans for PV Inverters: Testing, Modeling, and Case Studies

The reliability of photovoltaic (PV) inverters is critical for long-term solar system performance, with cooling fan failures frequently leading to costly downtime. While much research exists on general cooling fan reliability, little attention has been given to fans operating within PV inverters and their unique environmental challenges. Here, this article proposes a comprehensive methodology to address this gap. First, a failure mode and effects analysis is performed on fans to identify the key failure mechanisms in PV applications, their corresponding stressors, and the models necessary for lifetime prediction. Second, an accelerated life test is designed and conducted to collect valuable experimental data for PV inverter fans in a reasonable amount of time. Third, a mathematical conversion of dynamic mission profiles into effective constant stress levels is derived. Fourth, case studies are given, showcasing lifetime estimates that account for geographic variations in mission profile data. The results demonstrate that this integrated approach leads to an accurate reliability assessment for PV inverter cooling fans.

accelerated life testing (ALT)

Floating photovoltaic power plants: A review of energy yield, reliability, and operation and maintenance

Photovoltaic (PV) systems are essential for the transition to sustainable energy, reducing fossil fuel dependence and mitigating climate change. Although PV requires minimal land area — PV can meet the European Union's energy needs using only 0.26% of its land — space for deployment is often scarce in densely populated regions. Floating photovoltaics (FPV) offer an effective solution to land-use challenges by installing PV systems on floating structures in water bodies. FPV is a growing niche within PV with a cumulative installed capacity reaching 7.7 GW globally by 2023. Almost 90% of the installed FPV capacity is in Asia, with close to 50% of in China alone, while the Netherlands and France are the largest markets outside Asia. FPV shows strong potential to support climate targets, but still faces challenges like regulatory barriers, cost competitiveness compared to ground-based PV (GPV), and uncertainties about environmental impacts and system reliability. FPV systems are currently installed mainly on sheltered inland waters, such as quarry lakes, irrigation ponds and reservoirs. FPV technical standards are still being developed. Guidelines have been published by the World Bank, DNV, and Solar Power Europe, and emerging national standards from South Korea, China, and Singapore address design, components, and safety. The International Electrotechnical Commission (IEC) is working on formal standards for floats, mooring systems, and electrical connectors. However, the published best practices lack quantitative guidance for yield modelling and reliability, which this report aims to address. It provides data-driven insights, models, and parameters essential for accurate energy yield, reliability, and maintenance predictions over FPV systems' lifetimes.

14 SOLAR ENERGY

Challenges of standard halo models in constraining galaxy properties from cosmic infrared background anisotropies

The halo model, combined with halo occupation distribution (HOD) prescriptions, is widely used to interpret cosmic infrared background (CIB) anisotropies and extract physical information about star-forming galaxies and their connection to large-scale structures. Recent CIB-specific implementations of the halo model have adopted more physical parameterizations. However, the extent to which these models can reliably recover meaningful physical parameters remains uncertain. We assessed whether the current parameterization of CIB halo models is sufficient to recover astrophysical quantities, such as star formation efficiency, η(M h , z), and halo mass at which the peak of star formation efficiency occurs, M max , when fit to mock data. We also assessed whether discrepancies arise from assumptions about galaxy emission (the HOD ingredients) or from more fundamental components in the halo model, such as bias and matter clustering. We fit the M21 CIB HOD model, implemented within the halo model framework, to mock CIB power spectra and star formation rate density (SFRD) data generated from the SIDES-Uchuu simulation, and compared the best-fit parameters to the known simulation inputs. We then repeated the analysis using a simplified version of the simulation (SSU), explicitly designed to match the HOD assumptions. A detailed comparison of model and simulation outputs was carried out to trace the origin of observed discrepancies. While the M21 HOD model provides a good fit to the mock data, it failed to recover the intrinsic parameters accurately, particularly the halo mass at which star formation efficiency peaks. This mismatch persists even when fitting data generated with the same model assumptions. We find strong agreement (within 5%) in the emission-related components (SFRD, emissivity), but observe a scale- and redshift-dependent offset exceeding 20% in the two-halo term of the CIB power spectrum. This likely arises from limitations in the treatment of halo bias and matter clustering within the linear approximation. Additionally, incorporating scatter in the SFR–halo mass relation and the spectral energy distribution (SED) templates significantly affects the shot noise (∼50%), but has only a modest impact (less than 10%) on the clustered component. These results suggest that recovering physical parameters from CIB clustering requires improvements to the cosmological ingredients of the halo model framework, such as adopting scale-dependent halo bias and nonlinear matter power spectra in addition to careful modeling of emission physics.

cosmic background radiation

Closing the loop: model-predictive control for a closed-circuit reverse osmosis system

This article presents a model-predictive controller (MPC) for the maximization of the energy efficiency of a closed-circuit desalination reverse osmosis (CCRO) system. CCRO is a process for producing drinking water that is based on a cyclic operation with the following two phases: (a) filtration and (b) drain. In this article, we test model predictive control for optimal control of this process. The most important features of our approach are as follows: (a) the selection of a model structure that enables reliable forecasts of the filtration phase (up to 3 h), (b) an on-line model calibration strategy that ensures model forecast reliability, and (c) the satisfaction of equipment safety and operational constraints on the selected setpoints. We challenge this through deliberate introduction of changes in the unmeasured feed concentration and the applied constraints. Our results indicate that frequent model parameter updates are critical to maintain model reliability for MPC purposes. In addition, we illustrate that parameter identifiability is not guaranteed and that deliberate variation in flow rates is necessary even though the process never operates in steady state. Finally, MPC can compute flow rate setpoints that maximize the energy efficiency of the CCRO process while satisfying the applicable equipment and safety constraints.

closed-circuit reverse osmosis

Predictive links between microbial communities and biological oxygen utilization in the Arctic Ocean

Microbial metabolism influences rates of net community production (NCP), exerting a direct biological control on marine oxygen and carbon fluxes. In the Arctic, it is increasingly important to understand and quantify this process, as ecological and oceanographic conditions shift due to changing climate. Here, we describe potential ecological links between pelagic microbial diversity and an NCP precursor, biological oxygen utilization, using machine learning and paired observations of community structure and metabolic activity from a seasonally and spatially variable transect of the Arctic Ocean (2019–2020 MOSAiC Expedition). Community structure was determined using 16S (prokaryotic) and 18S (eukaryotic) rRNA gene amplicon sequencing, and metabolic activity was derived from ΔO 2 /Ar. Using self-organizing maps, we identified clear successional patterns in observed microbial community structure that were seasonally driven in the upper ocean and vertically stratified with depth. Metabolic activity was also stratified, with a primarily net heterotrophic water column (median −1.5% biological oxygen saturation), excepting periodic oxygen supersaturation (maximum: 13.6%) within the mixed layer. Using DNA sequences as predictor variables, we then constructed a random forest regression model that reliably reconstructed biological oxygen concentrations (root mean squared error = 4.14 μmol kg −1 ). Top predictors from this model were from heterotrophic (bacteria) or potentially mixotrophic (dinoflagellate) taxa. These analyses highlight biologically driven diagnostic tools that can be used to expand biogeochemical datasets and improve the microbial perspectives and metabolisms represented in ecological models of net productivity and carbon flux in a changing Arctic Ocean.

Chamberlain, Emelia J. [Univ. of San Diego, San Di

Deep-learning-derived planetary boundary layer height from conventional meteorological measurements

Abstract. The planetary boundary layer (PBL) height (PBLH) is an important parameter for various meteorological and climate studies. This study presents a multi-structure deep neural network (DNN) model, which can estimate PBLH by integrating the morning temperature profiles and surface meteorological observations. The DNN model is developed by leveraging a rich dataset of PBLH derived from long-standing radiosonde records augmented with high-resolution micro-pulse lidar and Doppler lidar observations. We access the performance of the DNN with an ensemble of 10 members, each featuring distinct hidden-layer structures, which collectively yield a robust 27-year PBLH dataset over the southern Great Plains from 1994 to 2020. The influence of various meteorological factors on PBLH is rigorously analyzed through the importance test. Moreover, the DNN model's accuracy is evaluated against radiosonde observations and juxtaposed with conventional remote sensing methodologies, including Doppler lidar, ceilometer, Raman lidar, and micro-pulse lidar. The DNN model exhibits reliable performance across diverse conditions and demonstrates lower biases relative to remote sensing methods. In addition, the DNN model, originally trained over a plain region, demonstrates remarkable adaptability when applied to the heterogeneous terrains and climates encountered during the GoAmazon (Green Ocean Amazon; tropical rainforest) and CACTI (Cloud, Aerosol, and Complex Terrain Interactions; middle-latitude mountain) campaigns. These findings demonstrate the effectiveness of deep learning models in estimating PBLH, enhancing our understanding of boundary layer processes with implications for improving the representation of PBL in weather forecasting and climate modeling.

54 ENVIRONMENTAL SCIENCES

AI Model Benchmarking for Nonproliferation Applications: Steel Thread Benchmarking Task Force Technical Report (Rev. 2)

Steel Thread is a NA-22 venture that seeks to build trustworthy, reliable AI models that can be used in a wide variety of nonproliferation tasks. A key aspect of building these models is developing appropriate benchmarks and evaluation methods, which will enable the venture to identify and adapt models to provide the most value in the nonproliferation domain. Benchmarks must be relevant to key tasks in this domain, such as question answering, information retrieval, document summarization and classification, consensus analysis, and image and data analysis. This report 1) provides an overview of benchmark design, evaluation, and challenges; 2) reviews a variety of open benchmarks, with a focus on language models and tasks; and 3) identifies benchmarks that are most relevant to Steel Thread. This report is intended to serve as a basis for further efforts to classify and evaluate benchmarks and their correlation with success on nonproliferation-specific tasks. The Steel Thread venture has defined benchmarks to be a particular combination of a dataset (or datasets) and a metric (or metrics) conceptualized as representing one or more specific tasks or sets of abilities for a specific modality. It is adopted by a research community as a shared framework for comparing methods.1 It includes 1) Data: Labeled (a designated subset not used for training, which could be all the data), 2) Metric: A way to quantify performance, 3) Task/Ability: What the benchmark is testing, 4) Protocol: A structured and repeatable evaluation process, 5) Baseline/Reference Model: For comparison; could be statistical, rule-based, SME-derived, or another model, and 6) Maintenance Plan: to update with new information over time; important for long-term utility. For further clarity, the definition includes what a benchmark, in this context, is not. It is not a corpus of training data, specific to a model (it is intended to apply to a range of models), a universal evaluation of performance, a guarantee that the ‘top’ model on the leaderboard will be the best fit for every specific use case, an all-encompassing proof of a model’s universal quality, nor is it a one-size-fits-all measure of success. It does not cover every real-world constraint (like operational, ethical, or cost considerations), a systems integration test, or a unit test. This definition was inspired by and resulted from discussions within the Steel Thread Benchmarking Task Force. This group was formed to define what we would mean as a benchmark within Steel Thread but persisted as the need to develop a thorough understanding of the large and expanding existing benchmarking space. This technical report is a result of the group’s divide and conquer approach to exploring this space. The release of benchmarks might not be progressing as quickly as model development, but it is moving very fast, as many benchmarks quickly become saturated, when state-of-the-art models score so close to the benchmark’s ceiling that their results are virtually indistinguishable. At that point, the test no longer differentiates between new systems, so researchers usually stop reporting scores as the benchmark no longer informs about improvements from the next generation of models. In the OpenAI announcement of GPT-5, they reported results on six flagship public benchmarks (AIME 2025, SWE-bench Verified, Aider Polyglot, MMMU, HealthBench Hard, GPQA) but the full system-card covers roughly thirty-five separate evaluations, comprising hundreds of test task items in total. There have been some efforts to summarize benchmarks in specific fields, like for text-to-image generation, but these surveys have had a narrow methodology scope. Therefore, a comprehensive survey of all benchmarks or even all benchmarks that could be relevant to Steel Thread is outside of the scope of this report. We chose some specific benchmarks to investigate in detail.

97 MATHEMATICS AND COMPUTING

Heterogeneous energetic material damage simulator (HEDS): A deep learning approach to simulate damage–sensitivity linkages

Damage in the microstructures of energetic materials (EMs), such as propellants and plastic bonded explosives (PBXs), can significantly alter their response to external loads. Both sensitization and desensitization can occur, causing concerns with safety and performance in the field; predictive models that connect damage and the sensitivity of EMs can enable design and provide confidence in their robustness and reliability. However, modeling of damage evolution is challenging for real microstructures of EMs; samples of damaged EMs are difficult to obtain, thereby hindering experiments and direct numerical simulations to determine the sensitivity of EMs at various stages of damage. Here, we develop an approach to generate synthetic, i.e., in silico produced, damaged microstructures for use in simulations to connect damage levels to sensitivity. The development of the present workflow to generate and impose varying levels of damage in microstructures, known as HEDS (Heterogeneous Energetic Material Damage Simulator), begins with a small set of images of damaged PBXs and combines a collection of deep neural network techniques to generate microstructures with varying levels of damage. By making the synthetic microstructures conform closely to those observed in available real, imaged microstructures, we develop an ensemble of damaged microstructures that can be used for in silico shock experiments. HEDS develops these microstructure ensembles as level set fields, which are directly employed in a sharp interface Eulerian hydrocode where shock simulations are performed to quantify the energy release rate from hotspot fields generated in the microstructure. These capabilities can be useful for the analysis and assessment of changes in the sensitivity of EMs and to design formulations that are less susceptible to damage-induced changes in sensitivity and performance.

Fang, Irene (ORCID:0009000844557122)