Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33

End-To-End Uncertainty Quantification with Analytical Derivatives for Design Under Uncertainty

Uncertainty quantification (UQ) is a rapidly growing and evolving discipline, especially within the aerospace community. Performing analysis with UQ can provide decision makers with a wealth of information about a candidate design. However, the value of UQ is fully realized when the information gained during UQ analysis is leveraged in a feedback loop of a design optimization process, often referred to as design under uncertainty. Although design under uncertainty can be a powerful risk mitigation technique, there are a number of roadblocks that prevent its implementation. Two primary factors are computational costs and added complexity of the analysis. High fidelity simulations on the order tens of uncertain variables quickly become computationally infeasible. Also, implementing UQ into an existing multidisciplinary design and optimization (MDO) process often requires extensive knowledge of the UQ methods and careful treatment of the problem formulation. The objective of this work is to address these two primary roadblocks and enable practitioners to efficiently perform design under uncertainty with limited knowledge of the UQ discipline. Methods outlined in this paper demonstrate MDO incorporating UQ into the design process, leveraging an analytic derivative tool chain through the entire optimization. The proposed approach leverages machine learning techniques to generate a differentiable confidence interval output from polynomial chaos models. This technique, coupled with the incorporation of analytical derivatives through the Polynomial Chaos Expansion (PCE) process, eliminates the need to estimate derivatives which are usually obtained from finite difference, complex step, or similar methods. Developing a differentiable confidence interval allows mixed uncertainty problems (both epistemic and aleatory) to be modeled. Without such modeling, these problems cannot accurately predict objective functions containing statistical quantities such as mean and variance. The addition of analytic derivatives to a polynomial chaos-based UQ method decreases the computational costs of performing design under uncertainty by orders of magnitude in comparison with methods such as complex step. The method and codes developed are modular in nature and are a drop-in solution for design under uncertainty within existing MDO problems. A low-fidelity analytical multidisciplinary optimization under uncertainty for a wing design in OpenMDAO is detailed in this paper. This demonstration case will include both objective functions and constraints which are influenced by uncertain parameters.

Ben Phillips↗

Decoding substance use disorder severity from clinical notes using a large language model

Substance use disorder (SUD) poses a major concern due to its detrimental effects on health and society. SUD identification and treatment depend on a variety of factors such as severity, co-determinants (e.g., withdrawal symptoms), and social determinants of health. Existing diagnostic coding systems used by insurance providers, like the International Classification of Diseases (ICD-10), lack granularity for certain diagnoses, but American clinicians will add this granularity (as that found within the Diagnostic and Statistical Manual of Mental Disorders classification or DSM-5) as supplemental unstructured text in clinical notes. Traditional natural language processing (NLP) methods face limitations in accurately parsing such diverse clinical language. Large language models (LLMs) offer promise in overcoming these challenges by adapting to diverse language patterns. This study investigates the application of LLMs for extracting severity-related information for various SUD diagnoses from clinical notes. We propose a workflow employing zero-shot learning of LLMs with carefully crafted prompts and post-processing techniques. Through experimentation with Flan-T5, an open-source LLM, we demonstrate its superior recall compared to the rule-based approach. Focusing on 11 categories of SUD diagnoses, we show the effectiveness of LLMs in extracting severity information, contributing to improved risk assessment and treatment planning for SUD patients.

60 APPLIED LIFE SCIENCES↗

High-precision Galaxy Clustering Predictions from Small-volume Hydrodynamical Simulations via Control Variates

Abstract Cosmological simulations of galaxy formation are an invaluable tool for understanding galaxy formation and its impact on cosmological parameter inference from large-scale structures. However, their high computational cost is a significant obstacle for running simulations that probe cosmological volumes comparable to those analyzed by contemporary large-scale structure experiments. In this work, we explore the possibility of obtaining high-precision galaxy clustering predictions from small-volume hydrodynamical simulations such as MillenniumTNG and FLAMINGO via control variates. In this approach, the hydrodynamical full-physics simulation is paired with a matched low-resolution gravity-only simulation. By learning the galaxy–halo connection from the hydrodynamical simulation and applying it to the gravity-only counterpart, one obtains a galaxy population that closely mimics the one in the more expensive simulation. One can then construct an estimator of galaxy clustering that combines the clustering amplitudes in the small-volume hydrodynamical and gravity-only simulations with clustering amplitudes in a large-volume gravity-only simulation. Depending on the galaxy sample, clustering statistic, and scale, this galaxy clustering estimator can have an effective volume of up to around 100 times the volume of the original hydrodynamical simulation in the nonlinear regime. With this approach, we can construct galaxy clustering predictions from existing simulations that are precise enough for mock analyses of next-generation large-scale structure surveys such as the Dark Energy Spectroscopic Instrument and the Legacy Survey of Space and Time.

Doytcheva, Alexandra (ORCID:0009000111254888)↗

Breaking the curse of dimensionality: Solving configurational integrals for crystalline solids by tensor networks

Accurately evaluating configurational integrals for dense solids remains a central and difficult challenge in the statistical mechanics of condensed systems. Here, we present a tensor network approach that reformulates the high-dimensional configurational integral for identical-particle crystals into a sequence of computationally efficient summations. We represent the integrand as a high-dimensional tensor and apply tensor-train (TT) decomposition together with a custom TT-cross interpolation. This approach circumvents the need to explicitly construct the full tensor. We introduce tailored rank-1 and rank-2 schemes optimized for sharply peaked Boltzmann probability densities, typical for identical-particle crystals. When applied to the calculation of internal energy and pressure-temperature curves for crystalline Cu and Ar at high (GPa) pressures, as well as the alpha-to-beta phase transition diagram of Sn, our method accurately reproduces molecular dynamics simulation results using tight-binding, machine learning, hierarchical interacting particle–neural network, and modified embedded atom method potentials,all within seconds of computation time.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Reduced‐Order Modeling for Linearized Representations of Microphysical Process Rates

Abstract Representing cloud microphysical processes in large scale atmospheric models is challenging because many processes depend on the details of the droplet size distribution (DSD, the spectrum of droplets with different sizes in a cloud). While full or partial statistical moments of droplet size distributions are the typical variables used in bulk models, prognostic moments are limited in their ability to represent microphysical processes across the range of conditions experienced in the atmosphere. Microphysical parameterizations employing prognostic moments are known to suffer from structural uncertainty in their representations of inherently higher dimensional cloud processes, which limit model fidelity and lead to forecasting errors. Here we investigate how data‐driven reduced‐order modeling can be used to learn predictors for microphysical process rates in bulk microphysics schemes in an unsupervised manner from higher dimensional bin distributions. Using simulations characteristic of marine stratiform clouds, we simultaneously learn lower dimensional representations of droplet size distributions and predict the evolution of the microphysical state of the system. Droplet collision‐coalescence, the main process for generating warm rain, is estimated to have an intrinsic dimension of three. This intrinsic dimension provides a lower limit on the number of degrees of freedom needed to accurately represent collision‐coalescence in models. We demonstrate how deep learning based reduced‐order modeling can be used to discover intrinsic coordinates describing the microphysical state of the system, where process rates such as collision‐coalescence are globally linearized. These implicitly learned representations of the DSD retain more information about the DSD than typical moment‐based representations.

54 ENVIRONMENTAL SCIENCES↗

The 3D Heliosphere: What Can We Learn from STEREO?

Many techniques have been used to study the 3D heliosphere, with the earliest probably being the analysis of comet tails. I will list most of these and mention a few, focusing on existing multi-point studies. The result, from more than 50 years of study, Is that a lot is known. This has led to a good picture of the quasi-steady heliosphere and its relation to the 3D Corona. But, there are also some large gaps and STEREO is designed to address one of these: the timing, size, geometry, mass, speed, direction, and 3D propagation of Corona[ mass ejections (CMEs). In spite of the statistical analysis of a large data archive, Imaginative use of in situ and remote measurements, and extensive modeling, these properties of CMES are poorly known. I will outline an example of how STEREO instruments might work together to develop a far better 30 description of CMEs In the 3D heliosphere and note that other examples are described in the Science Definition Team report and in the Science Objectives given by the four instrument teams. Since the two STEREO spacecraft are not intended to work in isolation, I will also outline how they might be used In combination With ground-based and other spacecraft observations.

Suess, S. T.↗

Lessons from Climate Modeling on the Design and Use of Ensembles for Crop Modeling

Working with ensembles of crop models is a recent but important development in crop modeling which promises to lead to better uncertainty estimates for model projections and predictions, better predictions using the ensemble mean or median, and closer collaboration within the modeling community. There are numerous open questions about the best way to create and analyze such ensembles. Much can be learned from the field of climate modeling, given its much longer experience with ensembles. We draw on that experience to identify questions and make propositions that should help make ensemble modeling with crop models more rigorous and informative. The propositions include defining criteria for acceptance of models in a crop MME, exploring criteria for evaluating the degree of relatedness of models in a MME, studying the effect of number of models in the ensemble, development of a statistical model of model sampling, creation of a repository for MME results, studies of possible differential weighting of models in an ensemble, creation of single model ensembles based on sampling from the uncertainty distribution of parameter values or inputs specifically oriented toward uncertainty estimation, the creation of super ensembles that sample more than one source of uncertainty, the analysis of super ensemble results to obtain information on total uncertainty and the separate contributions of different sources of uncertainty and finally further investigation of the use of the multi-model mean or median as a predictor.

Model ensembles↗

A High Performance Computing Approach to Tree Cover Delineation in 1-m NAIP Imagery Using a Probabilistic Learning Framework

Tree cover delineation is a useful instrument in deriving Above Ground Biomass (AGB) density estimates from Very High Resolution (VHR) airborne imagery data. Numerous algorithms have been designed to address this problem, but most of them do not scale to these datasets, which are of the order of terabytes. In this paper, we present a semi-automated probabilistic framework for the segmentation and classification of 1-m National Agriculture Imagery Program (NAIP) for tree-cover delineation for the whole of Continental United States, using a High Performance Computing Architecture. Classification is performed using a multi-layer Feedforward Backpropagation Neural Network and segmentation is performed using a Statistical Region Merging algorithm. The results from the classification and segmentation algorithms are then consolidated into a structured prediction framework using a discriminative undirected probabilistic graphical model based on Conditional Random Field, which helps in capturing the higher order contextual dependencies between neighboring pixels. Once the final probability maps are generated, the framework is updated and re-trained by relabeling misclassified image patches. This leads to a significant improvement in the true positive rates and reduction in false positive rates. The tree cover maps were generated for the whole state of California, spanning a total of 11,095 NAIP tiles covering a total geographical area of 163,696 sq. miles. The framework produced true positive rates of around 88% for fragmented forests and 74% for urban tree cover areas, with false positive rates lower than 2% for both landscapes. Comparative studies with the National Land Cover Data (NLCD) algorithm and the LiDAR canopy height model (CHM) showed the effectiveness of our framework for generating accurate high-resolution tree-cover maps.

Segments↗

A general mechanistic framework for cross-scale understanding of hot spots and hot moments in carbon and water fluxes

Semi-arid ecosystems, like those in the American Southwest, exert a massive impact on the interannual variability of carbon and water cycling. Unfortunately, these carbon and water fluxes are notoriously difficult to predict due to their high spatial and temporal variability, which is poorly captured by the current generation of vegetation models. Indeed, this region is exemplified by the ‘hot spots and hot moments’ concept, which states that small areas in space (‘hot spots’) and transient moments in time (‘hot moments’) exert an outsized influence on biogeochemical cycling. However, the factors that regulate these pulses in biogeochemical activity are unknown, as is their variability across space and time. These uncertainties severely limit efforts to better represent hot spots and hot moments in models. Here, we seek to develop a generalized method for detecting and quantifying the importance of hot spots and hot moments from individual plant to regional scales. Underpinning this method is our recently developed statistical approach for identifying hot spots and hot moments. By applying this method to semi-continuous measurements of plant water status, a depth profile of soil water potential, and ecosystem fluxes via eddy covariance, we will track the fate of water through the soil-plant-atmosphere continuum and identify the mechanistic drivers of these transient pulses in biogeochemical activity. Then, we will expand this approach across a broad network of Ameriflux towers, and apply a machine learning approach that will allow us to upscale measurements of hot spots and hot moments across the American Southwest and quantify their impact on carbon and water cycles. These products will allow us to identify hot spots and hot moments across spatio-temporal scales and will serve as crucial data sources for validating a new generation of models that can better capture highly dynamic carbon and water fluxes. The proposed method will be easily transferable across biomes and will serve as a framework for future research on hot spots and hot moments across the plant ecophysiology, biometeorology, and vegetation modeling communities.

54 ENVIRONMENTAL SCIENCES↗

A novel methodology for gamma-ray spectra dataset procurement over varying standoff distances and source activities

The adoption of machine learning approaches for gamma-ray spectroscopy has received considerable attention in the literature. Many studies have investigated the deployment of various algorithm architectures to a specific task. However, little attention has been afforded to the development of the datasets leveraged to train the models. Such training datasets typically span a set of environmental or detector parameters to encompass a problem space of interest to a user. Variations in these measurement parameters will also induce fluctuations in the detector response, including expected pile-up and ground scatter effects. Fundamental to this work is the understanding that 1) the underlying spectral shape varies as the measurement parameters change and 2) the statistical uncertainties associated with two spectra impact their level of similarity. While previous studies attribute some arbitrary discretization to the measurement parameters for the generation of their synthetic training data, this work introduces a principled methodology for efficient spectral-based discretization of a problem space. A signal-to-noise ratio (SNR) respective spectral comparison measure and a Gaussian Process Regression (GPR) model are used to predict the spectral similarity across a range of measurement parameters. This innovative approach effectively showcased its capability by dividing a problem space, ranging from 5 cm to 100 cm standoff distances and 5 μCi–100 μCi of 137 Cs, into three unique combinations of measurement parameters. The findings from this work will aid in creating more robust datasets, which incorporate many possible measurement scenarios, reduce the number of required experimental test set measurements, and possibly enable experimental training data collection for gamma-ray spectroscopy.

data science↗

Performance and Evaluation of the Global Modeling and Assimilation Office Observing System Simulation Experiment

The National Aeronautics and Space Administration Global Modeling and Assimilation Office (NASA/GMAO) has spent more than a decade developing and implementing a global Observing System Simulation Experiment framework for use in evaluting both new observation types as well as the behavior of data assimilation systems. The NASA/GMAO OSSE has constantly evolved to relect changes in the Gridpoint Statistical Interpolation data assimiation system, the Global Earth Observing System model, version 5 (GEOS-5), and the real world observational network. Software and observational datasets for the GMAO OSSE are publicly available, along with a technical report. Substantial modifications have recently been made to the NASA/GMAO OSSE framework, including the character of synthetic observation errors, new instrument types, and more sophisticated atmospheric wind vectors. These improvements will be described, along with the overall performance of the current OSSE. Lessons learned from investigations into correlated errors and model error will be discussed.

OSS↗

Application of Machine Learning Algorithms to the Study of Noise Artifacts in Gravitational-Wave Data

The sensitivity of searches for astrophysical transients in data from the Laser Interferometer Gravitationalwave Observatory (LIGO) is generally limited by the presence of transient, non-Gaussian noise artifacts, which occur at a high-enough rate such that accidental coincidence across multiple detectors is non-negligible. Furthermore, non-Gaussian noise artifacts typically dominate over the background contributed from stationary noise. These "glitches" can easily be confused for transient gravitational-wave signals, and their robust identification and removal will help any search for astrophysical gravitational-waves. We apply Machine Learning Algorithms (MLAs) to the problem, using data from auxiliary channels within the LIGO detectors that monitor degrees of freedom unaffected by astrophysical signals. Terrestrial noise sources may manifest characteristic disturbances in these auxiliary channels, inducing non-trivial correlations with glitches in the gravitational-wave data. The number of auxiliary-channel parameters describing these disturbances may also be extremely large; high dimensionality is an area where MLAs are particularly well-suited. We demonstrate the feasibility and applicability of three very different MLAs: Artificial Neural Networks, Support Vector Machines, and Random Forests. These classifiers identify and remove a substantial fraction of the glitches present in two very different data sets: four weeks of LIGO's fourth science run and one week of LIGO's sixth science run. We observe that all three algorithms agree on which events are glitches to within 10% for the sixth science run data, and support this by showing that the different optimization criteria used by each classifier generate the same decision surface, based on a likelihood-ratio statistic. Furthermore, we find that all classifiers obtain similar limiting performance, suggesting that most of the useful information currently contained in the auxiliary channel parameters we extract is already being used. Future performance gains are thus likely to involve additional sources of information, rather than improvements in the MLAs themselves.

gravitational-wave data↗

Mammalian Vestibular Macular Synaptic Plasticity: Results from SLS-2 Spaceflight

The effects of exposure to microgravity were studied in rat utricular maculas collected inflight (IF, day 13), post-flight on day of orbiter landing (day 14, R+O) and after 14 days (R+ML). Controls were collected at corresponding times. The objectives were 1) to learn whether hair cell ribbon synapses counts would be higher in tissues collected in space than in tissues collected postflight during or after readaptation to Earth's gravity; and 2) to compare results with those of SLS-1. Maculas were fixed by immersion, micro-dissected, dehydrated and prepared for ultrastructural study by usual methods. Synapses were counted in 100 serial sections 150 nm thick and were located to specific hair cells in montages of every 7th section. Counts were analyzed for statistical significance using analysis of variance. Results in maculas of IF dissected rats, one 13 day control (IFC), and one R + 0 rat have been analyzed. Study of an R+ML macula is nearly completed. For type I cells, IF mean is 2.3 +/-1.6; IFC mean is 1.6 +/-1.0; R+O mean is 2.3 +/- 1.6. For type II cells, IF mean is 11.4 +/- 17.1; IFC mean is 5.5 +/-3.5; R+O mean is 10.1 +/- 7.4. The difference between IF and IFC means for type I cells is statistically significant (p less than 0.0464). For type It cells, IF compared to IFC means, p less than 0.0003; and for IFC to R+O means, p less than 0.0139. Shifts toward spheres (p less than 0.0001) and pairs (p less than 0.0139) were significant in type II cells of IF rats. The results are largely replicating findings from SLS-1 and indicate that spaceflight affects synaptic number, form and distribution, particularly in type II hair cells. The increases in synaptic number and in sphere-like ribbons are interpreted to improve synaptic efficacy, to help return afferent discharges to a more normal state. Findings indicate that a great capacity for synaptic plasticity exists in mammalian gravity sensors, and that this plasticity is more dominant in the local circuitry. The local circuit includes type II cells and is interpreted to be responsible for shaping the final output of the system.

Ross, Muriel D.D.↗

Reconstruction and Selection of Neutrino Interactions in MicroBooNE using Deep Convolutional Neural Networks

In this document, we describe a new reconstruction workflow developed for the MicroBooNE experiment. It features the use of Deep Convolutional Neural Networks trained to recognize key structures within the data sufficient for the 3D reconstruction of neutrino interactions within the detector. As a test of the reconstruction utility, the products of the reconstruction workflow are used to select inclusive charged-current (CC) $\nu_e$ and $\nu_\mu$ interactions in both simulated and real MicroBooNE data. In simulation, our $\nu_e$ and $\nu_\mu$ selections achieve an efficiency of 57% and 68\%, respectively, with a purity of 91% and 96%, respectively. We find that these selections are competitive with the inclusive selections used for the most recent MicroBooNE LEE searches. In particular, the CC-$\nu_e$ inclusive selection efficiency improves by over 20% while also improving sample purity. As a first step in quantifying potential bias, the data and Monte Carlo expectati ons are compared for both selections using the MicroBooNE open data. Within statistical and systematic uncertainties, both the electron and muon CC-inclusive event samples agree. A comparison of the real data events chosen by our work and another reconstruction framework shows that the two analyses each identify a sizeable fraction of events the other does not. This suggests that future analyses integrating the strengths of each could lead to combined gains. This work demonstrates, for the first time on real LArTPC data, state-of-the-art neutrino interaction reconstruction centered around deep learning algorithms.

43 PARTICLE ACCELERATORS↗

Big-data Efficient and Automated Science Transfer (BEAST): An Open-Source Software Architecture for Arc Jet Data Management, Modeling, and Automation

Big-data Efficient and Automated Science Transfer (BEAST) is a facility data management application developed for the NASA Ames arc jet facilities. The current decentralized data management practices limit statistical tracking, synchronization between video/time series, search capability, data throughput, and data processing speed/efficiency. Consequently, BEAST was developed to provide a new data infrastructure with streamlined data collection, processing, transfer, and analysis. This new framework also seeks to implement the FAIR principles of data stewardship: Findable, Accessible, Interoperable, and Reusable. The BEAST framework is based on a combination of the Python Django web framework and the Python data stack to provide a monolithic, open-source platform for data management, automation, and machine learning. This architecture was chosen for maintainability and scalability for a small, in-house development team. This paper will describe the application framework, deployment, and discuss the benefits and future plans for the system.

Data management↗

Search for top squarks in final states with many light-flavor jets and 0, 1, or 2 charged leptons in proton-proton collisions at $\sqrt{s}=13$ TeV

Several new physics models including versions of supersymmetry (SUSY) characterized by R-parity violation (RPV) or with additional hidden sectors predict the production of events with top quarks, low missing transverse momentum, and many additional quarks or gluons. The results of a search for top squarks decaying to two top quarks and six additional light-flavor quarks or gluons are reported. The search employs a novel machine learning method for background estimation from control samples in data using decorrelated discriminators. The search is performed using events with 0, 1, or 2 electrons or muons in conjunction with at least six jets. No requirement is placed on the magnitude of the missing transverse momentum. The result is based on a sample of proton-proton collisions at $\sqrt{s}=13$ TeV corresponding to 138 fb −1 of integrated luminosity collected with the CMS detector at the LHC in 2016–2018. With no statistically significant excess of events observed beyond the expected contributions from the standard model, the data are used to determine upper limits on the top squark pair production cross section in the frameworks of RPV and stealth SUSY. Models with top squark masses less than 700 (930) GeV are excluded at 95% confidence level for RPV (stealth) SUSY scenarios.

Hadron-Hadron Scattering↗

SENTRA: A Modular Computational Graph Framework for Critical Mineral and Materials Supply Chains: Part I: Network Construction Latent-Quantity Estimation, and Temporal Graph Forecasting

Global supply chains for critical minerals and materials are complex, evolving networks of countries, products, production stages, and trade relationships. Existing analytical approaches are limited by fragmented data and static network representations that do not capture the dynamic production dependencies linking raw materials, intermediate products, and final goods across multiple countries. Trade and production statistics provide only a partial view of domestic production, inventories, and material flows, making it difficult to identify indirect sourcing pathways, hidden dependencies, and embedded foreign exposures. This paper introduces the Supply Chain Exposure Network Tracking and Risk Assessment (SENTRA) framework, a modular graph-based computational framework for constructing, analyzing, and forecasting dynamic supply chain networks. As the first paper in a three-part methodological series, it establishes the computational foundation of SENTRA by constructing a temporal attributed multi-relational graph whose nodes represent product–country pairs and whose edges encode observed trade and within-country value-chain relationships. Statistical estimation and constrained optimization recover latent production, final demand, and product input dependency coefficients while enforcing economic accounting constraints. Graph-derived exposure measures quantify direct, transshipment, value-chain, and multi-hop supply chain dependencies independently of the forecasting model. A temporal graph forecasting architecture based on a relational graph neural network then forecasts the evolution of the graph under mass-balance constraints with distribution-free conformal uncertainty quantification. Validation on the global aluminum supply chain shows that the learned graph representations recover economically meaningful supply chain structure, accurately forecast out-of-sample trade relationships, and produce well-calibrated prediction intervals. Subsequent papers apply this computational foundation to exposure assessment, disruption analysis, and scenario-based policy analysis, and extend the framework to multimaterial supply chain modeling and decision support.

36 MATERIALS SCIENCE↗

Moving beyond post hoc explainable artificial intelligence: a perspective paper on lessons learned from dynamical climate modeling

AI models are criticized as being black boxes, potentially subjecting climate science to greater uncertainty. Explainable artificial intelligence (XAI) has been proposed to probe AI models and increase trust. In this review and perspective paper, we suggest that, in addition to using XAI methods, AI researchers in climate science can learn from past successes in the development of physics-based dynamical climate models. Dynamical models are complex but have gained trust because their successes and failures can sometimes be attributed to specific components or sub-models, such as when model bias is explained by pointing to a particular parameterization. We propose three types of understanding as a basis to evaluate trust in dynamical and AI models alike: (1) instrumental understanding, which is obtained when a model has passed a functional test; (2) statistical understanding, obtained when researchers can make sense of the modeling results using statistical techniques to identify input–output relationships; and (3) component-level understanding, which refers to modelers' ability to point to specific model components or parts in the model architecture as the culprit for erratic model behaviors or as the crucial reason why the model functions well. We demonstrate how component-level understanding has been sought and achieved via climate model intercomparison projects over the past several decades. Such component-level understanding routinely leads to model improvements and may also serve as a template for thinking about AI-driven climate science. Currently, XAI methods can help explain the behaviors of AI models by focusing on the mapping between input and output, thereby increasing the statistical understanding of AI models. Yet, to further increase our understanding of AI models, we will have to build AI models that have interpretable components amenable to component-level understanding. We give recent examples from the AI climate science literature to highlight some recent, albeit limited, successes in achieving component-level understanding and thereby explaining model behavior. The merit of such interpretable AI models is that they serve as a stronger basis for trust in climate modeling and, by extension, downstream uses of climate model data.

54 ENVIRONMENTAL SCIENCES↗