Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Robust Design Under Uncertainty in Quantum Error Mitigation

Error mitigation techniques are crucial to achieving near-term quantum advantage. Classical postprocessing of quantum computation outcomes is a popular approach for error mitigation, which includes methods, such as zero noise extrapolation, virtual distillation, and learning-based error mitigation. However, these techniques have limitations due to the propagation of uncertainty resulting from the finite shot number of a quantum measurement. In this work, we introduce general and unbiased methods for quantifying the uncertainty and error of error-mitigated observables based on the strategic sampling of error mitigation outcomes. We then extend our approach to demonstrate the optimization of performance and robustness of error mitigation under uncertainty. To illustrate our methods, we apply them to zero noise extrapolation and Clifford date regression in the ground state of the XY model simulated using depolarizing and International Business Machines Corporation (IBM) Toronto noise models, respectively. In particular, we optimize the choice of noise levels and the allocation of shots for zero noise extrapolation and the distribution of the training circuits for Clifford data regression. While our methods are readily applicable to any postprocessing-based error mitigation approach, in practice they must not be prohibitively expensive—even though they perform optimizations of the error mitigation hyperparameters requiring sampling of a statistical distribution of error mitigation outcomes. By leveraging surrogate-based optimization, we show that our methods can efficiently perform optimal design for a zero noise extrapolation implementation. We then further demonstrate the transferability of learned zero noise extrapolation hyperparameters to other similar circuits.

97 MATHEMATICS AND COMPUTING↗

Evaluation of global terrestrial evapotranspiration using state-of-the-art approaches in remote sensing, machine learning and land surface modeling

Evapotranspiration (ET) is critical in linking global water, carbon and energy cycles. However, direct measurement of global terrestrial ET is not feasible. Here, we first reviewed the basic theory and state-of-the-art approaches for estimating global terrestrial ET, including remote-sensing-based physical models, machine-learning algorithms and land surface models (LSMs). We then utilized 4 remote-sensing-based physical models, 2 machine-learning algorithms and 14 LSMs to analyze the spatial and temporal variations in global terrestrial ET. The results showed that the ensemble means of annual global terrestrial ET estimated by these three categories of approaches agreed well, with values ranging from 589.6 mm/yr (6.56×10^4 cu.km/yr) to 617.1 mm/yr (6.87×10^4 cu.km/yr). For the period from 1982 to 2011, both the ensembles of remote-sensing-based physical models and machine-learning algorithms suggested increasing trends in global terrestrial ET (0.62 mm/sq.yr with a significance level of p<0.05 and 0.38 mm yr−2 with a significance level of p<0.05, respectively). In contrast, the ensemble mean of the LSMs showed no statistically significant change (0.23 mm/sq.yr, p>0.05), although many of the individual LSMs reproduced an increasing trend. Nevertheless, all 20 models used in this study showed that anthropogenic Earth greening had a positive role in increasing terrestrial ET. The concurrent small interannual variability, i.e., relative stability, found in all estimates of global terrestrial ET, suggests that a potential planetary boundary exists in regulating global terrestrial ET, with the value of this boundary being around 600 mm/yr. Uncertainties among approaches were identified in specific regions, particularly in the Amazon Basin and arid/semiarid regions. Improvements in parameterizing water stress and canopy dynamics, the utilization of new available satellite retrievals and deep-learning methods, and model–data fusion will advance our predictive understanding of global terrestrial ET.

surface modeling↗

Revisiting point defect thermodynamics in group IVB and VB transition metal carbides

We present a comprehensive re-examination of point defect thermodynamics in group IVB and VB transition metal carbides (TMCs) with the rocksalt structure using a combination of density functional theory (DFT) calculations and a statistical mechanical Wagner-Schottky model within the canonical ensemble. The most stable configurations of point defects were discovered using basin-hopping global optimization, driven by either a machine learning interatomic potential (MLIP) or DFT. A key finding is the identification of previously unreported dicarbon antisites—a C–C dimer occupying a metal site—as the structural (constitutional) defects on the carbon-rich side of stoichiometry in all group IVB and VB TMCs except TaC. Furthermore, dicarbon antisite-containing thermal defect complexes, such as quadruple and interbranch defects, can dominate in TMCs under specific stoichiometric and temperature conditions. In conclusion, by incorporating dicarbon antisites into the defect landscape, this work provides a revised understanding of the thermodynamics of point defects in TMCs.

Carbides↗

Data Agnostic Feature-Target Analysis & Ranking Machine Learning Pipeline (DAFTAR-ML) v0.1.0

DAFTAR-ML is a specialized machine-learning pipeline that identifies relevant features based on their relationship to a target variable. Many ML pipelines focus solely on prediction, and feature ranking is often absent or lacks robust statistical methods. DAFTAR-ML performs its tasks with this outcome in mind. Model training is robust, using nested cross-validation and hyperparameter tuning. Instead of relying on native feature-importance scores, it employs SHAP (SHapley Additive exPlanations) to quantify feature importance. The pipeline also produces comprehensive results, including publication-quality visualizations.

Melie, Tina [Lawrence Berkeley National Laboratory↗

Improving the Parameterization of Cloud and Rain Microphysics in E3SM using Novel Observationally-Constrained Bayesian Approach (Final Technical Report)

In this project, we sought to develop new cloud and rain microphysics frameworks within the Energy Exascale Earth System Model (E3SM). This work encompassed two primary avenues of research: 1) Further development of a Bayesian-based scheme called BOSS (Bayesian Observationally-constrained Statistical-physical Scheme) to represent cloud and rain microphysics, testing it in realistic high-resolution cloud models, and implementing it in E3SM; 2) Development of a methodology utilizing machine learning to enable computationally tractable use of tractable use of Markov chain Monte Carlo sampling for Bayesian parameter estimation in Earth system and cloud models. In this project, we adapted the BOSS microphysics scheme, originally formulated for rain-only, to include all liquid-phase microphysical processes for cloud and rain, in particular the processes that mediate between these two categories, for example the conversion from cloud to rain through collision and coalescence of drops. We constrained the scheme via comparison and testing against a detailed model that explicitly represents the evolution of cloud and rain particles, called a bin microphysics scheme.

54 ENVIRONMENTAL SCIENCES↗

Evaluation of fluxon synapse device based on superconducting loops for energy efficient neuromorphic computing

With Moore’s law nearing its end due to the physical scaling limitations of CMOS technology, alternative computing approaches have gained considerable attention as ways to improve computing performance. Here, we evaluate performance prospects of a new approach based on disordered superconducting loops with Josephson-junctions for energy efficient neuromorphic computing. Synaptic weights can be stored as internal trapped fluxon states of three superconducting loops connected with multiple Josephson-junctions (JJ) and modulated by input signals applied in the form of discrete fluxons (quantized flux) in a controlled manner. The stable trapped fluxon state directs the incoming flux through different pathways with the flow statistics representing different synaptic weights. We explore implementation of matrix–vector-multiplication (MVM) operations using arrays of these fluxon synapse devices. We investigate the energy efficiency of online-learning of MNIST dataset. Our results suggest that the fluxon synapse array can provide ~100× reduction in energy consumption compared to other state-of-the-art synaptic devices. This work presents a proof-of-concept that will pave the way for development of high-speed and highly energy efficient neuromorphic computing systems based on superconducting materials.

42 ENGINEERING↗

Methods for semi-automated indexing for high precision information retrieval

OBJECTIVE: To evaluate a new system, ISAID (Internet-based Semi-automated Indexing of Documents), and to generate textbook indexes that are more detailed and more useful to readers. DESIGN: Pilot evaluation: simple, nonrandomized trial comparing ISAID with manual indexing methods. Methods evaluation: randomized, cross-over trial comparing three versions of ISAID and usability survey. PARTICIPANTS: Pilot evaluation: two physicians. Methods evaluation: twelve physicians, each of whom used three different versions of the system for a total of 36 indexing sessions. MEASUREMENTS: Total index term tuples generated per document per minute (TPM), with and without adjustment for concordance with other subjects; inter-indexer consistency; ratings of the usability of the ISAID indexing system. RESULTS: Compared with manual methods, ISAID decreased indexing times greatly. Using three versions of ISAID, inter-indexer consistency ranged from 15% to 65% with a mean of 41%, 31%, and 40% for each of three documents. Subjects using the full version of ISAID were faster (average TPM: 5.6) and had higher rates of concordant index generation. There were substantial learning effects, despite our use of a training/run-in phase. Subjects using the full version of ISAID were much faster by the third indexing session (average TPM: 9.1). There was a statistically significant increase in three-subject concordant indexing rate using the full version of ISAID during the second indexing session (p < 0.05). SUMMARY: Users of the ISAID indexing system create complex, precise, and accurate indexing for full-text documents much faster than users of manual methods. Furthermore, the natural language processing methods that ISAID uses to suggest indexes contributes substantially to increased indexing speed and accuracy.

Evaluation Studies↗

CAHS: Context-Aware Homology Search

Protein homology search is foundational to bioinformatics: it supports annotation transfer, structure/function inference, and evolutionary analysis over rapidly expanding sequence repositories (e.g., UniProtKB). Profile hidden Markov models (pHMMs), as implemented in HMMER, remain the most widely trusted approach because they provide statistically calibrated E-values; however, their gap behavior is fixed once a profile is trained, despite biological evidence that insertion/deletion tolerance varies across flexible loops and intrinsically disordered regions. We present CAHS (Context-Aware Homology Search), a lightweight query-time adapter for pHMM search that incorporates learned and biologically motivated signals without changing HMMER's downstream search pipeline or its calibrated E-value reporting. Given a query sequence, CAHS computes per-residue representations from a protein language model and a disorder predictor, maps these to profile coordinates, and modulates only match-state transition rows (gap-open and gap-extension probabilities) while preserving Plan7 constraints. We comprehensively evaluate CAHS across six structurally diverse protein families and multi-domain architectures against a 570k-sequence target corpus. CAHS expands detection capability, retrieving thousands of additional remote homologs at relaxed thresholds by maintaining alignment quality through flexible regions. For multi-domain proteins, context-aware modulation resolves 94% of fragmented alignments. Crucially, CAHS preserves hit-set invariance at stringent operating points (E<10-10), demonstrating increased statistical confidence without inflating false positives. Furthermore, sharper statistical distinction between homologs and background noise during early filter stages yields up to a 3.87× acceleration in end-to-end wall-clock time on high-performance computing clusters. Overall, CAHS illustrates a practical AI-for-science design pattern: augmenting a trusted probabilistic model with query-specific learned signals to improve interpretable, reproducible inference in data-rich biology.

Bhattaram, Swethasree [Georgia Institute of Techno↗

Towards Validation of an Adaptive Flight Control Simulation Using Statistical Emulation

Traditional validation of flight control systems is based primarily upon empirical testing. Empirical testing is sufficient for simple systems in which a.) the behavior is approximately linear and b.) humans are in-the-loop and responsible for off-nominal flight regimes. A different possible concept of operation is to use adaptive flight control systems with online learning neural networks (OLNNs) in combination with a human pilot for off-nominal flight behavior (such as when a plane has been damaged). Validating these systems is difficult because the controller is changing during the flight in a nonlinear way, and because the pilot and the control system have the potential to co-adapt in adverse ways traditional empirical methods are unlikely to provide any guarantees in this case. Additionally, the time it takes to find unsafe regions within the flight envelope using empirical testing means that the time between adaptive controller design iterations is large. This paper describes a new concept for validating adaptive control systems using methods based on Bayesian statistics. This validation framework allows the analyst to build nonlinear models with modal behavior, and to have an uncertainty estimate for the difference between the behaviors of the model and system under test.

He, Yuning↗

Zentropy Theory for Transformative Functionalities of Magnetic and Superconducting Materials

The proposed research developed the zentropy theory through applications to complex magnetic materials and superconductors under the hypothesis that the emergent properties of complex magnetic materials and superconductors can be predicted by statistical mechanics of ergodic microstates with their partition functions computed from DFT-predicted free energies. The key objective is to develop approaches to systematically determine the types and number of microstates and the supercell size in DFT-based calculations through convergency of macroscopic functionalities, with the incorporation of our mixed-space approach accounting for the interactions between periodic supercells. In addition to use scientific intuitions to guide the design of important microstates, the key innovation of the proposed research is to integrate the domain knowledge and the material-property-descriptor database (MPDD) with 4 million microstates, which is supported by our deep neural network machine learning models (SIPFENN: structure-informed prediction of formation energy using neural networks) and integrated with our high throughput DFT Tool Kit (DFTTK). For complex magnetic materials, one of the objectives is to develop approaches to calculate short-range ordering from the statistical distribution of each microstate. For superconductors, the divergency of quasiparticle effective mass at a quantum critical point will be investigated, and the superconducting and non-superconducting microstates will be delineated through analysis of electronic band structure, density of states, charge density, and Fermi surface.

36 MATERIALS SCIENCE↗

Moment extraction using an unfolding protocol without binning

Deconvolving (“unfolding”) detector distortions is a critical step in the comparison of cross-section measurements with theoretical predictions in particle and nuclear physics. However, most existing approaches require histogram binning while many theoretical predictions are at the level of statistical moments. We develop a new approach to directly unfold distribution moments as a function of another observable without having to first discretize the data. Our moment unfolding technique uses machine learning and is inspired by Boltzmann weight factors and generative adversarial networks (GANs). We demonstrate the performance of this approach using jet substructure measurements in collider physics. With this illustrative example, we find that our moment unfolding protocol is more precise than bin-based approaches and is as or more precise than completely unbinned methods.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Towards a data-driven model of hadronization using normalizing flows

We introduce a model of hadronization based on invertible neural networks that faithfully reproduces a simplified version of the Lund string model for meson hadronization. Additionally, we introduce a new training method for normalizing flows, termed MAGIC, that improves the agreement between simulated and experimental distributions of high-level (macroscopic) observables by adjusting single-emission (microscopic) dynamics. Our results constitute an important step toward realizing a machine-learning based model of hadronization that utilizes experimental data during training. Finally, we demonstrate how a Bayesian extension to this normalizing-flow architecture can be used to provide analysis of statistical and modeling uncertainties on the generated observable distributions.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The DECam MAGIC Survey: Spectroscopic Follow-up of the Most Metal-poor Stars in the Distant Milky Way Halo *

In this work, we present high-resolution spectroscopic observations for six metal-poor stars with [Fe/H] < –3 (including one with [Fe/H] < –4), selected using narrowband Ca ii HK photometry from the DECam MAGIC Survey. The spectroscopic data confirm the accuracy of the photometric metallicities and allow for the determination of chemical abundances for 16 elements, from carbon to barium. The program stars have chemical abundances consistent with the [Fe/H] < –3 range. A kinematic/dynamical analysis suggests that all program stars belong to the distant Milky Way halo population (heliocentric distances 35 < d helio /kpc ≲ 55), including three with high-energy orbits that might have been associated with the Magellanic system and one, J0026−5445, having parameters consistent with being a member of the Sagittarius stream. The remaining two stars show kinematics consistent with the Gaia-Sausage/Enceladus dwarf galaxy merger. J0433−5548, with [Fe/H] = –4.12, is a carbon-enhanced ultra metal-poor star, with [C/Fe] = +1.73. This star is believed to be a bona fide second-generation star, and its chemical abundance pattern was compared with yields from metal-free supernova models. Results suggest that J0433−5548 could have been formed from a gas cloud enriched by a single supernova explosion from an ∼11 M ⊙ star in the early Universe. The successful identification of such objects demonstrates the reliability of photometric metallicity estimates, which can be used for target selection and statistical studies of faint targets in the Milky Way and its satellite population. These discoveries illustrate the power of measuring chemical abundances of metal-poor Milky Way halo stars to learn more about early galaxy formation and evolution.

CEMP stars↗

Sensing a Changing Chemical Mixture Using an Electronic Nose

A method of using an electronic nose to detect an airborne mixture of known chemical compounds and measure the temporally varying concentrations of the individual compounds is undergoing development. In a typical intended application, the method would be used to monitor the air in an inhabited space (e.g., the interior of a building) for the release of solvents, toxic fumes, and other compounds that are regarded as contaminants. At the present state of development, the method affords a capability for identifying and quantitating one or two compounds that are members of a set of some number (typically of the order of a dozen) known compounds. In principle, the method could be extended to enable monitoring of more than two compounds. An electronic nose consists of an array of sensors, typically made from polymer carbon composites, the electrical resistances of which change upon exposure to a variety of chemicals. By design, each sensor is unique in its responses to these chemicals: some or all of the sensitivities of a given sensor to the various vapors differ from the corresponding sensitivities of other sensors. In general, the responses of the sensors are nonlinear functions of the concentrations of the chemicals. Hence, mathematically, the monitoring problem is to solve the set of time-dependent nonlinear equations for the sensor responses to obtain the time dependent concentrations of individual compounds. In the present developmental method, successive approximations of the solution are generated by a learning algorithm based on independent-component analysis (ICA) an established information theoretic approach for transforming a vector of observed interdependent signals into a set of signals that are as nearly statistically independent as possible.

Duong, Tuan↗

Improving the SMAP Level-4 Soil Moisture Product

The NASA Soil Moisture Active Passive (SMAP) mission generates, among other data sets, the Level 4 Soil Moisture (L4 SM) product. The L4 SM algorithm is based on the assimilation of SMAP radiometer brightness temperature observations into the NASA Catchment land surface model using a spatially distributed ensemble Kalman filter. The L4 SM data are published with a mean latency of approx. 2.5 days from the time of observation and provide global, three-hourly, 9 km resolution estimates of surface and root-zone soil moisture and related land surface states and fluxes. In 2018, the product was upgraded from Version 3 to Version 4. Underlying the new version is a revised modeling system that includes improved input parameter datasets for land cover, topography, and vegetation height that are based on recent, high quality space-borne remote sensing observations. Land cover inputs were updated to the GlobCover2009 product, which is based on satellite observations from the Medium Resolution Imaging Spectrometer. Topographic statistics now rely on observations from the Shuttle Radar Topography Mission. Finally, vegetation height inputs are derived from space-borne Lidar measurements. Additionally, SMAP Level-2 soil moisture retrievals and in situ soil moisture measurements were used to calibrate a particular Catchment model parameter that governs the recharge of soil moisture from the models root-zone excess reservoir into the surface excess reservoir. Specifically, the replenishment of soil moisture near the surface from below under non-equilibrium conditions was substantially reduced, which brings the models surface soil moisture more in line with the SMAP Level 2 and in situ soil moisture. Finally, the calibration of the assimilated SMAP brightness temperatures changed substantially from Version 3 to Version 4. Considerable effort went into the version upgrade, creating an expectation that the new version is improved over the old version. Indeed, some aspects of the new version are clearly better. However, other aspects are not, and on balance, the overall improvement is modest at best. In this presentation we summarize the skill of the new and old versions vs. independent in situ measurements and in terms of data assimilation diagnostics, including, for example, the statistics of the (soil moisture) analysis increments and the observation minus forecast (brightness temperatures) residuals. We share our experience with trying to improve to the L4 SM product and the lessons learned from the effort.

Reichle, Rolf↗

Flight Mechanics Modeling and Simulation of the Earth Entry System

Introduction: The Mars Sample Return (MSR) Campaign being planned by NASA and ESA has the ambitious goal to return Mars samples back to Earth. This international collaboration had developed a concept of operations that included a ESA-designed Earth Return Orbiter (ERO) and NASA-designed Capture, Containment, and Return System (CCRS). The Earth Entry System (EES), consisting of a protective aeroshell that houses the samples as well as sample containment vessels, would conduct entry, descent, and landing (EDL) on a direct Earth trajectory. The EES would enter on a spin-stabilized ballistic trajectory with the goal to passively achieve aerodynamic stability throughout all regions of flight. The EDL sequence would end with the EES impacting the soft playa soil of the Utah Test and Training Range (UTTR). As of the submission of this abstract, the MSR campaign is undergoing a re-architecture leading to a pause in EES development. However, the novel approaches developed in flight mechanics modeling and simulation can significantly benefit the greater IPPW community in the development of Earth return vehicles. This paper will present the latest state of EES flight mechanics modeling and simulation. The paper will highlight the simulation architecture developed and key lessons learned from understanding of EDL trajectory sensitivities. Modeling and Simulation: Figure 1 provides a high-level concept of operations for the approach, entry, descent, and landing (AEDL) phase of the CCRS-portion of MSR. The objective of EES flight mechanics is to model and simulate the EES trajectory from ERO separation to ground impact at UTTR. A variety of flight mechanics simulation models were utilized to model both exo-atmopsheric and atmospheric portions of flight. 42, a 6-DOF simulation developed at Goddard Space Flight Center, is utilized for propagating the attitude of EES during exo-atmospheric flight. 42 allows for a variety of spin eject mechanism scenarios to be simulated for analysis. 10 minutes prior to entry, the 42 states are handed off to the EDL sims. The prime EDL sim utilized by EES is the Program to Optimize Simulated Trajectories II (POST2), a 6-DOF sim developed at Langley Research Center, and the independent verification and validation EDL sim utilized is DSENDS, a 6-DOF sim developed at Jet Propulsion Laboratory. Figure 2 provides a visualization of the flight mechanics simulation model flow through various points in the AEDL phase. Due to the existence of a variety of sim models, the EES flight mechanics team developed processes for data hand-off. These processes included the development of a centralized coordinate frame document, utilization of a single, centralized simulation input document for all sims to reference, and hand-off files containing both the technical data to be ingested by other flight mechanics sims as well as annotations of modeling assumptions utilized to generate the data. Figure~\ref{fig:post2simarchitecture} provides an overview of the POST2 sim architecture wherein POST2 ingests numerous subsystem models and input files. The dispersed state file generated by MONTE provides the position/velocity state of the trajectory while the 42 Handoff file provides the attitude. The aerodynamics database, delivered by the EES aeroscience team, is utilized to simulate the aerodynamic forces and moments experienced during EDL. A custom atmosphere model, developed by EES atmosphere team, is utilized to simulate the anticipated atmosphere environment around the region of Earth through which the EES trajectory flys. These inputs and subsystem models can be varied depending on the AEDL flight mechanics scenario being simulated. Monte Carlo simulations are utilized to generate statistical AEDL performance metrics in the form of scorecards and violin plots. Furthermore, outputs from the POST2 simulation are utilized for follow-on analyses including aerothermal and landing performance. \section{Flight Mechanics Lessons Learned} Though the EES flight mechanics team uncovered a variety of lessons learned through the analysis conducted to support CCRS through preliminary design review, this paper will highlight the most important lessons. A key AEDL performance goal is to ensure the landing footprint of EES remains on the UTTR south range. A common modeling strategy used in EDL analysis is One-Variable-At-a-Time (OVAT). OVAT analysis provides insight into the key drivers that affect AEDL performance metrics. Figure 3 shows the landing ellipses for single dispersion sources as compared to the baseline aggregate of all dispersions. The figure shows that atmosphere winds alone dominate the size of the footprint ellipse (note: EES does not use a parachute unlike previous Earth-return missions and is in wind-driven free fall for ~5min). The significance of the wind led the EES flight mechanics team to pursue the development of a Custom Atmosphere Model [4], in lieu of EarthGRAM [1], built on actual radiosonde wind measurements around the UTTR-region. This decision was driven by the realism in the generated footprint ellipses and lessons-learned from Stardust [5]. These findings will be invaluable for future Earth-return missions in providing an early understanding of the key drivers affecting footprint size and modeling considerations for which to account. Another lesson learned is tied to the AEDL performance goal of achieving passive stability throughout all regions of flight. It is well understood that blunt-body aeroshells are less stable as they transition from supersonic to subsonic. Eliminating a backshell does help improvestability; however, other phenomena such as roll-induced instability during terminal descent can still arise. The EES flight mechanics team developed stability metrics as tools to better understand the causes of and better predict the onset of dynamic instability. These tools were built upon analytical models developed by Jaffe [3] and Murphy [2]. The tools were shown to both be very accurate in correlation with actual unstable cases and useful in developing stability margin policies based on the vehicle design and simulation considerations (e.g. sphere-cone angle change, mass change, wind turbulence). These tools allowed for the current EES design to demonstrate the ability to achieve passive stability and can be an invaluable tool for consideration in the design of parachute-less Earth-return vehicles.

Rohan Deshmukh↗

Describing Point Defect Topology in 2D Energy Materials Through Computer Vision

Point defects such as vacancies and impurity atoms strongly impact the performance of 2D materials. Traditional efforts often rely on manual detection, a process that is time-intensive, prone to human error, and challenging to scale. Here we leverage machine learning (ML) methods to identify and quantify vacancies within 2D transition metal carbides (Ti3C2, MXenes), aiming to expedite detection while improving accuracy. MXenes exhibit valuable defect-defined electrochemical properties, but we currently lack statistical understanding of defect topology needed to fully harness these materials. Here we employ a convolutional neural network for semantic segmentation of experimental MXene images, opening an opportunity to conduct a rigorous statistical study on defect hierarchy while investigating local relaxation in the lattice. We show how the integration of ML can yield fundamental insight into point defects, providing a powerful tool that will play an increasingly crucial role in the future of materials science. ML is often not just a matter of straightforward application, and pretrained models proved ineffective in this case. Instead, we trained our own neural network (NN) and applied data augmentation techniques and fine-tuning to the training dataset. Since labeled microscopy data is often scarce, we developed training data from a previously published wide-frame MXene image, using customized Gaussian fitting to locate atomic positions. Our trained model was then applied to a large dataset of experimental images, enabling a statistical study of defect configurations across three samples prepared with different HF etchant concentrations (5%, 9.1%, and 12.5%), as shown in Fig. 1. This also allowed us to investigate local strain around vacancies, though we find that we are limited by the precision of measurements using high-angle annular dark field (HAADF) images, as shown in Fig. 2. This study demonstrates how ML enables large-scale, quantitative analysis of atomic defects - an otherwise infeasible task with traditional methods. While our NN was specialized for Ti3C2 MXenes, the pipeline we developed provides a foundation for future ML models tailored to other materials. Ultimately, we envision embedding the NN onto the microscope to give real-time feedback to the user. To make this a reality, continued work is necessary to fully understand the NN's capabilities and limitations. This study gets one step closer to our goals of automated experimentation moving away from traditional methods of manual labeling. As ML capabilities advance, we hope to continue adapting and applying these techniques in microscopy.

2D materials↗

Dynamic data-driven multiscale modeling for predicting the degradation of a 316L stainless steel nuclear cladding material

Here, we have developed a long short-term memory stacked ensemble (LSTM-SE) surrogate modeling approach that can provide rapid predictions of microstructural evolution and the resultant mechanical properties of American Iron and Steel Institute (AISI) 316L series stainless steel (316LSS) fuel cladding under conditions of varying temperature and radiation dose rate. To acquire training data, we developed and implemented a kinetic Monte Carlo (KMC) model to simulate precipitation kinetics of M 23 C 6 , γ', and G phases within SS316L cladding. Experimentally reported precipitation kinetics of SS316L in literature were linked to the kinetic parameters of the simulated precipitation in our KMC model. The model was then used to simulate microstructure evolution under synthetically generated treatments of varying temperature and radiation dose rate, for periods of up to 3000 hours. Changes in volume fraction, number density, and particle size of precipitates were recorded, and particle area fractions were correlated using statistical methods to develop the surrogate model. Simultaneously, the mechanical properties of the simulated microstructures were evaluated using microstructure-based finite element method (FEM) analysis to determine the elastic modulus, yield stress, ultimate tensile strength, and elongation to failure of the aged microstructures. Using this approach, our surrogate model can predict precipitation behavior within 0.25% volume fraction and mechanical properties within 6% relative error from the values predicted by the KMC and FEM models using 50 training simulations as input. The trained recurrent neural network-based model can return estimations of precipitation kinetics and mechanical properties ~1000 times faster than the physics-based codes. This work demonstrates, as a proof of concept, that reactor material service lifetimes under variable service conditions can be predicted for a statistics-based model from a practicably obtainable dataset.

36 MATERIALS SCIENCE↗