Search NASA⌕ Search

SEARCH · Search NASA

Results for “Classifications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

LABQ3: Bayesian method for quantification of mineral compositions and nano-scale elemental mapping of 3D synchrotron XCT data

Quantitative analysis of mineral compositions is essential in understanding geochemical, mineralogical and environmental processes. Fine-resolution 3D imaging is widely done using synchrotron X-ray computed tomography (XCT), but existing analyses are limited to visualization and segmentation. This paper presents a new method, Linear Attenuation Bayesian Quantitative 3D-mapper (LABQ3), based on the linearity of X-ray attenuation with respect to elemental concentrations. To address the random variability in attenuation measurements, LABQ3 employs Bayesian decision theory to minimize classification error, using reference attenuation distributions from scans of pure mineral standards. To demonstrate LABQ3 and test its performance, we studied precipitated carbonate samples. XCT scans were done at multiple energies using the transmission X-ray microscope (TXM) at beamline 32-ID-C of the Advanced Photon Source at Argonne National Laboratory. The reconstructed 3D images have a voxel size of 20 nm. Analyses revealed rich nano-scale compositional heterogeneity within individual particles. A mixture of calcium and cadmium produced an overall stoichiometric composition of (Ca 0.78 ,Cd 0.22 )CO 3 , with some voxels containing nearly pure CdCO 3 . The addition of zinc led to an overall stoichiometric composition of 33% Ca, 28% Cd, 39% Zn, with a nearly pure CaCO 3 core and compositional zonation through the rim. These compositional gradients are related to temporal sequences of carbonate mineral formation where Cd precipitated at the beginning in (Ca,Cd)CO 3 , while Cd and Zn precipitated at the end in (Ca, Cd,Zn)CO 3 . Results differ from bulk analyses using Inductively Coupled Plasma-Mass Spectrometry (ICP-MS), showing that LABQ3 provides particle-specific insights. LABQ3 distinguishes itself by quantifying chemical compositions along a continuum, making it different from XCT analyses based on segmentation. LABQ3 allows simultaneous acquisition of morphology and chemical composition in 3D, facilitating the interpretation of chemical gradients of trace elements, quantification of solid solution compositions, inferences about temporal sequences of mineral precipitation, and addressing other concerns about solid-phase chemistry.

58 GEOSCIENCES↗

A high-throughput workflow to analyze sequence-conformation relationships and explore hydrophobic patterning in disordered peptoids

Understanding how a macromolecule’s primary sequence governs its conformational landscape is crucial for elucidating its function, yet these design principles are still emerging for macromolecules with intrinsic disorder. Herein, we introduce a high-throughput workflow that implements a practical colorimetric conformational assay, introduces a semi-automated sequencing protocol using matrix-assisted laser desorption/ionization and tandem mass spectrometry (MALDI-MS/MS), and develops a generalizable sequence-structure algorithm. Using a model system of 20mer peptidomimetics containing polar glycine and hydrophobic N-butylglycine residues, we identified nine classifications of conformational disorder and isolated 122 unique sequences across varied compositions and conformations. Conformational distributions of three compositionally identical library sequences were corroborated through atomistic simulations and ion mobility spectrometry coupled with liquid chromatography. A data-driven strategy was developed using existing sequence variables and data-derived “motifs” to inform a machine-learning algorithm toward conformation prediction. Here, this multifaceted approach enhances our understanding of sequence-conformation relationships and offers a powerful tool for accelerating the discovery of materials with conformational control.

data-driven analysis↗

An adaptive and stability-promoting layerwise training approach for sparse deep neural network architecture

This work presents a two-stage adaptive framework for progressively developing deep neural network (DNN) architectures that generalize well for a given training data set. In the first stage, a layerwise training approach is adopted where a new layer is added each time and trained independently by freezing parameters in the previous layers. We impose desirable structures on the DNN by employing manifold regularization, sparsity regularization, and physics-informed terms. We introduce a ε – δ – stability-promoting concept as a desirable property for a learning algorithm and show that employing manifold regularization yields a ε – δ stability-promoting algorithm. Further, we also derive the necessary conditions for the trainability of a newly added layer and investigate the training saturation problem. In the second stage of the algorithm (post-processing), a sequence of shallow networks is employed to extract information from the residual produced in the first stage, thereby improving the prediction accuracy. Numerical investigations on prototype regression and classification problems demonstrate that the proposed approach can outperform fully connected DNNs of the same size. Moreover, by equipping the physics-informed neural network (PINN) with the proposed adaptive architecture strategy to solve partial differential equations, we numerically show that adaptive PINNs not only are superior to standard PINNs but also produce interpretable hidden layers with provable stability. As a result, we also apply our architecture design strategy to solve inverse problems governed by elliptic partial differential equations.

42 ENGINEERING↗

NanoPSD: A software for automatic detection of Nano-Particle Shape Distribution in electron microscopy images

Accurate quantification of the size and morphology of nanoparticles from electron microscopy (EM) images is essential to understand growth mechanisms, surface reactivity, and functional behavior in nanoscale materials. Manual analysis remains slow, subjective, and difficult to reproduce in large datasets. We introduce NanoPSD (Nano-Particle Shape Distribution), an open-source and fully automated framework for quantitative particle detection and morphology analysis from EM images. NanoPSD integrates adaptive contrast enhancement, polarity-agnostic scale-bar detection, Optical Character Recognition (OCR)-based calibration, and classical segmentation via Otsu thresholding with morphological refinement. Particle contours are used to extract geometric descriptors, including equivalent circular diameter, aspect ratio, circularity, and solidity, enabling automated classification into spherical, rod-like, and aggregate morphologies. The framework supports both single-image and batch processing, generating publication-quality visualizations, LaTeX-ready tables, and structured comma-separated values (CSV) datasets. As a demonstration, we applied NanoPSD to plasma-synthesized nanoparticle samples diagnosed via transmission electron microscopy (TEM). The code produced statistically robust size and morphology distributions spanning a few to tens of nanometers with minimal user supervision. The pipeline demonstrates high reproducibility and scalability, processing large image collections with consistent calibration and output formatting. Its modular design enables seamless integration of future deep-learning-based segmentation models, providing a pathway toward intelligent, data-driven electron microscopy analysis.

36 MATERIALS SCIENCE↗

Microstructure prediction for Ti-22Al-25Nb in laser powder bed fusion

This work presents a physics-informed framework for predicting solidification morphology and defect susceptibility in additively manufactured Ti–22Al–25Nb across a broad processing space. The framework integrates solidification microstructure selection (SMS) analysis with a single-track defect-based printability map to establish a unified methodology linking processing parameters to both interfacial morphology and manufacturability. Thermal gradients G and solidification rates R are first computed using the Thermo-Calc Additive Manufacturing (TC-AM) module, a finite-interface-dissipation (FID) phase-field (PF) model coupled with CALPHAD method is then employed to systematically distinguish planar and dendritic regimes as functions of $G$ and $R$. By superimposing the printability map onto the morphology projections, a comprehensive process–structure framework is obtained. Across most processing conditions, the predicted microstructure is predominantly dendritic, while planar growth emerges only under selected laser power $P$ and scan speed $v$ combinations. In addition to morphology classification, the framework quantifies the dendritic area fraction and introduces a width-based morphology descriptor to characterize the spatial extent of planar/dendritic regions within the melt pool. It provides mechanistic insight into the interplay between solidification physics and defect formation, offering practical guidance for parameter selection and microstructural control in Ti–22Al–25Nb additive manufacturing (AM).

36 MATERIALS SCIENCE↗

PlasmoData.jl — A Julia framework for modeling and analyzing complex data as graphs

Datasets encountered in scientific and engineering applications appear in complex formats (e.g., images, multivariate time series, molecules, video, text strings, networks). Graph theory provides a unifying framework to model such datasets and enables the use of powerful tools that can help analyze, visualize, and extract value from data. In this work, we present PlasmoData.jl, an open-source, Julia framework that uses concepts of graph theory to facilitate the modeling and analysis of complex datasets. The core of our framework is a general data modeling abstraction, which we call a DataGraph. We show how the abstraction and software implementation can be used to represent diverse data objects as graphs and to enable the use of tools from topology, graph theory, and machine learning (e.g., graph neural networks) to conduct a variety of tasks. We illustrate the versatility of the framework by using real datasets: (i) an image classification problem using topological data analysis to extract features from the graph model to train machine learning models; (ii) a disease outbreak problem where we model multivariate time series as graphs to detect abnormal events; and (iii) a technology pathway analysis problem where we highlight how we can use graphs to navigate connectivity. Further, our discussion also highlights how PlasmoData.jl leverages native Julia capabilities to enable compact syntax, scalable computations, and interfaces with diverse packages. Overall, we show that the DataGraph abstraction and PlasmoData.jl Julia package are able to model data within graphs and enable useful analysis.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Maximizing dynamic range and performance of anatase TiO 2 ECRAM through structure and programming

Here, in this study, we investigate the structure-dependent modulation characteristics of all-solid-state three-terminal electrochemical random-access memory (ECRAM) based on an anatase Li x TiO 2 channel. By directly comparing “asymmetric” and “symmetric” ECRAM device architectures, we reveal significant insight into the impact of a non-zero gate-drain open-circuit voltage and its influence on voltage vs. current-controlled gating. We also explore the impact of potentiation/depression write parameters on the symmetry, linearity, and dynamic range of the device response. Together, initial results from optimizing structure and programming approaches yielded unprecedented G max /G min ratios of >1,000 for ECRAM and hundreds of tunable memory states with excellent linearity and symmetry. Simulations based on these ECRAM devices further illustrate the promise of this analog memory technology, achieving near 2% classification error in the MNIST digit recognition benchmark for a range of training parameters compared to a theoretical best of 1.66% and outperforming other device models extracted from the literature.

AIHWKit↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

Defining an industrial heat pump: A review & synthesis

Clear and understandable definitions of advanced energy technologies are not common and often confuse those who are not experts. This review paper addresses a critical issue in the field of heat pumps: the ambiguity surrounding the term “industrial heat pump.” The term, as discussed in various literature and media, often leads to misconceptions about the capabilities and applications of heat pumps, particularly in industrial contexts. This review paper seeks to refine and unify the definition of industrial heat pumps by analyzing existing academic and commercial literature, industry reports, and classifications. Unique identifiers explored for the definition of an industrial heat pump include the use of waste heat, operating temperature ranges, size specifications, and specific end-use applications. The study reveals differences in how industrial heat pumps are defined in academic and commercial contexts, highlighting the need for a standardized definition to facilitate better communication among manufacturers, policymakers, and end users. This is especially important considering recent legislative efforts such as the HEAT Act, which seeks to promote industrial heat pump adoption. The proposed definition aims to provide clarity while supporting the effective deployment of industrial heat pumps. With proper adoption, the definition will help promote the role of industrial heat pumps in achieving energy efficiency goals in the industrial sector. The intended impact is for the definition to serve as a foundational reference for future research and policy development.

Heat pump characteristics↗

A spatiotemporally explicit and scalable indicator of intact lands across the conterminous United States, 1986–2023

Globally, ecologically intact areas are increasingly scarce. Agricultural expansion into previously uncultivated areas drives the loss of intact lands that might otherwise exhibit high levels of ecological integrity. Thus, the absence of cultivation can be an indicator of intact lands as measured from remote sensing data and thematic maps. Our objective for this study was to develop and compare tractable approaches based on remotely sensed satellite data to map spatial patterns of potentially intact lands across the conterminous U.S. (CONUS). Using annual cultivation probabilities derived from satellite observations, we classified and mapped potentially intact lands across CONUS from 1986 to 2023 at 30 m resolution. We created three maps, first by applying a constant cultivation probability threshold across CONUS, second by varying the threshold state-by-state to maximize state-level overall accuracies, and third by equalizing the state-level user's and producer's accuracies to minimize classification bias. Validation against 800,000+ independent ground samples resulted in CONUS-level overall accuracies ≥85% for the roughly 660 million ha of potentially intact land. Map accuracy varied with the proportion of potentially intact lands across regions, with the Pacific-Mountain and Great Plains regions exhibiting the highest accuracies, while Eastern CONUS exhibited a greater mix of potentially intact and non-intact lands and more moderate map accuracies. These novel maps and approaches can be adapted to different spatiotemporal extents to support conservation and production decisions ranging from species and ecosystems protection to reducing land conversion and climate mitigation.

agriculture↗

Analysis of US Industrial Assessment Centers (IACs) implementation

Industrial energy assessments are a fundamental action toward developing a decarbonization strategy at any level. They provide an understanding of where energy efficiency opportunities exist and help to make informed business decisions about the costs and benefits of implementing sustainable policies and practices. This paper examines the effectiveness of the Department of Energy's Industrial Assessment Centers Program at providing useful energy-efficiency audits as well as the barriers faced by the program and plans for future growth and improvement. This paper presents an analysis of the program between 1981 and 2022, covering 20,290 industrial assessments and 151,198 recommendations, with $2.6 billion of recommended savings. The analysis includes a breakdown of the IAC recommendations and implemented projects based on both the industrial subsectors according to the Standard Industrial Classification code, energy type, and systems evaluated. The results include a 47 % implementation rate, a gap analysis to understand the missed opportunities, and a discussion about reasons for recommendation rejection. A comparison of the IAC Program with other country-level energy auditing or assessment programs was conducted, and suggestions for improving implementation rates were mentioned.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Risk of Mortality in Family Members of Men Seeking Fertility Assessment

Objective: To assess mortality in family members of men seeking fertility assessment. Subfertility serves as a biomarker for overall somatic health, and poor semen quality is associated with increased risk of hospitalization and mortality from chronic conditions. However, it is unclear if these risks extend to family members of men with low sperm count. Design: Retrospective cohort study. Subjects: Family members, up to third-degree relatives, of men in the Subfertility, Health and Assisted Reproduction and the Environment cohort who underwent a semen analysis as part of a fertility assessment 1996–2017. Relatives of men with a recorded total sperm count who lived in Utah for ≥1 year 1904–2017 were included in the analysis (N = 22,280 families). Exposure: Individuals were classified by family membership. Families were classified as relatives of azoospermic (0M), oligozoospermic (<39M), or normozoospermic (≥39M) men. The average total sperm count of the proband (male relative) with fertility assessment was also included as a continuous exposure measure. Main Outcome Measures: The main outcomes were all-cause and cause-specific mortality risk by sex, age, and degree of relation: first-, second-, and third-degree. Cox proportional hazard models were used to test the association between fertility classification and mortality, controlling for sex, race/ethnicity, and birth year. Results: A total of 666,437 relatives of men with fertility assessment (N deaths = 183,974) were included in the analysis. Relative to normozoospermia families, all-cause mortality risk increased in oligozoospermia families (hazard ratio [HR] oligozoospermia , 1.03; 95% confidence interval [CI], 1.01–1.05). Close relatives, first- (HR oligozoospermia , 1.17; 95% CI, 1.07–1.28) and second-degree relatives (HR azoospermia , 1.11; 95% CI, 1.04–1.20; HR oligozoospermia , 1.05; 95% CI,1.01–1.09), of azoospermic and oligozoospermic men had the highest all-cause and cause-specific mortality risk, including death attributed to cardiovascular disease or congenital birth conditions. Conclusion: Our results suggest that familial all-cause and cause-specific mortality risk differ by fertility phenotype. Families of azoospermic and oligozoospermic men showed significantly increased risk, particularly for close relatives. This study provides further evidence that shared genetic and/or environmental factors could influence both fertility and somatic health.

Male fertility↗

Efficient distributed continual learning for steering experiments in real-time

Deep learning has emerged as a powerful method for extracting valuable information from large volumes of data. However, when new training data arrives continuously (i.e., is not fully available from the beginning), incremental training suffers from catastrophic forgetting (i.e., new patterns are reinforced at the expense of previously acquired knowledge). Training from scratch each time new training data becomes available would result in extremely long training times and massive data accumulation. Rehearsal-based continual learning has shown promise for addressing the catastrophic forgetting challenge, but research to date has not addressed performance and scalability. To fill this gap, we propose an approach based on a distributed rehearsal buffer that efficiently complements data-parallel training on multiple GPUs to achieve high accuracy, short runtime, and scalability. It leverages a set of buffers (local to each GPU) and uses several asynchronous techniques for updating these local buffers in an embarrassingly parallel fashion, all while handling the communication overheads necessary to augment input minibatches using unbiased, global sampling. We further propose a generalization of rehearsal buffers to support both classification and generative learning tasks, as well as more advanced rehearsal strategies (notably Dark Experience Replay, leveraging knowledge distillation). We illustrate this approach with a real-life HPC streaming application from the domain of ptychographic image reconstruction. Furthermore, we run extensive experiments on up to 128 GPUs of the ThetaGPU supercomputer to compare our approach with baselines representative of training-from-scratch (the upper bound in terms of accuracy) and incremental training (the lower bound). Results show that rehearsal-based continual learning achieves a top-5 validation accuracy close to the upper bound, while simultaneously exhibiting a runtime close to the lower bound.

Asynchronous data management↗

A terminology for scientific workflow systems

The term “scientific workflow” has evolved over the last two decades to encompass a broad range of compositions of interdependent compute tasks and data movements. It has also become an umbrella term for processing in modern scientific applications. Today, many scientific applications can be considered as workflows made of multiple dependent steps, and hundreds of workflow systems have been developed to manage and run these scientific workflows. However, no turnkey solution has emerged from the field to address the diversity of scientific processes and the infrastructure on which they are supposed to be implemented. Instead, new research problems requiring the execution of scientific workflows with some novel feature often lead to the development of an entirely new workflow system. A direct consequence of this situation is that many existing workflow management systems (WMSs) share some salient features, offer similar functionalities, and can manage the same categories of workflows but at the same time also have some distinct capabilities that can be important for specific applications. This situation makes researchers who develop workflows face the complex question of selecting a WMS. This selection can be driven by technical considerations, to find the system that is the most appropriate for their application and for the computing and storage resources available to them, or other factors such as reputation, adoption, strong community support, or long-term sustainability. To address this problem, a group of WMS developers and practitioners joined their efforts to produce a community-based terminology of WMSs. This paper summarizes their findings and introduces this new terminology to characterize WMSs. Furthermore, this terminology is composed of fives axes: workflow structure and characteristics, composition, orchestration, data management, and metadata capture. Each axis comprises several concepts that capture the prominent features of WMSs. Based on this terminology, this paper also presents a classification of 23 existing WMSs according to the proposed axes and terms.

Community-based terminology↗

Simulating water dynamics related to pedogenesis across space and time: Implications for four-dimensional digital soil mapping

Digital soil mapping (DSM) relies on machine-learning and geostatistics to represent soil property observations across space. DSM techniques are powerful but often empirical, being limited to the quality and density of point samples. Water dynamics are closely related to soil variability, and the physics that govern water movement are well known. Hydrological properties can hence be simulated by physical models through space and time, unveiling key characteristics about soils. We propose the use of hydrologic models to map soils across the surface (2D), depth (1D), and time (1D)–which provides a 4D approach to digital soil mapping (4DSM). The Distributed Hydrology Soil Vegetation Model (DHSVM) was applied to a watershed currently under pasture. Moisture sensors and wells were installed at different depths in the watershed on summit, sideslope and toeslope positions to validate the model. DHSVM simulations of soil moisture distribution and depth to saturation were performed during the hydrological year (October 2008-September 2009). Clusters of similar pixels based on soil moisture values were determined using Dynamic Time Warping (DTW) to align temporal data and K-means. Clustering was performed both seasonally and for the entire year. Temporal patterns simulated by DHSVM matched measurements given by moisture sensors and wells. Seasonal clusters differed from the annual cluster. Distinct clusters were observed for each season and with depth, showing that spatiotemporal soil variability is lost when statically assessing soils. Spatiotemporal clusters corroborated field observations of fragipan occurrence not explicitly spatially mapped by Soil Survey Geographic Database (SSURGO). If a connection can be made between water and soils, static and dynamic soil variability can be predicted using physically based hydrologic models. Hydrologic models can benefit soil mapping by enabling reliable 4D simulation of water dynamics, which are fundamental to soil variability and soil classification and directly relate to biological, physical and chemical soil processes not captured by typical soil sampling protocols.

54 ENVIRONMENTAL SCIENCES↗

Deep-learning-enhanced assessment of wellbore barrier effectiveness in geologic storage systems with intermediate aquifers

For geologic systems where carbon dioxide (CO 2 ) is injected underground, existing wells represent potential pathways for fluid migration. Here, this study introduces a novel deep learning model to quantify the likelihood and potential magnitude of fluid migration through wellbores at sites with intermediate aquifers or thief zones between the injection units and underground drinking water sources. Synthetic datasets, generated using reservoir simulations, captured a wide range of subsurface conditions, well attributes, operational parameters, and fluid migration scenarios. Among the regression models developed to predict brine and CO 2 leakage rates and CO 2 saturations along leaky wellbores, convolutional neural network (CNN) outperformed both Light Gradient Boosting Machine and deep neural network. Additionally, a CNN-based classification model was created to predict whether brine and CO 2 would leak along a wellbore, further improving performance over regression alone. The best models were integrated into the National Risk Assessment Partnership Open-source Integrated Assessment Model for rapid, stochastic assessment of storage system containment and leakage risks. A case study demonstrated the model’s ability to simulate fluid migration through existing wells with multiple intermediate aquifers. This computationally efficient wellbore model offers value in support of site performance evaluation and risk-informed decision making by stakeholders.

CO2 leakage↗

Deep learning model for fast, science-based forecasting of fluid migration along faults in geologic carbon storage scenarios

Effective long-term geologic storage depends on robust site selection and credible, science-based forecasting of subsurface behavior to ensure storage integrity. For this work, we develop a deep learning–based reduced-order model (ROM) to quantify potential carbon dioxide (CO₂) and brine migration through geological faults. The ROM combines a Transformer model for binary classification and a Stacked Ensemble for regression, trained on a comprehensive dataset generated from 1400 physics-based reservoir simulations. Key geologic and operational parameters—including fault geometry, reservoir structure, and injection conditions—were systematically varied to capture a wide range of fluid migration scenarios. The ROM accurately predicts the onset of migration, cumulative migration volumes of both CO₂ and brine, and associated migration rates, as compared to an independent set of validation simulations, while significantly reducing computational cost compared to traditional simulation methods. Model performance was evaluated across diverse fault configurations, revealing that shallow reservoir geometry and fault angle are among the most influential factors governing migration behavior. Sensitivity analysis using SHapley Additive exPlanations (SHAP) provided interpretability, revealing distinct patterns in how geological and operational features drive transient versus cumulative migration outcomes. The ROM’s ability to rapidly simulate fault migration scenarios enables efficient sensitivity analyses, scenario evaluations, and decision support for site selection and monitoring design. This approach enhances the safety, scalability, and long-term operational performance of geologic carbon storage (GCS) systems by providing a robust, interpretable tool for predicting subsurface fluid migration and assessing fault-related migration potential.

42 ENGINEERING↗

Predicting U 3 O 8 powder processing conditions: An AI/ML approach analyzing deep learning embeddings of SEM micrographs

High-resolution SEM images of uranium-oxide powders encode micro- and nanoscale clues to their synthesis route and calcination temperature. We trained a ResNet-50 model on 11 commercial-scale U₃O₈ classes, ammonium diuranate (ADU) or uranyl peroxide (H₂O₂) precursors calcined at temperatures ranging from 400 to 750 °C and added a 256-D projection head before the classifier to analyze the learned representation. The best of eight seeds reached 92.4 % accuracy on reserved testing data, but our focus is the structure of the embedding space rather than the accuracy and labels. We quantify class relatedness in the original 256-D space using centroid similarity and distributional distances, and we use Uniform Manifold Approximation Projection (UMAP) for visualization. ‘Unknown’ images from different preparation methods, SEM operators, and from the literature localized near the expected classes under a nearest-centroid analysis without retraining, as well as clustered in similar UMAP space. In conclusion, this embedding-centered workflow complements black-box classification by providing quantitative, similarity-based comparisons of U₃O₈ morphologies and reduces storage space by up to 98 % for image data used in millisecond vector search comparisons.

36 MATERIALS SCIENCE↗