Search NASASearch

SEARCH · Search NASA

Results for “feature importance analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Plasma confinement state classification via FPP relevant microwave diagnostics

We present a parsimonious and robust machine learning approach for identifying plasma confinement states in fusion power plants (FPPs) where reliable identification of the low-confinement and high-confinement regimes is critical for safe and efficient operation. Unlike research-oriented devices, FPPs must operate with a severely constrained set of diagnostics. To address this challenge, we demonstrate that a minimalist model, using only electron cyclotron emission (ECE) signals, can achieve accurate and reliable state classification. ECE provides electron temperature profiles without the engineering or survivability issues of in-vessel probes, making it a primary candidate for FPP-relevant diagnostics. Our framework employs ECE as input, extracts features using radial basis functions, and applies a gradient boosting classifier, achieving a test accuracy of 96% (correct predictions). Robustness analysis and feature importance analyzes confirm the approach’s reliability. These results demonstrate that state-of-the-art performance is attainable from a restricted diagnostic set, paving the way for minimalist yet resilient plasma control architectures for FPPs.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

IDENTIFICATION OF POTENTIAL SUPERCONDUCTOR QUENCH PRECURSORS USING FREQUENCY DOMAIN FEATURE ANALYSIS

Superconducting magnets are important pieces of technology in the world of particle accelerators, allowing researchers to study atomic and subatomic phenomena, among other things. In some instances, superconductors can lose this non-resistive property in a phenomenon known as quenching, which can cause damage to the magnets. This potential danger prompts the introduction of systems to predict when a quench is imminent; one such implementation is through the use of acoustic sensors that detect vibrations within the magnet. Within these acoustic sensor signals, significantly above-noise disturbances (referred to as ”events”) can be identified. Our research applies the statistical framework of a permutation test to features calculated from the power spectral density (PSD) to distinguish between events far from the quench at the end of the signal to events at the start of the signal. We found that dividing the PSD into frequency bands produced a feature capable of distinguishing between events early in the signal and late in the signal leading up the quench, providing a promising starting place for future quench prediction systems.

Roehrig, Benjamin [Northern Illinois U.]

Interpretable machine learning-guided design of Fe-based soft magnetic alloys

Here, we present a machine learning (ML) guided approach to predict saturation magnetization (𝑀 S ) and coercivity (𝐻 C ) in Fe-rich soft magnetic alloys, particularly Fe-Si-B systems. ML models trained on experimental data reveal that increasing Si and B content reduces 𝑀 S from 1.81 T (DFT ≈ 2.04 T) to ≈1.54 T (DFT ≈ 1.56T) in Fe-Si-B, which is attributed to decreased magnetic density and structural modifications. Experimental validation of ML predicted magnetic saturation on Fe-1Si-1B (2.09 T), Fe-5Si-5B (2.01 T), and Fe-10Si-10B (1.54 T) alloy compositions further supports our findings. These trends are consistent with density functional theory predictions, which link increased electronic disorder and band broadening to lower 𝑀 S values. Experimental validation on selected alloys confirms the predictive accuracy of the ML model, with good agreement across compositions. Beyond predictive accuracy, detailed uncertainty quantification and model interpretability including through feature importance and partial dependence analysis reveal that 𝑀 S is governed by a nonlinear interplay between Fe content and early transition metal ratios, while 𝐻 C is more sensitive to processing conditions such as ribbon thickness and thermal treatment windows. The ML framework was further applied to Fe-Si-B/Cr/Cu/Zr/Nb alloys in a pseudoquaternary compositional space, which shows comparable magnetic properties to NANOMET (Fe 84.8 ⁢Si 0.5 ⁢B 9.4 ⁢Cu 0.8⁢ P 3.5 ⁢C 1 ), FINEMET (Fe 73.5 ⁢Si 13.5 ⁢B 9 Cu 1 ⁢Nb 3 ), NANOPERM (Fe 88 ⁢Zr 7⁢ B 4 ⁢Cu 1 ), and HITPERM (Fe 44 ⁢Co 44 ⁢Zr 7⁢ B 4 ⁢Cu 1 . Our findings demonstrate the potential of the ML framework for accelerated search of high-performance soft magnetic materials.

density functional theory

On the minimum number of radiation field parameters to specify gas cooling and heating functions

Fast and accurate approximations of gas cooling and heating functions are needed for hydrodynamic galaxy simulations. We use machine learning to analyze atomic gas cooling and heating functions in the presence of a generalized incident local radiation field computed by Cloudy. We characterize the radiation field through binned radiation field intensities instead of the photoionization rates used in our previous work. We find a set of 6 energy bins whose intensities exhibit relatively low correlation. We use these bins as features to train machine learning models to predict Cloudy cooling and heating functions at fixed metallicity. We compare the relative SHapley Additive exPlanation (SHAP) value importance of the features. From the SHAP analysis, we identify a feature subset of 3 energy bins (0.5-1, 1-4, and 13-16Ry) with the largest importance and train additional models on this subset. We compare the mean squared errors and distribution of errors on both the entire training data table and a randomly selected 20% test set withheld from model training. The machine learning models trained with 3 and 6 bins, as well as 3 and 4 photoionization rates, have comparable accuracy everywhere, with errors ≳10 times smaller than for the interpolation table of Gnedin and Hollon (2012). We conclude that 3 energy bins (or 3 analogous photoionization rates: molecular hydrogen photodissociation, neutral hydrogen HI, and fully ionized carbon CVI) are sufficient to characterize the dependence of the gas cooling and heating functions on our assumed incident radiation field model.

79 ASTRONOMY AND ASTROPHYSICS

On the minimum number of radiation field parameters to specify gas cooling and heating functions

Fast and accurate approximations of gas cooling and heating functions are needed for hydrodynamic galaxy simulations. We use machine learning to analyze atomic gas cooling and heating functions computed by Cloudy in the presence of a generalized incident local radiation field. We characterize the radiation field through binned radiation field intensities instead of the photoionization rates used in our previous work. We find a set of 6 energy bins whose intensities exhibit relatively low correlation. We use these bins as features to train machine learning models to predict Cloudy cooling and heating functions at fixed metallicity. We compare the relative SHapley Additive exPlanation (SHAP) value importance of the features. From the SHAP analysis, we identify a feature subset of 3 energy bins ($0.5-1, 1-4$, and $13-16 \, \mathrm{Ry}$) with the largest importance and train additional models on this subset. We compare the mean squared errors and distribution of errors on both the entire training data table and a randomly selected 20% test set withheld from model training. The machine learning models trained with 3 and 6 bins, as well as 3 and 4 photoionization rates, have comparable accuracy everywhere, with errors $\gtrsim 10$ times smaller than for the interpolation table of Gnedin and Hollon (2012). We conclude that 3 energy bins (or 3 analogous photoionization rates: molecular hydrogen photodissociation, neutral hydrogen HI, and fully ionized carbon CVI) are sufficient to characterize the dependence of the gas cooling and heating functions on our assumed incident radiation field model.

79 ASTRONOMY AND ASTROPHYSICS

Mitigating Algorithmic Bias in Cancer Site Classification Models

Purpose Integrating artificial intelligence in cancer diagnostics has improved tumor classification beyond rule-based systems. Despite these advancements, these models may still encode demographic biases. We conducted a large-scale, applied bias-probing study of a deep learning–based cancer site classifier to quantify race information encoded in document embeddings. We then evaluated how performance changes when race-correlated embedding dimensions are removed in a post-training sensitivity analysis. Methods The cancer site classifier was trained using 3.5 million electronic cancer pathology reports from six of the National Cancer Institute's SEER registries. We trained a hierarchical self-attention network to generate 400-dimensional document embeddings. These embeddings were used to train two downstream, gradient-boosted decision tree classifiers: one to classify the cancer sites and another to predict racial categories. We identified overlapping features by intersecting the top 50 feature-importance rankings from the site and race models and computed their cumulative feature importance in each model. As a post hoc sensitivity analysis, we progressively pruned these overlapping dimensions, retrained the site model, and compared overall macro-F1 and accuracy, race-stratified macro-F1, and group fairness metrics on the basis of demographic parity and equalized odds before and after pruning. Results The analysis revealed minimal feature overlap between the cancer site and race prediction models, and the cumulative importance scores indicated a negligible influence of racial information on clinical predictions. Post-training pruning of overlapping features did not compromise the models' diagnostic accuracy, with a 0.07% loss in accuracy. Conclusion Our findings demonstrate that HiSAN-generated embeddings from SEER data can be used effectively in cancer site classification without significant demographic bias influencing the outcomes. Post-training pruning therefore functions as a practical audit and sensitivity check.

Shivanna, Abhishek [ORNL] (ORCID:0009000665228593)

Geographical Insights into Suicide Mortality Through Spatial Machine Learning

Suicide mortality is a leading cause of death in the United States, with an upward trend that emphasizes its significance as a public health issue. Previous research has employed global models like ordinary least squares (OLS) regression and local models such as geographically weighted regression (GWR). While local models are useful for analyzing spatial variations in suicide mortality, they share limitations with traditional global models, particularly about their inability to handle multi-collinearity and non-linear relationships. Machine learning approaches, like random forests (RF), can address some of these limitations but often fail to account for spatial variability. This gap highlights the need for spatial ML models specifically designed to tackle suicide mortality. This research seeks to fill this void by using a geographically weighted random forest model (GWRF) to examine the associations between county-level suicide mortality in the U.S. from 2010 to 2020 and various social and environmental determinants of health. A key aspect of our methodology is disciplined feature selection, which reduces the pool of explanatory variables by about 90%. This refinement enhances the explanatory power of both global (R2 improved from 0.59 to 0.67) and local (R2 improved from 0.64 to 0.67) RF models while reducing their run times. An analysis of the importance scores for these selected features reveals that the drivers of suicide mortality vary by context. Thus, to effectively address regional disparities and inform targeted public health interventions, a holistic approach that incorporates multiple county-level characteristics is essential.

Lebakula, Viswadeep [ORNL] (ORCID:0000000152935914

Predicting Initial Trans-Membrane Pressure for Optimized Operations in UF Unit Using Random Forest

With the growing scarcity of freshwater, innovative process design mechanisms like Reverse Osmosis (RO) are increasingly gaining attention among water treatment utilities to address the rising demand. Ensuring reliable water production necessitates efficient resource utilization, minimizing downtime in (ultra-filtration) UF systems. Recent advancements in machine learning (ML) have enabled the development of accurate data-driven models for Model Predictive Control (MPC), often requiring minimal prior knowledge of underlying physical processes. In this study, we present predictive regression models based on Random Forest (RF) and Auto-Regressive (AR) approaches to forecast the initial Trans-Membrane Pressure (TMP) for each filtration cycle in data generated by Direct Potable Reuse (DPR) systems. The proposed RF-based model demonstrates superior performance compared to baseline methods, including historical mean, Last Observation Carried Forward (LOCF), and naïve AR models, across various forecasting horizons in terms of root mean square error (RMSE) metric. To evaluate how different classes of process variables contribute to TMP dynamics over time, we examine the feature importance of independent covariates across multiple forecast horizons. This analysis provides insight into the temporal relevance of operational and sensor-derived features, guiding control and monitoring strategies. Additionally, the impact of hyperparameter tuning on TMP prediction performance is studied for both direct and recursive RF modelling approaches across increasing forecast horizons. Accurate prediction of initial TMP is critical for optimizing RO operations, as it enables the development of robust modelling frameworks by accurately estimating membrane fouling trends, thereby enhancing process efficiency and long-term reliability. The demonstrated efficacy of the RF-based approach highlights its potential as a tool for real-time decision-making in water treatment systems, paving the way for advanced process optimization and sustainable water resource management.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)

Predicting initial trans-membrane pressure across cycles in the ultrafiltration process using random forest

With growing freshwater scarcity, direct potable reuse (DPR) systems that reclaim wastewater for drinking are becoming increasingly important for sustainable water supply. Reliable operation requires minimizing downtime in ultrafiltration (UF) units, where membrane fouling leads to elevated trans-membrane pressure (TMP). This study develops data-driven regression models based on random forest (RF) and autoregressive (AR) approaches to forecast the initial TMP at the start of each UF filtration cycle in a pilot-scale DPR system. The RF model consistently outperforms baseline methods, including historical mean, last observation carried forward, and AR models, across multiple forecast horizons, achieving the lowest root mean square error. To evaluate how different classes of process variables contribute to TMP dynamics over time, we examine the feature importance of independent input variables across multiple forecast horizons. This analysis provides insight into the temporal relevance of operational and sensor-derived features, guiding control and monitoring strategies. Additionally, the impact of hyperparameter tuning on TMP prediction performance is assessed for both direct and recursive RF modelling approaches. The proposed RF framework establishes a robust foundation for predictive monitoring and real-time optimization of UF operations, supporting sustainable and reliable water reuse.

direct potable reuse

Dataset: "Widespread Drought-driven Declines in Streamflows and Water quality in the Upper Colorado River Basin (1998-2022)"

This data package contains the associated data and scripts for Nagamoto, E., Ombadi, M., Ciulla, F. et al. Widespread drought-driven declines in streamflows and water quality in the Upper Colorado River Basin during 1998-2022. Commun Earth Environ 7, 734 (2026). https://doi.org/10.1038/s43247-026-03890-5. This purpose of this study was to investigate the impact of the 21st century drought on water quantity and quality at catchments throughout the Upper Colorado River Basin (UCRB). We used stream flow, water temperature, specific conductance, air temperature, precipitation, and catchment attribute data for over 200 sites in the UCRB, collected from the National Water Information System using Basin3D (Varadharajan, 2023), GAGESII (Falcone, 2010), and the Google Earth Engine. We identified years of severe drought between 1998 and 2022 using the Standardized Precipitation Evaporation Index (SPEI), then calculated the relative change percentage of the stream flow, water temperature, and specific conductance from drought versus non-drought years. We used the attribute information from GAGESII to investigate what physical traits of catchments are associated streamflow vulnerability (greater relative change) or resilience to drought. We used land cover data from the National Land Cover Database (USGS, 2024) to assess any changes to physical attributes that may not be represented in the static attributes information in GAGESII. To increase data availability, we modeled stream temperature using methods from Willard, 2023. While the study period is water years 1998 to 2022, the raw water quantity and quality data extends to 1950 and the meteorological data extends to 1980. The data and code can be downloaded via the UCRB_drought.zip. Within the zip, the files are organized as follows: - INPUTS: Contains all input data used in UCRB_Drought_Workflow.ipynb - OUTPUTS: Contains all intermediate data created from UCRB_Drought_Workflow.ipynb as well as final products including the calculated Standardized Evapotranspiration Index (SPEI) - climatic_variables: The code used to collect meteorologic data from Google Earth Engine - feature_importance: The code used for the catchment attributes analysis - preprocessing: Code used in UCRB_Drought_Workflow_Preprocessing.ipynb - pyeto: Code used in UCRB_Drought_Workflow_Preprocessing.ipynb - calculations: Code used in UCRB_Drought_Workflow_Impacts.ipynb - plotting: Code used in UCRB_Drought_Workflow_Impacts.ipynb - README.md - UCRB_Drought_Workflow_Preprocessing.ipynb: The code used to prep raw data for the analysis - UCRB_Drought_Workflow_Impact.ipynb: The code which uses the prepped raw data for analysis, and plots all figures - requirements_ucrb-drought_v2.yml: The requirements file to create a virtual environment and Jupyter Lab kernel to run the code The INPUTS folder is organized into the following major directories and sub-directories. The "RDC_WT_SC_RAW" folder contains raw data for streamflow, water temperature, and specific conductance in a ".h5" file. The "NLCD_RAW" folder contains ".csv" files with annual land cover percentages for counties within the UCRB. The "MET_RAW" folder contains a ".csv" file with monthly meteorological data (air temperature and precipitation) for the sites in the UCRB which was obtained from code in the climatic_variables folder. The "GAGESII" folder contains ".csv" files with physical catchment attribute variables for catchments across the country. The "WT_LSTM_data" folder contains ".csv" files with calculated WT (Willard, 2023) and the associated RMSEs. The "Upper_Colorado_River_Basin_Boundary" folder contains geographic data including a shapefile for plotting in the UCRB_Drought_Workflow.ipynb. The "RESERVOIRS_RAW" folder contains ".csv" files for each reservoir in the UCRB with daily reservoir storage. There are also two files in the INPUTS folder that have combined reservoir storage data and reservoir metadata. The OUTPUTS folder is organized into the following major directories and sub-directories. The "RDC_WT_SC_data" folder contains a folder "Water_year" with the associated cleaned data, metadata, and data availability information in ".csv" files, a folder "Median_Relchange" with the relative change comparing drought to non-drought years in ".csv" files, and a folder "Peak95_Min5_Relchange" that has ".csv" files for the relative change in peak (95th %) and minimum (5th %) variables. The "NLCD_data" folder contains the difference in land cover from the beginning to end of the study period and the percentage of the county that is within UCRB bounds can be found in Nagamoto et al (2025)). The "MET_data" folder contains separated monthly air temperature and precipitation data and the calculated PET in ".csv" files. The "SPEI_data" folder contains ".csv" files with calculated SPEI values (one restricted to the study period and the other with information from the entire MET data period). The "Paper_Tables" folder contains two ".csv" files containing site information and data availability and information about the GAGESII trait aggregated categories. The base directory includes the file “flmd.csv” for a list and description of all files and the file “dd.csv” for data dictionaries. Scripts for preprocessing, analysis, and figure generation are located in the associated GitHub repository found at [https://github.com/iNAIADS/drought-impacts/tree/develop/UCRB-drought]. UPDATE 1: Title and code file updated to match submitted manuscript 10-15-2025. UPDATE 2: Code and data files updated to match revised manuscript 3-4-2026. UPDATE 3: Code and data files updated to match revised manuscript 6-7-2026. ** NOTE: DD and FLMD have not been updated yet. UPDATE 4: Added associated Manuscript information and DD and FLMD have been updated. To cite this code, please use the following BibTeX: @misc{nagamoto2025drought, author = {Emily Nagamoto and Fabio Ciulla and Mohammad Ombadi and Jared Willard and Rosemary Carroll and Charuleka Varadharajan}, title = {Dataset: "Widespread Drought-driven Declines in Streamflows and Water quality in the Upper Colorado River Basin (1998-2022)"}, year = {2025}, doi = {10.15485/2551894}, publisher = {ESS-DIVE Repository}, url = {https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2551894} }

54 ENVIRONMENTAL SCIENCES

Detecting Important Drivers of Gridded Population Modeling With Machine Learning

High-resolution population datasets have been lever-aged across a broad swath of domains, such as climate change, public policy, humanitarian aid, and rescue operations, among others. Machine learning methods were adopted to generate high-resolution or gridded population estimates by using various geospatial input features such as buildings, roads, and nighttime lights. In this study, we evaluate the importance of population features using Random Forest models across three levels of analysis, utilizing permutation measures. Our research aims to address key questions to enhance our understanding of high-resolution population modeling, such as: Are certain features globally (10 countries collectively) more important than others? Do optimal features vary by country? Within each country, do feature importance differ across administrative units? What similarities exist in feature importance at the global, country, and administrative unit levels? To answer these questions, we leverage the Kneedle algorithm to automate the selection of optimum features. We find that there are patterns displayed by features across spatial boundaries, evidenced by the same feature being the most important indicator of population across 7 of the 10 countries modeled. Our findings indicate that while important features may vary across geographies, certain features consistently hold greater importance than others agnostic of geography.

Lebakula, Viswadeep [ORNL] (ORCID:0000000152935914

Synthetic spectra for Lyman- α forest analysis in the Dark Energy Spectroscopic Instrument

Synthetic data sets are used in cosmology to test analysis procedures, to verify that systematic errors are well understood and to demonstrate that measurements are unbiased. In this work we describe the methods used to generate synthetic datasets of Lyman-α quasar spectra aimed for studies with the Dark Energy Spectroscopic Instrument (DESI). In particular, we focus on demonstrating that our simulations reproduces important features of real samples, making them suitable to test the analysis methods to be used in DESI and to place limits on systematic effects on measurements of Baryon Acoustic Oscillations (BAO). We present a set of mocks that reproduce the statistical properties of the DESI early data set with good agreement. Additionally, we use a synthetic dataset to forecast the BAO scale constraining power of the completed DESI survey through the Lyman-α forest.

79 ASTRONOMY AND ASTROPHYSICS

Machine Learning-Assisted Recovery of Delicate Kinetic Information from Transient Reactor Experiments

Identifying active sites and their roles in chemical reaction steps remains a vital challenge in heterogeneous catalysis. Transient experiments offer a unique way to probe active sites and distinguish subtle kinetic features. Although physics-based analysis methods may be well-developed, they can be highly susceptible to experimental noise, and smoothing methods may erase or even distort important features; a smooth curve is not always the best curve. We demonstrate a new workflow for the direct interpretation of intrinsic kinetic information from exit flux curves measured in transient reactor experiments. This workflow contains three artificial neural networks (ANNs), including a noise reducer, a concentration predictor, and a rate predictor to analyze experimental data, followed by the virtual TAP (VTAP) physics-based reactor model and density functional theory (DFT) calculations of adsorption energies on specific sites. We use this workflow to analyze the data from experiments titrating Pt/Al 2 O 3 and Pt/SiO 2 catalysts with carbon monoxide (CO) in the temporal analysis of products (TAP) reactor. Our workflow separates the time-evolving chemical reaction and mass transfer information contained in the TAP pulse response. The existence of strong- and weak-binding sites on the Pt/Al 2 O 3 catalyst is observed in the catalyst titration experiment in the transient reactor. The structures of the strong- and weak-binding sites are then identified by using DFT calculations. We find that the Pt/SiO 2 catalyst has only strong-binding sites, which aligns with the inactive support effect of SiO 2 . We demonstrate how machine learning methods provide unique insights with high-resolution data analysis that cannot be achieved by using state-of-the-art physics-based methods.

Adsorption

Design of diverse, functional mitochondrial targeting sequences across eukaryotic organisms using variational autoencoder

Mitochondria play a key role in energy production and metabolism, making them a promising target for metabolic engineering and disease treatment. However, despite the known influence of passenger proteins on localization efficiency, only a few protein-localization tags have been characterized for mitochondrial targeting. To address this limitation, we leverage a Variational Autoencoder to design novel mitochondrial targeting sequences. In silico analysis reveals that a high fraction of the generated peptides (90.14%) are functional and possess features important for mitochondrial targeting. We characterize artificial peptides in four eukaryotic organisms and, as a proof-of-concept, demonstrate their utility in increasing 3-hydroxypropionic acid titers through pathway compartmentalization and improving 5-aminolevulinate synthase delivery by 1.62-fold and 4.76-fold, respectively. Moreover, we employ latent space interpolation to shed light on the evolutionary origins of dual-targeting sequences. Overall, our work demonstrates the potential of generative artificial intelligence for both fundamental research and practical applications in mitochondrial biology.

59 BASIC BIOLOGICAL SCIENCES

Scalable Computation of Topological Abstractions for Scalar Data

Topological data analysis has become an important tool for large scale scalar data analysis and visualization, efficiently extracting the inherent structure and features of interest of the data. However, with growing dataset sizes and complexity, it is increasingly becoming infeasible to compute topological abstractions of interest in serial and on single machines. This paper presents the state of the art in the scalable computation of topological abstractions on scalar data, in shared memory parallel on single machines, and in distributed memory parallel on multiple machines. We highlight results for set‐based, graph‐based and complex‐based abstractions and organize the state of the art based on this taxonomy. The paper identifies parallelization and distribution techniques common in topological algorithms and highlights further areas of interest with underdeveloped efforts.

97 MATHEMATICS AND COMPUTING

Unlocking Scholarly Insights: Leveraging Machine Learning Approaches for Citation Analysis and Intent Classification

Publicly funded organizations, notably institutions like the Los Alamos National Laboratory (LANL), are deeply vested in acquiring robust productivity metrics to gauge the entirety of their research output. Motivated by the imperative to enhance institutional productivity assessment, this study investigates the utilization of Large Language Models (LLM) such as BERT-based models, as well as local LLaMa-30b-instruct and Mixtral-8x7b-instruct architectures for classifying type of URL referenced resources in academic papers such as software, dataset, as well as authorship intent. Challenges in discerning resource types from context are highlighted, along with the potential of BERT and LLMs to address these challenges. Through comprehensive analysis, this research unveils a notable surge in documents featuring URL citations, indicative of the escalating importance of digital resources in scholarly publications. Moreover, citations to datasets and software demonstrate consistent growth over time, underscoring their increasing significance. Our findings also reveal that LANL authors contribute substantially to accessible science, comprising about 10% of dataset and software mentions in LANL

Large Language Models, BERT, citation classificati

Fundamental Insights into Cathode Stability: Linking Compositional Tuning and Local Coordination in Complex Metal Oxides under Aqueous Transformations

Compositional tuning of complex metal oxides in Li-ion battery materials influences their performance as well as their end-of-life behavior, in particular, the tendency to release toxic metal cations in aqueous solution. We modeled ternary variants of a parent LiCoO 2 delafossite structure by varying the metal identity and relative amounts. This yielded ten model formulations of Li(A 4/6 B 1/6 C 1/6 )O 2 , where the material is enriched with the A metal and doped with B and C, with Ni, Mn, Co, Fe, Al, V, and Ti as constituent metals. To assess their stability in aqueous conditions, metal release energetics were calculated using a combination of Density Functional Theory calculations and thermodynamics. Metal release in ternary oxides is dictated by subtle variations in the coordination environment of the leaving group. To identify governing chemical features across diverse compositions with varying local coordination environments, we leverage random forest regression and descriptor importance analysis. A key result is that metal–oxygen orbital hybridization, quantified using a projected density-of-states-derived descriptor, H d/p , provides a physically grounded measure of interaction strength that governs metal release energetics. This refined perspective goes beyond conventional oxidation state considerations and offers more robust insights for materials science. Finally, we model defect surface-bound O 2 dimer formation as a proxy for reactive oxygen species (ROS) generation. The results show that Ni-rich compositions more readily stabilize spin-polarized O 2 dimers, corroborating experimental reports of an increased ROS-driven biological response. In conclusion, our results establish a compositional and electronic basis for metal release and surface oxygen reactivity that form a rationale for complex metal oxide design principles.

36 MATERIALS SCIENCE

Raman spectroscopic investigation of ianthinite [U$_2^{4+}$(UO$_2$)$_4$O$_6$(OH)$_4$(H$_2$O)$_4$]·$5$H$_2$O, a rare mixed-valence uranium oxide hydrate

Ianthinite ([[U$_2^{4+}$(UO$_2$)$_4$O$_6$(OH)$_4$(H$_2$O)$_4$]·$5$H$_2$O) is an exotic mineral that possesses U in both tetravalent and hexavalent oxidation states and is structurally related to the U 3 O 8 polymorphs, which are commonly encountered technogenic materials in the nuclear fuel cycle. Despite the similarities between U 3 O 8 and ianthinite, and the importance of ianthinite in U paragenesis, no Raman spectra have been reported for this mineral. Here, to gain a more complete understanding of how structural attributes of ianthinite give rise to observable spectroscopic features and how these may relate to important materials in the nuclear fuel cycle, we provide, for the first time, Raman spectra of ianthinite. Ianthinite readily oxidizes at ambient conditions, complicating analysis of phase-pure material. Several analytical methods are employed herein to decouple the Raman features of ianthinite from its alteration product(s). First, a simple difference spectrum is presented, then results of Raman spectroscopic mapping are employed, and finally, we use a novel processing and analysis method. Each analysis method provides different insight into structural features that are unique to ianthinite, in particular, features that are attributable to U(IV) in distorted octahedral coordination in both ianthinite and U 3 O 8 phases.

Spano, Tyler L. [Oak Ridge National Laboratory (OR