Search NASASearch

SEARCH · Search NASA

Results for “Feature selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Machine Learning–Augmented Laser-Induced Breakdown Spectroscopy for Spectral Discrimination of Iron Oxalates

Enhanced characterization and phase identification of post-PUREX Pu Oxalates (PuOXA) are pivotal for nonproliferation and pre-detonation nuclear forensics. Despite significant advances in the characterization of PuO 2 samples, little is known about the impact of both the chemical structure and oxidation states of PuOXA (i.e., Pu(III) and Pu(IV)) have on optical emission signatures. Here, we demonstrate the analytical capabilities of laser-induced breakdown spectroscopy (LIBS) applied to Fe(II) and Fe(III) oxalate samples as surrogates for PuOXA, highlighting the discriminating features in the LIBS emission spectra arising from differences in the oxidation states within mixed FeOXA samples. We report the enhancement of spectral feature selection using Principal Component Analysis (PCA), which enables the analytical superiority of machine learning algorithms such as Linear Discriminant Analysis (LDA), Quadratic Discriminant Analysis (QDA), Partial Least Squares Regression (PLSR), Support Vector Regression (SVR), and Random Forest Regression (RFR) over conventional univariate techniques for phase discrimination and chemometric analysis. Cluster analysis revealed how both matrix effects and laser ablation influence cluster separability by introducing spectral artifacts that misdirect the maximization of variance. PCA-selected emission lines were used in the regression models, demonstrating that both univariate and multivariate linear regression models (i.e., PLSR and SVR) can achieve acceptable performance, with machine learning models outperforming conventional calibration regressions. Furthermore, the application of non-linearly activated PCA-selected emission lines illustrates how simplifying the data while retaining captured variance enables the use of less complex and more computationally efficient models. Furthermore, this is particularly evident in the underperformance of RFR, which suffers from increased computational costs and overfitting owing to its high complexity.

Oxalates

Enhancing dimensionality prediction in hybrid metal halides via feature engineering and class-imbalance mitigation

We present a machine learning (ML) framework for predicting the structural dimensionality of hybrid metal halides (HMHs), including organic-inorganic perovskites, using a combination of chemically-informed feature engineering and advanced class-imbalance handling techniques. This study is motivated by the small and highly imbalanced nature of experimentally available HMH datasets, which limits the applicability and reliability of conventional ML approaches. The dataset, consisting of 494 HMH structures, is highly imbalanced across dimensionality classes (0D, 1D, 2D, 3D), posing significant challenges to predictive modeling. To mitigate this limitation, the dataset was augmented to 1336 samples using the synthetic minority oversampling technique, enabling improved learning of underrepresented dimensionality classes while preserving chemically meaningful feature relationships. We developed interaction-based descriptors designed to capture coupled steric and polarity effects relevant to dimensionality prediction, which are not readily captured by standard single-parameter or composition-only descriptors. These descriptors are integrated into a multi-stage workflow combining feature selection, ensemble stacking, and performance optimization. Our approach significantly improves F1-scores for underrepresented classes, achieving robust cross-validation performance across all dimensionalities. This work demonstrates a generalizable strategy for extracting reliable and interpretable structure–dimensionality relationships from limited experimental data, enabling pre-synthesis screening of organic cations and providing a practical blueprint for small-data ML in hybrid materials systems.

36 MATERIALS SCIENCE

Switchgrass Steroidal Saponins Reduce Fungal Disease but Decrease Yeast Fermentation Yield

Increasing the production of bioproducts from lignocellulosic feedstocks requires improvement in both field production and biorefinery efficiency. When plant traits arise that improve field production but decrease biofuel yield, these trade-offs can represent challenges in the entire production process. To examine trade-offs between field and production traits, we examined factors underlying switchgrass resistance to fungal rust pathogens in field conditions and factors that impede yeast fermentation in the lab using repeated measurements on a switchgrass genetic diversity panel. We found that the same switchgrass genotypes that showed high fungal pathogen resistance also showed recalcitrance to yeast fermentation. These switchgrass genotypes were mostly from the Atlantic genetic group, which had high levels of specialized metabolites of the saponin class. Among 1589 metabolites identified through metabolomics, we found that saponins were among the most likely to explain variation in both rust infection and fermentation yield using random forest feature selection, and that only four of these were sufficient to explain 57.9% of the variation in rust susceptibility. Through follow-up testing in recalcitrant biomass, we found that the bacterium Zymomonas mobilis does not suffer the same inhibition as the yeast Saccharomyces cerevisiae, and that the addition of ergosterol (thought to be the fungal cellular target of saponin inhibition) rescues yeast fermentation. Several lines of evidence point to a central role for saponins as key metabolites protecting switchgrass from fungal pathogens and interfering with yeast fermentation, underscoring an ongoing need for collaboration between plant breeders and biofuel production scientists.

VanWallendael, Acer [North Carolina State Universi

A reproducible study design for the MIMIC-IV in-hospital mortality task

Open, tabular electronic health record (EHR) datasets such as MIMIC-III and MIMIC-IV have become critical resources for developing machine learning (ML) models addressing clinical prediction tasks, including hospital readmission, length of stay, and in-hospital mortality (IHM). While MIMIC-III has benefited from well-established preprocessing pipelines and standardized feature sets, MIMIC-IV remains comparatively challenging to work with because there are no standardized benchmarks to support reproducibility and comparability across studies. To address this limitation, we present a rigorously curated MIMIC-IV custom feature set optimized for IHM prediction, constructed through a reproducible preprocessing pipeline and feature selection strategy.

97 MATHEMATICS AND COMPUTING

Analysis of Waste Material Feedstocks Using Laser-Induced Breakdown Spectroscopy and Machine Learning

Predicting properties such as heating value, ash fusion temperature, and mineral ash composition from Laser-Induced Breakdown Spectroscopy (LIBS) data can make gasifiers more flexible to different feedstocks. Understanding these feedstock properties in-situ improves feedstock conversion modelling methods that allow for consistent operation, higher carbon conversion, and reduced fouling and erosion rates. The purpose of this study is to demonstrate methods for model creation that take LIBS data as predictor features and estimate higher order material properties as a function of feedstock material properties. Six samples were chosen to represent a mixture of abundant and carbon rich waste materials. LIBS measurements were performed on these samples for elemental wavelengths and intensity values. Laboratory analytical results were obtained for each sample’s heating value, proximate and ultimate analysis, mineral ash composition, ash fusion temperatures, and viscosity temperatures. Thermal conductivity was measured using a HotDisk TPS 2500S. LIBS measurements were processed and used as predictor features for machine learning (ML) models to predict the sample’s material properties. Predictor feature selection algorithms, particularly minimum redundancy maximum relevance (mRMR), reduced the dimensionality of ML models. Many modelling methods such as Gaussian process regression (GPR), regression tree, neural networks (NN), and support vector machines (SVM) were demonstrated to be effective at predicting higher order properties; however, mRMR with GPR stood out as a clear winning combination.

01 COAL, LIGNITE, AND PEAT

Identification of Differential Equations by Dynamics-Guided Weighted Weak Form with Voting

In the identification of differential equations from data, significant progresses have been made with the weak/integral formulation. In this paper, we explore the direction of finding more efficient and robust test functions adaptively given the observed data. While this is a difficult task, we propose weighting a collection of localized test functions for better identification of differential equations from a single trajectory of noisy observations on the differential equation. We find that using high dynamic regions is effective in finding the equation as well as the coefficients, and propose a dynamics indicator per differential term and weight the weak form accordingly. For stable identification against noise, we further introduce a voting strategy to identify the active features from an ensemble of recovered results by selecting the features that frequently occur in different weighting of test functions. Systematic numerical experiments are provided to demonstrate the robustness of our method.

97 MATHEMATICS AND COMPUTING

SCULPT (Supervised Clustering and Uncovering Latent Patterns with Training) v1

SCULPT (Supervised Clustering and Uncovering Latent Patterns with Training) is a comprehensive data visualization and analysis application focused on working with COLTRIMS (COLd Target Recoil Ion Momentum Spectroscopy) data, which is used in atomic and molecular physics experiments. The application offers several powerful features: - Data uploading and processing capabilities for COLTRIMS files - Multiple visualization methods using UMAP (Uniform Manifold Approximation and Projection) for dimensionality reduction - Interactive selection of data points across multiple views - Feature engineering through various methods: - Manual feature selection from calculated physics parameters - Deep autoencoder for dimension reduction - Genetic programming for discovering meaningful features - Mutual information-based feature selection - Multiple clustering approaches (DBSCAN, KMeans, Agglomerative) - Quality metrics for evaluating clustering results - Export capabilities for selections and generated features

Daoud, Hazem [Lawrence Berkeley National Laborator

Laser-induced selective local patterning of vanadium oxide phases

The same elements can form different compounds with widely different physical properties. Synthesis of a single-phase material is commonly achieved by controlling experimental conditions. Synthesizing materials that incorporate multiple specific spatially distributed chemical phases is often challenging, especially if different phases must be organized into well-defined spatial patterns. Here, we present an efficient solid reaction laser annealing (SRLA) approach to directly write regions of different local chemical compositions. We demonstrate the practical utility of our approach by locally writing microscale patterns of distinct chemical phases in vanadium oxide thin films. Specifically, we achieved the controlled local recrystallization of a uniform V 2 O 3 matrix into VO 2 , V 3 O 5 , and V 4 O 7 regions exhibiting sharp 1st- and 2nd-order metal–insulator phase transitions over a wide range of critical temperatures, i.e., a characteristic feature of select vanadium oxides that is extremely sensitive to even minute structural or compositional imperfections. We utilized the local chemical phase writing to pattern spiking oscillators with distinct electrical behavior directly in the thin film sample without employing elaborate lithography fabrication. Our laser tuning local chemical composition opens a pathway to synthesize a wide range of artificially micropatterned composite materials, with precision and control unattainable in conventional material synthesis methods.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND

Influence of particle size on NIR spectroscopic characterization of sorghum biomass for the biofuel industry

NIR spectroscopy is a rapid and accurate green technology for high-throughput biomass characterization, including sorghum (Sorghum bicolor), a promising energy crop for the biofuel industry. This study assessed the influence of particle size on NIR spectroscopic analysis (wavelength range: 867–2535 nm) of sorghum biomass composition. Grown under field conditions, a total of 113 types of genetically diverse sorghum accessions were dried, ground, and sieved (<250, 250–600, 600–850, and > 850 µm particle size) for developing partial least square regression (PLSR) prediction models for moisture, ash, extractive, glucan, xylan, acid-soluble lignin (ASL), acid-insoluble lignin (AIL), and total lignin (ASL + AIL). Overall, smaller particle sizes provided better model performance, while no single particle size provided the best performance for all the selected components. With only 9 selected bands and 4 latent variables (LVs), the best PLSR model was obtained for moisture with particle size of 600–850 µm with the square root of the coefficient of determination (R) of 0.85, the ratio of prediction to deviation (RPD) of 2.2, and the root mean square error (RMSE) of 0.46 % in external validation. Similar model performances were also obtained for ash, extractive, glucan, and xylan. This study showed that size reduction could effectively improve NIR spectroscopic analysis for lipid-producing sorghum biomass for the biofuel industry.

09 BIOMASS FUELS

Manipulating the Assembly and Architecture of Fibrillar Silk

Silk is a unique and exceptionally strong biological material. However, no synthetic method has yet come close to replicating the properties of natural silk. This shortfall is attributed to an insufficient understanding of both silk nanofibril structure and the mechanism of formation. Here in situ atomic force microscopy (AFM) and photo-induced force microscopy (PiFM) is utilized to investigate the formation process and define the basic structural paradigm of individual silk nanofibrils. By visualizing the multistage process of silk nanofibril formation, the importance of conformational transformations along the assembly pathway is revealed. Unfolded silk structures initially accumulate into amorphous clusters, which then evolve into crystal nuclei via conformational transformation into β-crystallites. Nanofibril elongation then occurs through the attachment of silk molecules at a single end of the nanofibril tip; this is facilitated through the formation of a new amorphous cluster that then repeats the aforementioned conformational transformation. However, enzymatic digestion of the amorphous regions leads to direct, rapid elongation of β-crystalline fibers. These findings imply that the energy landscape is characterized by shallow minima associated with intermediate states, which can be eliminated by introducing β-crystallites, and motivate research into the directed modification of the silk assembly pathway to select for features beneficial to specific applications.

36 MATERIALS SCIENCE

A Rhenium Bis -tetramethylphenanthroline Catalyst for CO 2 Reduction to Formate

Catalytic CO 2 reduction reactions featuring high selectivity toward formate are relatively rare. In some homogeneous molecular CO 2 -reducing electrocatalysis, using triethylamine (TEA) and isopropanol (IPA) as additives improves catalytic performance in producing formate. In this paper, we investigate whether the rhenium(I) bis-diimine dicarbonyl complexes, cis-[Re(N^N) 2 (CO) 2 ] + , where N^N is 2,2’-bipyridine ([1] + ) or 3,4,7,8-tetramethyl-1,10-phenanthroline ([2] + ), are capable of electrocatalytically reducing CO 2 to formate in acetonitrile containing TEA and IPA. Catalyst [1] + was ineffective at CO 2 reduction, yielding formate quantities comparable to those produced in experiments without the catalyst. Catalyst [2] + , however, is a promising electrocatalyst for the CO 2 reduction reaction in the presence of TEA and IPA, with formate being produced in millimolar concentrations (10.5 mM), as detected by 1 H NMR spectroscopy after 6 h electrolysis (formate Faradaic efficiency = 11%, with the major balance going to H 2 ). Upon more detailed examination, [2] + exhibited a turnover frequency (TOF) of 12 s –1 for formate, comparable to other leading molecular catalysts that competently execute this reduction. Combinations of spectroscopy, electrochemistry, and theory were used to better understand the mechanism of CO 2 reduction by [2] + . Fourier transform infrared spectroelectrochemical (FTIR-SEC) data provided no evidence for CO ligand dissociation or substitution upon one- and two-electron reduction of [2] + , suggesting that a mechanism distinct from one that is metal-hydride-based is operative in catalysis. Computational studies guide mechanistic investigations toward the proposed formation of a hydrophenanthroline-based intermediate responsible for hydride transfer to CO 2 and electrocatalytic formate production from [2] + .

Beverages

Targeted genetic manipulation and yeast-like evolutionary genomics in the green alga Auxenochlorella

Auxenochlorella spp. are diploid oleaginous green algae whose streamlined genomes can be readily manipulated by homologous recombination, making them highly amenable to discovery research and bioengineering. Vegetatively diploid organisms experience specific evolutionary phenomena, including allodiploid hybridization, mitotic recombination, loss-of-heterozygosity, and aneuploidy; however, studies of these forces have largely focused on yeasts. Here, we present a telomere-to-telomere phased diploid genome assembly of Auxenochlorella UTEX 250-A (haploid length 22 Mb) and introduce a genetic toolkit for site-specific manipulation of the nuclear genome in multiple strains, featuring several selectable markers, inducible promoters, and fluorescent reporters for protein localization. UTEX 250-A is an allodiploid hybrid of Auxenochlorella protothecoides and Auxenochlorella symbiontica, two species differentiated by extensive chromosomal rearrangements. UTEX 250-A haplotypes are a mosaic of each parental species following mitotic recombination, and two chromosomes are trisomic. Loss-of-heterozygosity events are pervasive across Auxenochlorella and can evolve rapidly in the laboratory. High-quality structural annotation yielded ∼7,500 genes per haplotype. Auxenochlorella have experienced gene family loss and reduction, including core photosynthesis genes, and exhibit periodic adenine and cytosine methylation at promoters and gene bodies, respectively. Approximately 10% of genes, especially those involved in DNA repair and sex, overlap antisense long noncoding RNAs, which may participate in a regulatory mechanism. We demonstrate the utility of Auxenochlorella for fundamental research by knockout of a chlorophyll biosynthesis enzyme, and confirm one trisomy by allele-specific transformation. These results demonstrate the generality of several evolutionary forces associated with vegetative diploidy and provide a foundation for the use of Auxenochlorella as a reference organism.

CHL27

Deep Factorization Machine Learning for Disaggregation of Transmission Load Profiles with High Penetration of Behind-The-Meter Solar

The ever-growing integration of distributed energy resources (DERs), especially behind-the-meter (BTM) solar generations, poses imperative operational challenges to system operators such as regional transmission organizations (RTOs). It is important for RTOs to effectively and accurately extract actual load profiles at the transmission level for a single node with significant BTM solar injection. This paper first illustrates the necessity of disaggregating the daily actual load profile of a single node. Furthermore, by segmenting nodes with selected timeseries features, nodes with significant BTM solar generation are identified. Lastly, a bi-level framework is proposed, comprising reference node disaggregation and DeepFM nodal disaggregation, aimed at disaggregating the nodal load profiles from which system operators require more information. By adopting a hybrid Deep Factorization Machine (DeepFM) model, the model achieve accurate results by extracting both linear and nonlinear relations between nodes in the same region and the zonal load and nodal load profile. To overcome the lack of ground truth, this paper segments the load profile into daytime, nighttime, and zero-crossing points and utilizes the latter two for evaluation purposes. The proposed disaggregation procedure is validated using real world, minute-level, normalized, and anonymized nodal data in the PJM service territory.

42 ENGINEERING

EOSPAC User's Manual: Version 6.5 Second Edition (Rev. 3)

The EOSPAC utility package is a collection of interface routines, which can be used to access the SESAME data library and perform various data adjustments and interpolations on the SESAME data. The SESAME data library contains both thermodynamic (e.g., equation of state) and transport coefficients (e.g., opacity and conductivity). Note, for simplicity, the term EOS (equation of state) used herein includes both thermodynamic variables and transport coefficients. The EOSPAC utility package is designed to be used by physics codes (henceforth ”host codes”) written in multiple languages and on multiple platforms. The remainder of this manual is organized into several sections. Chapter 2 discusses conventions such as data organization and routine names. Chapter 3 provides a general overview of basic theory and models implemented within EOSPAC. Chapter 4 provides a general overview of how to use the EOSPAC interface library. Chapters 5 to 7 describe the public interfaces of EOSPAC in detail. Chapter 8 provides a brief introduction to some related tools, which may be of use to the user. Chapter 9 provides details related to some selected numerical features of EOSPAC. Chapter 10 gives examples for using the interface routines described in chapters 5 to 7. Chapter 11 provides technical support contact information. Chapter 12 contains a brief set of acknowledgments. Chapter 13 contains a list of referenced documents. Finally, chapter 14 lists the “table types: mnemonic conventions”, “table types: grouped by category, sorted by name”, “table types: eospac version 5 cross reference”, “options: setup phase”, “data information parameters”, “meta-data information parameters”, “options: interpolation phase”, and the “error codes”.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Balancing Trade-offs: Adaptive Differential Privacy in Interpretable Machine Learning Models

In the advancing field of machine learning, balancing accuracy, interpretability, and privacy represents a significant challenge. The problem is exacerbated by the widespread deployment of pre-trained models locally in diverse applications, which could lead to various amounts of privacy leakage. Conventional Differential Privacy strategies, in which uniform noises are applied to model gradients, guarantee data privacy at the expense of accuracy and interpretability. This paper introduces a Feature-Sensitive Adaptive Differential Privacy (FADP) framework with a unique noise-adding strategy. Noises are adaptively added based on feature importance clustering, where important features are considered for interpretability. By employing a unique masking technique, FADP selectively preserves crucial features with minimal noise interference, maintaining accuracy while enhancing interpretability. The FADP framework addresses the limitations of traditional DP methods by preserving critical channels and improving interpretability — a vital requirement in machine learning applications that demand transparency in model decisions. Through comprehensive testing, FADP is shown to balance the trade-offs among accuracy, privacy, and interpretability, marking a substantial advancement in the field of privacy-preserving machine learning.

Farhad Riya, Farhin [University of Tennessee, Knox

Detecting Important Drivers of Gridded Population Modeling With Machine Learning

High-resolution population datasets have been lever-aged across a broad swath of domains, such as climate change, public policy, humanitarian aid, and rescue operations, among others. Machine learning methods were adopted to generate high-resolution or gridded population estimates by using various geospatial input features such as buildings, roads, and nighttime lights. In this study, we evaluate the importance of population features using Random Forest models across three levels of analysis, utilizing permutation measures. Our research aims to address key questions to enhance our understanding of high-resolution population modeling, such as: Are certain features globally (10 countries collectively) more important than others? Do optimal features vary by country? Within each country, do feature importance differ across administrative units? What similarities exist in feature importance at the global, country, and administrative unit levels? To answer these questions, we leverage the Kneedle algorithm to automate the selection of optimum features. We find that there are patterns displayed by features across spatial boundaries, evidenced by the same feature being the most important indicator of population across 7 of the 10 countries modeled. Our findings indicate that while important features may vary across geographies, certain features consistently hold greater importance than others agnostic of geography.

Lebakula, Viswadeep [ORNL] (ORCID:0000000152935914

Improving galaxy cluster selection with the outskirt stellar mass of galaxies

The number density and redshift evolution of optically selected galaxy clusters offer an independent measurement of the amplitude of matter fluctuations, 𝑆 8 . However, recent results have shown that clusters chosen by the redMaPPer algorithm show richness-dependent biases that affect the weak lensing signals and number densities of clusters, increasing uncertainty in the cluster mass calibration and reducing their constraining power. Here, in this work, we evaluate an alternative cluster proxy, outskirt stellar mass, 𝑀 out , defined as the total stellar mass within a [50, 100] kpc envelope centered on a massive galaxy. This proxy exhibits scatter comparable to redMaPPer richness, 𝜆, but is less likely to be subject to projection effects. We compare the Dark Energy Survey Year 3 redMaPPer cluster catalog with a 𝑀 out selected cluster sample from the Hyper-Suprime Camera survey. We use weak lensing measurements to quantify and compare the scatter of 𝑀 out and 𝜆 with halo mass. Our results show 𝑀 out has a scatter consistent with 𝜆, with a similar halo mass dependence, and that both proxies contain unique information about the underlying halo mass. We find 𝜆-selected samples introduce features into the measured Δ⁢Σ signal that are not well fit by a log-normal scatter only model, absent in 𝑀 out selected samples. Our findings suggest that 𝑀 out offers an alternative for cluster selection with more easily calibrated selection biases, at least at the generally lower richnesses probed here. Combining both proxies may yield a mass proxy with a lower scatter and more tractable selection biases, enabling the use of lower mass clusters in cosmology. Finally, we find the scatter and slope in the 𝜆 −𝑀 out scaling relation to be 0.49 ±0.02 and 0.38 ±0.09.

79 ASTRONOMY AND ASTROPHYSICS

WigglyRivers: A tool to characterize the multiscale nature of meandering channels

Channel sinuosity is ubiquitous along river networks, producing complex patterns that encapsulate and influence morphodynamic processes and ecosystem services. Accurately characterizing these patterns is challenging with traditional curvature-based algorithms. Here, in this study, we present WigglyRivers, a Python package that builds on existing wavelet-based methods to create an unsupervised meander identification and characterization tool. The package uses planimetric information the user provides or from the USGS’s High-Resolution National Hydrography Dataset to characterize individual reaches or entire river networks. WigglyRivers also includes a supervised river identification tool for manually selecting individual meandering features. Here, we provide examples of idealized river transects and show the capabilities of WigglyRivers. We also use the supervised identification tool to validate the unsupervised identification on river transects across the continental US. WigglyRivers is a tool to understand better the multiscale characteristics of river networks and the link between river geomorphology and river corridor connectivity.

54 ENVIRONMENTAL SCIENCES