Search NASA⌕ Search

SEARCH · Search NASA

Results for “Feature importance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Interpretable Machine Learning for Molecular Biosignatures: a Novel Single-Sample Feature Importance Method That Is Sensitive To Statistical Interactions

Isotope ratio mass spectrometry (IRMS) of volatiles (e.g., CO 2 ) promises to be a powerful tool for potential biosignature detection for future missions to ocean worlds (OW) such as Europa and Enceladus. Machine learning (ML) methods for IRMS data could enable science autonomy by onboard prediction of seawater chemistry and biosignature presence. However, ML models are likely to be complex and involve statistical interactions between features (variables), which can make predictions seem opaque and enigmatic. For ML predictions as significant as extraterrestrial biosignatures, we must place extraordinary confidence in models. It is therefore essential that these models make interpretable predictions (i.e., human-understandable) and include false-prediction diagnostics. We achieve high accuracy and interpretability in ML biosignature and seawater chemistry models for OW through a nearest-neighbors feature selection tool that detects statistical interactions between predictors, constructs interaction networks for visualization of selected features working together to make a prediction, and reports single-sample feature importance scores for false-detection diagnostics. Here we develop a novel single-sample nearest-neighbors projected distance regression(ssNPDR) feature selection method that improves upon existing single-sample algorithms through the inclusion of statistical interactions while providing false-prediction diagnostics for ML models.

geochemistry↗

Algorithmic Detection of Elemental Biosignatures

Machine learning models that classify a sample as indicative or non-indicative of life could play an important role in life-detection missions. Their predictions result from agnostic algorithms and thereby add redundancy to judgements resulting from human expertise. Additionally, their important features can reveal the most informative measurements within the operational constraints of a life-detection mission. The Ladder of Life Detection (Neveu 2018) identifies the need for an understanding of how combinations of multiple biosignatures affect overall confidence. The present work provides a starting point to answer this need, and future work will expand the data types to obtain even more predictive combinations of features. Elemental abundance was chosen as a starting set of features due to its availability in diverse sample types, which are needed to train a generalizable model. A standardized dataset was collected, including 35 non-indicative, e.g., lunar rock, basalt; 19 indicative mixed, e.g., seawater, agricultural soil; 46 indicative non-alive, e.g., coal, chalk; and 10 indicative alive, e.g., biofilm, bacteria. This dataset could be valuable for complementary biosignature research. The samples were standardized to the same limit of detection of a simulated mission scenario. Four classification models were used: k-nearest neighbors (KNN), logistic regression (LR), linear support vector machines (SVM), and Gaussian naïve Bayes (GNB). To obtain feature importances, KNN was run on three principal components of the training data and LR and SVM were run with L1 and L2 regularization. The performances and feature importances of the six model variants on 40:60 train to validation ratios were assessed with Monte Carlo simulations. ROC AUC and mean accuracy scores ranged between 82% - 94%, with sensitivity greater than specificity. For indicative of life predictors, all models had C and Ca as strong and Cl as medium; a majority of models had N, K, and P as medium. For non-indicative of life predictors, all models had Si as strong, and a majority of models had Mg, Al, and Ti as medium. Varied elements were Fe (slightly non-indicative), H (slightly indicative), O (widely varied), Na, Mn, and S. These results serve as a proof of concept and suggest important elemental signals beyond merely the CHNOPS of Earth-based life.

Algorithmic↗

Effects of preprocessing Landsat MSS data on derived features

Important to the use of multitemporal Landsat MSS data for earth resources monitoring, such as agricultural inventories, is the ability to minimize the effects of varying atmospheric and satellite viewing conditions, while extracting physically meaningful features from the data. In general, the approaches to the preprocessing problem have been derived from either physical or statistical models. This paper compares three proposed algorithms; XSTAR haze correction, Color Normalization, and Multiple Acquisition Mean Level Adjustment. These techniques represent physical, statistical, and hybrid physical-statistical models, respectively. The comparisons are made in the context of three feature extraction techniques; the Tasseled Cap, the Cate Color Cube. and Normalized Difference.

Parris, T. M.↗

Artificial Neural Network (ANN) Surface Longwave and Shortwave Fluxes Trained on CERES Observations

The Clouds and Earth’s Radiant Energy System (CERES) project provides satellite-based observations of the radiative fluxes and clouds systems. CERES climate quality data products typically take several months of calibration and validation before release to the public. The Fast Longwave and Shortwave Radiative Flux (FLASHFlux) data product was developed to provide key data for the applied sciences and educational users within a week of observation. FLASHFlux achieves this by using simplified calibration, an operational meteorological product from Global Modeling and Assimilation Office (GMAO), and its own surface parameterizations model. The CERES FLASHFlux provides two data products: 1) an hourly Level 2 Single Scanner Footprint (SSF) data separately for Terra and NOAA-20 observations, and 2) a daily Level 3 Time Interpolated and Spatially Averaged (TISA) 1o x 1o gridded data that combines Terra and NOAA-20 observations. Currently, FLASHFlux uses the Langley Parameterized Shortwave Algorithm (LPSA) and Langley Parameterized Longwave Algorithm (LPLA) to derive its surface fluxes (Kratz et al., 2010; Gupta et al, 2001). A new Machine Learning (ML) based approach using Artificial Neural Networks to derive Surface Longwave (LW) & Shortwave (SW) fluxes based on training data from the CERES Clouds Radiative Swath (CRS) product is being investigated to replace LPSA and LPLA in the SSF surface flux products. One of the biggest hurdles in training ML model is model fitting. To overcome the problem of overfitting we use feature engineering that helps in finding the important feature and remove features that are irrelevant to the model. In our training we employed the Leave-One-Feature-Out Importance (LOFO) to evaluate the significance of each feature in our training. We intercompare ANN fluxes against surface fluxes produced from the Fu-Liou model in CRS and the LPSA/LPLA in FLASHFlux SSF. Furthermore, we validated ANN derived fluxes to the Baseline Surface Radiation Network (BSRN).

P C Sawaengphokhai↗

[ ] or SUCCESS is Not Enough: Current Technology and Future Directions in Proof Presentation

Automated theorem provers for first order logic are now around for several decades. Over the last few years, their deductive power to solve hard problems has increased tremendously. The annual CASC system competitions [Se97] give a clear picture of this situation. However, today's automated theorem provers are restricted "more by general usability than by raw deductive power." As a result of this, there are only very few serious applications of automated theorem provers. There are numerous features which a theorem prover lacks for real-world applicability. An automated theorem prover (as it is currently seen) is nothing more than a fast and elaborate search procedure. In that sense, an ATP can compared to a formulated race car, cool and fast, but virtually unusable for shopping groceries around the corner. Many important features are missing, or are optimized for speed rather than for applicability. [Schol] identifies important features which are needed for practical usability like detection of non-theorems, handling of modal/inductive proof tasks, control of the prover, and proof output. In this paper, we will focus solely on the last point, the presentation of the ATP's result to the user. In the rest of this paper, we will first discuss the general importance of providing feedback to the user, then we will describe the system ExplainIt!, a part of the deductive synthesis system AMPHION/NAV. In the conclusions we will relate proof presentation to other ways of post-processing a proof found by an ATP and stress their role in the future of automated deduction.

Schumann, Johann↗

Large eddy simulation of incompressible turbulent channel flow

The three-dimensional, time-dependent primitive equations of motion were numerically integrated for the case of turbulent channel flow. A partially implicit numerical method was developed. An important feature of this scheme is that the equation of continuity is solved directly. The residual field motions were simulated through an eddy viscosity model, while the large-scale field was obtained directly from the solution of the governing equations. An important portion of the initial velocity field was obtained from the solution of the linearized Navier-Stokes equations. The pseudospectral method was used for numerical differentiation in the horizontal directions, and second-order finite-difference schemes were used in the direction normal to the walls. The large eddy simulation technique is capable of reproducing some of the important features of wall-bounded turbulent flows. The resolvable portions of the root-mean square wall pressure fluctuations, pressure velocity-gradient correlations, and velocity pressure-gradient correlations are documented.

Moin, P.↗

Penetration of Magnetosheath Plasma into Dayside Magnetosphere: 1. Density, Velocity, and Rotation

In this study, we examine a large number of plasma structures (filaments), observed with the Cluster spacecraft during 2 years (2007-2008) in the dayside magnetosphere but consisting of magnetosheath plasma. To reduce the effects observed in the cusp regions and on magnetosphere flanks, we consider these events predominantly inside the narrow cone less than 30 about the subsolar point. Two important features of these filaments are (i) their stable antisunward (earthward) motion inside the magnetosphere, whereas the ambient magnetospheric plasma moves usually in the opposite direction (sunward), and (ii) between these filaments and the magnetopause, there is a region of magnetospheric plasma, which separates these filaments from the magnetosheath. The stable earthward motion of these magnetopause show the possible disconnection of these filaments from the magnetosheath, as suggested earlier by many researchers. The results also show that these events cannot be a result of back-and-forth motions of magnetopause position or surface waves propagating on the magnetopause. Another important feature of these filaments is their rotation about the filament axis, which might be a result of their passage through the velocity shear on magnetopause boundary. After crossing the velocity shear, the filaments get a rotational velocity, which has opposite directions in the noon-dusk and noon-dawn sectors. This rotation velocity may be an important factor, supporting the stability of these filaments and providing their motion into the magnetosphere.

Lyatsky, Wladislaw↗

Models of symbiotic stars

One of the most important features of symbiotic stars is the coexistence of a cool spectral component that is apparently very similar to the spectrum of a cool giant, with at least one hot continuum, and emission lines from very different stages of ionization. The cool component dominates the infrared spectrum of S-type symbiotics; it tends to be veiled in this wavelength range by what appears to be excess emission in D-type symbiotics, this excess usually being attributed to circumstellar dust. The hot continuum (or continua) dominates the ultraviolet. X-rays have sometimes also been observed. Another important feature of symbiotic stars that needs to be explained is the variability. Different forms occur, some variability being periodic. This type of variability can, in a few cases, strongly suggest the presence of eclipses of a binary system. One of the most characteristic forms of variability is that characterizing the active phases. This basic form of variation is traditionally associated in the optical with the veiling of the cool spectrum and the disappearance of high-ionization emission lines, the latter progressively appearing (in classical cases, reappearing) later. Such spectral changes recall those of novae, but spectroscopic signatures of the high-ejection velocities observed for novae are not usually detected in symbiotic stars. However, the light curves of the 'symbiotic nova' subclass recall those of novae. We may also mention in this connection that radio observations (or, in a few cases, optical observations) of nebulae indicate ejection from symbiotic stars, with deviations from spherical symmetry. We shall give a historical overview of the proposed models for symbiotic stars and make a critical analysis in the light of the observations of symbiotic stars. We describe the empirical approach to models and use the observational data to diagnose the physical conditions in the symbiotics stars. Finally, we compare the results of this empirical approach with existing models and discuss unresolved problems requiring new observational and theoretical work.

Friedjung, Michael↗

Some important geometrical features of conic-section-generated offset reflector antennas

Geometrical characteristics of conic-section-generated offset reflectors are studied in a unified fashion. Some unique geometrical features of the reflector rim constructed from the intersection of the reflector surface and a cone or cylinder are explored in detail. It is found that the intersection curve (rim) of the rotationally generated conic-section reflector surface and a circular cone with its tip at the focal point is always a planar curve and has a circular projection on the focal plane only for the offset parabolic reflector. Furthermore, in this case, the line going through the center of the circle, parallel to the focal axis, and the central axis of the cone do not intersect the reflector surface at the same point. Numerical results are presented to demonstrate some unique features of offset parabolic reflectors.

Jamnejad-Dailami, V.↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning (ML) Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Direct VLBI detection of the magnetosphere surrounding the young star S1 in Rho Ophiuchi

VLBI 6-mm data are presently used to investigate the circularly polarized radio core previously identified around the young B3 star S1 in Rho Ophiuchi. The measured angular diameter and brightness temperature are found to be consistent with gyrosynchrotrom radiation emission from mildly relativistic electrons. A simple model based on a pole-on dipolar magnetic field of about 2 kG at the stellar surface suggests itself as consistent with the main observed features of the S1 magnetosphere; an important feature of the model is its taking the influence of the X-ray-emitting plasma into account. S1 may represent a new type of young stellar object, characterized by very extended magnetic fields.

Andre, Philippe↗

Bayes classification of interferometric TOPSAR data

We report the Bayes classification of terrain types at different sites using airborne interferometric synthetic aperture radar (INSAR) data. A Gaussian maximum likelihood classifier was applied on multidimensional observations derived from the SAR intensity, the terrain elevation model, and the magnitude of the interferometric correlation. Training sets for forested, urban, agricultural, or bare areas were obtained either by selecting samples with known ground truth, or by k-means clustering of random sets of samples uniformly distributed across all sites, and subsequent assignments of these clusters using ground truth. The accuracy of the classifier was used to optimize the discriminating efficiency of the set of features that was chosen. The most important features include the SAR intensity, a canopy penetration depth model, and the terrain slope. We demonstrate the classifier's performance across sites using a unique set of training classes for the four main terrain categories. The scenes examined include San Francisco (CA) (predominantly urban and water), Mount Adams (WA) (forested with clear cuts), Pasadena (CA) (urban with mountains), and Antioch Hills (CA) (water, swamps, fields). Issues related to the effects of image calibration and the robustness of the classification to calibration errors are explored. The relative performance of single polarization Interferometric data classification is contrasted against classification schemes based on polarimetric SAR data.

Michel, T. R.↗

The Thematic Mapper Tasseled Cap - A preliminary formulation

A transformation is described which rotates Thematic Mapper data (excluding the thermal band) in a manner analogous to that used in the MSS Tasseled Cap Transformation, thus providing a direct view of the planes of data dispersion and a direct association of spectral features with physical scene characteristics. This TM Tasseled Cap Transformation includes MSS-equivalent Grenness and Brightness features, as well as at least one additional important feature. The new feature, tentatively termed 'Wetness', offers promise of enhanced ability to assess soil conditions, monitor vegetative development, and delineate cover classes. Relationships between scene characteristics and spectral variation in the transformed data space are discussed based on both simulated and actual TM data.

Crist, E. P.↗