Search NASASearch

SEARCH · Search NASA

Results for “Feature selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Support Vector Machines for Hyperspectral Remote Sensing Classification

The Support Vector Machine provides a new way to design classification algorithms which learn from examples (supervised learning) and generalize when applied to new data. We demonstrate its success on a difficult classification problem from hyperspectral remote sensing, where we obtain performances of 96%, and 87% correct for a 4 class problem, and a 16 class problem respectively. These results are somewhat better than other recent results on the same data. A key feature of this classifier is its ability to use high-dimensional data without the usual recourse to a feature selection step to reduce the dimensionality of the data. For this application, this is important, as hyperspectral data consists of several hundred contiguous spectral channels for each exemplar. We provide an introduction to this new approach, and demonstrate its application to classification of an agriculture scene.

Gualtieri, J. Anthony

Software Assurance of PLCs Training Course

Being heavily visually-oriented, I am a firm believer in communication and conveying emotions through the art of color, motion, and transformation. A four-part online training course was created in PowerPoint and needed to be translated over into a Flash format. Issues with the Powerpoint were that the size of the files caused noticeable delays when placed online, there were compatibility issues, and from a composition and design perspective, color schemes and layout left much to be desired. High contrast, pixilated yellow text spiraling and flying on to a background of overly rich hues of blue with cheesy gradient patterns was just not appeasing to my eye, along with the menu directory buttons located at the top resembling blue pills of NyQuil on top of a stale gray border that had nothing to do with the background. The course itself is extremely broad and verbose, and will get monotonous very soon after starting. Moving about the course was very cumbersome as well. efficient by drastically reducing the size (The file size of all four parts of the actual course combined will ultimately not even be a fifth as big as one part of the original PowerPoint alone!); along with that, the course was made to be more interactive and user-friendly, as well as pleasing to the eye. Upon being viewed by fellow co-workers, nothing but positive feedback has been received. When beginning the presentation, onscreen comes a 3' 2" chubby, balding professor, who is a master in his knowledge of Programmable Logic Controllers (PLCs). He introduces himself and presents all of his vast knowledge over PLCs in a fun and innovative manner, making it much easier to acquire the information presented. A Scene Selection feature has been added making it a lot easier to jump from part to part and the back and forth arrows are much easier to utilize, and they are both less obtrusive than its PowerPoint predecessor. The user can also go at their pace, as the presentation pauses after at the end of each statement. computer animation, it was.. .still. ..animation done on a computer. I was able to incorporate my artistic talent and intuitive creativity into it, one thing I am very proficient at doing when it comes to what I do and what I will do in my profession as an artist/computer animator. At first, I felt that there was no place for an artist within a faculty of scientists, engineers, chemists, mathematicians, and programmers, but I managed to fit in quite successfully. My task was to convert the course into a Flash format, which would make it much more My project surprisingly somewhat dealt with my field of interest-though it was not

Threatt, Jamarr

Mapped Landmark Algorithm for Precision Landing

A report discusses a computer vision algorithm for position estimation to enable precision landing during planetary descent. The Descent Image Motion Estimation System for the Mars Exploration Rovers has been used as a starting point for creating code for precision, terrain-relative navigation during planetary landing. The algorithm is designed to be general because it handles images taken at different scales and resolutions relative to the map, and can produce mapped landmark matches for any planetary terrain of sufficient texture. These matches provide a measurement of horizontal position relative to a known landing site specified on the surface map. Multiple mapped landmarks generated per image allow for automatic detection and elimination of bad matches. Attitude and position can be generated from each image; this image-based attitude measurement can be used by the onboard navigation filter to improve the attitude estimate, which will improve the position estimates. The algorithm uses normalized correlation of grayscale images, producing precise, sub-pixel images. The algorithm has been broken into two sub-algorithms: (1) FFT Map Matching (see figure), which matches a single large template by correlation in the frequency domain, and (2) Mapped Landmark Refinement, which matches many small templates by correlation in the spatial domain. Each relies on feature selection, the homography transform, and 3D image correlation. The algorithm is implemented in C++ and is rated at Technology Readiness Level (TRL) 4.

Johnson, Andrew

Pattern Recognition for a Flight Dynamics Monte Carlo Simulation

The design, analysis, and verification and validation of a spacecraft relies heavily on Monte Carlo simulations. Modern computational techniques are able to generate large amounts of Monte Carlo data but flight dynamics engineers lack the time and resources to analyze it all. The growing amounts of data combined with the diminished available time of engineers motivates the need to automate the analysis process. Pattern recognition algorithms are an innovative way of analyzing flight dynamics data efficiently. They can search large data sets for specific patterns and highlight critical variables so analysts can focus their analysis efforts. This work combines a few tractable pattern recognition algorithms with basic flight dynamics concepts to build a practical analysis tool for Monte Carlo simulations. Current results show that this tool can quickly and automatically identify individual design parameters, and most importantly, specific combinations of parameters that should be avoided in order to prevent specific system failures. The current version uses a kernel density estimation algorithm and a sequential feature selection algorithm combined with a k-nearest neighbor classifier to find and rank important design parameters. This provides an increased level of confidence in the analysis and saves a significant amount of time.

Restrepo, Carolina

Optimizing Muscle Parameters in Musculoskeletal Modeling Using Monte Carlo Simulations

Astronauts assigned to long-duration missions experience bone and muscle atrophy in the lower limbs. The use of musculoskeletal simulation software has become a useful tool for modeling joint and muscle forces during human activity in reduced gravity as access to direct experimentation is limited. Knowledge of muscle and joint loads can better inform the design of exercise protocols and exercise countermeasure equipment. In this study, the LifeModeler(TM) (San Clemente, CA) biomechanics simulation software was used to model a squat exercise. The initial model using default parameters yielded physiologically reasonable hip-joint forces. However, no activation was predicted in some large muscles such as rectus femoris, which have been shown to be active in 1-g performance of the activity. Parametric testing was conducted using Monte Carlo methods and combinatorial reduction to find a muscle parameter set that more closely matched physiologically observed activation patterns during the squat exercise. Peak hip joint force using the default parameters was 2.96 times body weight (BW) and increased to 3.21 BW in an optimized, feature-selected test case. The rectus femoris was predicted to peak at 60.1% activation following muscle recruitment optimization, compared to 19.2% activation with default parameters. These results indicate the critical role that muscle parameters play in joint force estimation and the need for exploration of the solution space to achieve physiologically realistic muscle activation.

Hanson, Andrea

Optimized Algorithms for Prediction within Robotic Tele-Operative Interfaces

Robonaut, the humanoid robot developed at the Dexterous Robotics Laboratory at NASA Johnson Space Center serves as a testbed for human-robot collaboration research and development efforts. One of the primary efforts investigates how adjustable autonomy can provide for a safe and more effective completion of manipulation-based tasks. A predictive algorithm developed in previous work was deployed as part of a software interface that can be used for long-distance tele-operation. In this paper we provide the details of this algorithm, how to improve upon the methods via optimization, and also present viable alternatives to the original algorithmic approach. We show that all of the algorithms presented can be optimized to meet the specifications of the metrics shown as being useful for measuring the performance of the predictive methods. Judicious feature selection also plays a significant role in the conclusions drawn.

Martin, Rodney A.

Application of Machine Learning Techniques to Aviation Operations: NASA Case Studies

There is an increasing interest in applying methods based on Machine Learning Techniques(MLT) to problems in aviation operations. The current interest is based on developments in Cloud Computing, the availability of open software and the success of MLT in automation, consumer behavior and finance involving large database. Historically aviation operations have been analyzed using physics-based models and provide information for making operational decisions. This talk describes issues to be addressed in applying either model-driven or data-driven methods. Aviation operations involving many decision makers, multiple objectives, poor or unavailable physics-based models and a rich historical database are prime candidates for analysis using data-driven methods. The issues relating to data, feature selection and validation of the models are illustrated by examining case studies of the application of MLT to problems in air traffic management at NASA. Further research is needed in the application of MLT to critical aviation operations. As always, the best approach depends on the task, the physical understanding of the problem and the quality and quantity of the available data.

Sridhar, Banavar

Lessons Learned in the Application of Machine Learning Techniques to Air Traffic Management

There is an increasing interest in applying methods based on Machine Learning Techniques (MLT) to problems in Air Traffic Management (ATM). The current interest is based on developments in Cloud Computing, the availability of open software and the success of MLT in automation, consumer behavior and finance involving large databases. This paper reviews the current-state-of-the art in applying MLT to aviation operations, its promises and challenges. Historically aviation operations have been analyzed using physics-based models and provide information for making operational decisions. Aviation operations involving many decision makers, multiple objectives, poor or unavailable physics-based models and a rich historical database are prime candidates for analysis using data-driven methods. The promises and challenges in applying MLT to ATM is traced through three examples based on the authors’ experience, each separated by a decade, to show the influence of data and feature selection in the successful application of MLT to ATM. As always, the best approach depends on the task, the physical understanding of the problem and the quality and quantity of the available data.

Machine Learning Techniques

Intelligent multi-spectral IR image segmentation

We present a neural network based multi-­‐spectral image segmentation method. A neural network is trained on the selected features of both the objects and background in the longwave (LW) Infrared (IR) images. Multiple iterations of training are performed until the accuracy of the segmentation reaches satisfactory level. The segmentation boundary of the LW image is used to segment the midwave (MW) and shortwave (SW) IR images. A second neural network detects the local discontinuities and refines the accuracy of the local boundaries. The neural net based segmentation method is compared with Wavelet-­‐threshold and Grab-­‐Cut methods. Test results have shown increased accuracy and robustness of this segmentation scheme for multi-­‐spectral IR images.

Torres, Gilbert

Advancing Methodologies for Applying Machine Learning and Evaluating Spatiotemporal Models of Fine Particulate Matter (PM 2.5 ) Using Satellite Data Over Large Regions

Reconstructing the distribution of fine particulate matter (PM 2.5 ) in space and time, even far from ground monitoring sites, is an important exposure science contribution to epidemiologic analyses of PM 2.5 health impacts. Flexible statistical methods for prediction have demonstrated the integration of satellite observations with other predictors, yet these algorithms are susceptible to overfitting the spatiotemporal structure of the training datasets. We present a new approach for predicting PM 2.5 using machine-learning methods and evaluating prediction models for the goal of making predictions where they were not previously available. We apply extreme gradient boosting (XGBoost) modeling to predict daily PM 2.5 on a 1 x 1 km 2 resolution for a 13 state region in the Northeastern USA for the years 2000–2015 using satellite-derived aerosol optical depth and implement a recursive feature selection to develop a parsimonious model. We demonstrate excellent predictions of withheld observations but also contrast an RMSE of 3.11 μg/m 3 in our spatial cross-validation withholding nearby sites versus an overfit RMSE of 2.10 μg/m 3 using a more conventional random ten-fold splitting of the dataset. As the field of exposure science moves forward with the use of advanced machine-learning approaches for spatiotemporal modeling of air pollutants, our results show the importance of addressing data leakage in training, overfitting to spatiotemporal structure, and the impact of the predominance of ground monitoring sites in dense urban sub-networks on model evaluation. The strengths of our resultant modeling approach for exposure in epidemiologic studies of PM 2.5 include improved efficiency, parsimony, and interpretability with robust validation while still accommodating complex spatiotemporal relationships.

air pollution

Assistive Relative Pose Estimation for On-orbit Assembly using Convolutional Neural Networks

Accurate real-time pose estimation of spacecraft or object in space is a key capability necessary for on orbit spacecraft servicing and assembly tasks. Pose estimation of objects in space is more challenging than for objects on Earth due to space images containing widely varying illumination conditions, high contrast, and poor resolution in addition to power and mass constraints. In this paper, a convolutional neural network is leveraged to uniquely determine the translation and rotation of an object of interest relative to the camera. The main idea of using CNN model is to assist object tracker used in on space assembly tasks where only feature based method is always not sufficient. The simulation framework designed for assembly task is used to generate dataset for training the modified CNN models and, then results of different models are compared with measure of how accurately models are predicting the pose. Unlike many current approaches for spacecraft or object in space pose estimation, the model does not rely on hand-crafted object-specific features which makes this model more robust and easier to apply to other types of spacecraft. It is shown that the model performs comparable to the current feature-selection methods and can therefore be used in conjunction with them to provide more reliable estimates.

Sonawani, Shubham

Revealing the Mysteries of Venus: The DAVINCI Mission

The Deep Atmosphere Venus Investigation of Noble gases, Chemistry, and Imaging (DAVINCI) mission described herein has been selected for flight to Venus as part of the NASA Discovery Program. DAVINCI will be the first mission to Venus to incorporate science driven flybys and an instrumented descent sphere into a unified architecture. The anticipated scientific outcome will be a new understanding of the atmosphere, surface, and evolutionary path of Venus as a possibly once-habitable planet and analog to hot terrestrial exoplanets. The primary mission design for DAVINCI as selected features a preferred launch in summer/fall 2029, two flybys in 2030, and descent sphere atmospheric entry by the end of 2031. The in situ atmospheric descent phase subsequently delivers definitive chemical and isotopic composition of the Venus atmosphere during an atmospheric transect above Alpha Regio. These in situ investigations of the atmosphere and near infrared descent imaging of the surface will complement remote flyby observations of the dynamic atmosphere, cloud deck, and surface near infrared emissivity. The overall mission yield will be at least 60 Gbits (compressed) new data about the atmosphere and near surface, as well as the first unique characterization of the deep atmosphere environment and chemistry, including trace gases, key stable isotopes, oxygen fugacity, constraints on local rock compositions, and topography of a tessera.

James B Garvin

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), has typically limited machine learning (ML) in space studies and further study of radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNAseq) data from 6 mouse liver GeneLab datasets (GLDS) with a total of 113 spaceflight and ground-control samples to determine top features relevant to spaceflight including the effect of radiation exposure. Data was normalized within each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. The top MRMR features were used to predict spaceflight vs. ground-control samples using a Random Forest (RF) classifier with 5-fold cross validation (CV). The ML-based gene sets were further compared against differential gene expression results from individual GLDS. CV training using the top 100 MRMR genes show averages of 86% accuracy and 0.95 AUC value on the validation set over 5 folds (Figure 1A). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 811 or 68 DEGs overlapping between at least 2 or 3 studies, respectively (Figure 1B). Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism. Set analysis between the MRMR features and the DEGs showed 60 or 8 genes overlapping with at least 1 or 2 studies, respectively. MRMR feature selection and ensemble ML methods (e.g. RF) improve performance relative to a Naïve Bayes classifier when NGS data sets are analyzed. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise ratio. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from RNASeq analysis. Non-intersecting sets introduce opportunity to explore spaceflight relevant genes and implementing ML methods across existing NGS datasets may overcome sample size limitations. ML coupled with existing analytical methods enhances understanding of disease by revealing common underlying pathways across datasets.

Machine Learning

Machine Learning Emulators and Empirical Models Combining Climate and Global Crop Models for Seasonal Agricultural Production

We present results from several connected efforts to apply machine learning methods to estimates of seasonal agricultural production anomalies around the world. First, we apply the XGBoost Random Forest method to fit emulators that mimic global crop models participating in the Agricultural Model Intercomparison and Improvement Project (AgMIP) Global Gridded Crop Model Intercomparison (GGCMI). These are the same models used in the agricultural sector simulations of the Inter-Sectoral Impacts Model Intercomparison Project (ISIMIP). These emulators use 8 climate variables split across 5 sub-seasonal representations of the growing season for each ½ degree grid cell around the world for maize, wheat, rice and soybeans. Emulators are useful for estimating conditions that have not already been simulated by GGCMI (e.g., in a seasonal prediction model) and also to diagnose model differences and capabilities. For example, emulators of the pDSSAT maize model tend to be more reliant on mean temperatures than the LPJmL model, and few models have strong responses to cold extremes. Second, we use a similar XGBoost approach to fit empirical models for national production data for the top 20 producing countries according to the United Nations Food and Agricultural Organization (FAO). Models utilize both climate observations and the GGCM models as predictors, resulting in skillful models for many (but not all) top producing-countries. The patterns of climate and crop model features selected indicate regions and systems that are better or worse simulated by the GGCMs. For example, information in cold extreme predictors is often combined with GGCM output predictors to provide sensitivity that models may underrepresent.

machine learning

Analysis of Waste Material Feedstocks Using Laser-Induced Breakdown Spectroscopy and Machine Learning

Predicting properties such as heating value, ash fusion temperature, and mineral ash composition from Laser-Induced Breakdown Spectroscopy (LIBS) data can make gasifiers more flexible to different feedstocks. Understanding these feedstock properties in-situ improves feedstock conversion modelling methods that allow for consistent operation, higher carbon conversion, and reduced fouling and erosion rates. The purpose of this study is to demonstrate methods for model creation that take LIBS data as predictor features and estimate higher order material properties as a function of feedstock material properties. Six samples were chosen to represent a mixture of abundant and carbon rich waste materials. LIBS measurements were performed on these samples for elemental wavelengths and intensity values. Laboratory analytical results were obtained for each sample’s heating value, proximate and ultimate analysis, mineral ash composition, ash fusion temperatures, and viscosity temperatures. Thermal conductivity was measured using a HotDisk TPS 2500S. LIBS measurements were processed and used as predictor features for machine learning (ML) models to predict the sample’s material properties. Predictor feature selection algorithms, particularly minimum redundancy maximum relevance (mRMR), reduced the dimensionality of ML models. Many modelling methods such as Gaussian process regression (GPR), regression tree, neural networks (NN), and support vector machines (SVM) were demonstrated to be effective at predicting higher order properties; however, mRMR with GPR stood out as a clear winning combination.

01 COAL, LIGNITE, AND PEAT

Development of a quantitative basis for selection of spectral features in a vegetation monitoring system

The development of an objective methodology for evaluation of alternative Landsat data preprocessing options, spectral transform features for monitoring vegetation, and feature summarization algorithms is presented. Based on estimates of spectral separability between a target class and its confusion classes, analysis of variance techniques are used to evaluate potential design options for large scale vegetation monitoring systems. Case studies are presented for early season and through the season spring small grains separation and for barley/other spring small grains separation. It is concluded that a basis for efficient, objective selection among alternative feature extraction approaches has been established for the large scale vegetation mapping/inventory problem. Although the approach has been demonstrated for the unitemporal class separability case, extensions to the multitemporal case are under development.

Phinney, D. E.

An Evaluation of optional timing/synchronization features to support selection of an optimum design for the DCS digital communication network

The task was to evaluate the ability of a set of timing/synchronization subsystem features to provide a set of desirable characteristics for the evolving Defense Communications System digital communications network. The set of features related to the approaches by which timing/synchronization information could be disseminated throughout the network and the manner in which this information could be utilized to provide a synchronized network. These features, which could be utilized in a large number of different combinations, included mutual control, directed control, double ended reference links, independence of clock error measurement and correction, phase reference combining, and self organizing.

Bradley, D. B.

Measurement of the Splashback Feature Around SZ-Selected Galaxy Clusters With DES, SPT, and ACT

We present a detection of the splashback feature around galaxy clusters selected using the Sunyaev–Zel’dovich (SZ) signal. Recent measurements of the splashback feature around optically selected galaxy clusters have found that the splashback radius, rsp, is smaller than predicted by N-body simulations. A possible explanation for this discrepancy is that rsp inferred from the observed radial distribution of galaxies is affected by selection effects related to the optical cluster-finding algorithms. We test this possibility by measuring the splashback feature in clusters selected via the SZ effect in data from the South Pole Telescope SZ survey and the Atacama Cosmology Telescope Polarimeter survey. The measurement is accomplished by correlating these cluster samples with galaxies detected in the Dark Energy Survey Year 3data. The SZ observable used to select clusters in this analysis is expected to have a tighter correlation with halo mass and to be more immune to projection effects and aperture-induced biases, potentially ameliorating causes of systematic error for optically selected clusters. We find that the measured rsp for SZ-selected clusters is consistent with the expectations from simulations, although the small number of SZ-selected clusters makes a precise comparison difficult. In agreement with previous work, when using optically selected red MaPPer clusters with similar mass and redshift distributions,rspis∼2σsmaller than in the simulations. These results motivate detailed investigations of selection biases in optically selected cluster catalogues and exploration of the splashback feature around larger samples of SZ-selected clusters. Additionally, we investigate trends in the galaxy profile and splashback feature as a function of galaxy colour, finding that blue galaxies have profiles close to a power law with no discernible splashback feature, which is consistent with them being on their first in fall into the cluster.

T Shin