Search NASA⌕ Search

SEARCH · Search NASA

Results for “learning classifiers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Ask-The-Expert: Minimizing Human Review for Big Data Analytics Through Active Learning

In this CIF project, we worked toward semi-automating knowledge discovery from anomaly detection algorithms through the use of active learning. Active learning is an area of research within machine learning that uses an "expert in the loop" to learn from large data sets that have very few annotations or labels available, and where providing such labels is expensive. In our case, the task can be defined as the identification of safety events from flight operational data. Since traditional anomaly detection algorithms cannot differentiate between operationally relevant and irrelevant statistical anomalies, Subject Matter Experts (SMEs) have a lengthy and expensive burden of investigating every example identified by the detection algorithm, classifying and labeling them as relevant or irrelevant. Active learningidentifies the unlabeled example for which a label would most improve the classifier, asks the domain expert for a label, and repeats this process until there are no more resources (time, budget) available for labeling or a minimum required performance is reached. A positive label indicates an operationally significant safety event whereas a negative label indicates otherwise. Based on these few labels we propose to build an active learning system that utilizes the SME's time in the most effective manner by iteratively asking for labels for as few informative instances as possible. Our work was proposed to be a stepping stone toward implementation and deployment of the system with user interface to be pursued by the Aviation Operations and Safety Program (AOSP) given its interest in safety monitoring and discovery of safety incidents.

aviation safety↗

Photointerpretation of LANDSAT images

Learning objectives include: (1) developing a facility for applying conventional techniques of photointerpretation to small scale (satellite) imager; (2) promoting the ability to locate, identify, and interpret small natural and man made surface features in a LANDSAT image; (3) using supporting imagery, such as aerial and space photography, to conduct specific applications analyses; (4) learning to apply change detection techniques to recognize and explain transient and temporal events in individual or seasonal imagery; (5) producing photointerpretation maps that define major surface units, themes, or classes; (6) classifying or analyzing a scene for specific discipline applications in geology, agriculture, forestry, hyrology, coastal wetlands, and environmental pollution; and (7) evaluating both advantages and shortcomings in relying on the photointerpretive approach (rather than computer based analytical approach) for extracting information from LANDSAT data.

Source record↗

Quantum Leap: Evaluating the Feasibility of Quantum Machine Learning Using NASA Earth Observational Data

This study explores the feasibility of leveraging quantum machine learning (QML) to analyze NASA Earth Observational (EO) data for climate change research, with a particular focus on the phenomenon of ”crop frosting” which has become more prevalent due to climate change. We implemented and evaluated two QML models, the Variational Quantum Classifier (VQC) and Quantum Support Vector Classifier (QSVC), in both simulated and real quantum computing environments using a 127 qubit IBM quantum processor. Our study emphasizes the scientific rigor in comparing these quantum models with a classical Support Vector Machine (SVM) classifier, highlighting their performance in processing climate data. The results offer valuable insights into the potential scientific advantages, limitations, and scalability of QML for analyzing EO datasets, thus paving the way for more advanced climate modeling and predictive analytics using quantum computing. We showcased how Environmental Interaction Knowledge Graphs (EIKGs) and Digital Twins (DTs) can be integrated into this study. This research underscores the transformative potential of Classical and QML leveraging KGs and DT to address the multifaceted challenges posed by climate change.

Quantum Computing↗

System diagnostic builder

The System Diagnostic Builder (SDB) is an automated software verification and validation tool using state-of-the-art Artificial Intelligence (AI) technologies. The SDB is used extensively by project BURKE at NASA-JSC as one component of a software re-engineering toolkit. The SDB is applicable to any government or commercial organization which performs verification and validation tasks. The SDB has an X-window interface, which allows the user to 'train' a set of rules for use in a rule-based evaluator. The interface has a window that allows the user to plot up to five data parameters (attributes) at a time. Using these plots and a mouse, the user can identify and classify a particular behavior of the subject software. Once the user has identified the general behavior patterns of the software, he can train a set of rules to represent his knowledge of that behavior. The training process builds rules and fuzzy sets to use in the evaluator. The fuzzy sets classify those data points not clearly identified as a particular classification. Once an initial set of rules is trained, each additional data set given to the SDB will be used by a machine learning mechanism to refine the rules and fuzzy sets. This is a passive process and, therefore, it does not require any additional operator time. The evaluation component of the SDB can be used to validate a single software system using some number of different data sets, such as a simulator. Moreover, it can be used to validate software systems which have been re-engineered from one language and design methodology to a totally new implementation.

Nieten, Joseph L.↗

Feature Acquisition with Imbalanced Training Data

This work considers cost-sensitive feature acquisition that attempts to classify a candidate datapoint from incomplete information. In this task, an agent acquires features of the datapoint using one or more costly diagnostic tests, and eventually ascribes a classification label. A cost function describes both the penalties for feature acquisition, as well as misclassification errors. A common solution is a Cost Sensitive Decision Tree (CSDT), a branching sequence of tests with features acquired at interior decision points and class assignment at the leaves. CSDT's can incorporate a wide range of diagnostic tests and can reflect arbitrary cost structures. They are particularly useful for online applications due to their low computational overhead. In this innovation, CSDT's are applied to cost-sensitive feature acquisition where the goal is to recognize very rare or unique phenomena in real time. Example applications from this domain include four areas. In stream processing, one seeks unique events in a real time data stream that is too large to store. In fault protection, a system must adapt quickly to react to anticipated errors by triggering repair activities or follow- up diagnostics. With real-time sensor networks, one seeks to classify unique, new events as they occur. With observational sciences, a new generation of instrumentation seeks unique events through online analysis of large observational datasets. This work presents a solution based on transfer learning principles that permits principled CSDT learning while exploiting any prior knowledge of the designer to correct both between-class and withinclass imbalance. Training examples are adaptively reweighted based on a decomposition of the data attributes. The result is a new, nonparametric representation that matches the anticipated attribute distribution for the target events.

Thompson, David R.↗

Data Mining for Anomaly Detection

The Vehicle Integrated Prognostics Reasoner (VIPR) program describes methods for enhanced diagnostics as well as a prognostic extension to current state of art Aircraft Diagnostic and Maintenance System (ADMS). VIPR introduced a new anomaly detection function for discovering previously undetected and undocumented situations, where there are clear deviations from nominal behavior. Once a baseline (nominal model of operations) is established, the detection and analysis is split between on-aircraft outlier generation and off-aircraft expert analysis to characterize and classify events that may not have been anticipated by individual system providers. Offline expert analysis is supported by data curation and data mining algorithms that can be applied in the contexts of supervised learning methods and unsupervised learning. In this report, we discuss efficient methods to implement the Kolmogorov complexity measure using compression algorithms, and run a systematic empirical analysis to determine the best compression measure. Our experiments established that the combination of the DZIP compression algorithm and CiDM distance measure provides the best results for capturing relevant properties of time series data encountered in aircraft operations. This combination was used as the basis for developing an unsupervised learning algorithm to define "nominal" flight segments using historical flight segments.

Biswas, Gautam↗

Lessons Learned: Mechanical Component and Tribology Activities in Support of Return to Flight

The February 2003 loss of the Space Shuttle Columbia resulted in NASA Management revisiting every critical system onboard this very complex, reusable space vehicle in a an effort to Return to Flight. Many months after the disaster, contact between NASA Johnson Space Center and NASA Glenn Research Center evolved into an in-depth assessment of the actuator drive systems for the Rudder Speed Brake and Body Flap Systems. The actuators are CRIT 1-1 systems that classifies them as failure of any of the actuators could result in loss of crew and vehicle. Upon further evaluation of these actuator systems and the resulting issues uncovered, several research activities were initiated, conducted, and reported to the NASA Space Shuttle Program Management. The papers contained in this document are the contributions of many researchers from NASA Glenn Research Center and Marshall Space Flight Center as part of a Lessons Learned on mechanical actuation systems as used in space applications. Many of the findings contained in this document were used as a basis to safely Return to Flight for the remaining Space Shuttle Fleet until their retirement.

Tribology↗

Reese Sorenson's Individual Professional Page

The subject document is a World Wide Web (WWW) page entitled, "Reese Sorenson's Individual Professional Page." Its can be accessed at "http://george.arc.nasa.gov/~sorenson/personal/index.html". The purpose of this page is to make the reader aware of me, who I am, and what I do. It lists my work assignments, my computer experience, my place in the NASA hierarchy, publications by me, awards received by me, my education, and how to contact me. Writing this page was a learning experience, pursuant to an element in my Job Description which calls for me to be able to use the latest computers. This web page contains very little technical information, none of which is classified or sensitive.

Sorenson, Reese↗

Application of Machine Learning Techniques to Delay Tolerant Network Routing

This dissertation discusses several machine learning techniques to improve routing in delay tolerant networks (DTNs). These are networks in which there may be long one-way trip times, asymmetric links, high error rates, and deterministic as well as non-deterministic loss of contact between network nodes, such as interplanetary satellite networks, mobile ad hoc networks and wireless sensor networks. This work uses historical network statistics to train a multi-label classifier to predict reliable paths through the network. In addition, a clustering technique is used to predict future mobile node locations. Both of these techniques are used to reduce the consumption of resources such as network bandwidth, memory and data storage that is required by replication routing methods often used in opportunistic DTN environments. Thesis contributions include: an emulation tool chain developed to create a DTN test bed for machine learning, the network and software architecture for a machine learning based routing method, the development and implementation of classification and clustering techniques and performance evaluation in terms of machine learning and routing metrics.

Dudukovich, Rachel M.↗

To Boldly Go: America's Next Era in Space. The Plasma Universe

Dr. France Cordova, NASA's Chief Scientist, chaired this, the eighth seminar in the Administrator's Seminar Series. She introduced the NASA Administrator, Daniel S. Goldin, who, in turn, introduced the subject of plasma. Plasma, an ionized gas, is a function of temperature and density. We ve learned that, at Jupiter, the radiation is dense. But, Goldin asked, what else do we know? Dr. Cordova then introduced Dr. James Van Allen, for whom the Van Allen radiation belt was named. Dr. Van Allen, a member of the University of Iowa faculty, discussed the growing interest in practical applications of space physics, including radiation fields and particles, plasmas and ionospheres. He listed a hierarchy of magnetic fields, beginning at the top, as pulsars, the Sun, planets, interplanetary medium, and interstellar medium. He pointed out that we have investigated eight of the nine known planets,. He listed three basic energy sources as 1) kinetic energy from flowing plasma such as constitutional solar wind or interstellar wind; 2) rotational energy of the planet, and 3) orbital energy of satellites. He believes there are seven sources of energetic particles and five potential places where particles may go. The next speaker, Dr. Ian Axford of New Zealand, has been associated with the Max Planck Institut fuer Aeronomie and plasma physics. He has studied solar and galactic winds and clusters of galaxies of which there are several thousand. He believes that the solar wind temperature is in the millions of degrees. The final speaker was Dr. Roger Blanford of the California Institute of Technology. He classified extreme plasmas as lab plasmas and cosmic plasmas. Cosmic plasmas are from supernovae remnants. These have supplied us with heavy elements and may come via a shock front of 10(sup 15) electron volts. To understand the physics of plasma, one must learn about x-rays, the maximum energy of acceleration by supernova remnants, particle acceleration and composition of cosmic rays, maximum acceleration, and how fast protons are heated by ions. He asked questions about where high energy cosmic rays are made, what accelerates electrons, radiates gamma rays, makes electronpositron plasma, and finally noted that pulsars are good time keepers, but we need a better understanding of their mechanism and of plasmas, both cosmic and ground-based. In the discussion period, Goldin asked if NASA should put up an x-ray interferometer. The answer was no; gamma rays are of greater interest just now. Goldin also asked what the assembled scientists would like to see for a future mission? They expressed an interest in learning more about the origin of galaxies, cosmic rays, solar systems, planets, the existence of life "out there", gamma ray sources, the nature of gamma ray bursts, and the flow of gases around black holes. The discussion concluded with a suggestion that NASA should communicate to the general public more information regarding actual technological trials and tribulations involved in getting an experiment to work. The speakers thought that this would help non-scientists to better appreciate what it is that NASA does in connection with the benefits that are achieved.

Source record↗

Remote Sensing of Lineage Functional Types for Modeling and Monitoring Biodiversity

Hyperspectral remote sensing has the potential to continuously scale plant function and plant diversity information from landscape to global extents. Numerous studies have indicated that VSWIR (400-2500 nm) reflectance properties of vegetation capture evolutionarily conserved biochemical, structural, and other functional attributes of plant species. Spectral properties conserved in plants provide the opportunity to both 1) aggregate species into lineages with improved classification accuracy and 2) link those lineages directly to plant traits. Full realization of this goal will enable parameterization of Land Surface Models (LSMs) with remotely sensed information, e.g., canopy nitrogen, and better representations of biodiversity and functional diversity in biogeographic studies. In this study, we use hyperspectral AVIRIS data from the 2013 HyspIRI campaign over the Southern Sierra Nevada, California flight box to investigate the potential for incorporating evolutionary thinking into landcover classification. We link the airborne hyperspectral data with vegetation plot data from roughly 1372 surveys and a phylogeny representing 1361 species. We aggregate species into lineages ranging from species level groups down to similar number of Plant Functional Types as often used in LSMs. We assessed the ability of Random Forest and Partial Least Squares Discriminant Analysis to discriminate across these different phylogenetic scales and determine the optimal number of lineages to classify. Although there are some temporal and spatial differences in our training data, our best approaches achieved moderate classification accuracy (Kappa > 0.65). Given an optimal number of lineages, we explored approaches to improve classifications including machine learning and unmixing approaches. This work suggests that lineage-based methods may be a promising way to leverage the huge amounts of data that will come from high resolution and high return interval hyperspectral data planned for the Surface Biology and Geology mission with sparsely sampled existing ground-based ecological data.

Hyperspectral↗

Development and Testing of Data Mining Algorithms for Earth Observation

The new algorithms developed under this project included a principled procedure for classification of objects, events or circumstances according to a target variable when a very large number of potential predictor variables is available but the number of cases that can be used for training a classifier is relatively small. These "high dimensional" problems require finding a minimal set of variables -called the Markov Blanket-- sufficient for predicting the value of the target variable. An algorithm, the Markov Blanket Fan Search, was developed, implemented and tested on both simulated and real data in conjunction with a graphical model classifier, which was also implemented. Another algorithm developed and implemented in TETRAD IV for time series elaborated on work by C. Granger and N. Swanson, which in turn exploited some of our earlier work. The algorithms in question learn a linear time series model from data. Given such a time series, the simultaneous residual covariances, after factoring out time dependencies, may provide information about causal processes that occur more rapidly than the time series representation allow, so called simultaneous or contemporaneous causal processes. Working with A. Monetta, a graduate student from Italy, we produced the correct statistics for estimating the contemporaneous causal structure from time series data using the TETRAD IV suite of algorithms. Two economists, David Bessler and Kevin Hoover, have independently published applications using TETRAD style algorithms to the same purpose. These implementations and algorithmic developments were separately used in two kinds of studies of climate data: Short time series of geographically proximate climate variables predicting agricultural effects in California, and longer duration climate measurements of temperature teleconnections.

Glymour, Clark↗

Application of High-Dimensional Fuzzy K-Means Cluster Analysis to CALIOP/CALIPSO Version 4.1 Cloud-Aerosol Discrimination

This study applies fuzzy k-means (FKM) cluster analyses to a subset of the parameters reported in the CALIPSO lidar level 2 data products in order to classify the layers detected as either clouds or aerosols. The results obtained are used to assess the reliability of the cloud–aerosol discrimination (CAD) scores reported in the version 4.1 release of the CALIPSO data products. FKM is an unsupervised learning algorithm, whereas the CALIPSO operational CAD algorithm (COCA) takes a highly supervised approach. Despite these substantial computational and architectural differences, our statistical analyses show that the FKM classifications agree with the COCA classifications for more than 94 % of the cases in the troposphere. This high degree of similarity is achieved because the lidar-measured signatures of the majority of the clouds and the aerosols are naturally distinct, and hence objective methods can independently and effectively separate the two classes in most cases. Classification differences most often occur in complex scenes (e.g., evaporating water cloud filaments embedded in dense aerosol) or when observing diffuse features that occur only intermittently (e.g., volcanic ash in the tropical tropopause layer). The two methods examined in this study establish overall classification correctness boundaries due to their differing algorithm uncertainties. In addition to comparing the outputs from the two algorithms, analysis of sampling, data training, performance measurements, fuzzy linear discriminants, defuzzification, error propagation, and key parameters in feature type discrimination with the FKM method are further discussed in order to better understand the utility and limits of the application of clustering algorithms to space lidar measurements. In general, we find that both FKM and COCA classification uncertainties are only minimally affected by noise in the CALIPSO measurements, though both algorithms can be challenged by especially complex scenes containing mixtures of discrete layer types. Our analysis results show that attenuated backscatter and color ratio are the driving factors that separate water clouds from aerosols; backscatter intensity, depolarization, and mid-layer altitude are most useful in discriminating between aerosols and ice clouds; and the joint distribution of backscatter intensity and depolarization ratio is critically important for distinguishing ice clouds from water clouds.

Zeng, Shan↗

Multilayer perceptron, fuzzy sets, and classification

A fuzzy neural network model based on the multilayer perceptron, using the back-propagation algorithm, and capable of fuzzy classification of patterns is described. The input vector consists of membership values to linguistic properties while the output vector is defined in terms of fuzzy class membership values. This allows efficient modeling of fuzzy or uncertain patterns with appropriate weights being assigned to the backpropagated errors depending upon the membership values at the corresponding outputs. During training, the learning rate is gradually decreased in discrete steps until the network converges to a minimum error solution. The effectiveness of the algorithm is demonstrated on a speech recognition problem. The results are compared with those of the conventional MLP, the Bayes classifier, and the other related models.

Pal, Sankar K.↗

Automated Cardiovascular Pathology Assessment using Semantic Segmentation and Ensemble Learning

Cardiac magnetic resonance imaging provides high spatial resolution, enabling improved extraction of important functional and morphological features for cardiovascular disease staging. Segmentation of ventricular cavities and myocardium in cardiac cine sequencing provides a basis to quantify cardiac measures such as ejection fraction. A method is presented that curtails the expense and observer bias of manual cardiac evaluation by combining semantic segmentation and disease classification into a fully automatic processing pipeline. The initial processing element consists of a robust dilated convolutional neural network architecture for voxel-wise segmentation of the myocardium and ventricular cavities. The resulting comprehensive volumetric feature matrix captures diagnostic clinical procedure data and is utilized by the final processing element to model a cardiac pathology classifier. Our approach evaluated anonymized cardiac images from a training data set of 100 patients (4 pathology groups, 1 healthy group, 20 patients per group) examined at the University Hospital of Dijon. The top average Dice index scores achieved were 0.940, 0.886, 0.849 for structure segmentation of the left ventricle (LV), myocardium and right ventricle (RV) respectively. A 5-ary pathology classification accuracy of 90% was recorded on an independent test set using the trained model. Performance results demonstrate potential for advanced machine learning methods to deliver accurate, efficient and reproducible cardiac pathological assessment.

Semantic Segmentation↗

Design and performance of a large vocabulary discrete word recognition system. Volume 1: Technical report

The development, construction, and test of a 100-word vocabulary near real time word recognition system are reported. Included are reasonable replacement of any one or all 100 words in the vocabulary, rapid learning of a new speaker, storage and retrieval of training sets, verbal or manual single word deletion, continuous adaptation with verbal or manual error correction, on-line verification of vocabulary as spoken, system modes selectable via verification display keyboard, relationship of classified word to neighboring word, and a versatile input/output interface to accommodate a variety of applications.

Source record↗

Mapping Global Forest Age from Forest Inventories, Biomass and Climate Data

Forest age can determine the capacity of a forest to uptake carbon from the atmosphere. However, a lack of global diagnostics that reflect the forest stage and associated disturbance regimes hampers the quantification of age-related differences in forest carbon dynamics. This study provides a new global distribution of forest age circa 2010, estimated using a machine learning approach trained with more than 40 000 plots using forest inventory, biomass and climate data. First, an evaluation against the plot-level measurements of forest age reveals that the data-driven method has a relatively good predictive capacity of classifying old-growth vs. non-old-growth (precision = 0.81 and 0.99 for old-growth and non-old-growth, respectively) forests and estimating corresponding forest age estimates (NSE = 0.6 – Nash–Sutcliffe efficiency – and RMSE = 50 years – root-mean-square error). However, there are systematic biases of overestimation in young- and underestimation in old-forest stands, respectively. Globally, we find a large variability in forest age with the old-growth forests in the tropical regions of Amazon and Congo, young forests in China, and intermediate stands in Europe. Furthermore, we find that the regions with high rates of deforestation or forest degradation (e.g. the arc of deforestation in the Amazon) are composed mainly of younger stands. Assessment of forest age in the climate space shows that the old forests are either in cold and dry regions or warm and wet regions, while young–intermediate forests span a large climatic gradient. Finally, comparing the presented forest age estimates with a series of regional products reveals differences rooted in different approaches and different in situ observations and global-scale products. Despite showing robustness in cross-validation results, additional methodological insights on further developments should as much as possible harmonize data across the different approaches. The forest age dataset presented here provides additional insights into the global distribution of forest age to better understand the global dynamics in the forest water and carbon cycles. The forest age datasets are openly available at https://doi.org/10.17871/ForestAgeBGI.2021 (Besnard et al., 2021).

Simon Besnard↗

Fast Feature-Recognizing Optoelectronic System

Proposed optoelectronic system recognizes features or classifies images by processing outputs of photosensors rapidly, in parallel, through circuits developed in research on neural networks. Array of photoconductive elements serve as photomodulated connections in electronic neural network, which provides high speed data compression to generate feature vector. System able to "learn" new patterns for subsequent recognition. Potential applications in robotic vision systems and pattern recognition.

Thakoor, S.↗