Search NASASearch

SEARCH · Search NASA

Results for “classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Addressing the dynamic nature of reference data: a new nucleotide database for robust metagenomic classification

Accurate metagenomic classification relies on comprehensive, up-to-date, and validated reference databases. While the NCBI BLAST Nucleotide (nt) database, encompassing a vast collection of sequences from all domains of life, represents an invaluable resource, its massive size—currently exceeding 10 12 nucleotides—and exponential growth pose significant challenges for researchers seeking to maintain current nt-based indices for metagenomic classification. Recognizing that no current nt-based indices exist for the widely used Centrifuge classifier, and the last public version currently available was released in 2018, we addressed this critical gap by leveraging advanced high-performance computing resources. We present new Centrifuge-compatible nt databases, meticulously constructed using a novel pipeline incorporating different quality control measures, including reference decontamination and filtering. These measures demonstrably reduce spurious classifications, as shown through our reanalysis of published metagenomic data where Plasmodium annotations were dramatically reduced using our decontaminated database, highlighting how database quality can significantly impact research conclusions. Through temporal comparisons, we also reveal how our approach minimizes inconsistencies in taxonomic assignments stemming from asynchronous updates between public sequence and taxonomy databases. These discrepancies are particularly evident in taxa such as Listeria monocytogenes and Naegleria fowleri, where classification accuracy varied significantly across database versions. These new databases, made available as pre-built Centrifuge indexes, respond to the need for an open, robust, nt-based pipeline for taxonomic classification in metagenomics. Applications such as environmental metagenomics, forensics, and clinical metagenomics, which require comprehensive taxonomic coverage, will benefit from this resource. Our work highlights the importance of treating reference databases as dynamic entities, subject to ongoing quality control and validation akin to software development best practices. This approach is crucial for ensuring accuracy and reliability of metagenomic analysis, especially as databases continue to expand in size and complexity.

59 BASIC BIOLOGICAL SCIENCES

On the nature of global classification

Molecular sequencing technology has brought biology into the era of global (universal) classification. Methodologically and philosophically, global classification differs significantly from traditional, local classification. The need for uniformity requires that higher level taxa be defined on the molecular level in terms of universally homologous functions. A global classification should reflect both principal dimensions of the evolutionary process: genealogical relationship and quality and extent of divergence within a group. The ultimate purpose of a global classification is not simply information storage and retrieval; such a system should also function as an heuristic representation of the evolutionary paradigm that exerts a directing influence on the course of biology. The global system envisioned allows paraphyletic taxa. To retain maximal phylogenetic information in these cases, minor notational amendments in existing taxonomic conventions should be adopted.

NASA Program Exobiology

An Active Learning Framework for Hyperspectral Image Classification Using Hierarchical Segmentation

Augmenting spectral data with spatial information for image classification has recently gained significant attention, as classification accuracy can often be improved by extracting spatial information from neighboring pixels. In this paper, we propose a new framework in which active learning (AL) and hierarchical segmentation (HSeg) are combined for spectral-spatial classification of hyperspectral images. The spatial information is extracted from a best segmentation obtained by pruning the HSeg tree using a new supervised strategy. The best segmentation is updated at each iteration of the AL process, thus taking advantage of informative labeled samples provided by the user. The proposed strategy incorporates spatial information in two ways: 1) concatenating the extracted spatial features and the original spectral features into a stacked vector and 2) extending the training set using a self-learning-based semi-supervised learning (SSL) approach. Finally, the two strategies are combined within an AL framework. The proposed framework is validated with two benchmark hyperspectral datasets. Higher classification accuracies are obtained by the proposed framework with respect to five other state-of-the-art spectral-spatial classification approaches. Moreover, the effectiveness of the proposed pruning strategy is also demonstrated relative to the approaches based on a fixed segmentation.

classification

Modernizing NASA's Risk Classification System

NASA's risk classification system dates back to an era when every new NASA space mission was a one-of-a-kind build, and the only way to obtain reliability was as a by-product through a combination of reliability analyses, extensive and stringent quality requirements, and extensive testing. Originally, there were very limited commercial capabilities to develop systems to work reliably in space, so NASA considered its own homegrown approach the only recipe for success. This approach involved very detailed and prescriptive piece-part controls and no reliance on (and to some extent a rejection of) any type of commercial practices. Often risk was considered to be the lowest when NASA had the maximum amount of control and prescription, and the highest when commercial practices were largely employed, and these principles drove risk classification in the agency. Over time, however, commercial capabilities grew, and many products became standardized and commercialized, while the agency maintained its tried-and-true approach, paying little attention to the evolution of the commercial sector. In fact, the commercial sector was developing systems that have direct, proven reliability, established over time, while NASA still maintained the approach to ignore the reality of the commercialized aspects of standard products, label them as high risk, and attempt to change them to align with the agency's piece-part control practices. A table of mission classification vs lifetime for missions launched after 2000 indicates no correlation between lifetime and classification, with the few exceptions involving missions that have very limited objectives and no valid purpose to continue after they were met. This paper steps through some of the key historical elements in risk classification and NASA's overall approach to assurance, and presents some elements being brought forward to modernize the approach and take advantage of the growing capability in the commercial sector.

risk

Hierarchical Mixture of Experts for Advanced Air Mobility Flight Phase Classification

Advanced Air Mobility (AAM) and Urban Air Mobility (UAM) operations will have numerous vehicles and aircraft flying in the airspace, which poses safety and security concerns. Commercial airlines utilize Air Traffic Management (ATM) and Air Traffic Control (ATC) for real-time monitoring, surveillance, traffic coordination, and rerouting to maintain safe and efficient flight patterns. Transferring ATM and ATC architectures to AAM/UAM will be difficult to implement since AAM/UAM aircraft fly at lower altitudes, have more static and dynamic obstacles, operate in highly dense environments, and have several more aircraft to monitor for a given volume of the national airspace (NAS). Automatic flight phase classification will enhance efficiencies of ATM/ATC-like architectures for AAM/UAM. Classifying the main flight phases (takeoff, climb, cruise, descent, and landing) provides insight to ensure safe operations, provide situational awareness of the NAS, and monitor flights in case there are any emergencies. Typical flight phase classification methods are all-or-nothing, which will not capture or accurately classify the transitions between flight phases. Utilizing hierarchical mixture of experts (HME) provides a flight phase classification solution that includes transitions between the flight phases by assigning weights based on ground-based distributed sensor readings from cameras and radar. Adding the transitions between flight phases increases the fidelity of flight phase classification and provides deeper insight for flight phase classification by leveraging distributed sensing concepts.

distributed sensing

Are light curve classification metrics good proxies for SN Ia cosmological constraining power?

Context. When selecting a light curve classifier for use as part of a photometric supernova Ia (SN Ia) cosmological analysis, it is common to make decisions based on metrics of classification performance, such as the contamination within the photometrically classified SN Ia sample, rather than a measure of cosmological constraining power. If the former is an appropriate proxy for the latter, this practice would eliminate the computational expense of a full cosmology forecast in the analysis pipeline design process. Aims. This study tests the assumption that light curve classification metrics are an appropriate proxy for cosmology metrics. Methods. We emulated photometric SN Ia cosmology light curve samples with controlled contamination rates of individual contaminant classes and evaluated each of them under a set of classification metrics. We then derived cosmological parameter constraints from all samples under two common analysis approaches and quantified the impact of contamination by each contaminant class on the resulting cosmological parameter estimates. Results. We observe that cosmology metrics are sensitive to both the contamination rate and the class of the contaminating population, whereas the classification metrics are shown to be insensitive to the latter. Conclusions. Based on these findings, we discourage any exclusive reliance on light curve classification-based metrics for analysis design decisions, which (counterintuitively) include but are not limited to the classifier choice. Instead, we recommend optimising science analysis pipeline design choices using a metric of the information gained about the physical parameters of interest.

79 ASTRONOMY AND ASTROPHYSICS

AutoSourceID-Classifier: Star-galaxy classification using a convolutional neural network with spatial information

Aims.Traditional star-galaxy classification techniques often rely on feature estimation from catalogs, a process susceptible to introducing inaccuracies, thereby potentially jeopardizing the classification’s reliability. Certain galaxies, especially those not manifesting as extended sources, can be misclassified when their shape parameters and flux solely drive the inference. We aim to create a robust and accurate classification network for identifying stars and galaxies directly from astronomical images. Methods.The AutoSourceID-Classifier (ASID-C) algorithm developed for this work uses 32x32 pixel single filter band source cutouts generated by the previously developed AutoSourceID-Light (ASID-L) code. By leveraging convolutional neural networks (CNN) and additional information about the source position within the full-field image, ASID-C aims to accurately classify all stars and galaxies within a survey. Subsequently, we employed a modified Platt scaling calibration for the output of the CNN, ensuring that the derived probabilities were effectively calibrated, delivering precise and reliable results. Results.We show that ASID-C, trained on MeerLICHT telescope images and using the Dark Energy Camera Legacy Survey (DECaLS) morphological classification, is a robust classifier and outperforms similar codes such as SourceExtractor. To facilitate a rigorous comparison, we also trained an eXtreme Gradient Boosting (XGBoost) model on tabular features extracted by SourceExtractor. While this XGBoost model approaches ASID-C in performance metrics, it does not offer the computational efficiency and reduced error propagation inherent in ASID-C’s direct image-based classification approach. ASID-C excels in low signal-to-noise ratio and crowded scenarios, potentially aiding in transient host identification and advancing deep-sky astronomy.

Astronomy & Astrophysics

MicroFisher: Fungal taxonomic classification for metatranscriptomic and metagenomic data using multiple short hypervariable markers

AbstractProfiling the taxonomic and functional composition of microbes using metagenomic (MG) and metatranscriptomic (MT) sequencing is advancing our understanding of microbial functions. However, the sensitivity and accuracy of microbial classification using genome– or core protein-based approaches, especially the classification of eukaryotic organisms, is limited by the availability of genomes and the resolution of sequence databases. To address this, we propose the MicroFisher, a novel approach that applies multiple hypervariable marker genes to profile fungal communities from MGs and MTs. This approach utilizes the hypervariable regions of ITS and large subunit (LSU) rRNA genes for fungal identification with high sensitivity and resolution. Simultaneously, we propose a computational pipeline (MicroFisher) to optimize and integrate the results from classifications using multiple hypervariable markers. To test the performance of our method, we applied MicroFisher to the synthetic community profiling and found high performance in fungal prediction and abundance estimation. In addition, we also used MGs from forest soil and MTs of root eukaryotic microbes to test our method and the results showed that MicroFisher provided more accurate profiling of environmental microbiomes compared to other classification tools. Overall, MicroFisher serves as a novel pipeline for classification of fungal communities from MGs and MTs.

Wang, Haihua

Grid Edge Waveform Analytics Framework for Event Detection and Classification

This paper provides a grid edge waveform analytics framework for power system event detection and classification in the local as well as in the wide area. This framework overviews data excellence for event detection and classification. The data excellence describes the data acquisition process and requirements, data processing, data quality, and data integrity. Power system event detection in the local area based on different features such as energy-based, cyclostationary approach, template matching, and wavelet transform are also discussed. Furthermore, local area event detection and classification using approaches such as statistical, signal processing, artificial intelligence, and hybrid are also discussed. Moreover, an overview of wide-area event detection and classification along with several other aspects such as wide-area events, wide-area event detection approaches, event location and system performance, event pattern recognition, inter-area oscillation, and wide-area frequency response under variable deployment of inverter-based resources are also provided. The proposed framework is the first step toward the goal of developing appropriate tools and methodologies to detect and classify local as well as wide-area events using waveform analytics. The appropriate event detection and classification framework development is especially important now as more and more grid edge devices with communication capabilities are being deployed in the modern power grid than ever before.

Bhusal, Narayan

Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events

Danovo Energy Solution's presented its paper named: Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events at the 2026 Georgia Tech Fault & Disturbance Analysis Conference. The full paper can be found at OSTI ID# 3169150 Paper Abstract—Phasor Measurement Units (PMUs) stream time synchronized, high-resolution measurements from the grid, enabling data-driven techniques for event detection and classification. Accurate event classification improves grid reliability and stability. Events can be detected by varying numbers of PMUs and exhibit different durations depending on the event type. This variability challenges standard classifiers that require uniform input sizes. Moreover, multiple events may coincide, which increases classification complexity. Standard classifiers assign each instance to the class with the highest predicted probability, whereas overlapping events may exhibit comparable probabilities across multiple classes. In this study, to handle data size variability, we extract a wide range of time–frequency domain features from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, LightGBM, Support Vector Machine, and Multilayer Perceptron. To account for overlapping events, a probabilistic post-processing step is applied. For a given data instance, if multiple predicted class probabilities exceed 30% and the differences between them are less than 10%, the event is assigned to multiple classes. Experiments using real-world PMU data demonstrate that the Random Forest and XGBoost models achieve the highest accuracy, while the proposed post-processing method yields perfect classification performance on external unseen test sets.

Nematirad, Reza [Danova Energy Solutions]

The minimum distance approach to classification

The work to advance the state-of-the-art of miminum distance classification is reportd. This is accomplished through a combination of theoretical and comprehensive experimental investigations based on multispectral scanner data. A survey of the literature for suitable distance measures was conducted and the results of this survey are presented. It is shown that minimum distance classification, using density estimators and Kullback-Leibler numbers as the distance measure, is equivalent to a form of maximum likelihood sample classification. It is also shown that for the parametric case, minimum distance classification is equivalent to nearest neighbor classification in the parameter space.

Wacker, A. G.

Application of LANDSAT images to wetland study and land use classification in west Tennessee, part 1

The author has identified the following significant results. densitometric analysis was performed on LANDSAT data to permit numerical classification of objects observed in the imagery on the basis of measurements of optical density. Relative light transmission measurements were taken on four types of scene elements in each of three LANDSAT black and white bands in order to determine which classification could be distinguished. The analysis of band 6 determined forest and agricultural classifications, but not the urban and wetlands. Both bands 4 and 5 showed a significant difference existed between the confirmed classification of wetlands-agriculture, and urban areas. Therefore, the combination of band 6 with either 4 or 5 would permit the separation of the urban from the wetland classification. To enhance the urban and wetland boundaries, the LANDSAT black and white bands were combined in a multispectral additive color viewer. Several combinations of filters and light intensities were used to obtain maximum discrimination between points of interest. The best results for enhancing wetland boundaries and urban areas were achieved by using a color composite (a blue, green, and red filter on bands 4, 5 and 6 respectively).

Shahrokhi, F.

LANDSAT applications to wetlands classification in the upper Mississippi River Valley

A 25% improvement in average classification accuracy was realized by processing double-date vs. single-date data. Under the spectrally and spatially complex site conditions characterizing the geographical area used, further improvement in wetland classification accuracy is apparently precluded by the spectral and spatial resolution restrictions of the LANDSAT MSS. Full scene analysis of scanning densitometer data extracted from scale infrared photography failed to permit discrimination of many wetland and nonwetland cover types. When classification of photographic data was limited to wetland areas only, much more detailed and accurate classification could be made. The integration of conventional image interpretation (to simply delineate wetland boundaries) and machine assisted classification (to discriminate among cover types present within the wetland areas) appears to warrant further research to study the feasibility and cost of extending this methodology over a large area using LANDSAT and/or small scale photography.

Lillesand, T. M.

An initial analysis of LANDSAT 4 Thematic Mapper data for the classification of agricultural, forested wetland, and urban land covers

An initial analysis of LANDSAT 4 thematic mapper (TM) data for the delineation and classification of agricultural, forested wetland, and urban land covers was conducted. A study area in Poinsett County, Arkansas was used to evaluate a classification of agricultural lands derived from multitemporal LANDSAT multispectral scanner (MSS) data in comparison with a classification of TM data for the same area. Data over Reelfoot Lake in northwestern Tennessee were utilized to evaluate the TM for delineating forested wetland species. A classification of the study area was assessed for accuracy in discriminating five forested wetland categories. Finally, the TM data were used to identify urban features within a small city. A computer generated classification of Union City, Tennessee was analyzed for accuracy in delineating urban land covers. An evaluation of digitally enhanced TM data using principal components analysis to facilitate photointerpretation of urban features was also performed.

Quattrochi, D. A.

Improving crop classification through attention to the timing of airborne radar acquisitions

Radar remote sensors may provide valuable input to crop classification procedures because of (1) their independence of weather conditions and solar illumination, and (2) their ability to respond to differences in crop type. Manual classification of multidate synthetic aperture radar (SAR) imagery resulted in an overall accuracy of 83 percent for corn, forest, grain, and 'other' cover types. Forests and corn fields were identified with accuracies approaching or exceeding 90 percent. Grain fields and 'other' fields were often confused with each other, resulting in classification accuracies of 51 and 66 percent, respectively. The 83 percent correct classification represents a 10 percent improvement when compared to similar SAR data for the same area collected at alternate time periods in 1978. These results demonstrate that improvements in crop classification accuracy can be achieved with SAR data by synchronizing data collection times with crop growth stages in order to maximize differences in the geometric and dielectric properties of the cover types of interest.

Brisco, B.

Analysis of Landsat-4 Thematic Mapper data for classification of forest stands in Baldwin County, Alabama

A computer-implemented classification has been derived from Landsat-4 Thematic Mapper data acquired over Baldwin County, Alabama on January 15, 1983. One set of spectral signatures was developed from the data by utilizing a 3x3 pixel sliding window approach. An analysis of the classification produced from this technique identified forested areas. Additional information regarding only the forested areas. Additional information regarding only the forested areas was extracted by employing a pixel-by-pixel signature development program which derived spectral statistics only for pixels within the forested land covers. The spectral statistics from both approaches were integrated and the data classified. This classification was evaluated by comparing the spectral classes produced from the data against corresponding ground verification polygons. This iterative data analysis technique resulted in an overall classification accuracy of 88.4 percent correct for slash pine, young pine, loblolly pine, natural pine, and mixed hardwood-pine. An accuracy assessment matrix has been produced for the classification.

Hill, C. L.

Impact of Thematic Mapper Sensor Characteristics on Classification Accuracy

A fixed effect, three factor (two levels per factor) analysis of variance was used to quantitatively assess the significance of the improved spectral, spatial and radiometric resolution capabilities of the LANDSAT-4 thematic mapper sensor relative to the familiar MSS sensor. TM data acquired over the Washington, D.C. area were progressively degraded in spectral, spatial and radiometric characteristics to simulate the MSS, and classification accuracies were derived in a consistent manner for all eight treatments in the ANOVA design. Statistical testing of the significance of differences in classification accuracies between treatments indicated that the increased number of spectral bands and the improved quantization capabilities afforded by the TM sensor design would lead to significant improvements in classification accuracies attainable relative to MSS. In contrast, however, the improved spatial resolution provided by the TM sensor did not enhance classification accuracy. This latter result was felt to be more a function of the type of classification algorithms available.

Williams, D. L.

Effect of Landsat Thematic Mapper sensor parameters on land cover classification

Selected sensor parameter differences between TM and MSS were assessed through classification performance of a suburban/regional test site. Overall classification accuracy of a seven-band Landsat TM scene in comparison to MSS yielded an improvement in accuracy from 74.8 percent to 83.2 percent. To study the possible causes for the difference in classification performance, key sensor parameter differences between MSS and TM, including: (1) spatial resolution (30 m for TM versus 80 m for MSS), (2) quantization level (256 levels for TM versus 64 for MSS), and (3) spectral regions (seven bands in four major spectral regions for TM versus four bands in two regions for MSS), were evaluated. Landsat TM data were processed to stimulate all possible combinations of these MSS and TM parameters, yielding a three-factor design with two levels per factor. The results indicated that the added spectral regions (TM 1, TM 5, and TM 7) and to a lesser degree the increase in quantization level to eight bits produced the improved TM classification accuracy. However, in this study, the higher 30 m spatial resolution of TM contributed to a reduced classification accuracy from increased within-field variability or class heterogeneity.

Toll, D. L.