Search NASA⌕ Search

SEARCH · Search NASA

Results for “read classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Modularization of EDGE Workflows Using Nextflow: Improving the Efficiency and Maintainability of Bioinformatics Software

EDGE is a bioinformatics platform developed in 2016 by researchers at Los Alamos National Laboratory (LANL) to facilitate the analysis of next-generation sequencing data by researchers with varying levels of experience in bioinformatics (Li et al., 2017). Users with single-end, paired-end or long-read sequencing data can provide their reads as input to EDGE and select the combination of workflows to run that are most useful for their research (e.g., quality control of reads, genome assembly, or the taxonomic classification of input reads). Table 1 summarizes the modules available in EDGE. EDGE is available as a web platform at https://edgebioinformatics.org, as installable source code maintained on GitHub under a GPLv3 license, and as a publicly hosted Docker image.

59 BASIC BIOLOGICAL SCIENCES↗

DL-TODA: A Deep Learning Tool for Omics Data Analysis

Metagenomics is a technique for genome-wide profiling of microbiomes; this technique generates billions of DNA sequences called reads. Given the multiplication of metagenomic projects, computational tools are necessary to enable the efficient and accurate classification of metagenomic reads without needing to construct a reference database. The program DL-TODA presented here aims to classify metagenomic reads using a deep learning model trained on over 3000 bacterial species. A convolutional neural network architecture originally designed for computer vision was applied for the modeling of species-specific features. Using synthetic testing data simulated with 2454 genomes from 639 species, DL-TODA was shown to classify nearly 75% of the reads with high confidence. The classification accuracy of DL-TODA was over 0.98 at taxonomic ranks above the genus level, making it comparable with Kraken2 and Centrifuge, two state-of-the-art taxonomic classification tools. DL-TODA also achieved an accuracy of 0.97 at the species level, which is higher than 0.93 by Kraken2 and 0.85 by Centrifuge on the same test set. Application of DL-TODA to the human oral and cropland soil metagenomes further demonstrated its use in analyzing microbiomes from diverse environments. Compared to Centrifuge and Kraken2, DL-TODA predicted distinct relative abundance rankings and is less biased toward a single taxon.

59 BASIC BIOLOGICAL SCIENCES↗

Enhanced read resolution in reconfigurable memristive synapses for Spiking Neural Networks

Abstract The synapse is a key element circuit in any memristor-based neuromorphic computing system. A memristor is a two-terminal analog memory device. Memristive synapses suffer from various challenges including high voltage, SET or RESET failure, and READ margin issues that can degrade the distinguishability of stored weights. Enhancing READ resolution is very important to improving the reliability of memristive synapses. Usually, the READ resolution is very small for a memristive synapse with a 4-bit data precision. This work considers a step-by-step analysis to enhance the READ current resolution or the read current difference between two resistance levels for a current-controlled memristor-based synapse. An empirical model is used to characterize the $${\hbox {HfO}}_{2}$$ HfO 2 based memristive device. $$1\textrm{st}$$ 1 st and $$2\textrm{nd}$$ 2 nd stage device of our proposed synapse design can be scaled to enhance the READ current margin up to $$\sim$$ ∼ 4.3 $$\times$$ × and $$\sim$$ ∼ 21%, respectively. Moreover, READ current resolution can be enhanced with run-time adaptation techniques such as READ voltage scaling and body biasing. The READ voltage scaling and body biasing can improve the READ current resolution by about 46% and 15%, respectively. TENNLab’s neuromorphic computing framework is leveraged to evaluate the effect of READ current resolution on classification, control, and reservoir computing applications. Higher READ current resolution shows better accuracy than lower resolution even when facing different levels of read noise.

97 MATHEMATICS AND COMPUTING↗

A Decision Support System to Compile Environmental Mitigations from Hydropower Licensing Documents

The process of deciphering, extracting, and compiling information from texts dense with domain-specific terminology and technical jargon is a challenging endeavor. It demands considerable expertise and deep knowledge in the respective field, resulting in a labor-intensive process when executed by humans. Furthermore, the task of identifying multiple class labels in extensive texts presents a challenge due to intra- and inter-reader variability, making the process time-consuming and costly.We’re introducing a user-friendly graphical interface, fortified with a BERT model-powered decision support system. This advanced system aims to augment efficiency, curtail data collection time, and sustain high precision in data acquisition. It is instrumental in deciphering and synthesizing intricate texts teeming with a spectrum of expressions, even within similar mitigation categories. Such tasks traditionally demand substantial human effort and specialized knowledge in the domain.Our system is specifically engineered for the task of extracting environmental mitigation information to promote sustainable hydropower development from licenses issued by the Federal Energy Regulatory Commission (FERC). These license documents are comprehensive, each containing over 15,000 words and requiring the identification of 135 different class labels. We anticipate that our system will boost reading speed, improve the consistency of classification outputs among readers, and contribute to the development of a robust scientific database of environmental mitigations associated with the 2,000+ non-federal hydropower facilities licensed by FERC in the United States.

Yoon, Hong-Jun [ORNL] (ORCID:0000000254505878)↗

Crop classification using airborne radar and Landsat data

NASA 13.3 GHz airborne radar data from a soil moisture measurement analysis is used to investigate the statistical nature of the radar backscattering coefficient for bare ground and three different crop types, and to evaluate the crop classification rates using Landsat data alone or combined with the airborne survey. The scatterometer was a fan-beam Doppler system, VV polarized, and is considered only for 50 deg angles of incidence. A total of 36 fields were covered a week apart by the aircraft and Landsat, and Rayleigh statistics were used in the frequency averaging to eliminate fluctuations due to random fluctuations. Within-field variances were calculated for the Landsat and the radar data and used to design optimum crop classification procedures. The Landsat Band 4 readings were 67% accurate, and an increase in accuracy of 10% was achieved by the addition of the radar data.

Ulaby, F. T.↗

The Land Use and Land Cover Dichotomy: A Comparison of Two Land Classification Systems in Support of Urban Earth Science Applications

One is likely to read the terms 'land use' and 'land cover' in the same sentence, yet these concepts have different origins and different applications. Land cover is typically analyzed by earth scientists working with remotely sensed images. Land use is typically studied by urban planners who must prescribe solutions that could prevent future problems. This apparent dichotomy has led to different classification systems for land-based data. The works of earth scientists and urban planning practitioners are beginning to come together in the field of spatial analysis and in their common use of new spatial analysis technology. In this context, the technology can stimulate a common 'language' that allows a broader sharing of ideas. The increasing amount of land use and land cover change challenges the various efforts to classify in ways that are efficient, effective, and agreeable to all groups of users. If land cover and land uses can be identified by remote methods using aerial photography and satellites, then these ways are more efficient than field surveys of the same area. New technology, such as high-resolution satellite sensors, and new methods, such as more refined algorithms for image interpretation, are providing refined data to better identify the actual cover and apparent use of land, thus effectiveness is improved. However, the closer together and the more vertical the land uses are, the more difficult the task of identification is, and the greater is the need to supplement remotely sensed data with field study (in situ). Thus, a number of land classification methods were developed in order to organize the greatly expanding volume of data on land characteristics in ways useful to different groups. This paper distinguishes two land based classification systems, one developed primarily for remotely sensed data, and the other, a more comprehensive system requiring in situ collection methods. The intent is to look at how the two systems developed and how they can work together so that land based information can be shared among different users and compared over time.

McAllister, William K.↗

Hierarchical Mixture of Experts for Advanced Air Mobility Flight Phase Classification

Advanced Air Mobility (AAM) and Urban Air Mobility (UAM) operations will have numerous vehicles and aircraft flying in the airspace, which poses safety and security concerns. Commercial airlines utilize Air Traffic Management (ATM) and Air Traffic Control (ATC) for real-time monitoring, surveillance, traffic coordination, and rerouting to maintain safe and efficient flight patterns. Transferring ATM and ATC architectures to AAM/UAM will be difficult to implement since AAM/UAM aircraft fly at lower altitudes, have more static and dynamic obstacles, operate in highly dense environments, and have several more aircraft to monitor for a given volume of the national airspace (NAS). Automatic flight phase classification will enhance efficiencies of ATM/ATC-like architectures for AAM/UAM. Classifying the main flight phases (takeoff, climb, cruise, descent, and landing) provides insight to ensure safe operations, provide situational awareness of the NAS, and monitor flights in case there are any emergencies. Typical flight phase classification methods are all-or-nothing, which will not capture or accurately classify the transitions between flight phases. Utilizing hierarchical mixture of experts (HME) provides a flight phase classification solution that includes transitions between the flight phases by assigning weights based on ground-based distributed sensor readings from cameras and radar. Adding the transitions between flight phases increases the fidelity of flight phase classification and provides deeper insight for flight phase classification by leveraging distributed sensing concepts.

distributed sensing↗

Hierarchical Mixture of Experts for Advanced Air Mobility Flight Phase Classification

Advanced Air Mobility (AAM) and Urban Air Mobility (UAM) operations will have numerous vehicles and aircraft flying in the airspace, which poses safety and security concerns. Commercial airlines utilize Air Traffic Management (ATM) and Air Traffic Control (ATC) for real-time monitoring, surveillance, traffic coordination, and rerouting to maintain safe and efficient flight patterns. Transferring ATM and ATC architectures to AAM/UAM will be challenging to implement since AAM/UAM aircraft fly at lower altitudes, have more static and dynamic obstacles, operate in highly dense environments, and have several more aircraft to monitor for a given volume of the national airspace (NAS). Aircraft typically have the following flight phases: takeoff, climb, cruise, descent, and landing. Classifying these flight phases provides insight into ensuring safe operations, providing situational awareness of the NAS, and monitoring flights in emergencies. Automatic flight phase classification will enhance the efficiencies of ATM/ATC-like architectures for AAM/UAM, especially since numerous aircraft will be flying in highly dense urban environments. Typical flight phase classification methods are all-or-nothing, which will not capture or accurately classify the transitions between flight phases. Utilizing hierarchical mixture of experts (HME) provides a flight phase classification solution that includes transitions between the flight phases by assigning weights based on ground-based distributed sensor readings from cameras and radar. Adding the transitions between flight phases increases the fidelity of flight phase classification and provides deeper insight into flight phase classification by leveraging distributed sensing concepts. Simulation results and post-processed flight test results demonstrate the utility of HME for automatic and robust flight phase classification for real-time AAM operations.

distributed sensing↗

Coordination and establishment of centralized facilities and services for the University of Alaska ERTS survey of the Alaskan environment

The author has identified the following significant results. Specifications have been prepared for the engineering design and construction of a digital color display unit which will be used for automatic processing of ERTS data. The color display unit is a disk refresh memory with computer interfaced input and a color cathode ray tube output display. The system features both analog and digital post disk data manipulation and a versatile color coding device suitable for displaying not only images, but also computer generated graphics such as diagrams, maps, and overlays. Input is from IBM compatible 9 track, 800 BPI tapes, as generated by an IBM 360 computer. ERTS digital tapes are read into the 360, where various analyses such as maximum likelihood classification are performed and the results are written on a magnetic tape which is the input to the color display unit. The greatest versatility in the data manipulation area is provided by the minicomputer built into the color display unit, which is off-line from the main 360 computer. The minicomputer is able to read any line from the refresh disk and place it in its 4K-16 bit memory. Considerable flexibility is available for post-processing enhancement of images by the investigator.

Belon, A. E.↗

Flood Mapping of Recent Major Hurricane Events with Synthetic Aperture Radar, Commercial Imaging, and Aerial Observations

Floodwater mapping is an important remote sensing process that is used for disaster response, recovery, and damage assessment practices. Developing a system to read in Synthetic Aperture Radar (SAR) data and perform land cover classification will allow for the production of near real-time inundation mapping, enabling government and emergency response entities to get a preliminary idea of the situation. SAR is a unique remote sensing tool. Data in this project was obtained by NASA Jet Propulsion Laboratory’s Uninhabited Aerial Vehicle SAR (UAVSAR), an L-band radar mounted to a Gulfstream III jet. Data collected by UAVSAR is similar to what will be available from the NASA-Indian Space Research Organization (NISAR) mission starting in early 2022. Using Python and ArcGIS applications, a model was developed using training samples taken from NOAA post-event aerial photography and UAVSAR data gathered in the aftermath of Hurricane Florence in September 2018.

Melancon, Alexander M.↗

Enhancing and Archiving the APS Catalog of the POSS I

We have worked on two different projects: 1) Archiving the APS Catalog of the POSS I for distribution to NASA's NED at IPAC, SIMBAD in France, and individual astronomers and 2) The automated morphological classification of galaxies. We have completed archiving the Catalog into easily readable binary files. The database together with the software to read it has been distributed on DVD's to the national and international data centers and to individual astronomers. The archived Catalog contains more than 89 million objects in 632 fields in the first epoch Palomar Observatory Sky Survey. Additional image parameters not available in the original on-line version are also included in the archived version. The archived Catalog is also available and can be queried at the APS web site (URL: http://aps.umn.edu) which has been improved with a much faster and more efficient querying system. The Catalog can be downloaded as binary datafiles with the source code for reading it. It is also being integrated into the SkyQuery system which includes the Sloan Digital Sky Survey, 2MASS, and the FIRST radio sky survey. We experimented with different classification algorithms to automate the morphological classification of galaxies. This is an especially difficult problem because there are not only a large number of attributes or parameters and measurement uncertainties, but also the added complication of human disagreement about the adopted types. To solve this problem we used 837 galaxy images from nine POSS I fields at the North Galactic Pole classified by two independent astronomers for which they agree on the morphological types. The initial goal was to separate the galaxies into the three broad classes relevant to issues of large scale structure and galaxy formation and evolution: early (ellipticals and lenticulars), spirals, and late (irregulars) with an accuracy or success rate that rivals the best astronomer classifiers. We also needed to identify a set of parameters derived from the digitized images that separate the galaxies by type. The human eye can easily recognize complicated patterns in images such as spiral arms which can be spotty, blotchy affairs that are difficult for automated techniques. A galaxy image can potentially be described by hundreds of parameters, all of which may have some relation to the morphological type. In the set of initial experiments we used 624 such parameters, in two colors, blue and red. These parameters include the surface brightness and color measured at different radii, ratios of these parameters at different radii, concentration indices, Fourier transforms and wavelet decomposition coefficients. We experimented with three different classes of classification algorithms; decision trees, k-nearest neighbors, and support vector machines (SVM). A range of experiments were conducted and we eventually narrowed the parameters to 23 selected parameters. SVM consistently outperformed the other algorithms with both sets of features. By combining the results from the different algorithms in a weighted scheme we achieved an overall classification success of 86%.

Humphreys, Roberta M.↗

Automated RF Phase Adjustment for Beam Stabilization in the Fermilab Linac

The Fermilab Linac experiences longitudinal beam phase drift, leading to increased particle loss, conventionally corrected through labor-intensive manual RF adjustments. This project explores machine learning-based automation for drift correction, employing a prototype-based classification approach. Our model utilizes a 34-dimensional feature set (RF settings and BPM readings) and leverages a 7x27 response matrix for system modeling. To overcome limited real-world data, we generate synthetic data, enhancing model training and generalizability. Custom loss functions, including a surrogate energy-consistent loss and a temporal smoothness constraint, ensure physically plausible drift predictions. The goal is a robust system for autonomous phase adjustments, ensuring stable beam acceleration and reduced manual intervention.

Chichili, R. R. [Illinois U., Chicago]↗

Automated RF Phase Adjustment for Beam Stabilization in the Fermilab Linac

The Fermilab Linac experiences longitudinal beam phase drift, leading to increased particle loss, conventionally cor- rected through labor-intensive manual RF adjustments. This project explores machine learning-based automation for drift correction, employing a prototype-based classification approach. Our model utilizes a 34-dimensional feature set (RF settings and BPM readings) and leverages a 7x27 response matrix for system modeling. To overcome limited real-world data, we generate synthetic data, enhancing model training and generalizability. Custom loss functions, including a sur- rogate energy-consistent loss and a temporal smoothness constraint, ensure physically plausible drift predictions. The goal is a robust system for autonomous phase adjustments, ensuring stable beam acceleration and reduced manual intervention.

Chichili, R. R. [U. Illinois, Chicago]↗

Numerical classification of coding sequences

DNA sequences coding for protein may be represented by counts of nucleotides or codons. A complete reading frame may be abbreviated by its base count, e.g. A76C158G121T74, or with the corresponding codon table, e.g. (AAA)0(AAC)1(AAG)9 ... (TTT)0. We propose that these numerical designations be used to augment current methods of sequence annotation. Because base counts and codon tables do not require revision as knowledge of function evolves, they are well-suited to act as cross-references, for example to identify redundant GenBank entries. These descriptors may be compared, in place of DNA sequences, to extract homologous genes from large databases. This approach permits rapid searching with good selectivity.

Non-NASA Center↗

Two novel Patescibacteria: Phycocordibacter aenigmaticus gen. nov. sp. nov. and Minusculum obligatum gen. nov. sp. nov., both associated with microalgae optimized for carbon dioxide sequestration from flue gas

The functional roles of bacterial symbionts associated with microalgae remain understudied despite the importance of microalgae in biotechnology and environmental microbiology. 16S rRNA gene sequencing was conducted to analyze bacterial communities associated with two microalgae optimized for growth with flue gas containing 5%–10% CO 2 . Two dominant bacteria with no taxonomic classification beyond the class level (Paceibacteria) were discovered repeatedly in the most productive algal cultures. Long-read metagenomic sequencing was conducted to yield high-quality metagenomes, from which two novel species were discovered under the Seqcode (seqco.de/r:ywe1blo2), Phycocordibacter aenigmaticus gen. nov. sp. nov. and Minusculum obligatum gen. nov. sp. nov. The genus Phycocordibacter gen. nov. was proposed as the nomenclatural type of the family Phycocordibacteraceae fam. nov. and the order Phycocordibacterales ord. nov. Both bacteria possessed features typical of Patescibacteria such as reduced genomes (<800 kbp), lack of complete glycolysis and tricarboxylic acid (TCA) cycle pathways, and inability to synthesize amino acids. Instead, they rely on the reductive pentose phosphate pathway (Calvin cycle) for essential biosynthesis and redox balance. P. aenigmaticus may also rely on elemental sulfur oxidation (sdo), partial nitrite reduction (nirK), and sulfur-related amino acid metabolism (SAMe → SAH). Both bacteria were found in high relative abundance in cultures of Tetradesmus obliquus HTB1 (freshwater) and Nannochloropsis oceanica IMET1 (marine), suggesting a tight association with microalgae in various environments. The absence of full metabolic pathways for energy production suggests extreme metabolic limitations and obligate symbiosis, most likely with other bacteria associated with the microalgae.

54 ENVIRONMENTAL SCIENCES↗

The use of ERTS imagery for lake classification

The feasibility of using photographic representations of the ERTS imagery to classify lakes in the State of Wisconsin as to their trophic level was studied. Densitometric readings in band 5 of ERTS 70 mm imagery were taken for all the lakes in Wisconsin greater than 100 acres (approximately 1000 lakes). An algorithm has been developed from ground truth measurements to predict from satellite imagery an indicator of trophic status.

Scarpace, F. L.↗

Investigation of correlation classification techniques

A two-step classification algorithm for processing multispectral scanner data was developed and tested. The first step is a single pass clustering algorithm that assigns each pixel, based on its spectral signature, to a particular cluster. The output of that step is a cluster tape in which a single integer is associated with each pixel. The cluster tape is used as the input to the second step, where ground truth information is used to classify each cluster using an iterative method of potentials. Once the clusters have been assigned to classes the cluster tape is read pixel-by-pixel and an output tape is produced in which each pixel is assigned to its proper class. In addition to the digital classification programs, a method of using correlation clustering to process multispectral scanner data in real time by means of an interactive color video display is also described.

Haskell, R. E.↗

As-built design specification for PARCLS

The PARCLS program, part of the CLASFYG package, reads a parameter file created by the CLASFYG program and a pure pixel ground truth file in order to create to classification file of three separate crop categories in universal format.

Tompkins, M. A.↗