Search NASA⌕ Search

SEARCH · Search NASA

Results for “Algorithm Classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43

Wire Detection Algorithms for Navigation

In this research we addressed the problem of obstacle detection for low altitude rotorcraft flight. In particular, the problem of detecting thin wires in the presence of image clutter and noise was studied. Wires present a serious hazard to rotorcrafts. Since they are very thin, their detection early enough so that the pilot has enough time to take evasive action is difficult, as their images can be less than one or two pixels wide. Two approaches were explored for this purpose. The first approach involved a technique for sub-pixel edge detection and subsequent post processing, in order to reduce the false alarms. After reviewing the line detection literature, an algorithm for sub-pixel edge detection proposed by Steger was identified as having good potential to solve the considered task. The algorithm was tested using a set of images synthetically generated by combining real outdoor images with computer generated wire images. The performance of the algorithm was evaluated both, at the pixel and the wire levels. It was observed that the algorithm performs well, provided that the wires are not too thin (or distant) and that some post processing is performed to remove false alarms due to clutter. The second approach involved the use of an example-based learning scheme namely, Support Vector Machines. The purpose of this approach was to explore the feasibility of an example-based learning based approach for the task of detecting wires from their images. Support Vector Machines (SVMs) have emerged as a promising pattern classification tool and have been used in various applications. It was found that this approach is not suitable for very thin wires and of course, not suitable at all for sub-pixel thick wires. High dimensionality of the data as such does not present a major problem for SVMs. However it is desirable to have a large number of training examples especially for high dimensional data. The main difficulty in using SVMs (or any other example-based learning method) is the need for a very good set of positive and negative examples since the performance depends on the quality of the training set.

Kasturi, Rangachar↗

Towards an Aviation Large Language Model by Fine-tuning and Evaluating Transformers

In the aviation domain, there are many applications for machine learning and artificial intelligence tools that utilize natural language. For example, there is a desire to know the commonalities in written safety reports such as voluntary post incidents reports or aerial wildfire operations reports to better understand the risks present. Another use-case is the possibility of extracting airspace procedures and constraints currently written in documents such as Letters of Agreement. These applications can benefit from the use of state-of-the-art natural language processing techniques when adapted to the language/phraseology specific to the aviation domain. This paper evaluates the viability of adaptation of NLP tools to the aviation domain by fine-tuning transformer based models using aviation data sets. In 2018, a novel language model based on neural units (also called transformers) was created and became known as “Bidirectional Encoder Representations from Transformers” or BERT. This architecture combined with large amounts of English training data and innovative semi-supervised training tasks set the standard for what would later emerge as Large Language Models. The performance of these models was further improved by hyperparameter tuning and refinement of the semi-supervised training task and resulted in “Robustly Optimized BERT Pre-training Approach through hyperparameter tuning” or RoBERTa models. These pre-trained Large Language Models proved to be useful for a wide variety of natural language processing tasks such as text classification and question answering through a process called fine-tuning. The transformer architecture with pre-trained weights served as the basis with the last few layers replaced with layers fine-tuned to perform a new task e.g., a layer that provides a label for the entire input text. This process of fine-tuning can also be used to adapt the models to new domains; e.g., BioBERT started with the pre-trained BERT model and was completed by additional fine-tuning and training on biomedical documents. Transformer-based architectures can also be used to create rich representations of text called embeddings which can serve as the input to other machine learning models. This allows simpler algorithms such as logistic regression to use context-rich representations of the text while still remaining quick to train and evaluate. In the world of aviation, there is a growing demand for natural language processing and understanding but the domain presents unique challenges. Due to the technical content (and specialized language) of most aviation documents, fine-tuning pre-trained Large Language Models to specific tasks has not met the benchmark on natural language processing tasks set by simpler models trained from scratch on the data. To address this deficiency, this paper evaluates the improvements from fine-tuning a Large Language Model on a large set of aviation documents using the original semi-supervised training tasks before performing specific natural language tasks. In fine-tuning, a domain-specific dataset is used on the original training task but with the pre-trained Large Language Model instead of starting from a random initialization. This approach allows the model to be adapted to the specific domain language without discarding the information gained from training on general English data. This paper utilized two major dataset types to train and assess the RoBERTa fine-tuning performance. The first are 7,057 Letters of Agreement which are Federal Aviation Administration (FAA) documents that formalize airspace operations across the national airspace system. They contain many examples of ‘aviation English’ using domain specific terminology and phrasing which serves as a representative basis to perform the semi-supervised fine-tuning. The second type is the 494 document classification labels to be used for evaluation. This down-stream evaluation aims to show the performance of the fine-tuned model, better understand how much data is needed for an effective fine-tuning, and how fine-tuning can be adapted for different applications in-the domain. After semi-supervised training, evaluation begins by encoding the documents for classification using the fine-tuned RoBERTa model. Then a logistic regression classifier is trained to label the document type and compared against our ground truth labels. This currently leads to a 82.8% accuracy on 10-fold cross validation showing improvement over baseline RoBERTa which achieved 81.0%. We plan to measure the improvements on additional tasks and it is expected that these improvements will lead to more robust models that can tackle the natural language processing challenges present in aviation datasets.

ATM↗

Science Autonomy for Ocean Worlds Astrobiology: A Perspective

Astrobiology missions to ocean worlds in our solar system must overcome both scientific and technological challenges due to extreme temperature and radiation conditions, long communication times, and limited bandwidth. While such tools could not replace ground-based analysis by science and engineering teams, machine learning algorithms could enhance the science return of these missions through development of autonomous science capabilities. Examples of science autonomy include onboard data analysis and subsequent instrument optimization, data prioritization (for transmission), and real-time decision-making based on data analysis. Similar advances could be made to develop streamlined data processing software for rapid ground-based analyses. Here we discuss several ways machine learning and autonomy could be used for astrobiology missions, including landing site selection, prioritization and targeting of samples, classification of “features” (e.g., proposed biosignatures) and novelties (uncharacterized, “new” features, which may be of most interest to agnostic astrobiological investigations), and data transmission.

ocean worlds↗

Statistical properties of filaments in the cosmic web

ABSTRACT In the context of the cosmological and constrained Exploring the Local Universe with the reConstructed Initial Density field (ELUCID) simulation, this study explores the statistical characteristics of filaments within the cosmic web, focussing on aspects such as the distribution of filament lengths and their radial density profiles. Using the classification of the cosmic web environment through the Hessian matrix of the density field, our primary focus is on how cosmic structures react to the two variables $R_{\rm s}$ and $\lambda _{\rm th}$. The findings show that the volume fractions of knots, filaments, sheets, and voids are highly influenced by the threshold parameter $\lambda _{\rm th}$, with only a slight influence from the smoothing length $R_{\rm s}$. The central axis of the cylindrical filament is pinpointed using the medial-axis thinning algorithm of the COsmic Web Skeleton (COWS) method. It is observed that median filament lengths tend to increase as the smoothing lengths increase. Analysis of filament length functions at different values of $R_{\rm s}$ indicates a reduction in shorter filaments and an increase in longer filaments as $R_{\rm s}$ increases, peaking around $2.5R_{\rm s}$. The study also shows that the radial density profiles of filaments are markedly affected by the parameters $R_{\rm s}$ and $\lambda _{\rm th}$, showing a valley at approximately $2R_{\rm s}$, with increases in the threshold leading to higher amplitudes of the density profile. Moreover, shorter filaments tend to have denser profiles than their longer counterparts.

Zhang, Youcai (ORCID:0000000319674091)↗

Mapping small-scale vegetation changes in Mexico

This research attempts to map small-scale vegetation changes in Mexico. Forty-eight weeks of coarse resolution Advanced Very High Resolution Radiometer Normalized Difference Vegetation Index (NDVI), a digitized climax vegetation map, land cover samples from space shuttle photographs and actual vegetation samples were used as inputs. Principal components analyses and a clustering algorithm were applied to the NDVI data to generate a single layer that was stratified by the climax vegetation zones map. The purpose is to create a new layer that differentiates climax vegetation (hypothesized potential vegetation) from non-climax vegetation land covers. One of the keys to developing a present-day vegetation map was differentiating intrazone land covers based on the stratification; as great as 75% of the sampled land cover types differed from the climax vegetation. The present-day vegetation map achieved 80% classification accuracy when calculated from available ground reference data. About 55% of the temperate zones and 37% of the tropical zones were found to contain original climax vegetation. Most changes coincide with areas of major agricultural activity.

Turcotte, Kevin M.↗

The effects of cloud inhomogeneities upon radiative fluxes, and the supply of a cloud truth validation dataset

A series of cloud and sea ice retrieval algorithms are being developed in support of the Advanced Spaceborne Thermal Emission and Reflection Radiometer (ASTER) Science Team objectives. These retrievals include the following: cloud fractional area, cloud optical thickness, cloud phase (water or ice), cloud particle effective radius, cloud top heights, cloud base height, cloud top temperature, cloud emissivity, cloud 3-D structure, cloud field scales of organization, sea ice fractional area, sea ice temperature, sea ice albedo, and sea surface temperature. Due to the problems of accurately retrieving cloud properties over bright surfaces, an advanced cloud classification method was developed which is based upon spectral and textural features and artificial intelligence classifiers.

Welch, Ronald M.↗

Land Boundary Conditions for the Goddard Earth Observing System Model Version 5 (GEOS-5) Climate Modeling System: Recent Updates and Data File Descriptions

The Earths land surface boundary conditions in the Goddard Earth Observing System version 5 (GEOS-5) modeling system were updated using recent high spatial and temporal resolution global data products. The updates include: (i) construction of a global 10-arcsec land-ocean lakes-ice mask; (ii) incorporation of a 10-arcsec Globcover 2009 land cover dataset; (iii) implementation of Level 12 Pfafstetter hydrologic catchments; (iv) use of hybridized SRTM global topography data; (v) construction of the HWSDv1.21-STATSGO2 merged global 30 arc second soil mineral and carbon data in conjunction with a highly-refined soil classification system; (vi) production of diffuse visible and near-infrared 8-day MODIS albedo climatologies at 30-arcsec from the period 2001-2011; and (vii) production of the GEOLAND2 and MODIS merged 8-day LAI climatology at 30-arcsec for GEOS-5. The global data sets were preprocessed and used to construct global raster data files for the software (mkCatchParam) that computes parameters on catchment-tiles for various atmospheric grids. The updates also include a few bug fixes in mkCatchParam, as well as changes (improvements in algorithms, etc.) to mkCatchParam that allow it to produce tile-space parameters efficiently for high resolution AGCM grids. The update process also includes the construction of data files describing the vegetation type fractions, soil background albedo, nitrogen deposition and mean annual 2m air temperature to be used with the future Catchment CN model and the global stream channel network to be used with the future global runoff routing model. This report provides detailed descriptions of the data production process and data file format of each updated data set.

GEOS-5↗

Vegetation Classification of Coffea on Hawaii Island using Worldview-2 Satellite Imagery

Coffee is an important crop in tropical regions of the world; about 125 million people depend on coffee agriculture for their livelihoods. Understanding the spatial extent of coffee fields is useful for management and control of coffee pests such as Hypothenemus hampei and other pests that use coffee fruit as a host for immature stages such as the Mediterranean fruit fly, for economic planning, and for following changes in coffee agroecosystems over time. We present two methods for detecting Coffea arabica fields using remote sensing and geospatial technologies on WorldView-2 high-resolution spectral data of the Kona region of Hawaii Island. The first method, a pixel-based method using a maximum likelihood algorithm, attained 72% producer accuracy and 69% user accuracy (68% overall accuracy) based on analysis of 104 ground truth testing polygons. The second method, an object-based image analysis (OBIA) method, considered both spectral and textural information and improved accuracy, resulting in 76% producer accuracy and 94% user accuracy (81% overall accuracy) for the same testing areas. We conclude that the OBIA method is useful for detecting coffee fields grown in the open and use it to estimate the distribution of about 1050 hectares under coffee agriculture in the Kona region in 2012.

Gaertner, Julie↗

Machine Learning for Well Log Analysis in Uranium Mining

This project explores the use of Artificial Intelligence (AI) and Machine Learning (ML) techniques to automate well log analysis for uranium mining. Geophysical log data—spontaneous potential, resistivity, and gamma ray—were used to classify lithology, correlate well logs and identify roll front zonation patterns, which are critical for locating uranium ore bodies. Supervised ML algorithms such as eXtreme Gradient Boosting (XGBoost), Categorical Boosting (CatBoost), and Random Forest were trained to classify lithology with high accuracy. Gradient Boosting Machines (GBM), XGBoost, Random Forest, and Neural Networks were also used for role front zone identification. Moreover, a Fast Dynamic Time Warping (FastDTW) algorithm was employed for well log correlation. Additionally, sample lag was addressed using dynamic programming. Results demonstrate the potential of AI and ML to streamline well log analysis and enhance uranium exploration workflows.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Data‐Efficient Generation of Synthetic Microstructures of Polymer‐Bonded Energetic Material With Fine‐Tuned Stable Diffusion

Among current deep learning approaches for synthetic image generation, diffusion-based models stand out in terms of algorithmic stability and ability to retain high-fidelity image features with detailed resolution. Here, in this work, we employ Dreambooth, a method for fine-tuning Stable Diffusion, on X-ray CT images of microstructure of the polymer-bonded form (PBX) of a commonly used high explosive, Pentaerythritol tetranitrate (PETN), which yields generative models for creating synthetic PBX images. The models developed here represent five classes (or ‘lots’) of microstructures and demonstrate successful generation of images of each class with high fidelity, as verified by computed classification accuracy of ∼ 94% or higher. Data augmentation afforded by such image synthesis can be used to more reliably decipher underlying statistics, build processing-structure correlations, recognize off-normal structural anomalies, and identify age-related changes. Ideas related to converting image data into appropriate density mapping and performing mesoscale simulation or surrogate modeling of detonation are also discussed.

Dreambooth↗

Optimal band selection for dimensionality reduction of hyperspectral imagery

Hyperspectral images have many bands requiring significant computational power for machine interpretation. During image pre-processing, regions of interest that warrant full examination need to be identified quickly. One technique for speeding up the processing is to use only a small subset of bands to determine the 'interesting' regions. The problem addressed here is how to determine the fewest bands required to achieve a specified performance goal for pixel classification. The band selection problem has been addressed previously Chen et al., Ghassemian et al., Henderson et al., and Kim et al.. Some popular techniques for reducing the dimensionality of a feature space, such as principal components analysis, reduce dimensionality by computing new features that are linear combinations of the original features. However, such approaches require measuring and processing all the available bands before the dimensionality is reduced. Our approach, adapted from previous multidimensional signal analysis research, is simpler and achieves dimensionality reduction by selecting bands. Feature selection algorithms are used to determine which combination of bands has the lowest probability of pixel misclassification. Two elements required by this approach are a choice of objective function and a choice of search strategy.

Stearns, Stephen D.↗

An iterative decoupling solution method for large scale Lyapunov equations

A great deal of attention has been given to the numerical solution of the Lyapunov equation. A useful classification of the variety of solution techniques are the groupings of direct, transformation, and iterative methods. The paper summarizes those methods that are at least partly favorable numerically, giving special attention to two criteria: exploitation of a general sparse system matrix structure and efficiency in resolving the governing linear matrix equation for different matrices. An iterative decoupling solution method is proposed as a promising approach for solving large-scale Lyapunov equation when the system matrix exhibits a general sparse structure. A Fortran computer program that realizes the iterative decoupling algorithm is also discussed.

Athay, T. M.↗

Developing processing techniques for Skylab data

The author has identified the following significant results. The effects of misregistration and the scan-line-straightening algorithm on multispectral data were found to be: (1) there is greatly increased misregistration in scan-line-straightening data over conic data; (2) scanner caused misregistration between any pairs of channels may not be corrected for in scan-line-straightened data; and (3) this data will have few pure field center pixels than will conic data. A program SIMSIG was developed implementing the signature simulation model. Data processing stages of the experiment were carried out, and an analysis was made of the effects of spatial misregistration on field center classification accuracy. Fifteen signatures originally used for classifying the data were analyzed, showing the following breakdown: corn (4 signatures), trees (2), brush (1), grasses, weeds, etc. (5), bare soil (1), soybeans (1), and alfalfa (1).

Nalepka, R. F.↗

Refining image segmentation by polygon skeletonization

A skeletonization algorithm was encoded and applied to a test data set of land-use polygons taken from a USGS digital land use dataset at 1:250,000. The distance transform produced by this method was instrumental in the description of the shape, size, and level of generalization of the outlines of the polygons. A comparison of the topology of skeletons for forested wetlands and lakes indicated that some distinction based solely upon the shape properties of the areas is possible, and may be of use in an intelligent automated land cover classification system.

Clarke, Keith C.↗

ICAT: The Interactive Corpus Analysis Tool

The Interactive Corpus Analysis Tool (ICAT) is a Python library for creating dashboards to explore textual datasets and build simple binary classification models to help filter through them and focus on entries of interest. This tool uses a form of interactive machine learning (IML), a paradigm of “machine teaching” (Simard et al., 2017) that sits at the intersection of the fields of human computer interaction (HCI), visual analytics, and machine learning. The intent of ICAT is to allow subject matter experts (SME) with limited to no experience in machine learning to benefit from an iterative human-in-the-loop (HITL) approach to building their own model without needing to understand the details of the underlying algorithm. This interactivity is achieved by allowing the user to create features, label data points, and visually manipulate a representation of the features to manually cluster and investigate data, while a model is trained on the fly based on these actions. ICAT is built on top of the Panel (Holoviz, 2018) library, using a combination of Vega, a custom IPyWidget using D3, and ipyvuetify, and is intended to be used inside of a Jupyter environment.

Martindale, Nathan [Oak Ridge National Laboratory ↗

Recognizing Blazars Using Radio Morphology from the VLA Sky Survey

Abstract Blazars are radio-loud active galactic nuclei whose jets have a very small angle to our line of sight. Observationally, the radio emissions are mostly compact or compact-core with a one-sided jet. With 2.″5 resolution at 3 GHz, the Very Large Array Sky Survey (VLASS) enables us to resolve the structure of some blazar candidates in the sky north of decl. −40°. We introduce an algorithm to classify radio sources as either blazar-like or non-blazar-like based on their morphology in the VLASS images. We apply our algorithm to three existing catalogs, including one of the known blazars (Roma-BzCAT) and two blazar candidates identified by Wide-field Infrared Survey Explorer colors and radio emission (WIBRaLS, KDEBLLACS). We show that in all three catalogs, there are objects with morphologies inconsistent with being blazars. Considering all the catalogs, more than 12% of the candidates are unlikely to be blazars, based on this analysis. Notably, we show that 3% of the Roma-BzCATconfirmedblazars could be a misclassification based on their VLASS morphology. The resulting table with all sources and their radio morphological classification is available online.

Astronomy & Astrophysics↗

Improved assessment of mangrove forests in Sundarbans East Wildlife Sanctuary using WorldView 2 and TanDEM-X high resolution imagery

Recent developments of remote sensing techniques which can capture both the structure and function of the ecosystem provide a more representative view of the landscape. These unique Earth observations were used to help improve traditional forestry surveys by providing species-specific land cover classes for mangrove forests in the Sundarbans East Wildlife Sanctuary. By combining optical data from WorldView2 (WV2; 2 m pixel) and a canopy height model derived using radar data from TanDEM-X (TDX; 12 m pixel), we identified nine mangrove and five non-mangrove classes by following an Iterative Self-Organizing Data Analysis Algorithm. Three dominant mangrove species accounted for nearly 50% of the sanctuary. Heritieria fomes disproportionately covered the largest area at 43%, overturning previous field-based estimates of Excoecaria agallocha dominance. E. agallocha and Sonneratia apetala, covered 3% and 1.47% of the sanctuary, respectively. Four mixed species classes were also identified with clear vegetation zonation patterns that trended toward species homogeneity with increasing distance from shore. The overall land cover accuracy (WV2: 89.33%; WV2-TDX: 89.89%), the Kappa Coefficient (WV2:0.88; WV2-TDX: 0.89) and change statistics between WV2 and WV2-TDX landcover classifications indicate that the WV2 imagery can separate mangrove community types without structural data. The combination of the land cover classifications and the canopy height model indicated that H. fomes were not only the most dominant forest but also, on average, the tallest (12.3 m) among the other eight mangrove types. Our large-scale mapping with high resolution optical and radar platforms can capture subtle changes in mangrove vegetation and canopy structural gradients more accurately and be used to monitor biodiversity changes and Aichi Biodiversity Targets and Indicators, which would contribute to biodiversity policy updating.

Md Mizanur Rahman↗

Improving the CERES SYN Cloud and Flux Products by Identifying GOES-17 Scan Anomalies Using a Convolutional Neural Network

The NASA Clouds and the Earth’s Radiant Energy System (CERES) project relies on top-of-atmosphere (TOA) broadband fluxes derived from geostationary (GEO) satellite imagery to account for the diurnal flux variations between the CERES observation intervals, and thereby produce a synoptic gridded (SYN1deg) product based on continuous temporal observations. Consistent broadband flux derivation depends on accurate radiative property measurements and cloud retrievals, which largely determine the radiance-to-flux conversion process. Therefore, it is important to ensure a high quality of cloud property input in order to maintain a reliable broadband flux record. In Edition 4 of the CERES SYN1deg product, a robust automated image anomaly detection algorithm based on inter-line and inter-pixel differences, spatial variance, and 2-D Fourier analysis has been successful in identifying imagery with linear artifacts, but the line-by-line inspection and cleaning process must still be performed by a human. Therefore, further automation of this quality assurance process is warranted, especially considering the excessive amount of additional cleaning necessitated by the GOES-17 Advance Baseline Imager (ABI) cooling system anomaly. As such, this article highlights advancement of the CERES GEO image artifact cleaning approach based on a convolutional neural network (CNN) for classification of bad scanlines. Once trained, the CNN approach is a computationally inexpensive means to ensure greater consistency in cloud retrievals, and therefore broadband flux derivation, based on GOES-17 measurements.

Benjamin Scarino↗