Search NASA⌕ Search

SEARCH · Search NASA

Results for “neighbor selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Nearest-Neighbor Machine Learning Feature Selection for Interpretation of Microbial Molecular Signatures from Isotope Ratio Mass Spectrometry Data

Mass spectrometry (MS) promises to be a powerful tool for potential biosignature detection during astrobiological missions on ocean worlds in our solar system. Accurate and generalizable machine learning methods could enhance science return on investment by predicting seawater chemistry and classifying isotopic biosignatures, either as a signature consistent with microbial life (biotic) or as a novelty (unclassified/unique). However, machine learning models are likely to be complex and involve interactions between MS features, making biosignatures difficult to interpret. Feature selection methods provide biological and chemical context that help interpret the mechanisms of machine learning models, but these methods also need the ability to detect complex interactions. Previously, we developed a machine learning feature selection algorithm called nearest-neighbor projected distance regression (NPDR) that has the ability to identify important model features that involve complex interactions and automatically reduce correlation and the dimensionality in a high-dimensional variable space. The standard distance metrics used in NPDR – Manhattan and Euclidean – assume the multivariate data are isotropic, which is often violated in real data due to differences in the covariance between variables. Thus, we extend NPDR to include a random forest distance, and other anisotropic distance metrics, for computing nearest neighbors. We also augment the isotope-ratio MS data with time-series features from the raw MS signal to improve biotic classification. We test NPDR on our novel experimental ocean world seawater analog MS data. We measure isotope fractionations of volatile CO 2 that could be measured in exospheres or plumes. Samples include baseline abiotic conditions using a range of possible seawater chemistry consistent with Europa and Enceladus, and biotic samples that include microbes in these seawaters. We use penalized NPDR with random forest proximity to identify interpretable microbial molecular signatures. We compare features with random forest importance, and we train a classifier that discriminates between biotic and abiotic samples with high accuracy. These ML-trained ocean-world analog MS data could be used to assist in identifying biosignatures during future missions.

geochemistry↗

Interpretable Machine Learning for Molecular Biosignatures: a Novel Single-Sample Feature Importance Method That Is Sensitive To Statistical Interactions

Isotope ratio mass spectrometry (IRMS) of volatiles (e.g., CO 2 ) promises to be a powerful tool for potential biosignature detection for future missions to ocean worlds (OW) such as Europa and Enceladus. Machine learning (ML) methods for IRMS data could enable science autonomy by onboard prediction of seawater chemistry and biosignature presence. However, ML models are likely to be complex and involve statistical interactions between features (variables), which can make predictions seem opaque and enigmatic. For ML predictions as significant as extraterrestrial biosignatures, we must place extraordinary confidence in models. It is therefore essential that these models make interpretable predictions (i.e., human-understandable) and include false-prediction diagnostics. We achieve high accuracy and interpretability in ML biosignature and seawater chemistry models for OW through a nearest-neighbors feature selection tool that detects statistical interactions between predictors, constructs interaction networks for visualization of selected features working together to make a prediction, and reports single-sample feature importance scores for false-detection diagnostics. Here we develop a novel single-sample nearest-neighbors projected distance regression(ssNPDR) feature selection method that improves upon existing single-sample algorithms through the inclusion of statistical interactions while providing false-prediction diagnostics for ML models.

geochemistry↗

Generating a Simulated Fluid Flow Over an Aircraft Surface Using Anisotropic Diffusion

A fluid-flow simulation over a computer-generated aircraft surface is generated using a diffusion technique. The surface is comprised of a surface mesh of polygons. A boundary-layer fluid property is obtained for a subset of the polygons of the surface mesh. A pressure-gradient vector is determined for a selected polygon, the selected polygon belonging to the surface mesh but not one of the subset of polygons. A maximum and minimum diffusion rate is determined along directions determined using a pressure gradient vector corresponding to the selected polygon. A diffusion-path vector is defined between a point in the selected polygon and a neighboring point in a neighboring polygon. An updated fluid property is determined for the selected polygon using a variable diffusion rate, the variable diffusion rate based on the minimum diffusion rate, maximum diffusion rate, and angular difference between the diffusion-path vector and the pressure-gradient vector.

Rodriguez, David L.↗

Generating a Simulated Fluid Flow over a Surface Using Anisotropic Diffusion

A fluid-flow simulation over a computer-generated surface is generated using a diffusion technique. The surface is comprised of a surface mesh of polygons. A boundary-layer fluid property is obtained for a subset of the polygons of the surface mesh. A gradient vector is determined for a selected polygon, the selected polygon belonging to the surface mesh but not one of the subset of polygons. A maximum and minimum diffusion rate is determined along directions determined using the gradient vector corresponding to the selected polygon. A diffusion-path vector is defined between a point in the selected polygon and a neighboring point in a neighboring polygon. An updated fluid property is determined for the selected polygon using a variable diffusion rate, the variable diffusion rate based on the minimum diffusion rate, maximum diffusion rate, and the gradient vector.

Rodriguez, David L.↗

A solution to the problem of SAR range curvature

When synthetic aperture radar systems are pushed to attain finer resolution at larger ranges than was previously the case for remote sensing purposes, the geometric signal aberration known as range curvature arises. Known techniques for correcting range curvature are exact at only one selected range, thus forcing neighboring ranges to use the same correction as an approximation. A solution to the problem is proposed that is exact at all ranges, thus simplifying and improving the image processing for such systems.

Raney, R. K.↗

A development of grid generation procedure for multicomponent aerodynamic configuration

Two approaches for solving the transonic flow in a multi-block grid were explored. The first approach examines a method involving "zonal decomposition" wherein block boundaries are treated as true boundary surfaces separating interfacing grids. The issues investigated involve techniques for matching solutions at a block boundary. A feasibility study was completed and the results are presented. The second approach involves overlapping grids for differencing across a block boundary near an artificially induced coordinate singularity occurring at a fictitious corner. This approach selects a set of neighboring nodes for the fictitious corner such that the resulting physical cells for a node are topologically the same as any other node on the airfoil surface.

Chen, H. C.↗

Statistics of associations among IR galaxies

In the course of expanding the search of Kleinmann et. al. (1988) for distant, infrared-luminous objects, the authors noticed (as is often remarked) that a large number of infrared-selected galaxies have close neighbors or show merger characteristics (e.g., tidal tails, distorted disks). Because the sample size is large (567 infrared galaxies and 2182 field galaxies), this sample is ideal for statistically examining the importance of interactions among infrared galaxies. In particular, the authors compare the nearest-neighbor distribution and the two-point correlation function of their sample with that of a control sample of field galaxies.

Gallimore, Jack F.↗

Imaging spectrometer for high resolution measurements of stratospheric trace constituents in the ultraviolet

A high-resolution spectrometer has been developed for studies of minor constituents in the middle atmosphere at ultraviolet wavelengths. In particular, the instrument is intended for observations of upper stratospheric UV bands. The spectrometer has a slit width of 0.08 A obtained by means of an echelle grating and a cross-disperser grating. The image plane detector is an intensified CCD consisting of a high gain proximity focused image intensifier that is fiber optically coupled to a two-dimensional CCD array. An instantaneous bandwidth of 9.2 A is resolved across 488 pixels at 0.018 A/pixel, permitting simultaneous acquisition of multiple lines of selected OH bands and the neighboring background. The spectrometer and the approach have been successfully demonstrated as a technique for measuring the concentration of OH on two high-altitude balloon flights. This paper reports the instrument design and its achieved performance.

Torr, Marsha R.↗

Interpretable Machine Learning Models for Autonomous Characterization of Analogue Ocean World Seawater Chemistry and Biosignature Potential Using Isotope Ratio Data

Background: Future missions to ocean worlds, such as Enceladus and Europa, will attempt to characterize the subsurface seawater chemistry and assess the potential for life. Such missions will be equipped with capabilities to precisely measure volatile isotopes in plumes, atmospheres, and exospheres. Motivation: While large isotopic fractionations can indicate a biological source, there are signatures resulting from abiotic geochemical processes that mimic isotopic biosignatures. While machine learning (ML) has the potential to disentangle competing effects and biotic mimicry, high-dimensional isotope ratio mass spectrometry (IRMS) data is likely to contain noise/irrelevant features and involve complex statistical interactions that make human inference and interpretation difficult. Further, ML predictions with as far-reaching implications as an extraterrestrial biosignature on an ocean world requires the use of interpretable models (i.e., not “black box” models) with physically and mathematically meaningful feature spaces along with false positive diagnostics. Methods: We use volatile CO2 IRMS data of analogue ocean world seawaters to validate an ML approach to provide biogeochemical context for biosignature detection. We employ a feature selection method called nearest-neighbor projected distance regression (NPDR) that detects statistical interactions and helps elucidate the mechanisms of the Random Forest classification models. Results: We train and validate predictive ML models on volatile CO2 IRMS data of analogue ocean world seawaters to predict major salt components (e.g., MgSO4, NaHCO3), pH, ionic strength, and the presence of biosignatures. Features derived from IRMS measurements are augmented with extracted time-series features. Our results show high test accuracy and interpretability, which is increased by interaction network visualization, sample-wise variable importance scores, and single-sample class probability estimates. We demonstrate an ML mission software solution that triggers autonomous data transmission and biogeochemical sample prediction.

geochemistry↗

Efficient Implementation of an Optimal Interpolator for Large Spatial Data Sets

Scattered data interpolation is a problem of interest in numerous areas such as electronic imaging, smooth surface modeling, and computational geometry. Our motivation arises from applications in geology and mining, which often involve large scattered data sets and a demand for high accuracy. The method of choice is ordinary kriging. This is because it is a best unbiased estimator. Unfortunately, this interpolant is computationally very expensive to compute exactly. For n scattered data points, computing the value of a single interpolant involves solving a dense linear system of size roughly n x n. This is infeasible for large n. In practice, kriging is solved approximately by local approaches that are based on considering only a relatively small'number of points that lie close to the query point. There are many problems with this local approach, however. The first is that determining the proper neighborhood size is tricky, and is usually solved by ad hoc methods such as selecting a fixed number of nearest neighbors or all the points lying within a fixed radius. Such fixed neighborhood sizes may not work well for all query points, depending on local density of the point distribution. Local methods also suffer from the problem that the resulting interpolant is not continuous. Meyer showed that while kriging produces smooth continues surfaces, it has zero order continuity along its borders. Thus, at interface boundaries where the neighborhood changes, the interpolant behaves discontinuously. Therefore, it is important to consider and solve the global system for each interpolant. However, solving such large dense systems for each query point is impractical. Recently a more principled approach to approximating kriging has been proposed based on a technique called covariance tapering. The problems arise from the fact that the covariance functions that are used in kriging have global support. Our implementations combine, utilize, and enhance a number of different approaches that have been introduced in literature for solving large linear systems for interpolation of scattered data points. For very large systems, exact methods such as Gaussian elimination are impractical since they require 0(n(exp 3)) time and 0(n(exp 2)) storage. As Billings et al. suggested, we use an iterative approach. In particular, we use the SYMMLQ method, for solving the large but sparse ordinary kriging systems that result from tapering. The main technical issue that need to be overcome in our algorithmic solution is that the points' covariance matrix for kriging should be symmetric positive definite. The goal of tapering is to obtain a sparse approximate representation of the covariance matrix while maintaining its positive definiteness. Furrer et al. used tapering to obtain a sparse linear system of the form Ax = b, where A is the tapered symmetric positive definite covariance matrix. Thus, Cholesky factorization could be used to solve their linear systems. They implemented an efficient sparse Cholesky decomposition method. They also showed if these tapers are used for a limited class of covariance models, the solution of the system converges to the solution of the original system. Matrix A in the ordinary kriging system, while symmetric, is not positive definite. Thus, their approach is not applicable to the ordinary kriging system. Therefore, we use tapering only to obtain a sparse linear system. Then, we use SYMMLQ to solve the ordinary kriging system. We show that solving large kriging systems becomes practical via tapering and iterative methods, and results in lower estimation errors compared to traditional local approaches, and significant memory savings compared to the original global system. We also developed a more efficient variant of the sparse SYMMLQ method for large ordinary kriging systems. This approach adaptively finds the correct local neighborhood for each query point in the interpolation process.

Memarsadeghi, Nargess↗

Wind-Driven Erosion and Exposure Potential at Mars 2020 Rover Candidate-Landing Sites

Aeolian processes have likely been the predominant geomorphic agent for most of Mars’ history and have the potential to produce relatively young exposure ages for geologic units. Thus, identifying local evidence for aeolian erosion is highly relevant to the selection of landing sites for future missions, such as the Mars 2020 Rover mission that aims to explore astrobiologically relevant ancient environments. Here we investigate wind-driven activity at eight Mars 2020 candidate-landing sites to constrain erosion potential at these locations. To demonstrate our methods, we found that contemporary dune-derived abrasion rates were in agreement with rover-derived exhumation rates at Gale crater and could be employed elsewhere. The Holden crater candidate site was interpreted to have low contemporary erosion rates, based on the presence of a thick sand coverage of static ripples. Active ripples at the Eberswalde and southwest Melas sites may account for local erosion and the dearth of small craters. Moderate-flux regional dunes near Mawrth Vallis were deemed unrepresentative of the candidate site, which is interpreted to currently be experiencing low levels of erosion. The Nili Fossae site displayed the most unambiguous evidence for local sand transport and erosion, likely yielding relatively young exposure ages. The down selected Jezero crater and northeast Syrtis sites had high-flux neighboring dunes and exhibited substantial evidence for sediment pathways across their ellipses. Both sites had relatively high estimated abrasion rates, which would yield young exposure ages. The down selected Columbia Hills site lacked evidence for sand movement, and contemporary local erosion rates are estimated to be relatively low.

Matthew Chojnacki↗

A LANDSAT digital image rectification system

DIRS is a digital image rectification system for the geometric correction of LANDSAT multispectral scanner digital image data. DIRS removes spatial distortions from the data and brings it into conformance with the Universal Transverse Mercator (UTM) map projection. Scene data in the form of landmarks are used to drive the geometric correction algorithms. Two dimensional least squares polynominal and spacecraft attitude modeling techniques for geometric mapping are provided. Entire scenes or selected quadrilaterals may be rectified. Resampling through nearest neighbor or cubic convolution at user designated intervals is available. The output products are in the form of digital tape in band interleaved, single band or CCT format in a rotated UTM projection. The system was designed and implemented on large scale IBM 360 computers.

Vanwie, P.↗

A Landsat Digital Image Rectification System

DIRS is a Digital Image Rectification System for the geometric correction of Landsat Multispectral Scanner digital image data. DIRS removes spatial distortions from the data and brings it into conformance with the Universal Transverse Mercator (UTM) map projection. Scene data in the form of landmarks or Ground Control Points (GCPs) are used to drive the geometric correction algorithms. The system offers extensive capabilities for 'shade printing' to aid in the determination of GCPs. Affine, two dimensional least squares polynominal and spacecraft attitude modeling techniques for geometric mapping are provided. Entire scenes or selected quadralaterals may be rectified. Resampling through nearest neighbor or cubic convolution at user designated intervals is available. The output products are in the form of digital tape in band interleaved, single band or CCT format in a rotated UTM projection. The system was designed and implemented on large scale IBM 360 computers with at least 300-500K bytes of memory for user application programs and five nine track tapes plus direct access storage.

Van Wie, P.↗

Trellis coding with multidimensional QAM signal sets

Trellis coding using multidimensional QAM signal sets is investigated. Finite-size 2D signal sets are presented that have minimum average energy, are 90-deg rotationally symmetric, and have from 16 to 1024 points. The best trellis codes using the finite 16-QAM signal set with two, four, six, and eight dimensions are found by computer search (the multidimensional signal set is constructed from the 2D signal set). The best moderate complexity trellis codes for infinite lattices with two, four, six, and eight dimensions are also found. The minimum free squared Euclidean distance and number of nearest neighbors for these codes were used as the selection criteria. Many of the multidimensional codes are fully rotationally invariant and give asymptotic coding gains up to 6.0 dB. From the infinite lattice codes, the best codes for transmitting J, J + 1/4, J + 1/3, J + 1/2, J + 2/3, and J + 3/4 bit/sym (J an integer) are presented.

Pietrobon, Steven S.↗

Pattern Recognition for a Flight Dynamics Monte Carlo Simulation

The design, analysis, and verification and validation of a spacecraft relies heavily on Monte Carlo simulations. Modern computational techniques are able to generate large amounts of Monte Carlo data but flight dynamics engineers lack the time and resources to analyze it all. The growing amounts of data combined with the diminished available time of engineers motivates the need to automate the analysis process. Pattern recognition algorithms are an innovative way of analyzing flight dynamics data efficiently. They can search large data sets for specific patterns and highlight critical variables so analysts can focus their analysis efforts. This work combines a few tractable pattern recognition algorithms with basic flight dynamics concepts to build a practical analysis tool for Monte Carlo simulations. Current results show that this tool can quickly and automatically identify individual design parameters, and most importantly, specific combinations of parameters that should be avoided in order to prevent specific system failures. The current version uses a kernel density estimation algorithm and a sequential feature selection algorithm combined with a k-nearest neighbor classifier to find and rank important design parameters. This provides an increased level of confidence in the analysis and saves a significant amount of time.

Restrepo, Carolina↗

Science Diplomacy Through Cities: Applying NASA Earth Observations at the Urban Scale

NASA's scientific expertise and data products are enhancing cities' environmental monitoring activities by pioneering applications of remote sensing and model-based Earth Observations at the urban scale. The above activities have greatly benefitted from engaging stakeholders and city practitioners from the start. Further, NASA's collaborations with cities have: Advanced NASA science, in testing new products and validating of satellite datasets, while meeting the needs of city governments. Broadened Rio de Janeiro's regional viewpoint and strengthened its relationships with neighboring cities. Scientific collaborations with cities benefit from: Selecting city partners with a high level of technical capacity and willing to make strong investments in joint projects. Sustained communication and face-to-face interactions. Well-defined deliverables, with dedicated resources and personnel. Pairing global datasets and projections with in situ measurements and local knowledgeSensitivity to local working culture and politics.

urban environment↗

Design and performance of a large vocabulary discrete word recognition system. Volume 1: Technical report

The development, construction, and test of a 100-word vocabulary near real time word recognition system are reported. Included are reasonable replacement of any one or all 100 words in the vocabulary, rapid learning of a new speaker, storage and retrieval of training sets, verbal or manual single word deletion, continuous adaptation with verbal or manual error correction, on-line verification of vocabulary as spoken, system modes selectable via verification display keyboard, relationship of classified word to neighboring word, and a versatile input/output interface to accommodate a variety of applications.

Source record↗