Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical Algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

The Art of Automation: Translating Electron Microscopy Workflows Into Automated Processes

Acquiring data using a scanning transmission electron microscope (STEM) is a complex, multi-step process. The intricacy of the process depends on the type of sample, composition of the material, desired results of the experiment, resolution requirement and other experimental factors. Each experiment presents unique complications, such as sample drift and contamination, that the microscopist must consider when acquiring data. All these challenges are handled fluidly and expertly by experienced microscopists, but to reach new levels of innovation in material development, including greater reproducibility, throughput, and precision, the automation of these workflows is essential. The initial phase of this work involved translating intuition-based workflows into discrete, programmable steps. Some common key stages in STEM workflows are the initial tuning, scanning the sample for areas of interest, and then acquiring the data. Each stage can be broken further into specific parameter adjustments, such as aberration correction and dwell time optimization, depending on the experiment. When deconstructing various experiments each step was assessed for automation feasibility based on the amount of real time operator decisions. There are steps that lend themselves to automation more readily than others, such as course focusing and sample screening, but there is potential for full automation of all stages with time. As an initial step, an automated montage routine was developed, allowing for the efficient acquisition of large portions of the sample without requiring continuous intervention from the operator. The automation of this small process of the procedure demonstrates the value of this capability. A major challenge in automation arises from discrepancies between commanded, reported and actual stage movements. Using systematic tests, stage movement was quantified. This error can be corrected algorithmically for more accurate workflows in the future. Expanding automation capabilities would result in larger, more efficient data acquisition which allows for more robust statistical analysis. Additionally, this work lays the groundwork for a closed loop system where machine learning algorithms would intake automatically acquired data and make real time decisions. By progressively automating this instrument, this work establishes the foundation for fully automated experimentation in transmission electron microscopy.

97 MATHEMATICS AND COMPUTING↗

Algorithms for the Computation of Debris Risks

Determining the risks from space debris involve a number of statistical calculations. These calculations inevitably involve assumptions about geometry - including the physical geometry of orbits and the geometry of non-spherical satellites. A number of tools have been developed in NASA's Orbital Debris Program Office to handle these calculations; many of which have never been published before. These include algorithms that are used in NASA's Orbital Debris Engineering Model ORDEM 3.0, as well as other tools useful for computing orbital collision rates and ground casualty risks. This paper will present an introduction to these algorithms and the assumptions upon which they are based.

Matney, Mark↗

Evaluating Approaches Relating Ecosystem Productivity with Desis Spectral Information

Data from the DLR Earth Sensing Imaging Spectrometer (DESIS), mounted on the International Space Station (ISS), were used to develop and test algorithms for remotely retrieving ecosystem productivity. Twenty DESIS images were used from three widely separated forested study sites representing deciduous and conifer forests. Gross primary production (GPP) values from eddy covariance flux towers at the sites were matched with DESIS spectral reflectances collected on the same days. Multiple algorithms were successful relating spectral reflectance with GPP, including: spectral vegetation indices (SVI) sensitive to chlorophyll content, SVI used in a photosynthetic light-use efficiency model framework, spectral shape characteristics through spectral derivatives and absorption feature analysis, and statistical models leading to multiband hyperspectral indices from partial least squares regression. Successful algorithms were able to achieve R2 better than 0.7 using a diverse set of observations combining data from different sites from multiple years and at multiple times during the year. The demonstrated robustness of the algorithms provides some confidence in using DESIS imagery to map spatial patterns of GPP.

K F Huemmrich↗

Evaluating Approaches Relating Ecosystem Productivity with DESIS Spectral Information

Data from the DLR Earth Sensing Imaging Spectrometer (DESIS), mounted on the International Space Station (ISS), were used to develop and test algorithms for remotely retrieving ecosystem productivity. Twenty DESIS images were used from three widely separated forested study sites representing deciduous and conifer forests. Gross primary production (GPP) values from eddy covariance flux towers at the sites were matched with DESIS spectral reflectances collected on the same days. Multiple algorithms were successful relating spectral reflectance with GPP, including: spectral vegetation indices (SVI) sensitive to chlorophyll content, SVI used in a photosynthetic light-use efficiency model framework, spectral shape characteristics through spectral derivatives and absorption feature analysis, and statistical models leading to multiband hyperspectral indices from partial least squares regression. Successful algorithms were able to achieve R2 better than 0.7 using a diverse set of observations combining data from different sites from multiple years and at multiple times during the year. The demonstrated robustness of the algorithms provides some confidence in using DESIS imagery to map spatial patterns of GPP.

Gross Primary Productivity (GPP)↗

VEXT: A Virtual Observatory Exploration Toolkit

This final report consists of two main parts. The first is taken from a paper by the PiCA (Pittsburgh Computational Astrostatistics) Group which describes our ongoing work in fast computation of n-point correlation functions. We present here a new algorithm for the fast computation of N-point correlation functions in large astronomical data sets. The algorithm is based on kd-trees which are decorated with cached sufficient statistics thus allowing for orders of magnitude speed-ups over the naive non-tree-based implementation of correlation functions. We further discuss the use of controlled approximations within the computation which allows for further acceleration. In summary, our algorithm now makes it possible to compute exact, all-pairs, measurements of the two, three and four-point correlation functions for cosmological data sets like the Sloan Digital Sky Survey and the next generation of Cosmic Microwave Background experiments. The second part summarizes the progress made by the PiCA Group in this area through the AISR grant.

Schneider, Jeff↗

Magnetospheric Image Unfolding

The Grant was a three year grant funded under the Space Physics Supporting Research and Technology and Suborbital Program. Our objective was to develop automated techniques needed to unfold or "invert" global images of the magnetospheric ion populations obtained by the new magnetospheric imaging techniques (ENA, EUV) in anticipation of future missions such as the Magnetospheric Imager and, now, IMAGE. Our focus on the present three year grant is to determine the degree to which such images can quantitatively constrain the global electromagnetic properties of the magnetosphere. In a previous three year grant period we successfully automated a forward modeling inversion algorithm, demonstrated that these inversions are robust in the face of realistic instrumental considerations such as counting statistics and backgrounds, applied error analysis techniques to the extracted parameters using variational procedures, implemented very realistic magnetospheric test images to test the inversion algorithms using the Rice University Magnetospheric Specification Model, and began the process of generating parametric models with the flexibility to handle the realistic magnetospheric images (e.g. Roelof et al, 1992; 1993). Our plan for the present 3 year grant period was to complete the development of the inversion tools needed to handle realistic magnetospheric images, assess the degree to which global electrodynamics is quantitatively constrained by ENA images of the magnetosphere, and bring the inversion of EUV images up to the maturity that we will have achieved for the ENA imaging. Below the accomplishments of our three year effort are present followed by a list of our presentations and publications. The accomplishments of all three years are presented here, and thus some of these items appeared on interim progress reports.

Source record↗

Two datasets are better than one: method of double moments for 3D reconstruction in cryo-EM

Cryo-electron microscopy is a powerful imaging technique for reconstructing three-dimensional molecular structures from noisy tomographic projection images of randomly oriented particles. We introduce a new data fusion framework, termed the method of double moments, which reconstructs molecular structures from two instances of the second-order moment of projection images obtained under distinct orientation distributions: one uniform, the other non-uniform and unknown. We prove that these moments generically uniquely determine the underlying structure, up to a global rotation and reflection, and we develop a convex-relaxation-based algorithm that achieves accurate recovery using only second-order statistics. Our results demonstrate the advantage of collecting and modeling multiple datasets under different experimental conditions, illustrating that leveraging dataset diversity can substantially enhance reconstruction quality in computational imaging tasks.

Kam’s method↗

Coincident learning for unsupervised anomaly detection of scientific instruments

Abstract Anomaly detection is an important task for complex scientific experiments and other complex systems (e.g. industrial facilities, manufacturing), where failures in a sub-system can lead to lost data, poor performance, or even damage to components. While scientific facilities generate a wealth of data, labeled anomalies may be rare (or even nonexistent), and expensive to acquire. Unsupervised approaches are therefore common and typically search for anomalies either by distance or density of examples in the input feature space (or some associated low-dimensional representation). This paper presents a novel approach called coincident learning for anomaly detection (CoAD), which is specifically designed for multi-modal tasks and identifies anomalies based on coincident behavior across two different slices of the feature space. We define an unsupervised metric, F ^ β , out of analogy to the supervised classification F β statistic. CoAD uses F ^ β to train an anomaly detection algorithm on unlabeled data , based on the expectation that anomalous behavior in one feature slice is coincident with anomalous behavior in the other. The method is illustrated using a synthetic outlier data set and a MNIST-based image data set, and is compared to prior state-of-the-art on two real-world tasks: a metal milling data set and our motivating task of identifying RF station anomalies in a particle accelerator.

43 PARTICLE ACCELERATORS↗

Neuroelectric Virtual Devices

This paper presents recent results in neuroelectric pattern recognition of electromyographic (EMG) signals used to control virtual computer input devices. The devices are designed to substitute for the functions of both a traditional joystick and keyboard entry method. We demonstrate recognition accuracy through neuroelectric control of a 757 class simulation aircraft landing at San Francisco International Airport using a virtual joystick as shown. This is accomplished by a pilot closing his fist in empty air and performing control movements that are captured by a dry electrode array on the arm which are then analyzed and routed through a flight director permitting full pilot outer loop control of the simulation. We then demonstrate finer grain motor pattern recognition through a virtual keyboard by having a typist tap his traders on a typical desk in a touch typist position. The EMG signals are then translated to keyboard presses and displayed. The paper describes the bioelectric pattern recognition methodology common to both examples. Figure 2 depicts raw EMG data from typing, the numeral '8' and the numeral '9'. These two gestures are very close in appearance and statistical properties yet are distinguishable by our hidden Kharkov model algorithms. Extensions of this work to NASA emissions and robotic control are considered.

Wheeler, Kevin↗

A Fast Implementation of the ISOCLUS Algorithm

Unsupervised clustering is a fundamental tool in numerous image processing and remote sensing applications. For example, unsupervised clustering is often used to obtain vegetation maps of an area of interest. This approach is useful when reliable training data are either scarce or expensive, and when relatively little a priori information about the data is available. Unsupervised clustering methods play a significant role in the pursuit of unsupervised classification. One of the most popular and widely used clustering schemes for remote sensing applications is the ISOCLUS algorithm, which is based on the ISODATA method. The algorithm is given a set of n data points (or samples) in d-dimensional space, an integer k indicating the initial number of clusters, and a number of additional parameters. The general goal is to compute a set of cluster centers in d-space. Although there is no specific optimization criterion, the algorithm is similar in spirit to the well known k-means clustering method in which the objective is to minimize the average squared distance of each point to its nearest center, called the average distortion. One significant feature of ISOCLUS over k-means is that clusters may be merged or split, and so the final number of clusters may be different from the number k supplied as part of the input. This algorithm will be described in later in this paper. The ISOCLUS algorithm can run very slowly, particularly on large data sets. Given its wide use in remote sensing, its efficient computation is an important goal. We have developed a fast implementation of the ISOCLUS algorithm. Our improvement is based on a recent acceleration to the k-means algorithm, the filtering algorithm, by Kanungo et al.. They showed that, by storing the data in a kd-tree, it was possible to significantly reduce the running time of k-means. We have adapted this method for the ISOCLUS algorithm. For technical reasons, which are explained later, it is necessary to make a minor modification to the ISOCLUS specification. We provide empirical evidence, on both synthetic and Landsat image data sets, that our algorithm's performance is essentially the same as that of ISOCLUS, but with significantly lower running times. We show that our algorithm runs from 3 to 30 times faster than a straightforward implementation of ISOCLUS. Our adaptation of the filtering algorithm involves the efficient computation of a number of cluster statistics that are needed for ISOCLUS, but not for k-means.

Memarsadeghi, Nargess↗

Classifying Agnostic Biosignatures using Raman, VNIR, and Elemental Data

How can we use our current wealth of terrestrial data, encompassing biogenic and abiogenic systems, to determine the distinguishing properties of life? SCOBI (Statistical Classification of Biosignature Information) uses machine learning techniques to algorithmically identify combinations of measurements that are “indicative of life”. A set of ~1000 observations, comprising elemental abundance, isotopic fractionation, VNIR reflectance, and (in progress) Raman spectra, have been assembled from existing literature and databases. The observations cover systems classified as “indicative alive” (e.g., cells, vegetation), “indicative non-alive” (e.g., fossils, teeth), “mixed indicative” (e.g., soil, pond water), or “non-indicative” (e.g., rocks, meteorites). VNIR data was preprocessed by linear interpolation from 400-2100 nm and smoothed with a Savitzky-Golay filter. To limit the amount of Earth-biochemistry-specific (non-agnostic) information included, the first five spectral features extracted were number of peaks, number of troughs, mean reflectance, mean peak width, and broadest peak width. To help further emphasize agnostic biosignatures, Earth-specific features such as chlorophylls have been manually flagged so that feature importance with and without them can be compared. Classifiers including k-nearest neighbors (KNN), Gaussian Naïve Bayes (GNB), logistic regression (LR), random forest (RF), and support vector machine (SVM) were implemented, as was a combination voting classifier. Performance metrics included false positive rates, false negative rates, and AUC with 50-50 test/train splits (Monte Carlo simulations). Key takeaways from this stage, prior to the inclusion of Raman spectra, are (1) the overall success rate of 0.933 AUC was most heavily influenced by the elemental abundance data; and (2) VNIR reflectance had the lowest classification performance with 0.52 AUC (58% of objects correctly classified). The next steps are to complete integration of Raman spectral data and to improve the approach to pre-processing and feature extraction for both types of spectral data, such as automated baseline removal, whole spectrum matching, and dimensionality reduction.

Biosignatures↗

Passive Microwave Arctic Sea Ice Melt Onset Dates From the Advanced Horizontal Range Algorithm 1979 - 2022

The onset of the summer melt season is a key stage of the Arctic sea ice seasonal cycle and is an indicator of climate change. Surface melting of the bare or snow-covered sea ice is detected using passive microwave satellite observations. The data set presented here is a 44 year record of Arctic sea ice annual melt onset (MO) dates for 1979–2022 produced using an updated version of the Advanced Horizontal Range Algorithm (AHRA). This data product contains annual maps of the sea ice MO date and a set of descriptive statistics summarizing the data. This paper describes a new update of the AHRA methodology, now AHRA V5, including key changes to the algorithm starting date and sea ice mask methodology to improve estimates of early-season MO dates especially near the sea ice periphery. AHRA V5 data are suitable for monitoring trends in Arctic and regional sea ice MO dates and for process studies of atmosphere-sea ice interactions during the early spring and summer months.

Angela C. Bliss↗

Statistical properties of DNA sequences

We review evidence supporting the idea that the DNA sequence in genes containing non-coding regions is correlated, and that the correlation is remarkably long range--indeed, nucleotides thousands of base pairs distant are correlated. We do not find such a long-range correlation in the coding regions of the gene. We resolve the problem of the "non-stationarity" feature of the sequence of base pairs by applying a new algorithm called detrended fluctuation analysis (DFA). We address the claim of Voss that there is no difference in the statistical properties of coding and non-coding regions of DNA by systematically applying the DFA algorithm, as well as standard FFT analysis, to every DNA sequence (33301 coding and 29453 non-coding) in the entire GenBank database. Finally, we describe briefly some recent work showing that the non-coding sequences have certain statistical features in common with natural and artificial languages. Specifically, we adapt to DNA the Zipf approach to analyzing linguistic texts. These statistical properties of non-coding sequences support the possibility that non-coding regions of DNA may carry biological information.

Non-NASA Center↗

A statistical technique for determining rainfall over land employing Nimbus-6 ESMR measurements

At 37 GHz, the frequency at which the Nimbus 6 Electrically Scanning Microwave Radiometer (ESMR 6) measures upwelling radiance, it was shown theoretically that the atmospheric scattering and the relative independence on electromagnetic polarization of the radiances emerging from hydrometers make it possible to monitor remotely active rainfall over land. In order to verify experimentally these theoretical findings and to develop an algorithm to monitor rainfall over land, the digitized ESMR 6 measurements were examined statistically. Horizontally and vertically polarized brightness temperature pairs (TH, TV) from ESMR 6 were sampled for areas of rainfall over land as determined from the rain recording stations and the WSR 57 radar, and areas of wet and dry ground (whose thermodynamic temperatures were greater than 5 C) over the Southeastern United States. These three categories of brightness temperatures were found to be significantly different in the sense that the chances that the mean vectors of any two populations coincided were less than 1 in 100.

Rodgers, E.↗

Frequency-Modulated, Continuous-Wave Laser Ranging Using Photon-Counting Detectors

Optical ranging is a problem of estimating the round-trip flight time of a phase- or amplitude-modulated optical beam that reflects off of a target. Frequency- modulated, continuous-wave (FMCW) ranging systems obtain this estimate by performing an interferometric measurement between a local frequency- modulated laser beam and a delayed copy returning from the target. The range estimate is formed by mixing the target-return field with the local reference field on a beamsplitter and detecting the resultant beat modulation. In conventional FMCW ranging, the source modulation is linear in instantaneous frequency, the reference-arm field has many more photons than the target-return field, and the time-of-flight estimate is generated by balanced difference- detection of the beamsplitter output, followed by a frequency-domain peak search. This work focused on determining the maximum-likelihood (ML) estimation algorithm when continuous-time photoncounting detectors are used. It is founded on a rigorous statistical characterization of the (random) photoelectron emission times as a function of the incident optical field, including the deleterious effects caused by dark current and dead time. These statistics enable derivation of the Cramér-Rao lower bound (CRB) on the accuracy of FMCW ranging, and derivation of the ML estimator, whose performance approaches this bound at high photon flux. The estimation algorithm was developed, and its optimality properties were shown in simulation. Experimental data show that it performs better than the conventional estimation algorithms used. The demonstrated improvement is a factor of 1.414 over frequency-domainbased estimation. If the target interrogating photons and the local reference field photons are costed equally, the optimal allocation of photons between these two arms is to have them equally distributed. This is different than the state of the art, in which the local field is stronger than the target return. The optimal processing of the photocurrent processes at the outputs of the two detectors is to perform log-matched filtering followed by a summation and peak detection. This implies that neither difference detection, nor Fourier-domain peak detection, which are the staples of the state-of-the-art systems, is optimal when a weak local oscillator is employed.

Erkmen, Baris I.↗

Noise Reduction using Frequency Sub-Band Adaptive Spectral Subtraction

A frequency sub-band based adaptive spectral subtraction algorithm is developed to remove noise from noise-corrupted speech signals. A single microphone is used to obtain both the noise-corrupted speech and the estimate of the statistics of the noise. The statistics of the noise are estimated during time frames that do not contain speech. These statistics are used to determine if future time frames contain speech. During speech time frames, the algorithm determines which frequency sub-bands contain useful speech information and which frequency sub-bands contain only noise. The frequency sub-bands, which contain only noise, are subtracted off at a larger proportion so the noise does not compete with the speech information. Simulation results are presented.

Kozel, David↗

Data Selection Improvement For MicroBooNE

Data selection is an extremely important part of data analysis for any experiment. Finding a physics result is often the result of sifting through a massive amount of data, keeping data that we believe to be signal, and throwing out data we do not. This process is called data selection. Creating a selection algorithm is an intensive process that must balance keeping enough data to have statistics and maximizing the signal purity of that data. We also need to choose the right reconstruction method, a tool to take raw data from the detector and convert it into physics results. In this study, we used three different reconstruction tools, Pandora, WireCell, and LANTERN, for the MicroBooNE experiment in conjunction to improve the selection algorithm for analysis. For the case of this study, we look into the charged current N proton 0 pions (CCNp0$\pi$) interaction channel. This is the dominant channel for the Short Baseline Neutrino (SBN) program and is expected to be a large contributor to the Deep Underground Neutrino Experiment (DUNE). We first investigated each of the three tools to find out more about their strengths and weaknesses as reconstructions, and compared them to the truth information directly from the MicroBooNE simulation pipeline. We then put together a direct comparison of the three methods to find which method or combination of methods would return the best result for us. While the study is ongoing, we have learned a lot about data selection for the experiment and the differences between the reconstruction tools.

Dillon, Brayden [Fermilab]↗

Data Analysis with Graphical Models: Software Tools

Probabilistic graphical models (directed and undirected Markov fields, and combined in chain graphs) are used widely in expert systems, image processing and other areas as a framework for representing and reasoning with probabilities. They come with corresponding algorithms for performing probabilistic inference. This paper discusses an extension to these models by Spiegelhalter and Gilks, plates, used to graphically model the notion of a sample. This offers a graphical specification language for representing data analysis problems. When combined with general methods for statistical inference, this also offers a unifying framework for prototyping and/or generating data analysis algorithms from graphical specifications. This paper outlines the framework and then presents some basic tools for the task: a graphical version of the Pitman-Koopman Theorem for the exponential family, problem decomposition, and the calculation of exact Bayes factors. Other tools already developed, such as automatic differentiation, Gibbs sampling, and use of the EM algorithm, make this a broad basis for the generation of data analysis software.

Buntine, Wray L.↗