Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical Algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Retrieval of Aerosol Microphysical Properties from AERONET Photo-Polarimetric Measurements. 2: A New Research Algorithm and Case Demonstration

A new research algorithm is presented here as the second part of a two-part study to retrieve aerosol microphysical properties from the multispectral and multiangular photopolarimetric measurements taken by Aerosol Robotic Network's (AERONET's) new-generation Sun photometer. The algorithm uses an advanced UNified and Linearized Vector Radiative Transfer Model and incorporates a statistical optimization approach.While the new algorithmhas heritage from AERONET operational inversion algorithm in constraining a priori and retrieval smoothness, it has two new features. First, the new algorithmretrieves the effective radius, effective variance, and total volume of aerosols associated with a continuous bimodal particle size distribution (PSD) function, while the AERONET operational algorithm retrieves aerosol volume over 22 size bins. Second, our algorithm retrieves complex refractive indices for both fine and coarsemodes,while the AERONET operational algorithm assumes a size-independent aerosol refractive index. Mode-resolved refractive indices can improve the estimate of the single-scattering albedo (SSA) for each aerosol mode and thus facilitate the validation of satellite products and chemistry transport models. We applied the algorithm to a suite of real cases over Beijing_RADI site and found that our retrievals are overall consistent with AERONET operational inversions but can offer mode-resolved refractive index and SSA with acceptable accuracy for the aerosol composed by spherical particles. Along with the retrieval using both radiance and polarization, we also performed radiance-only retrieval to demonstrate the improvements by adding polarization in the inversion. Contrast analysis indicates that with polarization, retrieval error can be reduced by over 50% in PSD parameters, 10-30% in the refractive index, and 10-40% in SSA, which is consistent with theoretical analysis presented in the companion paper of this two-part study.

aerosol retrieval algorithm↗

Locally adaptive vector quantization: Data compression with feature preservation

A study of a locally adaptive vector quantization (LAVQ) algorithm for data compression is presented. This algorithm provides high-speed one-pass compression and is fully adaptable to any data source and does not require a priori knowledge of the source statistics. Therefore, LAVQ is a universal data compression algorithm. The basic algorithm and several modifications to improve performance are discussed. These modifications are nonlinear quantization, coarse quantization of the codebook, and lossless compression of the output. Performance of LAVQ on various images using irreversible (lossy) coding is comparable to that of the Linde-Buzo-Gray algorithm, but LAVQ has a much higher speed; thus this algorithm has potential for real-time video compression. Unlike most other image compression algorithms, LAVQ preserves fine detail in images. LAVQ's performance as a lossless data compression algorithm is comparable to that of Lempel-Ziv-based algorithms, but LAVQ uses far less memory during the coding process.

Cheung, K. M.↗

Analysis of Sting Balance Calibration Data Using Optimized Regression Models

Calibration data of a wind tunnel sting balance was processed using a candidate math model search algorithm that recommends an optimized regression model for the data analysis. During the calibration the normal force and the moment at the balance moment center were selected as independent calibration variables. The sting balance itself had two moment gages. Therefore, after analyzing the connection between calibration loads and gage outputs, it was decided to choose the difference and the sum of the gage outputs as the two responses that best describe the behavior of the balance. The math model search algorithm was applied to these two responses. An optimized regression model was obtained for each response. Classical strain gage balance load transformations and the equations of the deflection of a cantilever beam under load are used to show that the search algorithm s two optimized regression models are supported by a theoretical analysis of the relationship between the applied calibration loads and the measured gage outputs. The analysis of the sting balance calibration data set is a rare example of a situation when terms of a regression model of a balance can directly be derived from first principles of physics. In addition, it is interesting to note that the search algorithm recommended the correct regression model term combinations using only a set of statistical quality metrics that were applied to the experimental data during the algorithm s term selection process.

Ulbrich, N.↗

A Dark Target research aerosol algorithm for MODIS observations over eastern China: increasing coverage while maintaining accuracy at high aerosol loading

Satellite aerosol products such as the Dark Target (DT) produced from the MODerate resolution Imaging Spectroradiometer (MODIS) are useful for monitoring the progress of air pollution. Unfortunately, the DT often fails to retrieve during the heaviest aerosol events as well as the more moderate events in winter. Some of the literature at-tributes this lack of retrieval to the cloud mask. However, we found this lack of retrieval is mainly traced to thresholds used for masking of inland water and snow. Modifications to these two masks greatly increase 50 % of the retrievals of aerosol optical depth at 0.55 μm (AOD) greater than 1.0. The “extra”-high-AOD retrievals tend to be biased when com-pared with a ground-based sun photometer (AErosol RObotic NETwork, AERONET). Reducing bias in new retrievals re-quires two additional steps. One is an update to the assumed aerosol optical properties (aerosol model); the haze in this region is both less absorbing and lower in altitude than what is assumed in the global algorithm. The second is account-ing for the scale height of the aerosol, specifically that the heavy-aerosol events in the region are much closer to the surface than what is assumed by the global DT algorithm. The resulting combination of modified masking thresholds, new aerosol model, and lower aerosol layer scale height was applied to 3 months of MODIS observations (January–March2013) over eastern China. After these two additional steps are implemented, the significant increase in new retrievals introduces no overall bias at a high-AOD regime but does degrade other overall validation statistics. We also find that the research algorithm is able to identify additional pollution events that AERONET instruments may not due to different spatial sampling. Mean AOD retrieved from the re-search algorithm increases from 0.11 to 0.18 compared to values calculated from the operational DT algorithm during January to March of 2013 over the study area. But near Beijing, where the severe pollution occurs, the new algorithm increases AOD by as much as 3.0 for each 0.5°grid box over the previous operational-algorithm values.

Dark Target↗

Accuracy Assessment of Aqua-MODIS Aerosol Optical Depth Over Coastal Regions: Importance of Quality Flag and Sea Surface Wind Speed

Coastal regions around the globe are a major source for anthropogenic aerosols in the atmosphere, but the underlying surface characteristics are not favorable for the Moderate Resolution Imaging Spectroradiometer (MODIS) algorithms designed for retrieval of aerosols over dark land or open-ocean surfaces. Using data collected from 62 coastal stations worldwide from the Aerosol Robotic Network (AERONET) from approximately 2002-2010, accuracy assessments are made for coastal aerosol optical depth (AOD) retrieved from MODIS aboard Aqua satellite. It is found that coastal AODs (at 550 nm) characterized respectively by the MODIS Dark Land (hereafter Land) surface algorithm, the Open-Ocean (hereafter Ocean) algorithm, and AERONET all exhibit a log-normal distribution. After filtering by quality flags, the MODIS AODs respectively retrieved from the Land and Ocean algorithms are highly correlated with AERONET (with R(sup 2) is approximately equal to 0.8), but only the Land algorithm AODs fall within the expected error envelope greater than 66% of the time. Furthermore, the MODIS AODs from the Land algorithm, Ocean algorithm, and combined Land and Ocean product show statistically significant discrepancies from their respective counterparts from AERONET in terms of mean, probability density function, and cumulative density function, which suggest a need for future improvement in retrieval algorithms. Without filtering with quality flag, the MODIS Land and Ocean AOD dataset can be degraded by 30-50% in terms of mean bias. Overall, the MODIS Ocean algorithm overestimates the AERONET coastal AOD by 0.021 for AOD less than 0.25 and underestimates it by 0.029 for AOD greater than 0.25. This dichotomy is shown to be related to the ocean surface wind speed and cloud contamination effects on the satellite aerosol retrieval. The Modern Era Retrospective-Analysis for Research and Applications (MERRA) reveals that wind speeds over the global coastal region 25 (with a mean and median value of 2.94 meters per second and 2.66 meters per second, respectively) are often slower than 6 meters per second assumed in the MODIS Ocean algorithm. As a result of high correlation (R(sup 2) greater than 0.98) between the bias in binned MODIS AOD and the corresponding binned wind speed over the coastal sea surface, an empirical scheme for correcting the bias of AOD retrieved from the MODIS Ocean algorithm is formulated and is shown to be effective over the majority of the coastal AERONET stations, and hence can be used in future analysis of AOD trend and MODIS AOD data assimilation.

Anderson, J. C.↗

Fast Quantum Algorithms for Numerical Integrals and Stochastic Processes

We discuss quantum algorithms that calculate numerical integrals and descriptive statistics of stochastic processes. With either of two distinct approaches, one obtains an exponential speed increase in comparison to the fastest known classical deterministic algotithms and a quadratic speed increase incomparison to classical Monte Carlo methods.

quantum algorithms numerical integrals↗

Round-off errors in cutting plane algorithms based on the revised simplex procedure

This report statistically analyzes computational round-off errors associated with the cutting plane approach to solving linear integer programming problems. Cutting plane methods require that the inverse of a sequence of matrices be computed. The problem basically reduces to one of minimizing round-off errors in the sequence of inverses. Two procedures for minimizing this problem are presented, and their influence on error accumulation is statistically analyzed. One procedure employs a very small tolerance factor to round computed values to zero. The other procedure is a numerical analysis technique for reinverting or improving the approximate inverse of a matrix. The results indicated that round-off accumulation can be effectively minimized by employing a tolerance factor which reflects the number of significant digits carried for each calculation and by applying the reinversion procedure once to each computed inverse. If 18 significant digits plus an exponent are carried for each variable during computations, then a tolerance value of 0.1 x 10 to the minus 12th power is reasonable.

Moore, J. E.↗

Maximally Informative Statistics for Localization and Mapping

This paper presents an algorithm for localization and mapping for a mobile robot using monocular vision and odometry as its means of sensing. The approach uses the Variable State Dimension filtering (VSDF) framework to combine aspects of Extended Kalman filtering and nonlinear batch optimization. This paper describes two primary improvements to the VSDF. The first is to use an interpolation scheme based on Gaussian quadrature to linearize measurements rather than relying on analytic Jacobians. The second is to replace the inverse covariance matrix in the VSDF with its Cholesky factor to improve the computational complexity. Results of applying the filter to the problem of localization and mapping with omnidirectional vision are presented.

Deans, Matthew C.↗

Efficient Autonomous Learning for Statistical Pattern Recognition

We describe a neural network learning algorithm that implements differential learning in a generalized backpropagation framework. The algorithm regulates model complexity during the learning procedure, generating the best low-complexity approximation to the Bayer-optimal classifier allowed by the training sample.

Pattern↗

Parallel Climate Data Assimilation PSAS Package Achieves 18 GFLOPs on 512-Node Intel Paragon

Several algorithms were added to the Physical-space Statistical Analysis System (PSAS) from Goddard, which assimilates observational weather data by correcting for different levels of uncertainty about the data and different locations for mobile observation platforms. The new algorithms and use of the 512-node Intel Paragon allowed a hundred-fold decrease in processing time.

weather prediction climate modeling data assimilat↗

AutoBayes Program Synthesis System Users Manual

Program synthesis is the systematic, automatic construction of efficient executable code from high-level declarative specifications. AutoBayes is a fully automatic program synthesis system for the statistical data analysis domain; in particular, it solves parameter estimation problems. It has seen many successful applications at NASA and is currently being used, for example, to analyze simulation results for Orion. The input to AutoBayes is a concise description of a data analysis problem composed of a parameterized statistical model and a goal that is a probability term involving parameters and input data. The output is optimized and fully documented C/C++ code computing the values for those parameters that maximize the probability term. AutoBayes can solve many subproblems symbolically rather than having to rely on numeric approximation algorithms, thus yielding effective, efficient, and compact code. Statistical analysis is faster and more reliable, because effort can be focused on model development and validation rather than manual development of solution algorithms and code.

Schumann, Johann↗

Improved Subseasonal Forecasting of Extreme Polar Vortices Using Machine Learning

Our research was focused on forecasting the position and shape of the winter stratospheric polar vortex at a subseasonal timescale of 15 days in advance. To achieve this, we employed both statistical and neural network machine learning techniques. The analysis was performed on 42 winter seasons of reanalysis data provided by NASA giving us a total of 6,342 days of data. The state of the polar vortex for determined by using geometric moments to calculate the centroid latitude and the aspect ratio of an ellipse fit onto the vortex. Timeseries for thirty additional precursors were calculated to help improve the predictive capabilities of the algorithm. Feature importance of these precursors was performed using random forest to measure the predictive importance and the ideal number of precursors. Then, using the precursors identified as important, various statistical methods were tested for predictive accuracy with random forest and nearest neighbor performing the best. An echo state network, a type of recurrent neural network that features sparsely connected hidden layer and a reduced number of trainable parameters that allows for rapid training and testing, was also implemented for the forecasting problem. Hyperparameter tuning was performed for each methods using a subset of the training data. The algorithms were trained and tuned on the first 41 years of data, then tested for accuracy on the final year. In general, the centroid latitude of the polar vortex proved easier to predict than the aspect ratio across all algorithms. Random forest outperformed other statistical forecasting algorithms overall but struggled to predict extreme values. Forecasting from echo state network suggested a strong predictive capability past 15 days, but further work is required to fully realize the potential of recurrent neural network approaches.

54 ENVIRONMENTAL SCIENCES↗

Analysis of Multivariate Experimental Data Using A Simplified Regression Model Search Algorithm

A new regression model search algorithm was developed that may be applied to both general multivariate experimental data sets and wind tunnel strain-gage balance calibration data. The algorithm is a simplified version of a more complex algorithm that was originally developed for the NASA Ames Balance Calibration Laboratory. The new algorithm performs regression model term reduction to prevent overfitting of data. It has the advantage that it needs only about one tenth of the original algorithm's CPU time for the completion of a regression model search. In addition, extensive testing showed that the prediction accuracy of math models obtained from the simplified algorithm is similar to the prediction accuracy of math models obtained from the original algorithm. The simplified algorithm, however, cannot guarantee that search constraints related to a set of statistical quality requirements are always satisfied in the optimized regression model. Therefore, the simplified algorithm is not intended to replace the original algorithm. Instead, it may be used to generate an alternate optimized regression model of experimental data whenever the application of the original search algorithm fails or requires too much CPU time. Data from a machine calibration of NASA's MK40 force balance is used to illustrate the application of the new search algorithm.

Ulbrich, Norbert M.↗

CSTAR star catalogue development

The Continuous Stellar Tracking Attitude Reference (CSTAR) system is an in-house project for the Space Station to provide high accuracy, drift free attitude and angular rate information for the GN&C system. Constraints exist on the star catalogue incorporated in the system. These constraints include the following: mass memory allocated for catalogue storage, star tracker imaging sensitivity, the minimum resolvable separation angle between stars, the width of the field of view of the star tracker, and the desired number of stars to be tracked in a field of view. The Smithsonian Astrophysical Observatory (SAO) catalogue is the basis reference for this study. As it stands, the SAO does not meet the requirements of any of the above constraints. Star selection algorithms have been devised for catalogue optimization. Star distribution statistics have been obtained to aid in the development of these rules. VAX based software has been developed to implement the star selection algorithms. The software is modular and provides a design tool to tailor the catalogue to available star tracker technology. The SAO catalogue has been optimized for the requirements of the present CSTAR system.

Uhde-Lacovara, J. A.↗

Clustering Days with Similar Airport Weather Conditions

On any given day, traffic flow managers must often rely on past experience and intuition when developing traffic flow management initiatives that mitigate imbalances between the aircraft demand and the weather impacted airport capacity. The goal of this study was to build on recent efforts to apply data mining classification and clustering algorithms to vast archives of historical weather and air traffic data to identify patterns and past decisions that can ultimately inform day-of-operations decision-making. More specifically, this study identified similar weather impacted days at select U.S. airports, and analyzed the traffic management initiatives implemented on these representative days. The identification of the similar days was accomplished by applying a decision tree algorithm to the hourly Localized Aviation Model Output Statistics Program observations and the arrival delays for Newark Liberty International Airport. The branches from the trained decision tree were subsequently pruned to identify four weather conditions that resulted in medium to high delays for the arrivals scheduled to Newark in 2012. Using these weather conditions, four, daily airport-level Weather Impacted Traffic Index values were calculated using the Localized Aviation Model Output Statistics Program observations and the 2012 scheduled arrival counts from the FAAs Aviation System Performance Metric system. The four, daily Weather Impacted Traffic Index values for 2012 were subsequently clustered using an Expectation Maximization clustering algorithm, and nine unique types of weather days at Newark were identified. By far the most prominent type of day at Newark was a day associated with relatively good weather conditions, where there was little convective activity, winds were low, ceilings and visibility were high and there was little precipitation. Moderate levels of convective activity characterized the next most prominent type of day. Days with persistently high winds or low ceiling and visibility levels were relatively rare in 2012. Lastly, the frequency at which Ground Delay Programs, Ground Stops and Miles-in-Trail restrictions were implemented on each of the typical types of days at Newark were analyzed. Based on the results, it does appear as if the usage of Miles-in-Trail, Ground Delay Program and Ground Stop restrictions correlates well with the severity of the weather associated with each unique type of weather impacted day at Newark. Furthermore, the results demonstrate that it is feasible to use historical weather and air traffic archives to provide guidance on the types of traffic management restrictions to implement in response to the weather conditions impacting an airport.

traffic flow management↗

Clustering Days with Similar Airport Weather Conditions

On any given day, traffic flow managers must often rely on past experience and intuition when developing traffic flow management initiatives that mitigate imbalances between the aircraft demand and the weather impacted airport capacity. The goal of this study was to build on recent efforts to apply data mining classification and clustering algorithms to vast archives of historical weather and air traffic data to identify patterns and past decisions that can ultimately inform day-of-operations decision-making. More specifically, this study identified similar weather impacted days at select U.S. airports, and analyzed the traffic management initiatives implemented on these representative days. The identification of the similar days was accomplished by applying a decision tree algorithm to the hourly Localized Aviation Model Output Statistics Program observations and the arrival delays for Newark Liberty International Airport. The branches from the trained decision tree were subsequently pruned to identify four weather conditions that resulted in medium to high delays for the arrivals scheduled to Newark in 2012. Using these weather conditions, four, daily airport-level Weather Impacted Traffic Index values were calculated using the Localized Aviation Model Output Statistics Program observations and the 2012 scheduled arrival counts from the FAAs Aviation System Performance Metric system. The four, daily Weather Impacted Traffic Index values for 2012 were subsequently clustered using an Expectation Maximization clustering algorithm, and nine unique types of weather days at Newark were identified. By far the most prominent type of day at Newark was a day associated with relatively good weather conditions, where there was little convective activity, winds were low, ceilings and visibility were high and there was little precipitation. Moderate levels of convective activity characterized the next most prominent type of day. Days with persistently high winds or low ceiling and visibility levels were relatively rare in 2012. Lastly, the frequency at which Ground Delay Programs, Ground Stops and Miles-in-Trail restrictions were implemented on each of the typical types of days at Newark were analyzed. Based on the results, it does appear as if the usage of Miles-in-Trail, Ground Delay Program and Ground Stop restrictions correlates well with the severity of the weather associated with each unique type of weather impacted day at Newark. Furthermore, the results demonstrate that it is feasible to use historical weather and air traffic archives to provide guidance on the types of traffic management restrictions to implement in response to the weather conditions impacting an airport.

weather↗

Training and Validation of Spectral Gap Filling Algorithm for Cpf-Ceres Intercalibration

The Climate Absolute Radiance and Refractivity Observatory (CLARREO) Pathfinder (CPF) mission is set to launch an SI-traceable reflective solar (RS) spectrometer aboard the International Space Station to measure Earth-reflected solar radiation with a radiometric uncertainty of 0.3% (k=1). The CPF intercalibration team has devised a cutting-edge methodology to accurately transfer the benchmark CPF calibration reference to the shortwave (SW) channel (200-5000 nm) of the Clouds and the Earth’s Radiant Energy System (CERES) instrument. The spectral range of CPF measurements spans from 350-2300 nm, while the CERES SW channel measures the Earth-reflected broadband solar radiances between 200 nm to 5 μm. To conduct precise CPF-CERES intercalibration analysis, the CPF-like spectral radiances outside the CPF spectral range need to be estimated to match the CERES SW spectral range. In response, the team has developed a fast algorithm that leverages spectrally redundant information within the CPF-measured portion through principal component analysis (PCA) and utilizes pre-established spectral correlation relationships among wavelengths to extend the CPF spectrum below 350 nm and above 2300 nm. Our results show that the algorithm achieves excellent accuracy in generating the missing energy in the UV and IR portions. The RMS error in the UV region is less than 4.5x10-3 W/m2/sr/nm, while in the IR region, it is smaller than 8x10-5 W/m2/sr/nm. Our methodology was validated using measured EMIT radiance data, which covers the spectral range from 0.381 μm to 2.493 μm. We employed EMIT radiances within the wavelength range of 0.43 – 2.25 μm to generate radiances for both the shorter wavelength range (0.381 – 0.43 μm) and longer wavelength range (2.25 – 2.493 μm). The generated radiances agree very well with the measured EMIT radiances. The standard deviation in the integrated broadband radiances was about 0.1%, and the bias is less than 0.004% for over 1.5 million EMIT measured samples. These statistics show that the spectral gap filling algorithm is robust and effective in substantially reducing the spectral difference-induced uncertainty in the CPF-CERES intercalibration samples.

Qiguang Yang↗

Signature extension through the application of cluster matching algorithms to determine appropriate signature transformations

Signature extension is intended to increase the space-time range over which a set of training statistics can be used to classify data without significant loss of recognition accuracy. A first cluster matching algorithm MASC (Multiplicative and Additive Signature Correction) was developed at the Environmental Research Institute of Michigan to test the concept of using associations between training and recognition area cluster statistics to define an average signature transformation. A more recent signature extension module CROP-A (Cluster Regression Ordered on Principal Axis) has shown evidence of making significant associations between training and recognition area cluster statistics, with the clusters to be matched being selected automatically by the algorithm.

Lambeck, P. F.↗