Search NASASearch

SEARCH · Search NASA

Results for “Statistical Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Earth-space links and fade-duration statistics

In recent years, fade-duration statistics have been the subject of several experimental investigations. A good knowledge of the fade-duration distribution is important for the assessment of a satellite communication system's channel dynamics: What is a typical link outage duration? How often do link outages exceeding a given duration occur? Unfortunately there is yet no model that can universally answer the above questions. The available field measurements mainly come from temperate climatic zones and only from a few sites. Furthermore, the available statistics are also limited in the choice of frequency and path elevation angle. Yet, much can be learned from the available information. For example, we now know that the fade-duration distribution is approximately lognormal. Under certain conditions, we can even determine the median and other percentiles of the distribution. This paper reviews the available data obtained by several experimenters in different parts of the world. Areas of emphasis are mobile and fixed satellite links. Fades in mobile links are due to roadside-tree shadowing, whereas fades in fixed links are due to rain attenuation.

Davarian, Faramaz

Earth-Space Links and Fade-Duration Statistics

In recent years, fade-duration statistics have been the subject of several experimental investigations. A good knowledge of the fade-duration distribution is important for the assessment of a satellite communication system's channel dynamics: What is a typical link outage duration? How often do link outages exceeding a given duration occur? Unfortunately there is yet no model that can universally answer the above questions. The available field measurements mainly come from temperate climatic zones and only from a few sites. Furthermore, the available statistics are also limited in the choice of frequency and path elevation angle. Yet, much can be learned from the available information. For example, we now know that the fade-duration distribution is approximately lognormal. Under certain conditions, we can even determine the median and other percentiles of the distribution. This paper reviews the available data obtained by several experimenters in different parts of the world. Areas of emphasis are mobile and fixed satellite links. Fades in mobile links are due to roadside-tree shadowing, whereas fades in fixed links are due to rain attenuation.

Davarian, Faramaz

A simulator for evaluating methods for the detection of lesion-deficit associations

Although much has been learned about the functional organization of the human brain through lesion-deficit analysis, the variety of statistical and image-processing methods developed for this purpose precludes a closed-form analysis of the statistical power of these systems. Therefore, we developed a lesion-deficit simulator (LDS), which generates artificial subjects, each of which consists of a set of functional deficits, and a brain image with lesions; the deficits and lesions conform to predefined distributions. We used probability distributions to model the number, sizes, and spatial distribution of lesions, to model the structure-function associations, and to model registration error. We used the LDS to evaluate, as examples, the effects of the complexities and strengths of lesion-deficit associations, and of registration error, on the power of lesion-deficit analysis. We measured the numbers of recovered associations from these simulated data, as a function of the number of subjects analyzed, the strengths and number of associations in the statistical model, the number of structures associated with a particular function, and the prior probabilities of structures being abnormal. The number of subjects required to recover the simulated lesion-deficit associations was found to have an inverse relationship to the strength of associations, and to the smallest probability in the structure-function model. The number of structures associated with a particular function (i.e., the complexity of associations) had a much greater effect on the performance of the analysis method than did the total number of associations. We also found that registration error of 5 mm or less reduces the number of associations discovered by approximately 13% compared to perfect registration. The LDS provides a flexible framework for evaluating many aspects of lesion-deficit analysis.

NASA Program Biomedical Research and Countermeasur

Building a Real-Time Flood Prediction Model for Improving Early Warning Systems in Ellicott City, Maryland

As flood events in the United States grow in frequency and intensity, the use of applied remote sensing analyses is increasingly necessary for effective flood monitoring and warning systems. The NASA DEVELOP National Program partnered with the local government of Howard County, Maryland, to investigate the use of machine learning for advanced flood risk detection, and to test the feasibility of integrating this approach into the county’s flood early warning system. To strengthen the efforts of the Howard County Office of Emergency Management (OEM), the project developed a statistical model capable of hindcasting the two severe flash flood events that devastated Ellicott City and transitioned to a ‘Long Short-Term Memory’ based sequence-to-sequence deep learning model with 8-hour forecast capability. The team combined data inputs from public sources including river and precipitation gauges, NASA and NOAA Earth observations, and numerical weather model products using scripts written in the Google Colaboratory Python scripting environment. In addition to designing the deep learning architecture, the team trained and tested the model, and evaluated its performance using Nash-Sutcliffe Efficiency. The final product, the Sequentially Trained Real-time EstimAted Model (STREAM) predicts stage height for the Hudson Branch gauge in Ellicott City using data products available in near real-time, including the High-Resolution Rapid Refresh model’s accumulated precipitation forecasts supplemented by stream gauge data from the OEM and the U.S. Geological Survey. STREAM was incorporated into an online dashboard in a user-friendly interface capable of triggering the alarms that initiate emergency response protocols up to 8 hours in advance of a predicted severe flood event. The project demonstrated the potential for the integration of open data and Earth observations into a flood risk forecasting tool capable of informing near real-time decision making.

NASA DEVELOP

Algorithms for Learning Preferences for Sets of Objects

A method is being developed that provides for an artificial-intelligence system to learn a user's preferences for sets of objects and to thereafter automatically select subsets of objects according to those preferences. The method was originally intended to enable automated selection, from among large sets of images acquired by instruments aboard spacecraft, of image subsets considered to be scientifically valuable enough to justify use of limited communication resources for transmission to Earth. The method is also applicable to other sets of objects: examples of sets of objects considered in the development of the method include food menus, radio-station music playlists, and assortments of colored blocks for creating mosaics. The method does not require the user to perform the often-difficult task of quantitatively specifying preferences; instead, the user provides examples of preferred sets of objects. This method goes beyond related prior artificial-intelligence methods for learning which individual items are preferred by the user: this method supports a concept of setbased preferences, which include not only preferences for individual items but also preferences regarding types and degrees of diversity of items in a set. Consideration of diversity in this method involves recognition that members of a set may interact with each other in the sense that when considered together, they may be regarded as being complementary, redundant, or incompatible to various degrees. The effects of such interactions are loosely summarized in the term portfolio effect. The learning method relies on a preference representation language, denoted DD-PREF, to express set-based preferences. In DD-PREF, a preference is represented by a tuple that includes quality (depth) functions to estimate how desired a specific value is, weights for each feature preference, the desired diversity of feature values, and the relative importance of diversity versus depth. The system applies statistical concepts to estimate quantitative measures of the user s preferences from training examples (preferred subsets) specified by the user. Once preferences have been learned, the system uses those preferences to select preferred subsets from new sets. The method was found to be viable when tested in computational experiments on menus, music playlists, and rover images. Contemplated future development efforts include further tests on more diverse sets and development of a sub-method for (a) estimating the parameter that represents the relative importance of diversity versus depth, and (b) incorporating background knowledge about the nature of quality functions, which are special functions that specify depth preferences for features.

Wagstaff, Kiri L.

Advancing Methodologies for Applying Machine Learning and Evaluating Spatiotemporal Models of Fine Particulate Matter (PM 2.5 ) Using Satellite Data Over Large Regions

Reconstructing the distribution of fine particulate matter (PM 2.5 ) in space and time, even far from ground monitoring sites, is an important exposure science contribution to epidemiologic analyses of PM 2.5 health impacts. Flexible statistical methods for prediction have demonstrated the integration of satellite observations with other predictors, yet these algorithms are susceptible to overfitting the spatiotemporal structure of the training datasets. We present a new approach for predicting PM 2.5 using machine-learning methods and evaluating prediction models for the goal of making predictions where they were not previously available. We apply extreme gradient boosting (XGBoost) modeling to predict daily PM 2.5 on a 1 x 1 km 2 resolution for a 13 state region in the Northeastern USA for the years 2000–2015 using satellite-derived aerosol optical depth and implement a recursive feature selection to develop a parsimonious model. We demonstrate excellent predictions of withheld observations but also contrast an RMSE of 3.11 μg/m 3 in our spatial cross-validation withholding nearby sites versus an overfit RMSE of 2.10 μg/m 3 using a more conventional random ten-fold splitting of the dataset. As the field of exposure science moves forward with the use of advanced machine-learning approaches for spatiotemporal modeling of air pollutants, our results show the importance of addressing data leakage in training, overfitting to spatiotemporal structure, and the impact of the predominance of ground monitoring sites in dense urban sub-networks on model evaluation. The strengths of our resultant modeling approach for exposure in epidemiologic studies of PM 2.5 include improved efficiency, parsimony, and interpretability with robust validation while still accommodating complex spatiotemporal relationships.

air pollution

A Novel Machine Learning Method for Surface PM2.5 Estimations from Geostationary Satellites

Particulate matter (PM) with a diameter of less or equal to 2.5 μm, known as PM , affects human health as it penetrates the respiratory system. The Environmental Protection Agency (EPA) measures the atmospheric concentration of PM using air quality monitors stationed throughout the Continental United States (CONUS). Such measurements are points on a spatial domain and therefore, might not be representative of the air quality at nearby areas considering that the composition of the atmosphere is highly variable from place to place. Satellite based AOD permits a spatially uniform means of estimating PM and new geostationary satellites provide high temporal and spatial resolution estimation of AOD. However, the concentration of PM is non-linearly dependent on other atmospheric parameters that include relative humidity, temperature, and height of the planetary boundary layer. This information may be estimated at similar spatial and temporal resolutions as AOD from numerical modeling such as from the National Oceanic and Atmospheric Administration’s (NOAA) High Resolution Rapid Refresh (HRRR) model which resolves near real-time atmospheric conditions over the CONUS. The estimation of PM concentration is a multi-parametric problem that considers the effect of temporal dependencies among the different parameters. Deep learning approaches are appropriate for such complex estimation problems as they intrinsically capture relations among multiple non-linear parameters. This study compares deep-learning methods to traditional regression analysis to demonstrate the capabilities of these methods in predicting PM2.5 concentrations. Additionally, a novel ensemble learning approach is employed to identify scientific processes that could further improve the estimation of PM concentration. Utilizing Long Short-Term Memory (LSTM) neural networks, which are suitable for multivariate time series estimation problems as they are capable of learning long-term dependencies, individual models are created for each EPA station and trained on the aforementioned dataset collocated over each station. Individual station models are merged if the model's performance is improved by reducing the root mean squared error (RMSE) metric. This ensemble training method ultimately reduces the RMSE value. Evaluation of these results provide insights into physical processes and related observable parameters that may contribute to PM concentrations. Identified parameters evaluated to be statistically different between the merged and unmerged models are expected to improve overall performance. These new parameters are then utilized for reevaluation of the deep learning methods with an extreme gradient boosting model with an RMSE of 5.5 providing the best results.

George Priftis

Exact and Approximate Probabilistic Symbolic Execution

Probabilistic software analysis seeks to quantify the likelihood of reaching a target event under uncertain environments. Recent approaches compute probabilities of execution paths using symbolic execution, but do not support nondeterminism. Nondeterminism arises naturally when no suitable probabilistic model can capture a program behavior, e.g., for multithreading or distributed systems. In this work, we propose a technique, based on symbolic execution, to synthesize schedulers that resolve nondeterminism to maximize the probability of reaching a target event. To scale to large systems, we also introduce approximate algorithms to search for good schedulers, speeding up established random sampling and reinforcement learning results through the quantification of path probabilities based on symbolic execution. We implemented the techniques in Symbolic PathFinder and evaluated them on nondeterministic Java programs. We show that our algorithms significantly improve upon a state-of- the-art statistical model checking algorithm, originally developed for Markov Decision Processes.

Symbolic Execution

The effects of long delay and transmission errors on the performance of TP-4 implementations

A set of tools that allows us to measure and examine the effects of transmission delay and errors on the performance of TP-4 implementations has been developed. The tools give insight into both the large- and small-scale behaviors of an implementation. These tools have been systematically applied to a commercial implementation of TP-4. Measurements show, among other things, that a 2-second one-way transmission delay and an effective bit-error rate of 1 error per 100,000 bits can result in a 95 percent reduction in TP-4 throughput. The detailed statistics give insight into why transmission delay and errors affect this implementations so significantly and support a number of 'lessons learned' that could be applied to TP-4 implementations that operate more robustly across networks with long transmission delays and transmission errors.

Durst, Robert C.

Multivariate statistical analysis software technologies for astrophysical research involving large data bases

We developed a package to process and analyze the data from the digital version of the Second Palomar Sky Survey. This system, called SKICAT, incorporates the latest in machine learning and expert systems software technology, in order to classify the detected objects objectively and uniformly, and facilitate handling of the enormous data sets from digital sky surveys and other sources. The system provides a powerful, integrated environment for the manipulation and scientific investigation of catalogs from virtually any source. It serves three principal functions: image catalog construction, catalog management, and catalog analysis. Through use of the GID3* Decision Tree artificial induction software, SKICAT automates the process of classifying objects within CCD and digitized plate images. To exploit these catalogs, the system also provides tools to merge them into a large, complete database which may be easily queried and modified when new data or better methods of calibrating or classifying become available. The most innovative feature of SKICAT is the facility it provides to experiment with and apply the latest in machine learning technology to the tasks of catalog construction and analysis. SKICAT provides a unique environment for implementing these tools for any number of future scientific purposes. Initial scientific verification and performance tests have been made using galaxy counts and measurements of galaxy clustering from small subsets of the survey data, and a search for very high redshift quasars. All of the tests were successful, and produced new and interesting scientific results. Attachments to this report give detailed accounts of the technical aspects for multivariate statistical analysis of small and moderate-size data sets, called STATPROG. The package was tested extensively on a number of real scientific applications, and has produced real, published results.

Djorgovski, S. George

A Communication Channel Density Estimating Generative Adversarial Network

Autoencoder-based communication systems use neural network channel models to backwardly propagate message reconstruction error gradients across an approximation of the physical communication channel. In this work, we develop and test a new generative adversarial network (GAN) architecture for the purpose of training a stochastic channel approximating neural network. In previous research, investigators have focused on additive white Gaussian noise (AWGN) channels and/or simplified Rayleigh fading channels, both of which are linear and have well defined analytic solutions. Given that training a neural network is computationally expensive, channel approximation networks— and more generally the autoencoder systems—should be evaluated in communication environments that are traditionally difficult. To that end, our investigation focuses on channels that contain a combination of non-linear amplifier distortion, pulse shape filtering, intersymbol interference, frequency-dependent group delay, multipath, and non-Gaussian statistics. Each of our models are trained without any prior knowledge of the channel. We show that the trained models have learned to generalize over an arbitrary amplifier drive level and constellation alphabet. We demonstrate the versatility of our GAN architecture by comparing the marginal probability density function of several channel simulations with that of their corresponding neural network approximations

Smith, Aaron

The Johnson Space Center Management Information Systems (JSCMIS): An interface for organizational databases

The Management Information and Decision Support Environment (MIDSE) is a research activity to build and test a prototype of a generic human interface on the Johnson Space Center (JSC) Information Network (CIN). The existing interfaces were developed specifically to support operations rather than the type of data which management could use. The diversity of the many interfaces and their relative difficulty discouraged occasional users from attempting to use them for their purposes. The MIDSE activity approached this problem by designing and building an interface to one JSC data base - the personnel statistics tables of the NASA Personnel and Payroll System (NPPS). The interface was designed against the following requirements: generic (use with any relational NOMAD data base); easy to learn (intuitive operations for new users); easy to use (efficient operations for experienced users); self-documenting (help facility which informs users about the data base structure as well as the operation of the interface); and low maintenance (easy configuration to new applications). A prototype interface entitled the JSC Management Information Systems (JSCMIS) was produced. It resides on CIN/PROFS and is available to JSC management who request it. The interface has passed management review and is ready for early use. Three kinds of data are now available: personnel statistics, personnel register, and plan/actual cost.

Bishop, Peter C.

NASA Langley's Approach to the Sandia's Structural Dynamics Challenge Problem

The objective of this challenge is to develop a data-based probabilistic model of uncertainty to predict the behavior of subsystems (payloads) by themselves and while coupled to a primary (target) system. Although this type of analysis is routinely performed and representative of issues faced in real-world system design and integration, there are still several key technical challenges that must be addressed when analyzing uncertain interconnected systems. For example, one key technical challenge is related to the fact that there is limited data on target configurations. Moreover, it is typical to have multiple data sets from experiments conducted at the subsystem level, but often samples sizes are not sufficient to compute high confidence statistics. In this challenge problem additional constraints are placed as ground rules for the participants. One such rule is that mathematical models of the subsystem are limited to linear approximations of the nonlinear physics of the problem at hand. Also, participants are constrained to use these models and the multiple data sets to make predictions about the target system response under completely different input conditions. Our approach involved initially the screening of several different methods. Three of the ones considered are presented herein. The first one is based on the transformation of the modal data to an orthogonal space where the mean and covariance of the data are matched by the model. The other two approaches worked solutions in physical space where the uncertain parameter set is made of masses, stiffnesses and damping coefficients; one matches confidence intervals of low order moments of the statistics via optimization while the second one uses a Kernel density estimation approach. The paper will touch on all the approaches, lessons learned, validation 1 metrics and their comparison, data quantity restriction, and assumptions/limitations of each approach. Keywords: Probabilistic modeling, model validation, uncertainty quantification, kernel density

Horta, Lucas G.

An Unobtrusive System to Measure, Assess, and Predict Cognitive Workload in Real-World Environments

Across many careers, individuals face alternating periods of high and low attention and cognitive workload, which can result in impaired cognitive functioning and can be detrimental to job performance. For example, some professions (e.g., fire fighters, emergency medical personnel, doctors and nurses working in an emergency room, pilots) require long periods of low workload (boredom), followed by sudden, high-tempo operations during which they may be required to respond to an emergency and perform at peak cognitive levels. Conversely, other professions (e.g., air traffic controllers, market investors in financial industries, analysts) require long periods of high workload and multitasking during which the addition of just one more task results in cognitive overload resulting in mistakes. An unobtrusive system to measure, assess, and predict cognitive workload could warn individuals, their teammates, or their supervisors when steps should be taken to augment cognitive readiness. In this talk I will describe an approach to this problem that we have found to be successful across work domains including: (1) a suite of unobtrusive, field-ready neurophysiological, physiological, and behavioral sensors that are chosen to best suit the target environment; (2) custom algorithms and statistical techniques to process and time-align raw data originating from the sensor suite; (3) probabilistic and statistical models designed to interpret the data into the human state of interest (e.g., cognitive workload, attention, fatigue); (4) and machine-learning techniques to predict upcoming performance based on the current pattern of events, and (5) display of each piece of information depending on the needs of the target user who may or may not want to drill down into the functioning of the system to determine how conclusions about human state and performance are determined. I will then focus in on our experimental results from our custom functional near-infrared spectroscopy sensor, designed to operate in real-world environments to be worn comfortably (e.g., positioned into a baseball cap or a surgeons cap) to measure changes in brain blood oxygenation without adding burden to the individual being assessed.

workload collection

Application of High-Dimensional Fuzzy K-Means Cluster Analysis to CALIOP/CALIPSO Version 4.1 Cloud-Aerosol Discrimination

This study applies fuzzy k-means (FKM) cluster analyses to a subset of the parameters reported in the CALIPSO lidar level 2 data products in order to classify the layers detected as either clouds or aerosols. The results obtained are used to assess the reliability of the cloud–aerosol discrimination (CAD) scores reported in the version 4.1 release of the CALIPSO data products. FKM is an unsupervised learning algorithm, whereas the CALIPSO operational CAD algorithm (COCA) takes a highly supervised approach. Despite these substantial computational and architectural differences, our statistical analyses show that the FKM classifications agree with the COCA classifications for more than 94 % of the cases in the troposphere. This high degree of similarity is achieved because the lidar-measured signatures of the majority of the clouds and the aerosols are naturally distinct, and hence objective methods can independently and effectively separate the two classes in most cases. Classification differences most often occur in complex scenes (e.g., evaporating water cloud filaments embedded in dense aerosol) or when observing diffuse features that occur only intermittently (e.g., volcanic ash in the tropical tropopause layer). The two methods examined in this study establish overall classification correctness boundaries due to their differing algorithm uncertainties. In addition to comparing the outputs from the two algorithms, analysis of sampling, data training, performance measurements, fuzzy linear discriminants, defuzzification, error propagation, and key parameters in feature type discrimination with the FKM method are further discussed in order to better understand the utility and limits of the application of clustering algorithms to space lidar measurements. In general, we find that both FKM and COCA classification uncertainties are only minimally affected by noise in the CALIPSO measurements, though both algorithms can be challenged by especially complex scenes containing mixtures of discrete layer types. Our analysis results show that attenuated backscatter and color ratio are the driving factors that separate water clouds from aerosols; backscatter intensity, depolarization, and mid-layer altitude are most useful in discriminating between aerosols and ice clouds; and the joint distribution of backscatter intensity and depolarization ratio is critically important for distinguishing ice clouds from water clouds.

Zeng, Shan

Root Cause Classification of Breakup Events 1961-2018

This paper uses the updated NASA “History of On-Orbit Satellite Fragmentations 15th Edition,” to examine and categorize the root cause of historical breakup events to the greatest degree possible. Classes of debris progenitors have evolved, as many classes of Cold War-era spacecraft are now extinct, only to be replaced by new classes of payloads and rocket bodies statistically likely to experience debris-producing events. The efficacy of international debris mitigation implementation and root cause/fault tree analyses and lessons learned is examined in relation to the breakup of satellite classes or specific events. In select cases, the remaining on-orbit inventory of specific classes is identified in the context of possible future events. The environmental impact of these specific classes is examined and compared to nominal space environment projections. When appropriate, recommendations for debris remediation are made for specific satellite classes.

Anz-Meador, P.

Multivariate Statistical Analysis Software Technologies for Astrophysical Research Involving Large Data Bases

We developed a package to process and analyze the data from the digital version of the Second Palomar Sky Survey. This system, called SKICAT, incorporates the latest in machine learning and expert systems software technology, in order to classify the detected objects objectively and uniformly, and facilitate handling of the enormous data sets from digital sky surveys and other sources. The system provides a powerful, integrated environment for the manipulation and scientific investigation of catalogs from virtually any source. It serves three principal functions: image catalog construction, catalog management, and catalog analysis. Through use of the GID3* Decision Tree artificial induction software, SKICAT automates the process of classifying objects within CCD and digitized plate images. To exploit these catalogs, the system also provides tools to merge them into a large, complex database which may be easily queried and modified when new data or better methods of calibrating or classifying become available. The most innovative feature of SKICAT is the facility it provides to experiment with and apply the latest in machine learning technology to the tasks of catalog construction and analysis. SKICAT provides a unique environment for implementing these tools for any number of future scientific purposes. Initial scientific verification and performance tests have been made using galaxy counts and measurements of galaxy clustering from small subsets of the survey data, and a search for very high redshift quasars. All of the tests were successful and produced new and interesting scientific results. Attachments to this report give detailed accounts of the technical aspects of the SKICAT system, and of some of the scientific results achieved to date. We also developed a user-friendly package for multivariate statistical analysis of small and moderate-size data sets, called STATPROG. The package was tested extensively on a number of real scientific applications and has produced real, published results.

Djorgovski, S. G.

New Neural Network Cloud Mask Algorithm Based on Radiative Transfer Simulations

Cloud detection and screening constitute critically important first steps required to derive many satellite data products. Traditional threshold-based cloud mask algorithms require a complicated design process and fine tuning for each sensor, and they have difficulties over areas partially covered with snow/ice. Exploiting advances in machine learning techniques and radiative transfer modeling of coupled environmental systems, we have developed a new, threshold-free cloud mask algorithm based on a neural network classifier driven by extensive radiative transfer simulations. Statistical validation results obtained by using collocated CALIOP and MODIS data show that its performance is consistent over different ecosystems and significantly better than the MODIS Cloud Mask (MOD35 C6) during the winter seasons over snow-covered areas in the mid-latitudes. Simulations using a reduced number of satellite channels also show satisfactory results, indicating its flexibility to be configured for different sensors. Comparedto threshold-based methods and previous machine-learning approaches, this new cloud mask (i) does not rely on thresholds, (ii) needs fewer satellite channels, (iii) has superior performance during winter seasons in mid-latitude areas, and (iv) can easily be applied to different sensors.

cloud mask algorithms