Search NASA⌕ Search

SEARCH · Search NASA

Results for “STATISTICS”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Statistical Hypothesis Testing in Wavelet Analysis: Theoretical Developments and Applications to Indian Rainfall

Statistical hypothesis tests in wavelet analysis are methods that assess the degree to which a wavelet quantity(e.g., power and coherence) exceeds background noise. Commonly, a point-wise approach is adopted in which a wavelet quantity at every point in a wavelet spectrum is individually compared to the critical level of the point-wise test. However, because adjacent wavelet coefficients are correlated and wavelet spectra often contain many wavelet quantities, the point-wise test can produce many false positive results that occur in clusters or patches. To circumvent the point-wise test drawbacks, it is necessary to implement the recently developed area-wise, geometric, cumulative area-wise, and topological significance tests, which are reviewed and developed in this paper. To improve the computational efficiency of the cumulative area-wise test, a simplified version of the testing procedure is created based on the idea that its output is the mean of individual estimates of statistical significance calculated from the geometric test applied at a set of point-wise significance levels. Ideal examples are used to show that the geometric and cumulative area-wise tests are unable to differentiate wavelet spectral features arising from singularity-like structures from those associated with periodicities. A cumulative arc-wise test is therefore developed to strictly test for periodicities by using normalized arc length, which is defined as the number of points composing a cross section of a patch divided by the wavelet scale in question. A previously proposed topological significance test is formalized using persistent homology profiles (PHPs) measuring the number of patches and holes corresponding to the set of all point-wise significance values. Ideal examples show that the PHPs can be used to distinguish time series containing signal components from those that are purely noise. To demonstrate the practical uses of the existing and newly developed statistical methodologies, a first comprehensive wavelet analysis of Indian rainfall is also provided. An R software package has been written by the author to implement the various testing procedures.

Wind speed↗

A Global Model for Estimating Atmospheric Phase Scintillation Statistics

Since 2007, the National Aeronautics and Space Administration (NASA) has been collecting atmospheric phase turbulence data from various NASA ground stations throughout the world. The goal of these measurement campaigns has been to generate statistics to characterize the local site turbulence conditions and their impact on widely distributed ground based antenna arrays. This is of critical importance for the situation of uplink arraying, in which a priori knowledge of the fast varying turbulent conditions of water vapor in the troposphere may not be known, and will impact the power combining efficiency of ground based transmitting arrays. Therefore, the design of these type of systems will be dependent on the local climatology of the particular ground station site. Based on the 30+ station years of data collected characterizing atmospheric phase scintillation statistics at various sites, a global model is presented which attempts to predict the average phase statistics of a generic site based on local surface weather data, such as surface pressure, temperature, relative humidity, wind speed, and median wind direction. A model is proposed and based on a standard log power distribution similar to amplitude scintillation models trained on the existing data sets and shows reasonable accuracy against existing data sets.

propagation↗

Statistical learning framework for safety and failure analysis of a DNN-based autonomous aircraft system

Deep Neural Networks (DNNs) and Machine Learning technology is increasingly used for safety-critical applications in the Aerospace domain. To ensure safe operations, the DNN and the system must undergo rigorous verification and validation, including advanced statistical analyses. Performance and safety of the DNN and system behavior must not only be analyzed for the nominal case, but under numerous off-nominal and failure cases. In this paper we will describe how our statistical learning framework SYSAI can efficiently perform such analyses using the tool’s unique combination of advanced learning modeling and statistical analysis techniques. SYSAI can effectively explore the high-dimensional state and failure space of the system under test; geometrical shape detection of safety regions and boundaries support explainability of the results to the designer. In this paper, we report experiments and results obtained with a vision-based DNN control system (ACT) that is capable of autonomously steering an aircraft down a runway.

Yuning He↗

Truncated ARQ Statistical Link Analysis for Dynamic Links

The future deep space links are migrating towards higher frequency bands such as Ka band and optical. These links are susceptible to non Gaussian and non linear effects such as atmospheric turbulence, scintillation, antenna mis-pointing, jitter, etc. These dynamic links thus will experience various degrees of fading loss, and some of these link disruptions cannot be effectively mitigated by forward error correction coding and/or interleaving. One effective way to ensure reliable communication is by using Automatic Repeat Request (ARQ) protocol, where the receiver acknowledges to the transmitter whether or not a data unit is successfully received. If a data unit is not successfully received (such as after a pre-set time-out), the transmitter would then re-transmit the lost data unit to the receiver. In a previous paper, we derived a statistical link analysis method of finding the optimal operating Signal-to-Noise Ratio (SNR) and estimating the latency of an ARQ scheme. In a more recent paper, we demonstrated the above method using the SNR distribution constructed from the Ka-band (32 GHz) flight data. To simplify the discussion, we considered the academic approach that the ARQ scheme allows for an infinite number of retransmissions. In this paper, we consider the more practical case of a truncated ARQ scheme, where there is a limit on the number of retransmissions. We derive the error probability, the optimal SNR setting, and the latency statistics of the correctly received frames of the truncated ARQ schemes. We first discuss the truncated ARQ link analysis principles using the Gaussian assumption for SNR distribution with a large variance. Next, we demonstrate the statistical truncated ARQ link analysis using the SNR distribution constructed from the Ka-band flight data. The results in this paper can be applied in the design of reliable communication systems such as the Consultative Committee for Space Data System (CCSDS) File Transfer Protocol (CFTP) and the Delay Tolerant Network (DTN).

Morabito, David↗

Markovian Statistical Model of Cloud Optical Thickness. Part I: Theory and Examples

We present a generalization of the binary-value Markovian model previously used for statistical characterization of cloud masks to a continuous-value model describing 1D fields of cloud optical thickness (COT). This model has simple functional expressions and is specified by four parameters: the cloud fraction, the autocorrelation (scale) length, and the two parameters of the normalized probability density function of (non-zero) COT values (this PDF is assumed to have gamma-distribution form). Cloud masks derived from this model by separation between the values above and below some threshold in COT appear to have the same statistical properties as in binary-value model described in our previous publications. We demonstrate the ability of our model to generate examples of various cloud-field types by using it to statistically imitate actual cloud observations made by the Research Scanning Polarimeter (RSP) during two field experiments.

binary-value Markovian model↗

Interpretable Machine Learning for Molecular Biosignatures: a Novel Single-Sample Feature Importance Method That Is Sensitive To Statistical Interactions

Isotope ratio mass spectrometry (IRMS) of volatiles (e.g., CO 2 ) promises to be a powerful tool for potential biosignature detection for future missions to ocean worlds (OW) such as Europa and Enceladus. Machine learning (ML) methods for IRMS data could enable science autonomy by onboard prediction of seawater chemistry and biosignature presence. However, ML models are likely to be complex and involve statistical interactions between features (variables), which can make predictions seem opaque and enigmatic. For ML predictions as significant as extraterrestrial biosignatures, we must place extraordinary confidence in models. It is therefore essential that these models make interpretable predictions (i.e., human-understandable) and include false-prediction diagnostics. We achieve high accuracy and interpretability in ML biosignature and seawater chemistry models for OW through a nearest-neighbors feature selection tool that detects statistical interactions between predictors, constructs interaction networks for visualization of selected features working together to make a prediction, and reports single-sample feature importance scores for false-detection diagnostics. Here we develop a novel single-sample nearest-neighbors projected distance regression(ssNPDR) feature selection method that improves upon existing single-sample algorithms through the inclusion of statistical interactions while providing false-prediction diagnostics for ML models.

geochemistry↗

A statistical model for interpreting computerized dynamic posturography data

Computerized dynamic posturography (CDP) is widely used for assessment of altered balance control. CDP trials are quantified using the equilibrium score (ES), which ranges from zero to 100, as a decreasing function of peak sway angle. The problem of how best to model and analyze ESs from a controlled study is considered. The ES often exhibits a skewed distribution in repeated trials, which can lead to incorrect inference when applying standard regression or analysis of variance models. Furthermore, CDP trials are terminated when a patient loses balance. In these situations, the ES is not observable, but is assigned the lowest possible score--zero. As a result, the response variable has a mixed discrete-continuous distribution, further compromising inference obtained by standard statistical methods. Here, we develop alternative methodology for analyzing ESs under a stochastic model extending the ES to a continuous latent random variable that always exists, but is unobserved in the event of a fall. Loss of balance occurs conditionally, with probability depending on the realized latent ES. After fitting the model by a form of quasi-maximum-likelihood, one may perform statistical inference to assess the effects of explanatory variables. An example is provided, using data from the NIH/NIA Baltimore Longitudinal Study on Aging.

NASA Discipline Neuroscience↗

Statistical and linguistic features of DNA sequences

We present evidence supporting the idea that the DNA sequence in genes containing noncoding regions is correlated, and that the correlation is remarkably long range--indeed, base pairs thousands of base pairs distant are correlated. We do not find such a long-range correlation in the coding regions of the gene. We resolve the problem of the "non-stationary" feature of the sequence of base pairs by applying a new algorithm called Detrended Fluctuation Analysis (DFA). We address the claim of Voss that there is no difference in the statistical properties of coding and noncoding regions of DNA by systematically applying the DFA algorithm, as well as standard FFT analysis, to all eukaryotic DNA sequences (33 301 coding and 29 453 noncoding) in the entire GenBank database. We describe a simple model to account for the presence of long-range power-law correlations which is based upon a generalization of the classic Levy walk. Finally, we describe briefly some recent work showing that the noncoding sequences have certain statistical features in common with natural languages. Specifically, we adapt to DNA the Zipf approach to analyzing linguistic texts, and the Shannon approach to quantifying the "redundancy" of a linguistic text in terms of a measurable entropy function. We suggest that noncoding regions in plants and invertebrates may display a smaller entropy and larger redundancy than coding regions, further supporting the possibility that noncoding regions of DNA may carry biological information.

Non-NASA Center↗

A statistical model of the human core-temperature circadian rhythm

We formulate a statistical model of the human core-temperature circadian rhythm in which the circadian signal is modeled as a van der Pol oscillator, the thermoregulatory response is represented as a first-order autoregressive process, and the evoked effect of activity is modeled with a function specific for each circadian protocol. The new model directly links differential equation-based simulation models and harmonic regression analysis methods and permits statistical analysis of both static and dynamical properties of the circadian pacemaker from experimental data. We estimate the model parameters by using numerically efficient maximum likelihood algorithms and analyze human core-temperature data from forced desynchrony, free-run, and constant-routine protocols. By representing explicitly the dynamical effects of ambient light input to the human circadian pacemaker, the new model can estimate with high precision the correct intrinsic period of this oscillator ( approximately 24 h) from both free-run and forced desynchrony studies. Although the van der Pol model approximates well the dynamical features of the circadian pacemaker, the optimal dynamical model of the human biological clock may have a harmonic structure different from that of the van der Pol oscillator.

NASA Discipline Regulatory Physiology↗

Statistical Analysis of CFD Solutions From the Fifth AIAA Drag Prediction Workshop

A graphical framework is used for statistical analysis of the results from an extensive N-version test of a collection of Reynolds-averaged Navier-Stokes computational fluid dynamics codes. The solutions were obtained by code developers and users from North America, Europe, Asia, and South America using a common grid sequence and multiple turbulence models for the June 2012 fifth Drag Prediction Workshop sponsored by the AIAA Applied Aerodynamics Technical Committee. The aerodynamic configuration for this workshop was the Common Research Model subsonic transport wing-body previously used for the 4th Drag Prediction Workshop. This work continues the statistical analysis begun in the earlier workshops and compares the results from the grid convergence study of the most recent workshop with previous workshops.

Statistical analysis↗

Ensemble Statistical Post-Processing of the National Air Quality Forecast Capability: Enhancing Ozone Forecasts in Baltimore, Maryland

An ensemble statistical post-processor (ESP) is developed for the National Air Quality Forecast Capability (NAQFC) to address the unique challenges of forecasting surface ozone in Baltimore, MD. Air quality and meteorological data were collected from the eight monitors that constitute the Baltimore forecast region. These data were used to build the ESP using a moving-block bootstrap, regression tree models, and extreme-value theory. The ESP was evaluated using a 10-fold cross-validation to avoid evaluation with the same data used in the development process. Results indicate that the ESP is conditionally biased, likely due to slight overfitting while training the regression tree models. When viewed from the perspective of a decision-maker, the ESP provides a wealth of additional information previously not available through the NAQFC alone. The user is provided the freedom to tailor the forecast to the decision at hand by using decision-specific probability thresholds that define a forecast for an ozone exceedance. Taking advantage of the ESP, the user not only receives an increase in value over the NAQFC, but also receives value for An ensemble statistical post-processor (ESP) is developed for the National Air Quality Forecast Capability (NAQFC) to address the unique challenges of forecasting surface ozone in Baltimore, MD. Air quality and meteorological data were collected from the eight monitors that constitute the Baltimore forecast region. These data were used to build the ESP using a moving-block bootstrap, regression tree models, and extreme-value theory. The ESP was evaluated using a 10-fold cross-validation to avoid evaluation with the same data used in the development process. Results indicate that the ESP is conditionally biased, likely due to slight overfitting while training the regression tree models. When viewed from the perspective of a decision-maker, the ESP provides a wealth of additional information previously not available through the NAQFC alone. The user is provided the freedom to tailor the forecast to the decision at hand by using decision-specific probability thresholds that define a forecast for an ozone exceedance. Taking advantage of the ESP, the user not only receives an increase in value over the NAQFC, but also receives value for

ozone↗

Statistical Analysis of Instantaneous Frequency Scaling Factor as Derived From Optical Disdrometer Measurements At KQ Bands

The rain rate data and statistics of a location are often used in conjunction with models to predict rain attenuation. However, the true attenuation is a function not only of rain rate, but also of the drop size distribution (DSD). Generally, models utilize an average drop size distribution (Laws and Parsons or Marshall and Palmer [1]). However, individual rain events may deviate from these models significantly if their DSD is not well approximated by the average. Therefore, characterizing the relationship between the DSD and attenuation is valuable in improving modeled predictions of rain attenuation statistics. The DSD may also be used to derive the instantaneous frequency scaling factor and thus validate frequency scaling models. Since June of 2014, NASA Glenn Research Center (GRC) and the Politecnico di Milano (POLIMI) have jointly conducted a propagation study in Milan, Italy utilizing the 20 and 40 GHz beacon signals of the Alphasat TDP#5 Aldo Paraboni payload. The Ka- and Q-band beacon receivers provide a direct measurement of the signal attenuation while concurrent weather instrumentation provides measurements of the atmospheric conditions at the receiver. Among these instruments is a Thies Clima Laser Precipitation Monitor (optical disdrometer) which yields droplet size distributions (DSD); this DSD information can be used to derive a scaling factor that scales the measured 20 GHz data to expected 40 GHz attenuation. Given the capability to both predict and directly observe 40 GHz attenuation, this site is uniquely situated to assess and characterize such predictions. Previous work using this data has examined the relationship between the measured drop-size distribution and the measured attenuation of the link [2]. The focus of this paper now turns to a deeper analysis of the scaling factor, including the prediction error as a function of attenuation level, correlation between the scaling factor and the rain rate, and the temporal variability of the drop size distribution both within a given rain event and across different varieties of rain events. Index Terms-drop size distribution, frequency scaling, propagation losses, radiowave propagation.

Radio Attenuation↗

Statistical Analysis of Instantaneous Frequency Scaling Factor as Derived From Optical Disdrometer Measurements At KQ Bands

The rain rate data and statistics of a location are often used in conjunction with models to predict rain attenuation. However, the true attenuation is a function not only of rain rate, but also of the drop size distribution (DSD). Generally, models utilize an average drop size distribution (Laws and Parsons or Marshall and Palmer. However, individual rain events may deviate from these models significantly if their DSD is not well approximated by the average. Therefore, characterizing the relationship between the DSD and attenuation is valuable in improving modeled predictions of rain attenuation statistics. The DSD may also be used to derive the instantaneous frequency scaling factor and thus validate frequency scaling models. Since June of 2014, NASA Glenn Research Center (GRC) and the Politecnico di Milano (POLIMI) have jointly conducted a propagation study in Milan, Italy utilizing the 20 and 40 GHz beacon signals of the Alphasat TDP#5 Aldo Paraboni payload. The Ka- and Q-band beacon receivers provide a direct measurement of the signal attenuation while concurrent weather instrumentation provides measurements of the atmospheric conditions at the receiver. Among these instruments is a Thies Clima Laser Precipitation Monitor (optical disdrometer) which yields droplet size distributions (DSD); this DSD information can be used to derive a scaling factor that scales the measured 20 GHz data to expected 40 GHz attenuation. Given the capability to both predict and directly observe 40 GHz attenuation, this site is uniquely situated to assess and characterize such predictions. Previous work using this data has examined the relationship between the measured drop-size distribution and the measured attenuation of the link]. The focus of this paper now turns to a deeper analysis of the scaling factor, including the prediction error as a function of attenuation level, correlation between the scaling factor and the rain rate, and the temporal variability of the drop size distribution both within a given rain event and across different varieties of rain events. Index Terms-drop size distribution, frequency scaling, propagation losses, radiowave propagation.

Statistics↗

Statistical Analysis of CFD Solutions from the 6th AIAA CFD Drag Prediction Workshop

A graphical framework is used for statistical analysis of the results from an extensive N-version test of a collection of Reynolds-averaged Navier-Stokes computational fluid dynamics codes. The solutions were obtained by code developers and users from North America, Europe, Asia, and South America using both common and custom grid sequences as well as multiple turbulence models for the June 2016 6th AIAA CFD Drag Prediction Workshop sponsored by the AIAA Applied Aerodynamics Technical Committee. The aerodynamic configuration for this workshop was the Common Research Model subsonic transport wing-body previously used for both the 4th and 5th Drag Prediction Workshops. This work continues the statistical analysis begun in the earlier workshops and compares the results from the grid convergence study of the most recent workshop with previous workshops.

Statistical analysis↗

Relevance of the American Statistical Society's Warning on P-Values for Conjunction Assessment

On March 7, 2016, the American Statistical Association (ASA) issued an online editorial paper on the "context, process, and purpose of p-values." According to the paper, "the statement articulates in non-technical terms a few select principles that could improve the conductor interpretation of quantitative science, according to widespread consensus in the statistical community." This "consensus" statement was accompanied by 21 "commentaries" expressing a diversity of opinions among the panel of experts ASA convened. In the present work, we express our view that the ASA p-value warning has relevance to the space object Conjunction Assessment (CA) community.

p-values↗

Statistical Descriptors of Composite Fiber Aggregation

This study introduces a method of characterizing fiber aggregation and resin rich regions in composite microstructures. Microscale models of representative elements (RVE) need to be indicative of the extend of clustering (i.e. close fiber-to-fiber interaction) and resin rich “pools” which may impact the overall strength and performance of a composite structure. This algorithm was used to evaluate different unidirectional 2-D microstructure scans, which will be compared to their manufacturing method or any special treatment processes. These cluster and pool scan statistics can be used as criteria to judge statistical equivalency of artificially constructed microstructures.

Statistics↗

Artificial Generation of 2-D Fiber Reinforced Composite Microstructures with Statistically Equivalent Features

Fiber reinforced composites are used widely for their high strength and low weight advantages in various aerospace and automotive applications. While their use may be sought after, modeling of these material requires increasing fidelity at the lower scales to capture accurate material behavior under loading. The first steps in creating statistically equivalent models to real life cases is developing a method of rapid evaluation and artificial microstructure generation. The outlined work is capable of tracking microscale fiber positions and determining regions of localized volume fraction extrema (high and low end). Groupings of high and low volume fraction regions are called clusters and their geometry is used to characterize the microstructure. These cluster features can be evaluated for both artificial models and actual scans, allowing correlation to be established which can ultimately be used to regenerate statistically equivalent models. The results of this work show that if one feature is to be correlated, a model can be generated which matches almost exactly. But once more features are equally taken into account, the regeneration loses accuracy.

micromechanics↗