Search NASA⌕ Search

SEARCH · Search NASA

Results for “STATISTICS”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Predicting lettuce canopy photosynthesis with statistical and neural network models

An artificial neural network (NN) and a statistical regression model were developed to predict canopy photosynthetic rates (Pn) for 'Waldman's Green' leaf lettuce (Latuca sativa L.). All data used to develop and test the models were collected for crop stands grown hydroponically and under controlled-environment conditions. In the NN and regression models, canopy Pn was predicted as a function of three independent variables: shootzone CO2 concentration (600 to 1500 micromoles mol-1), photosynthetic photon flux (PPF) (600 to 1100 micromoles m-2 s-1), and canopy age (10 to 20 days after planting). The models were used to determine the combinations of CO2 and PPF setpoints required each day to maintain maximum canopy Pn. The statistical model (a third-order polynomial) predicted Pn more accurately than the simple NN (a three-layer, fully connected net). Over an 11-day validation period, average percent difference between predicted and actual Pn was 12.3% and 24.6% for the statistical and NN models, respectively. Both models lost considerable accuracy when used to determine relatively long-range Pn predictions (> or = 6 days into the future).

Non-NASA Center↗

Statistical properties of DNA sequences

We review evidence supporting the idea that the DNA sequence in genes containing non-coding regions is correlated, and that the correlation is remarkably long range--indeed, nucleotides thousands of base pairs distant are correlated. We do not find such a long-range correlation in the coding regions of the gene. We resolve the problem of the "non-stationarity" feature of the sequence of base pairs by applying a new algorithm called detrended fluctuation analysis (DFA). We address the claim of Voss that there is no difference in the statistical properties of coding and non-coding regions of DNA by systematically applying the DFA algorithm, as well as standard FFT analysis, to every DNA sequence (33301 coding and 29453 non-coding) in the entire GenBank database. Finally, we describe briefly some recent work showing that the non-coding sequences have certain statistical features in common with natural and artificial languages. Specifically, we adapt to DNA the Zipf approach to analyzing linguistic texts. These statistical properties of non-coding sequences support the possibility that non-coding regions of DNA may carry biological information.

Non-NASA Center↗

Systematic analysis of coding and noncoding DNA sequences using methods of statistical linguistics

We compare the statistical properties of coding and noncoding regions in eukaryotic and viral DNA sequences by adapting two tests developed for the analysis of natural languages and symbolic sequences. The data set comprises all 30 sequences of length above 50 000 base pairs in GenBank Release No. 81.0, as well as the recently published sequences of C. elegans chromosome III (2.2 Mbp) and yeast chromosome XI (661 Kbp). We find that for the three chromosomes we studied the statistical properties of noncoding regions appear to be closer to those observed in natural languages than those of coding regions. In particular, (i) a n-tuple Zipf analysis of noncoding regions reveals a regime close to power-law behavior while the coding regions show logarithmic behavior over a wide interval, while (ii) an n-gram entropy measurement shows that the noncoding regions have a lower n-gram entropy (and hence a larger "n-gram redundancy") than the coding regions. In contrast to the three chromosomes, we find that for vertebrates such as primates and rodents and for viral DNA, the difference between the statistical properties of coding and noncoding regions is not pronounced and therefore the results of the analyses of the investigated sequences are less conclusive. After noting the intrinsic limitations of the n-gram redundancy analysis, we also briefly discuss the failure of the zeroth- and first-order Markovian models or simple nucleotide repeats to account fully for these "linguistic" features of DNA. Finally, we emphasize that our results by no means prove the existence of a "language" in noncoding DNA.

NASA Discipline Number 14-10↗

Examination of two methods for statistical analysis of data with magnitude and direction emphasizing vestibular research applications

When the dependent (or response) variable response variable in an experiment has direction and magnitude, one approach that has been used for statistical analysis involves splitting magnitude and direction and applying univariate statistical techniques to the components. However, such treatment of quantities with direction and magnitude is not justifiable mathematically and can lead to incorrect conclusions about relationships among variables and, as a result, to flawed interpretations. This note discusses a problem with that practice and recommends mathematically correct procedures to be used with dependent variables that have direction and magnitude for 1) computation of mean values, 2) statistical contrasts of and confidence intervals for means, and 3) correlation methods.

Vestibule↗

Statistical Analysis of CFD Solutions from the Fourth AIAA Drag Prediction Workshop

A graphical framework is used for statistical analysis of the results from an extensive N-version test of a collection of Reynolds-averaged Navier-Stokes computational fluid dynamics codes. The solutions were obtained by code developers and users from the U.S., Europe, Asia, and Russia using a variety of grid systems and turbulence models for the June 2009 4th Drag Prediction Workshop sponsored by the AIAA Applied Aerodynamics Technical Committee. The aerodynamic configuration for this workshop was a new subsonic transport model, the Common Research Model, designed using a modern approach for the wing and included a horizontal tail. The fourth workshop focused on the prediction of both absolute and incremental drag levels for wing-body and wing-body-horizontal tail configurations. This work continues the statistical analysis begun in the earlier workshops and compares the results from the grid convergence study of the most recent workshop with earlier workshops using the statistical framework.

Statistical analysis↗

Validating an Air Traffic Management Concept of Operation Using Statistical Modeling

Validating a concept of operation for a complex, safety-critical system (like the National Airspace System) is challenging because of the high dimensionality of the controllable parameters and the infinite number of states of the system. In this paper, we use statistical modeling techniques to explore the behavior of a conflict detection and resolution algorithm designed for the terminal airspace. These techniques predict the robustness of the system simulation to both nominal and off-nominal behaviors within the overall airspace. They also can be used to evaluate the output of the simulation against recorded airspace data. Additionally, the techniques carry with them a mathematical value of the worth of each prediction-a statistical uncertainty for any robustness estimate. Uncertainty Quantification (UQ) is the process of quantitative characterization and ultimately a reduction of uncertainties in complex systems. UQ is important for understanding the influence of uncertainties on the behavior of a system and therefore is valuable for design, analysis, and verification and validation. In this paper, we apply advanced statistical modeling methodologies and techniques on an advanced air traffic management system, namely the Terminal Tactical Separation Assured Flight Environment (T-TSAFE). We show initial results for a parameter analysis and safety boundary (envelope) detection in the high-dimensional parameter space. For our boundary analysis, we developed a new sequential approach based upon the design of computer experiments, allowing us to incorporate knowledge from domain experts into our modeling and to determine the most likely boundary shapes and its parameters. We carried out the analysis on system parameters and describe an initial approach that will allow us to include time-series inputs, such as the radar track data, into the analysis

Statistical emulation↗

Bayesian Statistics and Uncertainty Quantification for Safety Boundary Analysis in Complex Systems

The analysis of a safety-critical system often requires detailed knowledge of safe regions and their highdimensional non-linear boundaries. We present a statistical approach to iteratively detect and characterize the boundaries, which are provided as parameterized shape candidates. Using methods from uncertainty quantification and active learning, we incrementally construct a statistical model from only few simulation runs and obtain statistically sound estimates of the shape parameters for safety boundaries.

Active Learning↗

Canonical Statistical Model for Maximum Expected Immission of Wire Conductor in an Aperture Enclosure

Prediction of the maximum expected electromagnetic pick-up of conductors inside a realistic shielding enclosure is an important canonical problem for system-level EMC design of space craft, launch vehicles, aircraft and automobiles. This paper introduces a simple statistical power balance model for prediction of the maximum expected current in a wire conductor inside an aperture enclosure. It calculates both the statistical mean and variance of the immission from the physical design parameters of the problem. Familiar probability density functions can then be used to predict the maximum expected immission for deign purposes. The statistical power balance model requires minimal EMC design information and solves orders of magnitude faster than existing numerical models, making it ultimately viable for scaled-up, full system-level modeling. Both experimental test results and full wave simulation results are used to validate the foundational model.

Computational electromagnetics↗

Chrono-Validation of Near-Real-Time Landslide Susceptibility Models via Plugin Statistical Simulations

The idea behind any validation scheme in landslide susceptibility studies is to test whether a model calibrated on a certain data can predict an unknown dataset of the same nature (landslide presences/absences and covariates). Almost the entirety of landslide susceptibility studies are validated by subsetting a single dataset into a training and test sets. This dataset usually corresponds either to event-specific or to historical inventories. Very rarely, a multi-temporal inventory is available and, in the few cases where this condition is met, the validation practices involve training a model on a specific landslide inventory, deriving a single predictive equation and validating it on a subsequent landslide inventory. This commonly leads landslide predictive studies, even those with a strong statistical rigor, to neglect the uncertainty estimation in their modeling scheme. In statistics, validation can also be performed via statistical simulations. This means that after fitting a given model, one can generate any number of predictive functions and test their predictive skills on any type and number of unknown datasets. In this work, we take a similar direction and we apply it to model and validate three separate co-seismic inventories, including an uncertainty estimation phase. We mapped these inventories within the same area in Indonesia, for three earthquakes occurred in 2012, 2017 and 2018. Specifically, we build three event-specific Bayesian Generalize Additive Models of the binomial family. From each model we then simulate 1000 predictive realizations over the remaining two inventories, by using a plug-in scheme where all the morphometric covariates are kept fixed and only the ground motion is replaced according to the prediction target. By doing so, we introduce a new analytical tool for near-real-time landslide predictive purposes, which is able to produce a probabilistic model which stands in between the definitions of susceptibility and hazard. In fact, our model is able to accurately estimate “where” and “when” - although not “how frequently” - landslide have occurred by featuring the multitemporal information of the trigger. In our findings, the simulations are quite similar to the fitted models; and the nine combinations we analyse produce excellent performance. This result confirms the assumption that “the past is the key to the future”, as we show that the relative contribution of each variable and their interactions in each probabilistic model remains practically the same across temporal replicates. This information is not trivial because it supports the routines implemented in global near-real-time applications.

Temporal validation↗

System and Safety Analysis with SysAI A Statistical Learning Framework

This is a tutorial on how to use the SYSAI (System Analysis using Statistical AI), a flexible statistical learning framework for the V&V and analysis of complex and high-dimensional Aerospace systems with DNN and AI components. SYSAI provides functionality for a variety of analyses and V&V tasks, including statistical data analysis, high dimensional safety-envelope and time-series analysis, property checking, as well as intelligent test-case generation. The tutorial will demonstrate SYSAI with our industrial partner’s Autonomous Centerline Tracking system, which uses a DNN to enable autonomous aircraft taxiing as an example. Video & Tutorial

Statistical V&V for Complex safety-critical system↗

Recent statistical methods for orientation data

The application of statistical methods for determining the areas of animal orientation and navigation are discussed. The method employed is limited to the two-dimensional case. Various tests for determining the validity of the statistical analysis are presented. Mathematical models are included to support the theoretical considerations and tables of data are developed to show the value of information obtained by statistical analysis.

Batschelet, E.↗

Evaluation of two statistical models using the shock structure problem.

The accuracy of two statistical models for the collision integral of the Boltzmann equation has been evaluated by applying the models to the solution of the problem of shock structure in a monatomic gas and then comparing the theoretical results with available ex perimental data. The two models considered here are the Bhatnagar-Gross-Krook and ellipsoidal statistical models. The Mach number range covered is 1.59-10.7 and profiles for density and, where available, temperature are compared. The method of numerical solution is the discrete ordinate technique which looks quite promising for application to more complicated models. The results indicate that the ellipsoidal statistical model, which gives a correct value for the Prandtl number, gives accurate results for a low Mach number shock. However, the accuracy degenerates as the Mach number increases. The Bhatnagar-Gross-Krook model gives poorer agreement with experimental data in all cases examined.

Giddens, D. P.↗

Site preferences of Ni/2+/ and Co/2+/ in clinopyroxene and olivine - Limitations of the statistical approach.

Criticism of a statistical approach used by Dasgupta (1972) in analyzing Snyder's (1959) chemical data for minerals from the Duluth Complex in Minnesota. Apart from obvious mathematical objections to citing correlation coefficients to four significant figures from Snyder's relatively inaccurate analytical data, several more fundamental criticisms are leveled at the statistical approach of Dasgupta. These relate to compositional zoning and disequilibrium in the minerals, inhomogeneities of the samples caused by inclusions and exsolved phases, measured site population data for the major cations in olivines and pyroxenes, and the importance of coupled substitutions in the crystal structures. It is concluded that the crystal field predictions of relative enrichments of Ni(2+) and Co(2+) ions in olivine and pyroxene structures have not been disproved by Dasgupta's statistical approach.

Burns, R. G.↗

Statistical analysis of close pairs of QSOs

The observation of close pairs of QSOs with very different redshifts has been suggested by some as evidence in support of the noncosmological redshift hypothesis. A method is described for determining the statistical significance of such pairs. As an example, it is shown that the statistical significance of the pair 1548+115a,b is not well defined and ranges from approximately 99% confidence to about 60%. If statistical methods are to be used in such cases, they must not be argued a posteriori.

Burbidge, E. M.↗

Statistical separability of spectral classes of blighted corn

A study was conducted to determine the statistical separability of multispectral measurements from corn having varying levels of southern corn leaf blight severity. Multispectral scanner data in twelve spectral channels in the wavelength range 0.4 to 11.7 microns were analyzed for ten selected flightlines of the 1971 Corn Blight Watch Experiment. A total of 168 corn fields having 18,804 sample points were analyzed. The blight rating information for these fields was available from ground observations. Maximum average transformed divergence between spectral classes of all possible pairs of blight levels, maximized over a subset of channels, was computed in each of one, two, three, and four spectral channels for each of ten flightlines. From the statistical analysis of the values of average transformed divergence, it was concluded that the greater the difference between the blight levels, the more statistically separable they are.

Kumar, R.↗

Statistical separability of agricultural cover types in subsets of one to twelve spectral channels

The purpose of this study was to determine the statistical separability of multispectral measurements from agricultural cover types: corn, soybeans, green forage (hay and pasture) and forest, in one to twelve spectral channels. Multispectral scanner data in twelve spectral channels in the wavelength range 0.4 to 11.7 microns, acquired for three flightlines were analysed by applying automatic pattern recognition techniques. The same analysis was performed for the data acquired a month later over the same three flightlines to investigate the effect of time on statistical separability of agricultural cover types. In the subsets of one to six spectral channels, the combination of wavelength regions (where V, N, M and T denote the visible, near infrared, middle infrared and thermal infrared wavelength regions, respectively): V, V M, V N M, V N M T, V V N M T, V V N M M T, respectively, were found to be the best choices for getting good overall statistical separability of the agricultural cover types for the data acquired.

Kumar, R.↗

Radar Derived Spatial Statistics of Summer Rain: Data Reduction and Analysis - Volume 2

Data reduction and analysis procedures are discussed along with the physical and statistical descriptors used. The statistical modeling techniques are outlined and examples of the derived statistical characterization of rain cells in terms of the several physical descriptors are presented. Recommendations concerning analyses which can be pursued using the data base collected during the experiment are included.

Konrad, T. G.↗

Statistical properties of the interplanetary microscale fluctuations

Results are reported for a statistical study of short-period fluctuations in the interplanetary plasma velocity and magnetic field. The data base used consists of measurements of the interplanetary plasma and magnetic field by Pioneer 6 with a time resolution of 72 sec and by Mariner 5 with a resolution of 5 min. The analysis is conducted to characterize the parent population from which all individual microscale events on these time scales are drawn. The microscale changes are grouped according to their energy densities relative to the energy density of the background magnetic field, and it is found that each grouping exhibits certain statistical properties within the limits of observational uncertainty. It is noted that these statistical properties are purely observational and independent of any physical model which may be used to interpret them. The results are discussed in the context of MHD discontinuity theory for a thermally anisotropic plasma.

Belcher, J. W.↗