Search NASA⌕ Search

SEARCH · Search NASA

Results for “data statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37

Earth Observations into Action: Systemic Integration of Earth Observation Applications into National Risk Reduction Decision Structures

As stated in the United Nations Global Assessment Report 2022 Concept Note, decision makers everywhere need data and statistics that are accurate, timely, sufficiently disaggregated, relevant, accessible, and easy to use. The purpose of this paper is to demonstrate scalable and replicable methods to advance and integrate the use of Earth observation, specifically ongoing efforts within the Group on Earth Observations Work Programme and the Committee on Earth Observation Satellites Work Plan, to support risk-informed decision making, based on documented national and subnational needs and requirements.

earth observations↗

Southern Rockies Western Slope Agriculture: Identifying Drivers of Rangeland Production for Drought Planning on the Western Slope of the Southern Rockies

Over the last decade, the southern Rocky Mountains of the United States have experienced increasingly severe and variable drought. Local ranchers and landowners have reported strain on their operations, citing decreasing forage production for their cattle and a need to adjust their business models, even considering abandoning their businesses altogether. The study identified Major Land Resource Area-48 (MLRA-48) and northwestern Colorado as the key region for analysis. NASA DEVELOP partnered with the BLM Colorado River Field Office, Colorado State University Extension, USDA Forest Service, and the National Drought Mitigation Center to address stakeholder concerns of the efficacy of existing remotely sensed rangeland production estimation platforms and explore possible early warning climatic indicators of drought. The study identified two key rangeland platforms, the Rangeland Production Monitoring Service (RPMS) and Rangeland Analysis Platform (RAP) and used in-situ data to statistically validate their efficacy. RAP outperformed RPMS in estimating in-situ biomass and was therefore used in our climate modeling. Our study performed a random forest analysis, sampling 1500 points across the study area, comparing monthly RAP biomass estimates to a variety of climatic variables, including mean precipitation, temperature, palmer drought severity index, snow water equivalent, wind speed and direction, and vapor pressure deficit. After analysis, our study determined that vapor pressure deficit is a key indicator in predicting forage production in MLRA-48. Our study recommends the use of RAP in estimating potential forage, with caution for its tendency to overestimate. Our climate analysis provided our partners with greater understanding of the influence of various climatic factors in determining forage production and allows them to assist landowners in planning for future drought.

Addie Gonzalez↗

Landslide Hazard is Projected to Increase Across High Mountain Asia

High Mountain Asia has long been known as a hotspot for landslide risk, and studies have suggested that landslide hazard is likely to increase in this region over the coming decades. Extreme precipitation may become more frequent, with a nonlinear response relative to increasing global temperatures. However, these changes are geographically varied. This article maps probable changes to landslide hazard, as shown by a landslide hazard indicator (LHI) derived from downscaled precipitation and temperature. In order to capture the nonlinear response of slopes to extreme precipitation, a simple machine-learning model was trained on a database of landslides across High Mountain Asia to develop a regional LHI. This model was applied to statistically downscaled data from the 30 members of the Seamless System for Prediction and Earth System Research large ensembles to produce a range of possible outcomes under the Shared Socioeconomic Pathways 2-4.5 and 5-8.5. The LHI reveals that landslide hazard will increase in most parts of High Mountain Asia. Absolute increases will be highest in already hazardous areas such as the Central Himalaya, but relative change is greatest on the Tibetan Plateau. Even in regions where landslide hazard declines by year 2100, it will increase prior to the mid-century mark. However, the seasonal cycle of landslide occurrence will not change greatly across High Mountain Asia. Although substantial uncertainty remains in these projections, the overall direction of change seems reliable. These findings highlight the importance of continued analysis to inform disaster risk reduction strategies for stakeholders across High Mountain Asia.

Thomas A Stanley↗

System Engineers and Decisions: It?s All about Knowledge

In order to guarantee that a system meets adequate levels of reliability and availability, system performances are continuously monitored and analyzed thanks to the technological advancements driving the Industry 4.0 revolution. An Industry 4.0 approach is typically based on advanced statistical, big data mining, machine learning, and internet-of-things methods designed to detect anomalies in the behavior of system, detect the most likely failure modes, and provide indications to system engineers on when maintenance activities should be performed before system performance are deemed unacceptable (which can be generated by diagnostic and prognostic methods). However, these analyses, which are designed to automatize and increase the efficacy of the system maintenance program, require large amount of data which can come in various forms: numeric, textual, images, sounds etc. Such data constitutes the historic knowledge benchmark to track system performances and support system engineer decisions. Here we claim that data is not sufficient to support this kind of analyses when applied to systems characterized by complex architectures and behaviors. Robust system engineer decisions require the ability to understand the system operational context that lies behind the observed data elements. In this respect, system models are in fact necessary to “put data in context” and capture relationships between data elements. Industry 4.0 methods require in fact contextual knowledge as a basis upon which hypotheses can be generated and assumptions tested. In our view, for complex systems, model-based system engineering (MBSE) models can afford this contextual knowledge, as they are typically used to describe systems architecture and dynamic behaviors. System knowledge is here intended as the blending of collected data and system architecture which takes the form of a “knowledge graph”. A knowledge graph is a database which consists of a large set of nodes (in our case an entity can be either a data or an MBSE element) which are linked to each other. The types of nodes and links follow a pre-defined topology, sometimes also refers as an ontology, that is designed to fit the actual decisions that needs to be performed. We show here how a knowledge graph can be defined to support system engineer maintenance decisions and how the same graph can be built based on system MBSE models and pre-processed data from numeric (through anomaly detections and diagnostic methods) and textual elements (through technical language processing TLP).

97 - MATHEMATICS AND COMPUTING↗

Application of Statistical Methods of Rain Rate Estimation to Data From The TRMM Precipitation Radar

The TRMM Precipitation Radar is well suited to statistical methods in that the measurements over any given region are sparsely sampled in time. Moreover, the instantaneous rain rate estimates are often of limited accuracy at high rain rates because of attenuation effects and at light rain rates because of receiver sensitivity. For the estimation of the time-averaged rain characteristics over an area both errors are relevant. By enlarging the space-time region over which the data are collected, the sampling error can be reduced. However. the bias and distortion of the estimated rain distribution generally will remain if estimates at the high and low rain rates are not corrected. In this paper we use the TRMM PR data to investigate the behavior of 2 statistical methods the purpose of which is to estimate the rain rate over large space-time domains. Examination of large-scale rain characteristics provides a useful starting point. The high correlation between the mean and standard deviation of rain rate implies that the conditional distribution of this quantity can be approximated by a one-parameter distribution. This property is used to explore the behavior of the area-time-integral (ATI) methods where fractional area above a threshold is related to the mean rain rate. In the usual application of the ATI method a correlation is established between these quantities. However, if a particular form of the rain rate distribution is assumed and if the ratio of the mean to standard deviation is known, then not only the mean but the full distribution can be extracted from a measurement of fractional area above a threshold. The second method is an extension of this idea where the distribution is estimated from data over a range of rain rates chosen in an intermediate range where the effects of attenuation and poor sensitivity can be neglected. The advantage of estimating the distribution itself rather than the mean value is that it yields the fraction of rain contributed by the light and heavy rain rates. This is useful in estimating the fraction of rainfall contributed by the rain rates that go undetected by the radar. The results at high rain rates provide a cross-check on the usual attenuation correction methods that are applied at the highest resolution of the instrument.

Meneghini, R.↗

Determination of monthly mean humidity in the atmospheric surface layer over oceans from satellite data

A simple statistical technique is described to determine monthly mean marine surface-layer humidity, which is essential in the specification of surface latent heat flux, from total water vapor in the atmospheric column measured by space-borne sensors. Good correlation between the two quantities was found in examining the humidity soundings from radiosonde reports of mid-ocean island stations and weather ships. The relation agrees with that obtained from satellite (Seasat) data and ship reports averaged over 2 deg areas and a 92-day period in the North Atlantic and in the tropical Pacific. The results demonstrate that, by using a local regression in the tropical Pacific, total water vapor can be used to determine monthly mean surface layer humidity to an accuracy of 0.4 g/kg. With a global regression, determination to an accuracy of 0.8 g/kg is possible. These accuracies correspond to approximately 10 to 20 W/sq m in the determination of latent heat flux with the bulk parameterization method, provided that other required parameters are known.

Liu, W. T.↗

Interrelationships Between Receiver/Relative Operating Characteristics Display, Binomial, Logit, and Bayes' Rule Probability of Detection Methodologies

Unknown risks are introduced into failure critical systems when probability of detection (POD) capabilities are accepted without a complete understanding of the statistical method applied and the interpretation of the statistical results. The presence of this risk in the nondestructive evaluation (NDE) community is revealed in common statements about POD. These statements are often interpreted in a variety of ways and therefore, the very existence of the statements identifies the need for a more comprehensive understanding of POD methodologies. Statistical methodologies have data requirements to be met, procedures to be followed, and requirements for validation or demonstration of adequacy of the POD estimates. Risks are further enhanced due to the wide range of statistical methodologies used for determining the POD capability. Receiver/Relative Operating Characteristics (ROC) Display, simple binomial, logistic regression, and Bayes' rule POD methodologies are widely used in determining POD capability. This work focuses on Hit-Miss data to reveal the framework of the interrelationships between Receiver/Relative Operating Characteristics Display, simple binomial, logistic regression, and Bayes' Rule methodologies as they are applied to POD. Knowledge of these interrelationships leads to an intuitive and global understanding of the statistical data, procedural and validation requirements for establishing credible POD estimates.

Generazio, Edward R.↗

Statistical Treatment of Earth Observing System Pyroshock Separation Test Data

The Earth Observing System (EOS) AM-1 spacecraft for NASA's Mission to Planet Earth is scheduled to be launched on an Atlas IIAS vehicle in June of 1998. One concern is that the instruments on the EOS spacecraft are sensitive to the shock-induced vibration produced when the spacecraft separates from the launch vehicle. By employing unique statistical analysis to the available ground test shock data, the NASA Lewis Research Center found that shock-induced vibrations would not be as great as the previously specified levels of Lockheed Martin. The EOS pyroshock separation testing, which was completed in 1997, produced a large quantity of accelerometer data to characterize the shock response levels at the launch vehicle/spacecraft interface. Thirteen pyroshock separation firings of the EOS and payload adapter configuration yielded 78 total measurements at the interface. The multiple firings were necessary to qualify the newly developed Lockheed Martin six-hardpoint separation system. Because of the unusually large amount of data acquired, Lewis developed a statistical methodology to predict the maximum expected shock levels at the interface between the EOS spacecraft and the launch vehicle. Then, this methodology, which is based on six shear plate accelerometer measurements per test firing at the spacecraft/launch vehicle interface, was used to determine the shock endurance specification for EOS. Each pyroshock separation test of the EOS spacecraft simulator produced its own set of interface accelerometer data. Probability distributions, histograms, the median, and higher order moments (skew and kurtosis) were analyzed. The data were found to be lognormally distributed, which is consistent with NASA pyroshock standards. Each set of lognormally transformed test data produced was analyzed to determine if the data should be combined statistically. Statistical testing of the data's standard deviations and means (F and t testing, respectively) determined if data sets were significantly different at a 95-percent confidence level. If two data sets were found to be significantly different, these families of data were not combined for statistical purposes. This methodology produced three separate statistical data families of shear plate data. For each population, a P99.1/50 (probability/confidence) per-separation-nut firing level was calculated. By using the binomial distribution, Lewis researchers determined that this pernut firing level was equivalent to a P95/50 per-flight confidence level. The overall envelope of the per-flight P95/50 levels led to Lewis' recommended EOS interface shock endurance specification. A similar methodology was used to develop Lewis' recommended EOS mission assurance levels. The available test data for the EOS mission are significantly larger than for a normal mission, thus increasing the confidence level in the calculated expected shock environment. Lewis significantly affected the EOS mission by properly employing statistical analysis to the data. This analysis prevented a costly requalification of the spacecraft's instruments, which otherwise would have been exposed to significantly higher test levels.

McNelis, Anne M.↗

The anomaly data base of screwworm information

Standard statistical processing of anomaly data in the screwworm eradication data system is possible from data compiled on magnetic tapes with the Univac 1108 computer. The format and organization of the data in the data base, which is also available on dedicated disc storage, are described.

Giddings, L. E.↗

Average ozone vertical distribution at Sodankyla based on the 1988-1991 ozone sounding data

The study presents the statistical analysis of ozone sonde data obtained at Sodankyla (67.4 deg N, 26.6 deg E) from the beginning of the sounding program on March 1988 to the end of December 1991. The Sodankyla sounding data offers the longest continuous record of the ozone vertical distribution in the European Arctic. In this paper, we present the average ozone partial pressures within each 1 km column obtained for different seasons during the almost four year long period. We believe that the data represented here are useful as an interim reference ozone atmosphere, especially considering the fact that northern Scandinavia has become a popular campaign site for the big international ozone experiments.

Kyro, Esko↗

Comparison of Grammar-Based and Statistical Language Models Trained on the Same Data

This paper presents a methodologically sound comparison of the performance of grammar-based (GLM) and statistical-based (SLM) recognizer architectures using data from the Clarissa procedure navigator domain. The Regulus open source packages make this possible with a method for constructing a grammar-based language model by training on a corpus. We construct grammar-based and statistical language models from the same corpus for comparison, and find that the grammar-based language models provide better performance in this domain. The best SLM version has a semantic error rate of 9.6%, while the best GLM version has an error rate of 6.0%. Part of this advantage is accounted for by the superior WER and Sentence Error Rate (SER) of the GLM (WER 7.42% versus 6.27%, and SER 12.41% versus 9.79%). The rest is most likely accounted for by the fact that the GLM architecture is able to use logical-form-based features, which permit tighter integration of recognition and semantic interpretation.

Hockey, Beth Ann↗

A Step Beyond Simple Keyword Searches: Services Enabled by a Full Content Digital Journal Archive

The problems of managing and searching large archives of scientific journal articles can potentially be addressed through data mining and statistical techniques matured primarily for quantitative scientific data analysis. A journal paper could be represented by a multivariate descriptor, e.g., the occurrence counts of a number key technical terms or phrases (keywords), perhaps derived from a controlled vocabulary ( e . g . , the American Meteorological Society's Glossary of Meteorology) or bootstrapped from the journal archive itself. With this technique, conventional statistical classification tools can be leveraged to address challenges faced by both scientists and professional societies in knowledge management. For example, cluster analyses can be used to find bundles of "most-related" papers, and address the issue of journal bifurcation (when is a new journal necessary, and what topics should it encompass). Similarly, neural networks can be trained to predict the optimal journal (within a society's collection) in which a newly submitted paper should be published. Comparable techniques could enable very powerful end-user tools for journal searches, all premised on the view of a paper as a data point in a multidimensional descriptor space, e.g.: "find papers most similar to the one I am reading", "build a personalized subscription service, based on the content of the papers I am interested in, rather than preselected keywords", "find suitable reviewers, based on the content of their own published works", etc. Such services may represent the next "quantum leap" beyond the rudimentary search interfaces currently provided to end-users, as well as a compelling value-added component needed to bridge the print-to-digital-medium gap, and help stabilize professional societies' revenue stream during the print-to-digital transition.

Boccippio, Dennis J.↗

Kepler Planet Detection Metrics: Statistical Bootstrap Test

This document describes the data produced by the Statistical Bootstrap Test over the final three Threshold Crossing Event (TCE) deliveries to NExScI: SOC 9.1 (Q1Q16)1 (Tenenbaum et al. 2014), SOC 9.2 (Q1Q17) aka DR242 (Seader et al. 2015), and SOC 9.3 (Q1Q17) aka DR253 (Twicken et al. 2016). The last few years have seen significant improvements in the SOC science data processing pipeline, leading to higher quality light curves and more sensitive transit searches. The statistical bootstrap analysis results presented here and the numerical results archived at NASAs Exoplanet Science Institute (NExScI) bear witness to these software improvements. This document attempts to introduce and describe the main features and differences between these three data sets as a consequence of the software changes.

Bootstrap↗

Research and operational efforts in support of Skylab experiment M093

The objectives of Skylab Experiment M093 were to measure electrocardiographic signals during spaceflight, to elucidate the electrophysiological basis for the changes observed, and to assess the effect of the change on the human cardiovascular system. Vectorcardiographic methods were used to quantitate changes, standardize data collection, and to facilitate reduction and statistical analysis of data. In this report the authors describe the M093 experiment design, the data transmission system, data reduction methods, and the analysis of data from the three Skylab missions. The report also includes clinical applications of the techniques developed for Skylab Experiment M093.

Smith, R. F.↗

Vectorcardiographic changes during extended space flight (M093): Observations at rest and during exercise

The objectives of Skylab Experiment M093 were to measure electrocardiographic signals during space flight, to elucidate the electrophysiological basis for the changes observed, and to assess the effect of the change on the human cardiovascular system. Vectorcardiographic methods were used to quantitate changes, standardize data collection, and to facilitate reduction and statistical analysis of data. Since the Skylab missions provided a unique opportunity to study the effects of prolonged weightlessness on human subjects, an effort was made to construct a data base that contained measurements taken with precision and in adequate number to enable conclusions to be made with a high degree of confidence. Standardized exercise loads were incorporated into the experiment protocol to increase the sensitivity of the electrocardiogram for effects of deconditioning and to detect susceptability for arrhythmias.

Smith, R. F.↗

Determining Monthly Mean Humidities From Satellite Data

Report describes statistical study to estimate monthly average humidity of marine surface layer of atmosphere from measurements by radiometers on satellites. Study part of continuing effort to determine flux density of latent heat due to evaporation at ocean surface. Such observations and measurements important because latent-heat flux affects weather and temperature and salinity of upper ocean layers.

Liu, W. Y. T.↗