Search NASASearch

SEARCH · Search NASA

Results for “Statistical techniques”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Misclassification in Workers’ Telecommuting Frequency Choices Using a Generalized Extreme Value Model

Telecommuting frequency is a response variable collected in travel surveys and is, therefore, prone to errors leading to mismeasurements or misclassification. Misclassification of explanatory variables is a common risk when using statistical modeling techniques. We define “misclassification” as a response reported or recorded in the wrong category; for example, a variable is recorded as a 1 when it should be 0. Here, in this context, this study aims to develop a statistical model to analyze telecommuting data which accounts for potential misclassification errors by building on existing literature in econometrics. The empirical analysis was undertaken using the 2017 National Household Travel Survey (NHTS) and the general extreme value (GEV) models available in the literature. Specifically, the frequency of telecommuting days was analyzed using the negative binomial (NB) model recast as the multinomial logit (MNL) model. By nature—and consistent with other studies—NHTS data are prone to errors that can be classified as intentional or unintentional misinformation provided by the person being interviewed. Ignoring these errors while modeling telecommuting frequencies using standard discrete count models can result in biased parameter estimates. The misclassification parameter was calculated for both over-reporting and under-reporting scenarios. The misclassification errors can be as high as 14% over-reported and 10% under-reported, particularly for the neighboring values. Statistical fit comparison between the models shows that models that ignore misclassification have worse data fit and biased parameter estimates with significant policy implications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Mars lander position estimation in the presence of ephemeris biases.

The process of estimating the location of a spacecraft landed on the surface of Mars is investigated through the application of statistical estimation techniques to earth-based radio tracking data. The spacecraft location and the tracking geometry and schedule are consistent with Viking-type mission constraints. With mission control requirements in mind, the investigation is restricted to analysis of a short data arc (approximately 3 days). Statistics of the spacecraft location are obtained through analysis of (direct-link) tracking data for the landed spacecraft and through simultaneous analysis of tracking data for both a landed and an orbiting spacecraft. These estimates include the effects of model uncertainties in the ephemeris of Mars, tracking station locations, the Mars rotational period, the Mars gravity field, and the orientation of Mars axis of rotation. The most significant of these effects is shown to be due to the Mars ephemeris uncertainty. A dual spacecraft tracking technique is presented for substantially reducing these ephemeris effects.

Blackshear, W. T.

Detection and Modeling of High-Dimensional Thresholds for Fault Detection and Diagnosis

Many Fault Detection and Diagnosis (FDD) systems use discrete models for detection and reasoning. To obtain categorical values like oil pressure too high, analog sensor values need to be discretized using a suitablethreshold. Time series of analog and discrete sensor readings are processed and discretized as they come in. This task isusually performed by the wrapper code'' of the FDD system, together with signal preprocessing and filtering. In practice,selecting the right threshold is very difficult, because it heavily influences the quality of diagnosis. If a threshold causesthe alarm trigger even in nominal situations, false alarms will be the consequence. On the other hand, if threshold settingdoes not trigger in case of an off-nominal condition, important alarms might be missed, potentially causing hazardoussituations. In this paper, we will in detail describe the underlying statistical modeling techniques and algorithm as well as the Bayesian method for selecting the most likely shape and its parameters. Our approach will be illustrated by several examples from the Aerospace domain.

Computer Systems

Operations for Learning with Graphical Models

This paper is a multidisciplinary review of empirical, statistical learning from a graphical model perspective. Well-known examples of graphical models include Bayesian net- works, directed graphs representing a Markov chain, and undirected networks representing a Markov field. These graphical models are extended to model data analysis and empirical learning using the notation of plates. Graphical operations for simplifying and manipulating a problem are provided including decomposition, differentiation, and the manipulation of probability models from the exponential family. These operations adapt existing techniques from statistics and automatic differentiation to graphs. Two standard algorithm schemes for learning are reviewed in a graphical framework: Gibbs sampling and the expectation maximization algorithm. Some algorithms are developed in this graphical framework including a generalized version of linear regression, techniques for feed-forward networks, and learning Gaussian and discrete Bayesian networks from data. The paper concludes by sketching some implications for data analysis and summarizing some popular algorithms that fall within the framework presented. The main original contributions here are the decomposition techniques and the demonstration that graphical models provide a framework for understanding and developing complex learning algorithms.

Buntine, Wray L.

A Framework for the Analysis of Deep Neural Networks in Autonomous Aerospace Applications using Bayesian Statistics

Deep Neural Networks (DNNs) are considered to be key components in many autonomous systems. Applications range from vision-based obstacle avoidance to intelligent/learning control and planning. Safety-critical applications as found in the aerospace domain require that the behavior of the DNN is validated and tested rigorously for safety of the autonomous system (AUS). In this paper, we present a framework to support testing of DNNs and the analysis of the network structure. Our framework employs techniques from statistical modeling and active learning to effectively generate test cases for DNN safety testing and performance analysis. We will present results of a case study on a physics-based Deep recurrent residual neural network (DR-RNN), which has been trained to emulate the aerodynamics behavior of a fixed-wing aircraft.

Deep Neural networks

Knowledge-based vision for space station object motion detection, recognition, and tracking

Computer vision, especially color image analysis and understanding, has much to offer in the area of the automation of Space Station tasks such as construction, satellite servicing, rendezvous and proximity operations, inspection, experiment monitoring, data management and training. Knowledge-based techniques improve the performance of vision algorithms for unstructured environments because of their ability to deal with imprecise a priori information or inaccurately estimated feature data and still produce useful results. Conventional techniques using statistical and purely model-based approaches lack flexibility in dealing with the variabilities anticipated in the unstructured viewing environment of space. Algorithms developed under NASA sponsorship for Space Station applications to demonstrate the value of a hypothesized architecture for a Video Image Processor (VIP) are presented. Approaches to the enhancement of the performance of these algorithms with knowledge-based techniques and the potential for deployment of highly-parallel multi-processor systems for these algorithms are discussed.

Symosek, P.

Application of Land Surface Data Assimilation to Simulations of Sea Breeze Circulations

A technique has been developed for assimilating GOES-derived skin temperature tendencies and insolation into the surface energy budget equation of a mesoscale model so that the simulated rate of temperature change closely agrees with the satellite observations. A critical assumption of the technique is that the availability of moisture (either from the soil or vegetation) is the least known term in the model's surface energy budget. Therefore, the simulated latent heat flux, which is a function of surface moisture availability, is adjusted based upon differences between the modeled and satellite- observed skin temperature tendencies. An advantage of this technique is that satellite temperature tendencies are assimilated in an energetically consistent manner that avoids energy imbalances and surface stability problems that arise from direct assimilation of surface shelter temperatures. The fact that the rate of change of the satellite skin temperature is used rather than the absolute temperature means that sensor calibration is not as critical. The sea/land breeze is a well-documented mesoscale circulation that affects many coastal areas of the world including the northern Gulf Coast of the United States. The focus of this paper is to examine how the satellite assimilation technique impacts the simulation of a sea breeze circulation observed along the Mississippi/Alabama coast in the spring of 2001. The technique is implemented within the PSUNCAR MM5 V3-5 and applied at spatial resolutions of 12- and 4-km. It is recognized that even 4-km grid spacing is too coarse to explicitly resolve the detailed, mesoscale structure of sea breezes. Nevertheless, the model can forecast certain characteristics of the observed sea breeze including a thermally direct circulation that results from differential low-level heating across the land-sea interface. Our intent is to determine the sensitivity of the circulation to the differential land surface forcing produced via the assimilation of GOES skin temperature tendencies. Results will be quantified through statistical verification techniques.

Mackaro, Scott

Land Surface Data Assimilation and the Northern Gulf Coast Land/Sea Breeze

A technique has been developed for assimilating GOES-derived skin temperature tendencies and insolation into the surface energy budget equation of a mesoscale model so that the simulated rate of temperature change closely agrees with the satellite observations. A critical assumption of the technique is that the availability of moisture (either from the soil or vegetation) is the least known term in the model's surface energy budget. Therefore, the simulated latent heat flux, which is a function of surface moisture availability, is adjusted based upon differences between the modeled and satellite observed skin temperature tendencies. An advantage of this technique is that satellite temperature tendencies are assimilated in an energetically consistent manner that avoids energy imbalances and surface stability problems that arise from direct assimilation of surface shelter temperatures. The fact that the rate of change of the satellite skin temperature is used rather than the absolute temperature means that sensor calibration is not as critical. The sea/land breeze is a well-documented mesoscale circulation that affects many coastal areas of the world including the northern Gulf Coast of the United States. The focus of this paper is to examine how the satellite assimilation technique impacts the simulation of a sea breeze circulation observed along the Mississippi/Alabama coast in the spring of 2001. The technique is implemented within the PSU/NCAR MM5 V3-4 and applied on a 4-km domain for this particular application. It is recognized that a 4-km grid spacing is too coarse to explicitly resolve the detailed, mesoscale structure of sea breezes. Nevertheless, the model can forecast certain characteristics of the observed sea breeze including a thermally direct circulation that results from differential low-level heating across the land-sea interface. Our intent is to determine the sensitivity of the circulation to the differential land surface forcing produced via the assimilation of GOES skin temperature tendencies. Results will be quantified through statistical verification techniques.

Lapenta, William M.

Uncertainty Assessment of the NASA Earth Exchange Global Daily Downscaled Climate Projections (NEX-GDDP) Dataset

The NASA Earth Exchange Global Daily Downscaled Projections (NEX-GDDP) dataset is comprised of downscaled climate projections that are derived from 21 General Circulation Model (GCM) runs conducted under the Coupled Model Intercomparison Project Phase 5 (CMIP5) and across two of the four greenhouse gas emissions scenarios (RCP4.5 and RCP8.5). Each of the climate projections includes daily maximum temperature, minimum temperature, and precipitation for the periods from 1950 through 2100 and the spatial resolution is 0.25 degrees (approximately 25 km x 25 km). The GDDP dataset has received warm welcome from the science community in conducting studies of climate change impacts at local to regional scales, but a comprehensive evaluation of its uncertainties is still missing. In this study, we apply the Perfect Model Experiment framework (Dixon et al. 2016) to quantify the key sources of uncertainties from the observational baseline dataset, the downscaling algorithm, and some intrinsic assumptions (e.g., the stationary assumption) inherent to the statistical downscaling techniques. We developed a set of metrics to evaluate downscaling errors resulted from bias-correction ("quantile-mapping"), spatial disaggregation, as well as the temporal-spatial non-stationarity of climate variability. Our results highlight the spatial disaggregation (or interpolation) errors, which dominate the overall uncertainties of the GDDP dataset, especially over heterogeneous and complex terrains (e.g., mountains and coastal area). In comparison, the temporal errors in the GDDP dataset tend to be more constrained. Our results also indicate that the downscaled daily precipitation also has relatively larger uncertainties than the temperature fields, reflecting the rather stochastic nature of precipitation in space. Therefore, our results provide insights in improving statistical downscaling algorithms and products in the future.

climate projection

Uncertainty Assessment of the NASA Earth Exchange Global Daily Downscaled Climate Projections (NEX-GDDP) Dataset

The NASA Earth Exchange Global Daily Downscaled Projections (NEX-GDDP) dataset is comprised of downscaled climate projections that are derived from 21 General Circulation Model (GCM) runs conducted under the Coupled Model Intercomparison Project Phase 5 (CMIP5) and across two of the four greenhouse gas emissions scenarios (RCP4.5 and RCP8.5). Each of the climate projections includes daily maximum temperature, minimum temperature, and precipitation for the periods from 1950 through 2100 and the spatial resolution is 0.25 degrees (approximately 25 km by 25 km). The GDDP dataset has received warm welcome from the science community in conducting studies of climate change impacts at local to regional scales, but a comprehensive evaluation of its uncertainties is still missing. In this study, we apply the Perfect Model Experiment framework (Dixon et al. 2016) to quantify the key sources of uncertainties from the observational baseline dataset, the downscaling algorithm, and some intrinsic assumptions (e.g., the stationary assumption) inherent to the statistical downscaling techniques. We developed a set of metrics to evaluate downscaling errors resulted from bias-correction ("quantile-mapping"), spatial disaggregation, as well as the temporal-spatial non-stationarity of climate variability. Our results highlight the spatial disaggregation (or interpolation) errors, which dominate the overall uncertainties of the GDDP dataset, especially over heterogeneous and complex terrains (e.g., mountains and coastal area). In comparison, the temporal errors in the GDDP dataset tend to be more constrained. Our results also indicate that the downscaled daily precipitation also has relatively larger uncertainties than the temperature fields, reflecting the rather stochastic nature of precipitation in space. Therefore, our results provide insights in improving statistical downscaling algorithms and products in the future.

general circulation model (GCM)

A statistical correlation method for the retrieval of atmospheric moisture profiles by microwave radiometry

A statistical correlation technique is applied to the retrieval of vertical moisture profiles under clear-sky conditions from down-looking radiometric measurements of atmospheric radiation at microwave wavelengths. For a given set of channels, the method selects the optimum radiometric channels for estimating water vapor at specific pressure levels between the surface and 300 mb. The water vapor mixing ratio at these pressure levels is then calculated from a linear combination of the selected channel brightness temperatures. To test its validity the algorithm was applied, in a numerical experiment, to fifty independent tropical radiosondes. The rms absolute deviation of the estimated moisture profiles from the actual profiles was comparable to that obtained using an iterative retrieval method reported earlier. The statistical method, however, requires several orders of magnitude less computer time than the iterative method; it is suitable for high speed processing of large amounts of data.

Kakar, R. K.

Non-linear Least Square Fitting Technique for the Determination of Field Line Resonance Frequency in Ground Magnetometer Data: Application to Remote Sensing of Plasmaspheric Mass Density

The accurate determination of the Field Line Resonance (FLR) frequency of a resonating geomagnetic field line is necessary to remotely monitor the plasmaspheric mass density during geomagnetic storms and quiet times alike. Under certain assumptions the plasmaspheric mass density at the equator is inversely proportional to the square of the FLR frequency. The most common techniques to determine the FLR frequency from ground magnetometer measurements are the amplitude ratio and phase difference techniques, both based on geomagnetic field observations at two latitudinally separated ground stations along the same magnetic meridian. Previously developed automated techniques have used statistical methods to pinpoint the FLR frequency using the amplitude ratio and phase difference calculations. We now introduce a physics-based automated technique, using non-linear least square fitting of the ground magnetometer data to the analytical resonant wave equations, that reproduces the wave characteristics on the ground, and from those determine the FLR frequency. One of the advantages of the new technique is the estimation of physics-based errors of the FLR frequency, and as a result of the equatorial plasmaspheric mass density. We present analytical results of the new technique, and test it using data from the Inner-Magnetospheric Array for Geospace Science (iMAGS) ground magnetometer chain along the coast of Chile and the east coast of the United States. We compare the results with the results of previously published statistical automated techniques.

A. Boudouridis

Dealing with Ion LET Uncertainties: An Application of Generalized Linear Models

Although most SEE rate estimation methods presume a fit to SEE cross section vs. LET, fitting SEE data is challenging because the data are not compatible with the assumptions of many common fitting techniques (e.g. linear regression. The difficulty of fitting such data is compounded when the LET of the ion responsible for an SEE is uncertain. We modify a Generalized Linear Model SEE data fitting methodology to accommodate uncertain LET and apply the method to the problem of backside heavy-ion SEE testing to demonstrate the utility of the method, explore the dependence of systematic errors that arise from improper treatment of LET uncertainty and develop guidelines for minimizing such systematic errors when proper treatment is not possible. Additional applications are suggested and assessed for suitability of treatment by the model.

Single-event effects

Quantifying the Importance of Selected Drought Indicators for the United States Drought Monitor

Using information theory, our study quantifies the importance of selected indicators for the U.S. Drought Monitor (USDM) maps. We use the technique of mutual information (MI) to measure the importance of any indicator to the USDM, and because MI is derived solely from the data, our findings are independent of any model structure (conceptual, physically-based, or empirical). We also compare these MIs against the drought representation effectiveness ratings in the North America Drought Indices and Indicators Assessment (NADIIA) survey for Koeppen climate zones. This reveals: [1] agreement between some ratings and our MI values (high for example indicators like Standardized Precipitation-Evapotranspiration Index or SPEI); [2] some divergences (for example, soil moisture has high ratings but near-zero MIs for ESA-CCI soil moisture in the Western U.S., indicating the need of another remotely sensed soil moisture source); and [3] new insights into the importance of variables such as Snow Water Equivalent (SWE) that are not included in sources like NADIIA. Further analysis of the MI results yields findings related to: [1] hydrological mechanisms (summertime SWE domination during individual drought events through snowmelt into the water-scarce soil); [2] hydroclimatic types (the top pair of inputs in the Western and non-Western regions are SPEIs and soil moistures respectively); and [3] predictability (high for the California 2012-2017 event, with longer-timescale indicators dominating). Finally, the high MIs between multiple indicators jointly and the USDM indicate potentially high drought forecasting accuracies achievable using only model-based inputs, and the potential for global drought monitoring using only remotely sensed inputs, especially for locations having insufficient in situ observations.

Drought

Multi-objective Bayesian active learning for MeV-ultrafast electron diffraction

Ultrafast electron diffraction using MeV energy beams(MeV-UED) has enabled unprecedented scientific opportunities in the study of ultrafast structural dynamics in a variety of gas, liquid and solid state systems. Broad scientific applications usually pose different requirements for electron probe properties. Due to the complex, nonlinear and correlated nature of accelerator systems, electron beam property optimization is a time-taking process and often relies on extensive hand-tuning by experienced human operators. Algorithm based efficient online tuning strategies are highly desired. Here, we demonstrate multi-objective Bayesian active learning for speeding up online beam tuning at the SLAC MeV-UED facility. The multi-objective Bayesian optimization algorithm was used for efficiently searching the parameter space and mapping out the Pareto Fronts which give the trade-offs between key beam properties. Such scheme enables an unprecedented overview of the global behavior of the experimental system and takes a significantly smaller number of measurements compared with traditional methods such as a grid scan. This methodology can be applied in other experimental scenarios that require simultaneously optimizing multiple objectives by explorations in high dimensional, nonlinear and correlated systems.

43 PARTICLE ACCELERATORS

White paper on light sterile neutrino searches and related phenomenology

This white paper provides a comprehensive review of our present understanding of experimental neutrino anomalies that remain unresolved, charting the progress achieved over the last decade at the experimental and phenomenological level, and sets the stage for future programmatic prospects in addressing those anomalies. It is purposed to serve as a guiding and motivational "encyclopedic" reference, with emphasis on needs and options for future exploration that may lead to the ultimate resolution of the anomalies. We see the main experimental, analysis, and theory-driven thrusts that will be essential to achieving this goal being: 1) Cover all anomaly sectors -- given the unresolved nature of all four canonical anomalies, it is imperative to support all pillars of a diverse experimental portfolio, source, reactor, decay-at-rest, decay-in-flight, and other methods/sources, to provide complementary probes of and increased precision for new physics explanations; 2) Pursue diverse signatures -- it is imperative that experiments make design and analysis choices that maximize sensitivity to as broad an array of these potential new physics signatures as possible; 3) Deepen theoretical engagement -- priority in the theory community should be placed on development of standard and beyond standard models relevant to all four short-baseline anomalies and the development of tools for efficient tests of these models with existing and future experimental datasets; 4) Openly share data -- Fluid communication between the experimental and theory communities will be required, which implies that both experimental data releases and theoretical calculations should be publicly available; and 5) Apply robust analysis techniques -- Appropriate statistical treatment is crucial to assess the compatibility of data sets within the context of any given model.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Forced Component Estimation Statistical Method Intercomparison Project (ForceSMIP)

Anthropogenic climate change is unfolding rapidly, yet its regional manifestation can be obscured by internal variability. A primary goal of climate science is to identify the externally forced climate response from among the noise of internal variability. Separating the forced response from internal variability can be addressed in climate models by using a large ensemble to average over different possible realizations of internal variability. However, with only one realization of the real world, it is a major challenge to isolate the forced response directly in observations. In the Forced Component Estimation Statistical Method Intercomparison Project (ForceSMIP), contributors used existing and newly developed statistical and machine learning methods to estimate the forced response over 1950–2022 within individual realizations of the climate system. Participants used neural networks, linear inverse models, fingerprinting methods, and low-frequency component analysis, among other approaches. These methods were trained using large ensembles from multiple climate models and then applied to observations. Here, we evaluate method performance within large ensembles and investigate the estimates of the forced response in observations. Our results show that many different types of methods are skillful for estimating the forced response in climate models, though the relative skill of individual methods varies depending on the variable and evaluation metric. Methods with comparable skill in models can give a wide range of estimates of the forced response pattern in observations, illustrating the epistemic uncertainty in forced response estimates. ForceSMIP gives new insights into the forced response in observations, its uncertainty, and methods for its estimation.

Climate attribution