Estimate accuracy and project selection models in industrial research.
Estimate accuracy and project selection models in industrial research, examining company data for miscellaneous, technical and commercial failure
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Estimate accuracy and project selection models in industrial research, examining company data for miscellaneous, technical and commercial failure
A statistical decision procedure called chain pooling had been developed for model selection in fitting the results of a two-level fixed-effects full or fractional factorial experiment not having replication. The basic strategy included the use of one nominal level of significance for a preliminary test and a second nominal level of significance for the final test. The subject has been reexamined from the point of view of using as many as three successive statistical model deletion procedures in fitting the results of a single experiment. The investigation consisted of random number studies intended to simulate the results of a proposed aircraft turbine-engine rotor-burst-protection experiment. As a conservative approach, population model coefficients were chosen to represent a saturated 2 to the 4th power experiment with a distribution of parameter values unfavorable to the decision procedures. Three model selection strategies were developed.
As many as three iterated statistical model deletion procedures are considered for an experiment. Population model coefficients were chosen to simulate a saturated factorial experiment having an unfavorable distribution of parameter values. Using random number studies, three model selection strategies were developed, namely, (1) a strategy to be used in anticipation of large coefficients of variation (neighborhood of 65 percent), (2) strategy to be used in anticipation of small coefficients of variation (4 percent or less), and (3) a security regret strategy to be used in the absence of such prior knowledge.
Multi-model Monte Carlo methods have been demonstrated to be an efficient and accurate alternative to standard Monte Carlo (MC) in the model-based propagation of uncertainty in entry, descent, and landing (EDL) applications. These multi-model MC methods fuse predictions from low-fidelity models with the high-fidelity EDL model of interest to produce unbiased statistics with a fraction of the computational cost. The accuracy and efficiency of the multi-model MC methods are dependent upon the magnitude of correlations of the low-fidelity models with the high-fidelity model, but also upon the correlation amongst the low-fidelity models, and their relative computational cost. Because of this layer of complexity, the question of how to optimally select the set of low-fidelity models has remained open. In this work, methods for optimal model construction and tuning are investigated as a means to increase the speed and precision of trajectory simulation for EDL. Specifically, the focus is on the inclusion of low-fidelity model tuning within the sample allocation optimization that accompanies multi-model MC methods. Preliminary results indicate that low-fidelity model tuning can significantly improve efficiency and precision of trajectory simulations and provide an increased edge to multi-model MC methods when compared to standard MC. The challenges and potential benefits to exploring a fully iterative and comprehensive optimization strategy in future work are highlighted.
Multi-model Monte Carlo methods have been demonstrated to be an efficient and accurate alternative to standard Monte Carlo (MC) in the model-based propagation of uncertainty in entry, descent, and landing (EDL) applications. These multi-model MC methods fuse predictions from low-fidelity models with the high-fidelity EDL model of interest to produce unbiased statistics with a fraction of the computational cost. The accuracy and efficiency of the multi-model MC methods are dependent upon the magnitude of correlations of the low-fidelity models with the high-fidelity model, but also upon the correlation among the low-fidelity models, and their relative computational cost. Because of this layer of complexity, the question of how to optimally select the set of low-fidelity models has remained open. In this work, methods for optimal model construction and tuning are investigated as a means to increase the speed and precision of trajectory simulation for EDL. Specifically, the focus is on the inclusion of low-fidelity model tuning within the sample allocation optimization that accompanies multi-model MC methods. Preliminary results indicate that low-fidelity model tuning can significantly improve efficiency and precision of trajectory simulations and provide an increased edge to multi-model MC methods when compared to standard MC. The challenges and potential benefits to exploring a fully iterative and comprehensive optimization strategy in future work are highlighted.
As many as three iterated statistical model deletion procedures were considered for an experiment. Population model coefficients were chosen to simulate a saturated 2 to the 4th power experiment having an unfavorable distribution of parameter values. Using random number studies, three model selection strategies were developed, namely, (1) a strategy to be used in anticipation of large coefficients of variation, approximately 65 percent, (2) a strategy to be sued in anticipation of small coefficients of variation, 4 percent or less, and (3) a security regret strategy to be used in the absence of such prior knowledge.
When performing symbolic regression using genetic programming, overfitting and bloat can negatively impact generalizability and interpretability of the resulting equations as well as increase computation times. A Bayesian fitness metric is introduced and its impact on bloat and overfitting during population evolution is studied and compared to common alternatives in the literature. The proposed approach was found to be more robust to noise and data sparsity in numerical experiments, guiding evolution to a level of complexity appropriate to the dataset. Further evolution of the population resulted not in overfitting or bloat, but rather in slight simplifications in model form. The ability to identify an equation of complexity appropriate to the scale of noise in the training data was also demonstrated. In general, the Bayesian model selection algorithm was shown to be an effective means of regularization which resulted in less bloat and overfitting when any amount of noise was present in the training data.
A computer program is described that performs a statistical multiple-decision procedure called chain pooling. It uses a number of mean squares assigned to error variance that is conditioned on the relative magnitudes of the mean squares. The model selection is done according to user-specified levels of type 1 or type 2 error probabilities.
When performing symbolic regression using genetic programming, overfitting and bloat can negatively impact generalizability and interpretability of the resulting equations as well as increase computation times. A Bayesian fitness metric is introduced and its impact on bloat and overfitting during population evolution is studied and compared to common alternatives in the literature. The proposed approach was found to be more robust to noise and data sparsity in numerical experiments, guiding evolution to a level of complexity appropriate to the dataset. Further evolution of the population resulted not in overfitting or bloat, but rather in slight simplifications in model form. The ability to identify an equation of complexity appropriate to the scale of noise in the training data was also demonstrated. In general, the Bayesian model selection algorithm was shown to be an effective means of regularization which resulted in less bloat and overfitting when any amount of noise was present in the training data.
Interpretable Machine Learning (IML) has performed well when tasked with deriving constitutive material models. However, IML has been shown to prefer models that overfit noise in data, which tends to lead to bloat and a decrease in interpretability. Due to these issues, the ability of IML to reliably derive models that fit the data and are both interpretable and generalizable is limited. A method developed recently has shown promise to improve upon traditional IML by using a Bayesian fitness definition for the evolution of free-form models with non-deterministic parameters. This framework was developed for genetic-programming-based symbolic regression(GPSR) and involves model parameter estimation using Sequential Monte Carlo sampling (SMC).The method has demonstrated a reduction in bloat when dealing with noisy data in comparison to conventional GPSR. The results of this framework applied to stress-strain data for copper show models that more effectively predict the experimental data better than was previously shown with GPSR.
As the development of the NAVSTAR Global Positioning System (GPS) continues, there will increasingly be the need for a software centered signal model. This model must accurately generate the observed pseudorange which would typically be encountered. The observed pseudorange varies from the true geometric (slant) range due to range measurement errors. Errors in range measurement stem from a variety of hardware and environment factors. These errors are classified as either deterministic or random and, where appropriate, their models are summarized. Of particular interest is the model for Selective Availability which is derived from actual GPS data. The procedure for the determination of this model, known as the System Identification Theory, is briefly outlined. The synthesis of these error sources into the final signal model is given along with simulation results.
A model is described for planning and decision making in research project selection. Evaluations of each project's direct and indirect benefits, uncertainty in achieving these benefits, and schedule priority with resource budget and program balance constraints are considered. The combination of the interactive effect of project selection, resource allocation and scheduling considerations into one model permits tradeoff alternatives to be studied. Clients' value judgments are used in evaluating the benefits from each proposed project. The model is applied to the NASA Civil Helicopter Technology Program. Research project priorities for this program are established, strengths and weaknesses of the model are discussed, and areas of future development are recommended.
Effect of gravitational model variations on accuracy of lunar orbit determination from data arcs
The analysis of current and future cosmological surveys of Type Ia supernovae (SNe Ia) at high redshift depends on the accuratephotometric classification of the SN events detected. Generating realistic simulations of photometric SN surveys constitutes anessential step for training and testing photometric classification algorithms, and for correcting biases introduced by selectioneffects and contamination arising from core-collapse SNe in the photometric SN Ia samples. We use published SN time-seriesspectrophotometric templates, rates, luminosity functions, and empirical relationships between SNe and their host galaxies toconstruct a framework for simulating photometric SN surveys. We present this framework in the context of the Dark EnergySurvey (DES) 5-yr photometric SN sample, comparing our simulations of DES with the observed DES transient populations.We demonstrate excellent agreement in many distributions, including Hubble residuals, between our simulations and data.We estimate the core collapse fraction expected in the DES SN sample after selection requirements are applied and beforephotometric classification. After testing different modelling choices and astrophysical assumptions underlying our simulation,we find that the predicted contamination varies from 7.2 to 11.7 per cent, with an average of 8.8 per cent and an r.m.s. of 1.1 percent. Our simulations are the first to reproduce the observed photometric SN and host galaxy properties in high-redshift surveyswithout fine-tuning the input parameters. The simulation methods presented here will be a critical component of the cosmologyanalysis of the DES photometric SN Ia sample: correcting for biases arising from contamination, and evaluating the associatedsystematic uncertainty.
Currently, there are several models in the literature, such as kinetic models, microstructural models, and mass transport models that describe a Li-air battery's discharge behavior. Many of these models are calibrated and tested at low current densities and cannot be easily transferred to high current densities. Even at low current densities, there is no quantitative method for a researcher to choose a reaction kinetic model such as classical Butler-Volmer and its derivatives, and modified Marcus-Hush-Chidsey, a resistance model for lithium peroxide such as electron transport via tunneling or linear resistivity, a surface coverage model (lithium peroxide growth) such as partial coverage or full coverage, and mass transport model (discussed in Ref. [1]). Also, it is time-consuming to test different models at high current density (1C) due to a lack of well-tested models and well-calibrated model parameters. For this presentation, we will develop an analytical model, which acts as a surrogate model for a sophisticated finite element model to predict discharge time and discharge voltage. Next, we use an uncertainty quantifying technique called reduced-order stochastic optimization [2, 3] to determine the uncertainty in model parameters for rate kinetics, lithium peroxide resistivity, and parasitic resistance. Finally, a finite element simulation is performed to determine the error introduced by the surrogate model and its influence on the uncertainty in the model parameters.
In this paper we examine the problem of monitoring and diagnosing noisy complex dynamical systems that are modeled as hybrid systems-models of continuous behavior, interleaved by discrete transitions. In particular, we examine continuous systems with embedded supervisory controllers that experience abrupt, partial or full failure of component devices. Building on our previous work in this area (MBCG99;MBCG00), our specific focus in this paper ins on the mathematical formulation of the hybrid monitoring and diagnosis task as a Bayesian model tracking algorithm. The nonlinear dynamics of many hybrid systems present challenges to probabilistic tracking. Further, probabilistic tracking of a system for the purposes of diagnosis is problematic because the models of the system corresponding to failure modes are numerous and generally very unlikely. To focus tracking on these unlikely models and to reduce the number of potential models under consideration, we exploit logic-based techniques for qualitative model-based diagnosis to conjecture a limited initial set of consistent candidate models. In this paper we discuss alternative tracking techniques that are relevant to different classes of hybrid systems, focusing specifically on a method for tracking multiple models of nonlinear behavior simultaneously using factored sampling and conditional density propagation. To illustrate and motivate the approach described in this paper we examine the problem of monitoring and diganosing NASA's Sprint AERCam, a small spherical robotic camera unit with 12 thrusters that enable both linear and rotational motion.
The ROCket Combustor Interactive Design (ROCCID) methodology is an interactive computer program that combines previously developed combustion analysis models to calculate the combustion performance and stability of liquid rocket engines. Test data from a 213 kN (48,000 lbf) Liquid Oxygen (LOX)/RP-1 combustor with a O-F-O (oxidizer-fuel-oxidizer) triplet injector were used to characterize the predictive capabilities of the ROCCID analysis models for this injector/propellant configuration. Thirteen combustion performance and stability models have been incorporated into ROCCID, and ten of them, which have options for triplet injectors, were examined in this study. Calculations using different combinations of analysis models, with little or no anchoring, were carried out on a test matrix of operating conditions matching those of the test program. Results of the computer analyses were compared to test data, and the ability of the model combinations to correctly predict combustion stability or instability was determined. For the best model combination(s), sensitivity of the calculations to fuel drop size and mixing efficiency was examined. Error in the stability calculations due to uncertainty in the pressure interaction index (N) was examined. The recommended model combinations for this O-F-O triplet LOX/RP-1 configuration are proposed.
The ROCket Combustor Interactive Design (ROCCID) methodology is an interactive computer program that combines previously developed combustion analysis models to calculate the combustion performance and stability of liquid rocket engines. Test data from 213 kN (48,000 lbf) Liquid Oxygen (LOX)/RP-1 combustor with an O-F-O (oxidizer-fuel-oxidizer) triplet injector were used to characterize the predictive capabilities of the ROCCID analysis models for this injector/propellant configuration. Thirteen combustion performance and stability models were incorporated into ROCCID, and ten of them, which have options for triplet injectors, were examined. Calculations using different combinations of analysis models, with little or no anchoring, were carried out on a test matrix of operating combinations matching those of the test program. Results of the computer analyses were compared to test data, and the ability of the model combinations to correctly predict combustion stability or instability was determined. For the best model combination(s), sensitivity of the calculations to fuel drop size and mixing efficiency was examined. Error in the stability calculations due to uncertainty in the pressure interaction index (N) was examined. The recommended model combinations for this O-F-O triplet LOX/RP-1 configuration are proposed.