Search NASASearch

Engineering topics

DiNicola, Michael

Publications and source records attributed to DiNicola, Michael.

Sampling Size Optimization for Bioburden Density Estimation in Planetary Protection

Planetary protection (PP) is a discipline that focuses on minimizing the biological contamination of spacecraft to ensure compliance with international policy. Precise estimation of bioburden - the total number of microbes in or on spacecraft hardware – and the bioburden density are of utmost importance for PP. Such estimation is the way concordance with requirements is demonstrated, and it is critical for quantifying the potential risk of inadvertently contaminating other planetary bodies. Although a suite of molecular techniques have been used to thoroughly characterize and profile the microbiome of various cleanroom environments and spacecraft, the gold standard remains the physical enumeration of microbes via culturing of samples directly taken from spacecraft and associated surfaces. However, due to technical, budgetary, and programmatic constraints, only a manageable portion (around 10%) of the entire spacecraft surface is directly sampled with cotton swabs or wipes. To generate the bioburden current best estimate (CBE) for components not directly verifiable, the accepted approach is to apply a NASA-defined bioburden estimate based on the components’ manufacturing or assembly environment. This approach utilizes a prespecified bioburden density estimation that applies a maximum value across the total surface area of the specified component. For hardware components that underwent similar assembly processes, an implied bioburden is adopted for all components, based on a direct verification of a representative component within the same lot. Once all components have a CBE, the bioburden estimates are generated. In previous publication [ 1], we have shown that statistical risks quantifying the accuracy of the estimates for sampled, prespecified, and implied components can be derived and ranked. For mean squared error (MSE) function, the risks are available analytically and hence a cost function can be obtained to optimize the risks with respect to the sampling area and sampling cost. Since the sampling area and sampling cost are two complimentary variables, their sum will have a well-defined minimum. This paper presents the multivariate optimization of the integrated risk of an empirical Bayes estimator to determine the optimal sampling schedule for a given number of components. It is assumed that given a number of components, N, the bioburden density for each component can either be sampled, implied, or prespecified. The multivariate optimization searches through different options to sample, imply or prespecify the bioburden density for a component, and account for the component’s surface area and cost of sampling. The idea of the optimization is based on the observation that the statistical risk of using an estimator is a monotonically decreasing function of the sampled area. The larger the sampled area, the lower the risk of using the estimator as the estimator becomes more and more accurate as the sampling area increases. On the other hand, the cost of sampling is monotonically increasing as the sampled surface grows. This makes the risk and total cost of sampling complimentary variables which can be counterbalanced to achieve an optimal overall value with respect to the sampled surface. In this paper, the integrated risk has been used to quantify the accuracy of the estimator. This risk has been selected because it depends on neither the true value of the parameter nor on the collected data. The cost of each sample was also available to obtain the total cost of sampling of N components. The paper will present the results based on computer-simulated data as well as the data collected during the InSight mission. The computer-simulated data have N components with randomly generated total areas and each component assigned to one of the three categories according to the method of estimating of bioburden density: sampled, implied, or prespecified. The cost of sampling is also available. The cost of sampling is estimated based on a cost model provided by the planetary protection group at JPL. For this paper, the overall cost was assumed to be a linear function of exposure. The optimization process finds the allocation of the components to the three categories that minimizes the tradeoff between integrated risk and total cost. For the InSight data, a set of components is selected representing all three categories, and optimization is performed to determine if the performed allocation was optimal or if a better allocation could have been obtained. To the best of our knowledge, this work is the first attempt not only perform an accurate estimation of bioburden density but also do it in an optimal way.

97 - MATHEMATICS AND COMPUTING

Regularized Differentiation for Bioburden Density Estimation in Planetary Protection

In this paper, we propose and investigate the performance of two novel shrinkage estimators for bioburden density estimation in planetary protection. The estimators are based on the regularized differentiation of a cumulative count of colony forming units collected throughout the data collecting session or the life cycle of the entire mission. The regularized differentiation recasts the problem of bioburden density estimation as a linear least squares problem. The least squares problem is then solved through regularization techniques, such as truncated singular value decomposition and penalized least squares. The regularization is necessary to avoid noise amplification during the differentiation of noisy data. The two regularization estimators are compared with four other commonly used estimators to simultaneously evaluate the means of multivariable independent Poisson distributions: the maximum likelihood, noninformative Bayes estimator with Jeffreys prior, Empirical Bayes using conjugate gamma-Poisson model with gamma parameters selected by method of moments, and the Clevenson-Zidek estimator. It is shown through computer-simulated data that the regularized differentiation based on ridge regression has the smallest mean-squared error among all estimators. The analysis of shrinkage mechanism implemented by regularized differentiation is performed, and it is shown that the regularized differentiation amounts to performing a weighted averaging of all the samples. The weights are determined by the regularization parameter automatically selected by the L-curve technique. Since the method of least squares makes no distributional assumptions about the data, it presents an attractive technique for bioburden density estimation when there are concerns about the misspecification of the distributional model. The paper concludes with the analysis of the bioburden data collected during InSight mission and directions for future work.

97 - MATHEMATICS AND COMPUTING

Salvaging Data Records with Missing Data: Data Imputation using the Multivariate t Distribution

When doing multivariate data analysis, one commonobstacle is the presence of incomplete observations, i.e., observationsfor which one or more key fields are blank. Missing datais often countered by deleting entire observations that containmissing data. The negative effects of deleting entire observationsare multiple: deleting observations reduces sample size andcan also result in biased inferences even if data is missing atrandom. In addition, knowledge contained within incompleteobservations is knowledge lost when they are deleted– and theeffort spent collecting that knowledge is effort wasted. Data imputationmethods, or methods of statistically “filling-in” missingdata, can help combat small sample sizes by using the existinginformation in partially complete observations with the end goalof producing less biased and higher confidence inferences. Whena sample from a multivariate normal population is only partiallycomplete, and the missing data meets appropriate assumptions(missing at random), robust data imputation of the missing datacan be implemented with monotone data augmentation (MDA)using the multivariate t distribution.Missing data imputation is applied to data from the NASA InstrumentCost Model (NICM) using the MDA algorithm underthe assumption of having a multivariate t distribution with fixeddegrees of freedom. A sensitivity analysis to the degrees offreedom parameter is presented to demonstrate robustness ofthe multivariate t distribution when dealing with small samplesas compared to the multivariate normal distribution.

DiNicola, Michael

Modeling Spacecraft Safe Mode Events

Spacecraft enter a ‘safe mode’ to protect the vehicle when a potentially harmful anomaly occurs. This minimally functioning state isolates faults, establishes contact with Earth, and orients the vehicle into a power positive attitude until operators intervene. Though ‘safings’ are inherently unpredictable, mission teams build in time margin during operations to determine root causes and restore functionality. Planning and managing this margin is both critical and enabling on mission architectures dependent on near-continuous operability – such as a low-thrust electric propulsion mission. To better quantify the occurrences and severity of safe mode anomalies, the Jet Propulsion Laboratory (JPL) has assembled a database of safings from past and active missions. Currently nearly 240 records are captured from 21 beyond-Earth missions, stemming from a collaboration between teams at JPL, Ames Research Center, Goddard Space Flight Center, and the Johns Hopkins University Applied Physics Laboratory. This paper discusses the event database, explores a statistical approach in modeling the occurrences and severity of safing events, presents a simulation technique, and details recommendations and future work to benefit future concepts.

Nicholas, Austin

Nasa Space Flight Instruments: Cost Time Trends

Are NASA’s space flight instruments becoming cheaper or more expensive as time marches forward? After analyzing the costs of hundreds of instruments launched over the last 30 years, the short answer to this question is no… and yes. This paper gives a visual analysis of the cost time trends for various NASA space flight instrument types, such as optical, particles detectors, fields detectors and microwave instruments. In addition to the statistical approaches utilized, such as significance tests, cluster analysis and principle components analysis (PCA), we will also discuss the intangibles which are likely at play, including technological progress, NASA policy and the luck of the draw associated with mission manifests. This analysis was performed as the main driver for the NASA Instrument Cost Model (NICM) recent cost estimating model redesign. Started in 2004, the first version of NICM was based off of instruments launched from 1985-2005, or 20 years’ worth of data. As NICM hit its 10-year anniversary, we wanted to know: should NICM continue to only use the most recent 20 years’ worth of data (1995-2015)? Are instruments becoming cheaper or more expensive as time marches forward? There is evidence in favor of a drop in the median dollar-per-kg value across some instrument types, but little in others. Whereas further research is needed to substantiate, Particles and Optical-Planetary instrument types show moderate to strong evidence of a downward trend in dollar-per-kg. Further research is required to study the nature of this trend (shift, taper, cyclic, etc.). Little evidence for a similar downward trend was detected for Fields or Microwave instruments, or Optical instruments on Earth Orbiting spacecraft. We presented evidence in favor of a drop in the median dollar-per-kg value for Particles and Optical-Planetary instrument types. While similar evidence was weak at best for Fields and Microwave instruments. We can speculate as to the causes for this effect, but we are also equipped to begin to rule out, or at least prioritize, some of the suspected drivers. We observed, for Particles and Optical-Planetary instruments, that perhaps a launch manifest effect was playing part of the role in the observed decrease in dollar-per-kg over the years, noting that the more flagship class missions, which have more money to spend on their instruments, were seen in the earlier years in our data, versus the later years which were dominated by less expensive class missions. However, if this were a dominating driver, would we not have seen the downward trend in the Fields and Microwave instruments as well, which were drawn from that same launch manifest? The fact that we did not observe this helps us rule out the launch manifest effect, and other drivers, such as advances in technology, that seem to be more likely suspects. In that case, however, why would technology advances be helping the Particles and Optical-Planetary instruments only? Why would it not be impacting Optical-Earth Orbiting instruments? Further suspects were looked at as well and ruled out, such as the “Faster, Better, Cheaper”era of NASA development which did not seem to actually impact trends by instrument type on a dollar-perkg scale. VI. Future Work A. Time Series Detailed Statistical Assessment The analysis discussed above sets the foundation for a more rigorous time series analysis of the data. Time series analysis will further explore evidence to-date of time trends for the instrument types which showed the strongest indicators for a decrease in dollar-per-kg: Optical (Planetary) and Particles instruments. More than providing evidence and top-level significance tests, time series analysis would help elucidate what kind of trend that exists in the data, their significance and allow statistically based forecasting (see Figure 10)

Mrozinksi, Joseph