Search NASA⌕ Search

SEARCH · Search NASA

Results for “regularized regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Kernel Partial Least Squares for Nonlinear Regression and Discrimination

This paper summarizes recent results on applying the method of partial least squares (PLS) in a reproducing kernel Hilbert space (RKHS). A previously proposed kernel PLS regression model was proven to be competitive with other regularized regression methods in RKHS. The family of nonlinear kernel-based PLS models is extended by considering the kernel PLS method for discrimination. Theoretical and experimental results on a two-class discrimination problem indicate usefulness of the method.

Rosipal, Roman↗

IN11B-1621: Quantifying How Climate Affects Vegetation in the Amazon Rainforest

Amazon droughts in 2005 and 2010 have raised serious concern about the future of the rainforest. Amazon forests are crucial because of their role as the largest carbon sink in the world which would effect the global warming phenomena with decreased photosynthesis activity. Especially, after a decline in plant growth in 1.68 million km2 forest area during the once-in-a-century severe drought in 2010, it is of primary importance to understand the relationship between different climatic variables and vegetation. In an earlier study, we have shown that non-linear models are better at capturing the relation dynamics of vegetation and climate variables such as temperature and precipitation, compared to linear models. In this research, we learn precise models between vegetation and climatic variables (temperature, precipitation) for normal conditions in the Amazon region using genetic programming based symbolic regression. This is done by removing high elevation and drought affected areas and also considering the slope of the region as one of the important factors while building the model. The model learned reveals new and interesting ways historical and current climate variables affect the vegetation at any location. MAIAC data has been used as a vegetation surrogate in our study. For temperature and precipitation, we have used TRMM and MODIS Land Surface Temperature data sets while learning the non-linear regression model. However, to generalize the model to make it independent of the data source, we perform transfer learning where we regress a regularized least squares to learn the parameters of the non-linear model using other data sources such as the precipitation and temperature from the Climatic Research Center (CRU). This new model is very similar in structure and performance compared to the original learned model and verifies the same claims about the nature of dependency between these climate variables and the vegetation in the Amazon region. As a result of this study, we are able to learn, for the very first time how exactly different climate factors influence vegetation at any location in the Amazon rainforests, independent of the specific sources from which the data has been obtained.

global warming↗

Soil temperature extrema recovery rates after precipitation cooling

From a one dimensional view of temperature alone variations at the Earth's surface manifest themselves in two cyclic patterns of diurnal and annual periods, due principally to the effects of diurnal and seasonal changes in solar heating as well as gains and losses of available moisture. Beside these two well known cyclic patterns, a third cycle has been identified which occurs in values of diurnal maxima and minima soil temperature extrema at 10 cm depth usually over a mesoscale period of roughly 3 to 14 days. This mesoscale period cycle starts with precipitation cooling of soil and is followed by a power curve temperature recovery. The temperature recovery clearly depends on solar heating of the soil with an increased soil moisture content from precipitation combined with evaporation cooling at soil temperatures lowered by precipitation cooling, but is quite regular and universal for vastly different geographical locations, and soil types and structures. The regularity of the power curve recovery allows a predictive model approach over the recovery period. Multivariable linear regression models alloy predictions of both the power of the temperature recovery curve as well as the total temperature recovery amplitude of the mesoscale temperature recovery, from data available one day after the temperature recovery begins.

Welker, J. E.↗

A Statistical Model to Predict the Extratropical Transition of Tropical Cyclones

This paper introduces a logistic regression model for the extratropical transition(ET) of tropical cyclones in the North Atlantic and the Western North Pacific, using elastic net regularization to select predictors and estimate coefficients.Predictors are chosen from the 1979-2017 best track and reanalysis datasets, and verification is done against the tropical/extratropical labels in the best track data. In an independent test set, the model skillfully predicts ET at lead times up to two days, with latitude and sea surface temperature as its most important predictors. At a lead time of 24 h, it predicts ET with a Matthews correlation coefficient of 0.4 in the North Atlantic, and 0.6 in the Western North Pacific. It identifies 80% of storms undergoing ET in the North Atlantic, and 92% of those in the Western North Pacific. 90% of transition time errors are less than 24 h. Select examples of the model's performance on individual storms illustrate its strengths and weaknesses.Two versions of the model are presented: an "operational model" that may provide baseline guidance for operational forecasts, and a "hazard model"that can be integrated into statistical TC risk models. As instantaneous diagnostics for tropical/extratropical status, both models' zero lead time predictions perform about as well as the widely used Cyclone Phase Space (CPS) in the Western North Pacific and better than the CPS in the North Atlantic, and predict the timings of the transitions better than CPS in both basins.

Melanie Bieli↗

Smoothing splines: Regression, derivatives and deconvolution

The statistical properties of a cubic smoothing spline and its derivative are analyzed. It is shown that unless unnatural boundary conditions hold, the integrated squared bias is dominated by local effects near the boundary. Similar effects are shown to occur in the regularized solution of a translation-kernel intergral equation. These results are derived by developing a Fourier representation for a smoothing spline.

Rice, J.↗

Algorithmic Detection of Elemental Biosignatures

Machine learning models that classify a sample as indicative or non-indicative of life could play an important role in life-detection missions. Their predictions result from agnostic algorithms and thereby add redundancy to judgements resulting from human expertise. Additionally, their important features can reveal the most informative measurements within the operational constraints of a life-detection mission. The Ladder of Life Detection (Neveu 2018) identifies the need for an understanding of how combinations of multiple biosignatures affect overall confidence. The present work provides a starting point to answer this need, and future work will expand the data types to obtain even more predictive combinations of features. Elemental abundance was chosen as a starting set of features due to its availability in diverse sample types, which are needed to train a generalizable model. A standardized dataset was collected, including 35 non-indicative, e.g., lunar rock, basalt; 19 indicative mixed, e.g., seawater, agricultural soil; 46 indicative non-alive, e.g., coal, chalk; and 10 indicative alive, e.g., biofilm, bacteria. This dataset could be valuable for complementary biosignature research. The samples were standardized to the same limit of detection of a simulated mission scenario. Four classification models were used: k-nearest neighbors (KNN), logistic regression (LR), linear support vector machines (SVM), and Gaussian naïve Bayes (GNB). To obtain feature importances, KNN was run on three principal components of the training data and LR and SVM were run with L1 and L2 regularization. The performances and feature importances of the six model variants on 40:60 train to validation ratios were assessed with Monte Carlo simulations. ROC AUC and mean accuracy scores ranged between 82% - 94%, with sensitivity greater than specificity. For indicative of life predictors, all models had C and Ca as strong and Cl as medium; a majority of models had N, K, and P as medium. For non-indicative of life predictors, all models had Si as strong, and a majority of models had Mg, Al, and Ti as medium. Varied elements were Fe (slightly non-indicative), H (slightly indicative), O (widely varied), Na, Mn, and S. These results serve as a proof of concept and suggest important elemental signals beyond merely the CHNOPS of Earth-based life.

Algorithmic↗

Bayesian Model Selection for Reducing Bloat and Overfitting in Genetic Programming for Symbolic Regression

When performing symbolic regression using genetic programming, overfitting and bloat can negatively impact generalizability and interpretability of the resulting equations as well as increase computation times. A Bayesian fitness metric is introduced and its impact on bloat and overfitting during population evolution is studied and compared to common alternatives in the literature. The proposed approach was found to be more robust to noise and data sparsity in numerical experiments, guiding evolution to a level of complexity appropriate to the dataset. Further evolution of the population resulted not in overfitting or bloat, but rather in slight simplifications in model form. The ability to identify an equation of complexity appropriate to the scale of noise in the training data was also demonstrated. In general, the Bayesian model selection algorithm was shown to be an effective means of regularization which resulted in less bloat and overfitting when any amount of noise was present in the training data.

Uncertainty quantification↗

Neural network uncertainty assessment using Bayesian statistics: a remote sensing application

Neural network (NN) techniques have proved successful for many regression problems, in particular for remote sensing; however, uncertainty estimates are rarely provided. In this article, a Bayesian technique to evaluate uncertainties of the NN parameters (i.e., synaptic weights) is first presented. In contrast to more traditional approaches based on point estimation of the NN weights, we assess uncertainties on such estimates to monitor the robustness of the NN model. These theoretical developments are illustrated by applying them to the problem of retrieving surface skin temperature, microwave surface emissivities, and integrated water vapor content from a combined analysis of satellite microwave and infrared observations over land. The weight uncertainty estimates are then used to compute analytically the uncertainties in the network outputs (i.e., error bars and correlation structure of these errors). Such quantities are very important for evaluating any application of an NN model. The uncertainties on the NN Jacobians are then considered in the third part of this article. Used for regression fitting, NN models can be used effectively to represent highly nonlinear, multivariate functions. In this situation, most emphasis is put on estimating the output errors, but almost no attention has been given to errors associated with the internal structure of the regression model. The complex structure of dependency inside the NN is the essence of the model, and assessing its quality, coherency, and physical character makes all the difference between a blackbox model with small output errors and a reliable, robust, and physically coherent model. Such dependency structures are described to the first order by the NN Jacobians: they indicate the sensitivity of one output with respect to the inputs of the model for given input data. We use a Monte Carlo integration procedure to estimate the robustness of the NN Jacobians. A regularization strategy based on principal component analysis is proposed to suppress the multicollinearities in order to make these Jacobians robust and physically meaningful.

Neural Networks (Computer)↗

Improvements of the Load Schedule for the Machine Calibration of a Strain-Gage Balance

The load schedule for the calibration of a six-component force balance in a calibration machine was improved. Now, single-component loads are repeated in regular intervals during the calibration. This approach has several advantages. First, the number of single-component loads increases to about twenty-two percent of all loads and load combinations. Consequently, more accurate numerical estimates of the primary bridge sensitivities can be obtained if global regression is used for the analysis of the calibration data. In addition, single-component repeats make it possible to track the stability of the applied loads during the calibration process. Finally, interactions of single-component repeats can be compared with interactions that are observed during the application of manual loads to the balance. Machine calibration and manual data sets of two force balances are used to illustrate benefits of the new load schedule. It is shown in the examples how differences between the observed interactions of machine calibration and manual data can be quantified. The suggested improvements can also be implemented in the load schedule for the machine calibration of a moment or direct-read balance as long as single-component loads are included that are described in the design load format of the balance.

wind tunnel test↗

MODIS Tree Cover Validation for the Circumpolar Taiga-Tundra Transition Zone

A validation of the 2005 500m MODIS vegetation continuous fields (VCF) tree cover product in the circumpolar taiga-tundra ecotone was performed using high resolution Quickbird imagery. Assessing the VCF's performance near the northern limits of the boreal forest can help quantify the accuracy of the product within this vegetation transition area. The circumpolar region was divided into longitudinal zones and validation sites were selected in areas of varying tree cover where Quickbird imagery is available in Google Earth. Each site was linked to the corresponding VCF pixel and overlaid with a regular dot grid within the VCF pixel's boundary to estimate percent tree crown cover in the area. Percent tree crown cover was estimated using Quickbird imagery for 396 sites throughout the circumpolar region and related to the VCF's estimates of canopy cover for 2000-2005. Regression results of VCF inter-annual comparisons (2000-2005) and VCF-Quickbird image-interpreted estimates indicate that: (1) Pixel-level, inter-annual comparisons of VCF estimates of percent canopy cover were linearly related (mean R(sup 2) = 0.77) and exhibited an average root mean square error (RMSE) of 10.1 % and an average root mean square difference (RMSD) of 7.3%. (2) A comparison of image-interpreted percent tree crown cover estimates based on dot counts on Quickbird color images by two different interpreters were more variable (R(sup 2) = 0.73, RMSE = 14.8%, RMSD = 18.7%) than VCF inter-annual comparisons. (3) Across the circumpolar boreal region, 2005 VCF-Quickbird comparisons were linearly related, with an R(sup 2) = 0.57, a RMSE = 13.4% and a RMSD = 21.3%, with a tendency to over-estimate areas of low percent tree cover and anomalous VCF results in Scandinavia. The relationship of the VCF estimates and ground reference indicate to potential users that the VCF's tree cover values for individual pixels, particularly those below 20% tree cover, may not be precise enough to monitor 500m pixel-level tree cover in the taiga-tundra transition zone.

Montesano, P. M.↗

In-Situ Measurement of Hall Thruster Erosion Using a Fiber Optic Regression Probe

One potential life-limiting mechanism in a Hall thruster is the erosion of the ceramic material comprising the discharge channel. This is especially true for missions that require long thrusting periods and can be problematic for lifetime qualification, especially when attempting to qualify a thruster by analysis rather than a test lasting the full duration of the mission. In addition to lifetime, several analytical and numerical models include electrode erosion as a mechanism contributing to enhanced transport properties. However, there is still a great deal of dispute over the importance of erosion to transport in Hall thrusters. The capability to perform an in-situ measurement of discharge channel erosion is useful in addressing both the lifetime and transport concerns. An in-situ measurement would allow for real-time data regarding the erosion rates at different operating points, providing a quick method for empirically anchoring any analysis geared towards lifetime qualification. Erosion rate data over a thruster s operating envelope would also be useful in the modeling of the detailed physics inside the discharge chamber. There are many different sensors and techniques that have been employed to quantify discharge channel erosion in Hall thrusters. Snapshots of the wear pattern can be obtained at regular shutdown intervals using laser profilometry. Many non-intrusive techniques of varying complexity and sensitivity have been employed to detect the time-varying presence of erosion products in the thruster plume. These include the use quartz crystal microbalances, emission spectroscopy, laser induced flourescence, and cavity ring-down spectroscopy. While these techniques can provide a very accurate picture of the level of eroded material in the thruster plume, it is more difficult to use them to determine the location from which the material was eroded. Furthermore, none of the methods cited provide a true in-situ measure of erosion at the channel surface while the thruster is in operation (i.e. none yield a continuous channel erosion measurement). A recent fundamental sensor development effort has led to a novel regression, erosion, and ablation sensor technology (REAST). The REAST sensor allows for measurement of real-time surface erosion rates at a discrete surface location. The sensor was tested using a linear Hall thruster geometry (see Fig. 1), which served as a means of producing plasma erosion of a ceramic discharge chamber. The mass flow rate, discharge voltage, and applied magnetic field strength could be varied, allowing for erosion measurements over a broad thruster operating envelope. Results are presented demonstrating the ability of the REAST sensor to capture not only the insulator erosion rates but also changes in these rates as a function of the discharge parameters.

Polzin, Kurt↗

A Method for Calculating the Probability of Successfully Completing a Rocket Propulsion Ground Test

Propulsion ground test facilities face the daily challenges of scheduling multiple customers into limited facility space and successfully completing their propulsion test projects. Due to budgetary and schedule constraints, NASA and industry customers are pushing to test more components, for less money, in a shorter period of time. As these new rocket engine component test programs are undertaken, the lack of technology maturity in the test articles, combined with pushing the test facilities capabilities to their limits, tends to lead to an increase in facility breakdowns and unsuccessful tests. Over the last five years Stennis Space Center's propulsion test facilities have performed hundreds of tests, collected thousands of seconds of test data, and broken numerous test facility and test article parts. While various initiatives have been implemented to provide better propulsion test techniques and improve the quality, reliability, and maintainability of goods and parts used in the propulsion test facilities, unexpected failures during testing still occur quite regularly due to the harsh environment in which the propulsion test facilities operate. Previous attempts at modeling the lifecycle of a propulsion component test project have met with little success. Each of the attempts suffered form incomplete or inconsistent data on which to base the models. By focusing on the actual test phase of the tests project rather than the formulation, design or construction phases of the test project, the quality and quantity of available data increases dramatically. A logistic regression model has been developed form the data collected over the last five years, allowing the probability of successfully completing a rocket propulsion component test to be calculated. A logistic regression model is a mathematical modeling approach that can be used to describe the relationship of several independent predictor variables X(sub 1), X(sub 2),..,X(sub k) to a binary or dichotomous dependent variable Y, where Y can only be one of two possible outcomes, in this case Success or Failure. Logistic regression has primarily been used in the fields of epidemiology and biomedical research, but lends itself to many other applications. As indicated the use of logistic regression is not new, however, modeling propulsion ground test facilities using logistic regression is both a new and unique application of the statistical technique. Results from the models provide project managers with insight and confidence into the affectivity of rocket engine component ground test projects. The initial success in modeling rocket propulsion ground test projects clears the way for more complex models to be developed in this area.

Messer, Bradley P.↗

Ride quality evaluation. IV - Models of subjective reaction to aircraft motion

The paper examines models of human reaction to the motions typically experienced on short-haul aircraft flights. Data are taken on the regularly scheduled flights of four commercial airlines - three airplanes and one helicopter. The data base consists of: (1) a series of motion recordings distributed over each flight, each including all six degrees of freedom of motion; temperature, pressure, and noise are also recorded; (2) ratings of perceived comfort and satisfaction from the passengers on each flight; (3) moment-by-moment comfort ratings from a test subject assigned to each airplane; and (4) overall comfort ratings for each flight from the test subjects. Regression models are obtained for prediction of rated comfort from rms values for six degrees of freedom of motion. It is shown that the model C = 2.1 + 17.1 T + 17.2 V (T = transverse acceleration, V = vertical acceleration) gives a good fit to the airplane data but is less acceptable for the helicopter data.

Jacobson, I. D.↗

Evaluating Combinations of Sentinel-2 Data and Machine-Learning Algorithms for Mangrove Mapping in West Africa

Creating a national baseline for natural resources, such as mangrove forests, and monitoring them regularly often requires a consistent and robust methodology. With freely available satellite data archives and cloud computing resources, it is now more accessible to conduct such large-scale monitoring and assessment. Yet, few studies examine the reproducibility of such mangrove monitoring frameworks, especially in terms of generating consistent spatial extent. Our objective was to evaluate a combination of image processing approaches to classify mangrove forests along the coast of Senegal and The Gambia. We used freely available global satellite data (Sentinel-2), and cloud computing platform (Google Earth Engine) to run two machine learning algorithms, random forest (RF), and classification and regression trees (CART). We calibrated and validated the algorithms using 800 reference points collected using high-resolution images. We further re-ran 10 iterations for each algorithm, utilizing unique subsets of the initial training data. While all iterations resulted in thematic mangrove maps with over 90% accuracy, the mangrove extent ranges between 827-2807 km2 for Senegal and 245-1271 km2 for The Gambia with one outlier for each country. We further report "Places of Agreement" (PoA) to identify areas where all iterations for both methods agree (506.6 km2 and 129.6 km2 for Senegal and The Gambia, respectively), thus have a high confidence in predicting mangrove extent. While we acknowledge the time- and cost-effectiveness of such methods for the landscape managers, we recommend utilizing them with utmost caution, as well as post-classification on-the-ground checks, especially for decision making.

Mondal, Pinki↗

Bayesian Model Selection for Reducing Bloat and Overfitting in Genetic Programming for Symbolic Regression

When performing symbolic regression using genetic programming, overfitting and bloat can negatively impact generalizability and interpretability of the resulting equations as well as increase computation times. A Bayesian fitness metric is introduced and its impact on bloat and overfitting during population evolution is studied and compared to common alternatives in the literature. The proposed approach was found to be more robust to noise and data sparsity in numerical experiments, guiding evolution to a level of complexity appropriate to the dataset. Further evolution of the population resulted not in overfitting or bloat, but rather in slight simplifications in model form. The ability to identify an equation of complexity appropriate to the scale of noise in the training data was also demonstrated. In general, the Bayesian model selection algorithm was shown to be an effective means of regularization which resulted in less bloat and overfitting when any amount of noise was present in the training data.

G F Bomarito↗

Estimation of Aerosol Optical Depth at Different Wavelengths by Multiple Regression Method

This study aims to investigate and establish a suitable model that can help to estimate aerosol optical depth (AOD) in order to monitor aerosol variations especially during non-retrieval time. The relationship between actual ground measurements (such as air pollution index, visibility, relative humidity, temperature, and pressure) and AOD obtained with a CIMEL sun photometer was determined through a series of statistical procedures to produce an AOD prediction model with reasonable accuracy. The AOD prediction model calibrated for each wavelength has a set of coefficients. The model was validated using a set of statistical tests. The validated model was then employed to calculate AOD at different wavelengths. The results show that the proposed model successfully predicted AOD at each studied wavelength ranging from 340 nm to 1020 nm. To illustrate the application of the model, the aerosol size determined using measure AOD data for Penang was compared with that determined using the model. This was done by examining the curvature in the ln [AOD]-ln [wavelength] plot. Consistency was obtained when it was concluded that Penang was dominated by fine mode aerosol in 2012 and 2013 using both measured and predicted AOD data. These results indicate that the proposed AOD prediction model using routine measurements as input is a promising tool for the regular monitoring of aerosol variation during non-retrieval time.

Tan, Fuyi↗

Cardiac interbeat interval dynamics from childhood to senescence : comparison of conventional and new measures based on fractals and chaos theory

BACKGROUND: New methods of R-R interval variability based on fractal scaling and nonlinear dynamics ("chaos theory") may give new insights into heart rate dynamics. The aims of this study were to (1) systematically characterize and quantify the effects of aging from early childhood to advanced age on 24-hour heart rate dynamics in healthy subjects; (2) compare age-related changes in conventional time- and frequency-domain measures with changes in newly derived measures based on fractal scaling and complexity (chaos) theory; and (3) further test the hypothesis that there is loss of complexity and altered fractal scaling of heart rate dynamics with advanced age. METHODS AND RESULTS: The relationship between age and cardiac interbeat (R-R) interval dynamics from childhood to senescence was studied in 114 healthy subjects (age range, 1 to 82 years) by measurement of the slope, beta, of the power-law regression line (log power-log frequency) of R-R interval variability (10(-4) to 10(-2) Hz), approximate entropy (ApEn), short-term (alpha(1)) and intermediate-term (alpha(2)) fractal scaling exponents obtained by detrended fluctuation analysis, and traditional time- and frequency-domain measures from 24-hour ECG recordings. Compared with young adults (<40 years old, n=29), children (<15 years old, n=27) showed similar complexity (ApEn) and fractal correlation properties (alpha(1), alpha(2), beta) of R-R interval dynamics despite lower spectral and time-domain measures. Progressive loss of complexity (decreased ApEn, r=-0.69, P<0.001) and alterations of long-term fractal-like heart rate behavior (increased alpha(2), r=0.63, decreased beta, r=-0.60, P<0.001 for both) were observed thereafter from middle age (40 to 60 years, n=29) to old age (>60 years, n=29). CONCLUSIONS: Cardiac interbeat interval dynamics change markedly from childhood to old age in healthy subjects. Children show complexity and fractal correlation properties of R-R interval time series comparable to those of young adults, despite lower overall heart rate variability. Healthy aging is associated with R-R interval dynamics showing higher regularity and altered fractal scaling consistent with a loss of complex variability.

NASA Discipline Cardiopulmonary↗

Sparse Regression as a Sparse Eigenvalue Problem

We extend the l0-norm "subspectral" algorithms for sparse-LDA [5] and sparse-PCA [6] to general quadratic costs such as MSE in linear (kernel) regression. The resulting "Sparse Least Squares" (SLS) problem is also NP-hard, by way of its equivalence to a rank-1 sparse eigenvalue problem (e.g., binary sparse-LDA [7]). Specifically, for a general quadratic cost we use a highly-efficient technique for direct eigenvalue computation using partitioned matrix inverses which leads to dramatic x103 speed-ups over standard eigenvalue decomposition. This increased efficiency mitigates the O(n4) scaling behaviour that up to now has limited the previous algorithms' utility for high-dimensional learning problems. Moreover, the new computation prioritizes the role of the less-myopic backward elimination stage which becomes more efficient than forward selection. Similarly, branch-and-bound search for Exact Sparse Least Squares (ESLS) also benefits from partitioned matrix inverse techniques. Our Greedy Sparse Least Squares (GSLS) generalizes Natarajan's algorithm [9] also known as Order-Recursive Matching Pursuit (ORMP). Specifically, the forward half of GSLS is exactly equivalent to ORMP but more efficient. By including the backward pass, which only doubles the computation, we can achieve lower MSE than ORMP. Experimental comparisons to the state-of-the-art LARS algorithm [3] show forward-GSLS is faster, more accurate and more flexible in terms of choice of regularization

Exact Sparse Least Squares (ESLS)↗