Search NASA⌕ Search

SEARCH · Search NASA

Results for “simple linear regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A Method for Assessing the Quality of Model-Based Estimates of Ground Temperature and Atmospheric Moisture Using Satellite Data

A method is developed for validating model-based estimates of atmospheric moisture and ground temperature using satellite data. The approach relates errors in estimates of clear-sky longwave fluxes at the top of the Earth-atmosphere system to errors in geophysical parameters. The fluxes include clear-sky outgoing longwave radiation (CLR) and radiative flux in the window region between 8 and 12 microns (RadWn). The approach capitalizes on the availability of satellite estimates of CLR and RadWn and other auxiliary satellite data, and multiple global four-dimensional data assimilation (4-DDA) products. The basic methodology employs off-line forward radiative transfer calculations to generate synthetic clear-sky longwave fluxes from two different 4-DDA data sets. Simple linear regression is used to relate the clear-sky longwave flux discrepancies to discrepancies in ground temperature ((delta)T(sub g)) and broad-layer integrated atmospheric precipitable water ((delta)pw). The slopes of the regression lines define sensitivity parameters which can be exploited to help interpret mismatches between satellite observations and model-based estimates of clear-sky longwave fluxes. For illustration we analyze the discrepancies in the clear-sky longwave fluxes between an early implementation of the Goddard Earth Observing System Data Assimilation System (GEOS2) and a recent operational version of the European Centre for Medium-Range Weather Forecasts data assimilation system. The analysis of the synthetic clear-sky flux data shows that simple linear regression employing (delta)T(sub g)) and broad layer (delta)pw provides a good approximation to the full radiative transfer calculations, typically explaining more thin 90% of the 6 hourly variance in the flux differences. These simple regression relations can be inverted to "retrieve" the errors in the geophysical parameters, Uncertainties (normalized by standard deviation) in the monthly mean retrieved parameters range from 7% for (delta)T(sub g) to approx. 20% for the lower tropospheric moisture between 500 hPa and surface. The regression relationships developed from the synthetic flux data, together with CLR and RadWn observed with the Clouds and Earth Radiant Energy System instrument, ire used to assess the quality of the GEOS2 T(sub g) and pw. Results showed that the GEOS2 T(sub g) is too cold over land, and pw in upper layers is too high over the tropical oceans and too low in the lower atmosphere.

Wu, Man Li C.↗

Estimates of Ground Temperature and Atmospheric Moisture from CERES Observations

A method is developed to retrieve surface ground temperature (Tg) and atmospheric moisture using clear sky fluxes (CSF) from CERES-TRMM observations. In general, the clear sky outgoing long-wave radiation (CLR) is sensitive to upper level moisture (q(sub h)) over wet regions and Tg over dry regions The clear sky window flux from 800 to 1200 /cm (RadWn) is sensitive to low level moisture (q(sub j)) and Tg. Combining these two measurements (CLR and RadWn), Tg and q(sub h) can be estimated over land, while q(sub h) and q(sub t) can be estimated over the oceans. The approach capitalizes on the availability of satellite estimates of CLR and RadWn and other auxiliary satellite data. The basic methodology employs off-line forward radiative transfer calculations to generate synthetic CSF data from two different global 4-dimensional data assimilation products. Simple linear regression is used to relate discrepancies in CSF to discrepancies in Tg, q(sub h) and q(sub t). The slopes of the regression lines define sensitivity parameters that can be exploited to help interpret mismatches between satellite observations and model-based estimates of CSF. For illustration, we analyze the discrepancies in the CSF between an early implementation of the Goddard Earth Observing System Data Assimilation System (GEOS-DAS) and a recent operational version of the European Center for Medium-Range Weather Prediction data assimilation system. In particular, our analysis of synthetic total and window region SCF differences (computed from two different assimilated data sets) shows that simple linear regression employing (Delta)Tg and broad layer (Delta)q(sub l) from 500 hPa to surface and (Delta)q(sub h) from 200 to 500 hPa provides a good approximation to the full radiative transfer calculations, typically explaining more than 90% of the 6-hourly variance in the flux differences. These simple regression relations can be inverted to "retrieve" the errors in the geophysical parameters. Uncertainties (normalized by standard deviation) in the monthly mean retrieved parameters range from 7% for (Delta)T to about 20% for (Delta)q(sub t). Our initial application of the methodology employed an early CERES-TRMM data set (CLR and Radwn) to assess the quality of the GEOS2 data. The results showed that over the tropical and subtropical oceans GEOS2 is, in general, too wet in the upper troposphere (mean bias of 0.99 mm) and too dry in the lower troposphere (mean bias of -4.7 mm). We note that these errors, as well as a cold bias in the Tg, have largely been corrected in the current version of GEOS-2 with the introduction of a land surface model, a moist turbulence scheme and the assimilation of SSTM/I total precipitable water.

Wu, Man Li C.↗

Estimates of Ground Temperature and Atmospheric Moisture from CERES Observations

A method is developed to retrieve surface ground temperature (T(sub g)) and atmospheric moisture using clear sky fluxes (CSF) from CERES-TRMM observations. In general, the clear sky outgoing longwave radiation (CLR) is sensitive to upper level moisture (q(sub l)) over wet regions and (T(sub g)) over dry regions The clear sky window flux from 800 to 1200/cm (RadWn) is sensitive to low level moisture (q(sub t)) and T(sub g). Combining these two measurements (CLR and RadWn), Tg and q(sub h) can be estimated over land, while q(sub h) and q(sub l) can be estimated over the oceans. The approach capitalizes on the availability of satellite estimates of CLR and RadWn and other auxiliary satellite data. The basic methodology employs off-line forward radiative transfer calculations to generate synthetic CSF data from two different global 4-dimensional data assimilation products. Simple linear regression is used to relate discrepancies in CSF to discrepancies in T(sub g), q(sub h) and q(sub l). The slopes of the regression lines define sensitivity parameters that can be exploited to help interpret mismatches between satellite observations and model-based estimates of CSF. For illustration, we analyze the discrepancies in the CSF between an early implementation of the Goddard Earth Observing System Data Assimilation System (GEOS-DAS) and a recent operational version of the European Center for Medium-Range Weather Prediction data assimilation system. In particular, our analysis of synthetic total and window region SCF differences (computed from two different assimilated data sets) shows that simple linear regression employing Delta(T(sub g)) and broad layer Delta(q(sub l) from .500 hPa to surface and Delta(q(sub h)) from 200 to .300 hPa provides a good approximation to the full radiative transfer calculations. typically explaining more than 90% of the 6-hourly variance in the flux differences. These simple regression relations can be inverted to "retrieve" the errors in the geophysical parameters. Uncertainties (normalized by standard deviation) in the monthly mean retrieved parameters range from 7% for Delta(T(sub g)) to about 20% for Delta(q(sub l)). Our initial application of the methodology employed an early CERES-TRMM data set (CLR and Radwn) to assess the quality of the GEOS2 data. The results showed that over the tropical and subtropical oceans GEOS2 is, in general, too wet in the upper troposphere (mean bias of 0.99 mm) and too dry in the lower troposphere (mean bias of -4.7 min). We note that these errors, as well as a cold bias in the T(sub g). have largely been corrected in the current version of GEOS-2 with the introduction of a land surface model, a moist turbulence scheme and the assimilation of SSM/I total precipitable water.

Wu, Man Li C.↗

High temperature oxidation of corrosion resistant alloys from machine learning

Parabolic rate constants, k p , were collected from published reports and calculated from corrosion product data (sample mass gain or corrosion product thickness) and tabulated for 75 alloys exposed to temperatures between ~800 and 2000 K (~500–1700 °C; 900–3000°F). Data were collected for environments including lab air, ambient and supercritical carbon dioxide, supercritical water, and steam. Materials studied include low- and high-Cr ferritic and austenitic steels, nickel superalloys, and aluminide materials. A combination of Arrhenius analysis, simple linear regression, supervised and unsupervised machine learning methods were used to investigate the relations between composition and oxidation kinetics. The supervised machine learning techniques produced the lowest mean standard errors. The most significant elements controlling oxidation kinetics were Ni, Cr, Al, and Fe, with Mo and Co composition also found to be significant features. The activation energies produced from the machine learning analysis were in the correct distributions for the diffusion constants for the oxide scales expected to dominate in each class.

Materials Science↗

The empirical accuracy of uncertain inference models

Uncertainty is a pervasive feature of the domains in which expert systems are designed to function. Research design to test uncertain inference methods for accuracy and robustness, in accordance with standard engineering practice is reviewed. Several studies were conducted to assess how well various methods perform on problems constructed so that correct answers are known, and to find out what underlying features of a problem cause strong or weak performance. For each method studied, situations were identified in which performance deteriorates dramatically. Over a broad range of problems, some well known methods do only about as well as a simple linear regression model, and often much worse than a simple independence probability model. The results indicate that some commercially available expert system shells should be used with caution, because the uncertain inference models that they implement can yield rather inaccurate results.

Vaughan, David S.↗

The influence of snow depth and surface air temperature on satellite-derived microwave brightness temperature

Areas of the steppes of central Russia, the high plains of Montana and North Dakota, and the high plains of Canada were studied in an effort to determine the relationship between passive microwave satellite brightness temperature, surface air temperature, and snow depth. Significant regression relationships were developed in each of these homogeneous areas. Results show that sq R values obtained for air temperature versus snow depth and the ratio of microwave brightness temperature and air temperature versus snow depth were not as the sq R values obtained by simply plotting microwave brightness temperature versus snow depth. Multiple regression analysis provided only marginal improvement over the results obtained by using simple linear regression.

Foster, J. L.↗

Application of thermal model for pan evaporation to the hydrology of a defined medium, the sponge

A technique is presented which estimates pan evaporation from the commonly observed values of daily maximum and minimum air temperatures. These two variables are transformed to saturation vapor pressure equivalents which are used in a simple linear regression model. The model provides reasonably accurate estimates of pan evaporation rates over a large geographic area. The derived evaporation algorithm is combined with precipitation to obtain a simple moisture variable. A hypothetical medium with a capacity of 8 inches of water is initialized at 4 inches. The medium behaves like a sponge: it absorbs all incident precipitation, with runoff or drainage occurring only after it is saturated. Water is lost from this simple system through evaporation just as from a Class A pan, but at a rate proportional to its degree of saturation. The contents of the sponge is a moisture index calculated from only the maximum and minium temperatures and precipitation.

Trenchard, M. H.↗

Empirical Estimation of Shortest Route Length along U.S. Interstate Highways Based on Great Circle Distance

In this study, 98 regression models were specified for easily estimating shortest distances based on great circle distances along the U.S. interstate highways nationwide and for each of the continental 48 states. This allows transportation professionals to quickly generate distance, or even distance matrix, without expending significant efforts on complicated shortest path calculations. For simple usage by all professionals, all models are present in the simple linear regression form. Only one explanatory variable, the great circle distance, is considered to calculate the route distance. For each geographic scope (i.e., the national or one of the states), two different models were considered, with and without the intercept. Based on the adjusted R-squared, it was observed that models without intercepts generally have better fitness. Additionally, all these models generally have good fitness with the linear regression relationship between the great circle distance and route distance. At the state level, significant variations in the slope coefficients between the state-level models were also observed. Furthermore, a preliminary analysis of the effect of highway density on this variation was conducted.

33 ADVANCED PROPULSION SYSTEMS↗

Using diverse potentials and scoring functions for the development of improved machine-learned models for protein–ligand affinity and docking pose prediction

The advent of computational drug discovery holds the promise of significantly reducing the effort of experimentalists, along with monetary cost. More generally, predicting the binding of small organic molecules to biological macromolecules has far-reaching implications for a range of problems, including metabolomics. However, problems such as predicting the bound structure of a protein–ligand complex along with its affinity have proven to be an enormous challenge. In recent years, machine learning-based methods have proven to be more accurate than older methods, many based on simple linear regression. Nonetheless, there remains room for improvement, as these methods are often trained on a small set of features, with a single functional form for any given physical effect, and often with little mention of the rationale behind choosing one functional form over another. Moreover, it is not entirely clear why one machine learning method is favored over another. Here, we endeavor to undertake a comprehensive effort towards developing high-accuracy, machine-learned scoring functions, systematically investigating the effects of machine learning method and choice of features, and, when possible, providing insights into the relevant physics using methods that assess feature importance. Here, we show synergism among disparate features, yielding adjusted R 2 with experimental binding affinities of up to 0.871 on an independent test set and enrichment for native bound structures of up to 0.913. When purely physical terms that model enthalpic and entropic effects are used in the training, we use feature importance assessments to probe the relevant physics and hopefully guide future investigators working on this and other computational chemistry problems.

59 BASIC BIOLOGICAL SCIENCES↗

Artificial intelligence-based predictive modeling for imaging neutral particle analyzers on the DIII-D tokamak

The Imaging Neutral Particle Analyzer (INPA) at DIII-D is a diagnostic system used to accurately resolve the energy and spatial distributions of fast ions in fusion plasmas. A novel artificial intelligence (AI) technique named INPA-net is based on Reservoir Computing Networks and developed here to predict active and passive signals produced by charge-exchange reactions from injected and edge-cold neutrals, respectively, in magnetically confined fusion plasmas. This model is trained using a set of 21 time domain signals between 0 s to 3.35 s that includes injected beam and thermal plasma information, and 6444 real 2D experimental images of the INPA in 12 plasma discharges at DIII-D. The trained neural network is able to forecast experimental images in real-time. The model achieves an R-squared value of 0.91, which is higher than the 0.83 value achieved by a simple linear regression model. This improvement highlights the model's enhanced predictive accuracy for measured images from the validation set. This AI approach is valuable due to its rapid response times and potential for integration into real-time plasma control systems. A version of this model capable of generating syntehic images would be useful for the real-time monitoring of fast-ion transport. A comprehensive sensitivity study reveals that INPA-net maintains high performance even with variations in the input parameters, indicating the model's robustness and reliability. While developed for the INPA, the underlying architecture is adaptable and may be applied to various 2D imaging diagnostics in fusion research.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Prediction of DIII-D Pedestal Structure From Externally Controllable Parameters

The sharp increase of pressure at the edge of a high confinement mode (H-mode) plasma, the pedestal, strongly impacts overall plasma performance. Predicting the pedestal is a necessity to control and optimize tokamak operations. Here, an experimental data-driven machine learning (ML) approach is presented that predicts the pedestal heights and widths of electron density (n e ) and electron temperature (T e ) profiles as well as the separatrix ne from externally controllable parameters such as the plasma shape, heating method and power, and gas puff rate and integrated gas puff. The OMFIT framework was used with DIII-D data to efficiently, robustly, and automatically build a database of pedestal parameters to train machine learning models. Database creation was enabled by the search engine tool for DIII-D data, TokSearch, which parallelizes data fetching, enabling fast searches through basic signals of thousands of DIII-D shots and selection of relevant time intervals. Principal Component Analysis (PCA) separated the database into three clusters that represent classes of plasma shapes that are regularly used in DIII-D. The most important parameters for setting the pedestal structure were plasma current (I p ), toroidal magnetic field (B Φ ), neutral beam heating power (P NBI ) and shaping quantities. The Deep Jointly Informed Neural Networks (DJINN) algorithm was applied to identify suitable neural network (NN) architectures that appropriately capture the features of the pedestal database. Separate NNs were implemented for each pedestal parameter, and ensembling methods were used to improve the prediction accuracy and allowed estimation of the prediction uncertainty. The pedestal predictions of the test dataset lie within the measurement uncertainties of the pedestal parameters. The NN outperformed simple Linear Regression (LR) analysis, indicating non-linear dependencies in the pedestal structure. The presented achievements illustrate a promising path for future research, using feature extraction to infer experimental trends and thereby improve pedestal models as well as deploying NN for a fast pedestal prediction in DIII-D scenario development.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Next-Level Energy Management in Manufacturing: Facility-Level Energy Digital Twin Framework Based on Machine Learning and Automated Data Collection

This research introduces an energy prediction framework at the facility level supported by automated data collection and machine learning models. It investigates whether reducing the prediction time scale allows for applying more complex machine learning techniques and if those techniques improve the prediction accuracy. The primary advantages of this framework lie in its automation of the energy prediction process and its provision of real-time energy data suitable for use in energy dashboards or digital twins. A sitewide dataset was created by combining 15 min energy and daily production data of five shops—assembly, battery, body (electric), body (gas), and paint—from a globally recognized electric vehicle manufacturer. Various machine learning models were evaluated on daily, weekly, and monthly datasets, including, in increasingly complex order: naïve, simple linear regression, net regularized generalized linear regression, principal component regression, k-nearest neighbor, random forest, and Bayesian regularized neural network. Compared to the current state-of-the-art energy consumption prediction for the industrial facility level, this research investigates more complex models and smaller time intervals for higher accuracy. The findings revealed that the more complex monthly models require a minimum of a year and a half of data to operate, while weekly models demand a year of data to achieve improved accuracy. Daily models can operate with only six months of data but exhibit poor performance due to reduced prediction accuracy of production. Key challenges identified include access to reliable, high-quality energy and production data and the initial demand for human labor.

digital twin↗

Prediction of DIII-D Pedestal Structure from Externally Controllable Parameters

The sharp increase of pressure at the edge of a high confinement mode (H-mode) plasma, the pedestal, strongly impacts overall plasma performance. Predicting the pedestal is a necessity to control and optimize tokamak operations. An experimental data-driven machine learning (ML) approach is presented that predicts the pedestal heights and widths of electron density (ne) and electron temperature (Te) profiles as well as the separatrix ne from externally controllable parameters such as the plasma shape, heating method and power, and gas puff rate and integrated gas puff. The OMFIT framework was used with DIII-D data to efficiently, robustly, and automatically build a database of pedestal parameters to train machine learning models. Database creation was enabled by the search engine tool for DIII-D data, TokSearch, which parallelizes data fetching, enabling fast searches through basic signals of thousands of DIII-D shots and selection of relevant time intervals. Principal Component Analysis (PCA) separated the database into three clusters that represent classes of plasma shapes that are regularly used in DIII-D. The most important parameters for setting the pedestal structure were plasma current (Ip), toroidal magnetic field (Bφ), neutral beam heating power (PNBI) and shaping quantities. The Deep Jointly Informed Neural Networks (DJINN) algorithm was applied to identify suitable neural network (NN) architectures that appropriately capture the features of the pedestal database. Separate NNs were implemented for each pedestal parameter, and ensembling methods were used to improve the prediction accuracy and allowed estimation of the prediction uncertainty. The pedestal predictions of the test dataset lie within the measurement uncertainties of the pedestal parameters. The NN outperformed simple Linear Regression (LR) analysis, indicating non-linear dependencies in the pedestal structure. The presented achievements illustrate a promising path for future research, using feature extraction to infer experimental trends and thereby improve pedestal models as well as deploying NN for a fast pedestal prediction in DIII-D scenario development.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

The application of dimensional analysis to the problem of solar wind-magnetosphere energy coupling

The constraints imposed by dimensional analysis are used to find how the solar wind-magnetosphere energy transfer rate depends upon interplanetary parameters. The analyses assume that only magnetohydrodynamic processes are important in controlling the rate of energy transfer. The study utilizes ISEE-3 solar wind observations, the AE index, and UT from three 10-day intervals during the International Magnetospheric Study. Simple linear regression and histogram techniques are used to find the value of the magnetohydrodynamic coupling exponent, alpha, which is consistent with observations of magnetospheric response. Once alpha is estimated, the form of the solar wind energy transfer rate is obtained by substitution into an equation of the interplanetary variables whose exponents depend upon alpha.

Bargatze, L. F.↗

Analysis of the lettuce data from the variable pressure growth chamber at NASA Johnson Space Center: A three-stage nested design model

A model of three-stage nested experimental design was applied to analyze the lettuce data obtained from the variable pressure growth chamber test bed at NASA-Johnson Space Center. From the results of an application of the analysis of variance and covariance on the data set, it was noted that all of the (uncontrollable) factors, Side, Zone, Height and (controllable) PAR (photosynthetically active radiation), had nonhomogeneous effects on the dry weight of the edible biomass of lettuce per pot. Incidentally, the variations accountable to the (uncontrollable) factorial heterogeneities are merely 9 percent and 17 percent of the total variation for both the first and second crop test, respectively. After adjusting for the PAR as a covariate in the no-intercept model, the accountable variations to all the four factors are 94 percent and 92 percent for the first and the second crop test, respectively. With the use of a no-intercept simple linear regression model, the accountable variations to the factor PAR are 92 percent and 90 percent for the first and the second crop test, respectively. Evidently, the (controllable) factor PAR is the dominating one.

Lee, Tze-San↗

Multiple Use One-Sided Hypotheses Testing in Univariate Linear Calibration

Consider a normally distributed response variable, related to an explanatory variable through the simple linear regression model. Data obtained on the response variable, corresponding to known values of the explanatory variable (i.e., calibration data), are to be used for testing hypotheses concerning unknown values of the explanatory variable. We consider the problem of testing an unlimited sequence of one sided hypotheses concerning the explanatory variable, using the corresponding sequence of values of the response variable and the same set of calibration data. This is the situation of multiple use of the calibration data. The tests derived in this context are characterized by two types of uncertainties: one uncertainty associated with the sequence of values of the response variable, and a second uncertainty associated with the calibration data. We derive tests based on a condition that incorporates both of these uncertainties. The solution has practical applications in the decision limit problem. We illustrate our results using an example dealing with the estimation of blood alcohol concentration based on breath estimates of the alcohol concentration. In the example, the problem is to test if the unknown blood alcohol concentration of an individual exceeds a threshold that is safe for driving.

Krishnamoorthy, K.↗

Instrumentation and methodology for quantifying GFP fluorescence in intact plant organs

The General Fluorescence Plant Meter (GFP-Meter) is a portable spectrofluorometer that utilizes a fiber-optic cable and a leaf clip to gather spectrofluorescence data. In contrast to traditional analytical systems, this instrument allows for the rapid detection and fluorescence measurement of proteins under field conditions with no damage to plant tissue. Here we discuss the methodology of gathering and standardizing spectrofluorescence data from tobacco and canola plants expressing GFP. Furthermore, we demonstrate the accuracy and effectiveness of the GFP-Meter. We first compared GFP fluorescence measurements taken by the GFP-Meter to those taken by a standard laboratory-based spectrofluorometer, the FluoroMax-2. Spectrofluorescence measurements were taken from the same location on intact leaves. When these measurements were tested by simple linear regression analysis, we found that there was a positive functional relationship between instruments. Finally, to exhibit that the GFP-Meter recorded accurate measurements over a span of time, we completed a time-course analysis of GFP fluorescence measurements. We found that only initial measurements were accurate; however, subsequent measurements could be used for qualitative purposes.

Evaluation Studies↗

Homozygous diploid deletion strains of Saccharomyces cerevisiae that determine lag phase and dehydration tolerance

This study identifies genes that determine length of lag phase, using the model eukaryotic organism, Saccharomyces cerevisiae. We report growth of a yeast deletion series following variations in the lag phase induced by variable storage times after drying-down yeast on filters. Using a homozygous diploid deletion pool, lag times ranging from 0 h to 90 h were associated with increased drop-out of mitochondrial genes and increased survival of nuclear genes. Simple linear regression (R2 analysis) shows that there are over 500 genes for which > 70% of the variation can be explained by lag alone. In the genes with a positive correlation, such that the gene abundance increases with lag and hence the deletion strain is suitable for survival during prolonged storage, there is a strong predominance of nucleonic genes. In the genes with a negative correlation, such that the gene abundance decreases with lag and hence the strain may be critical for getting yeast out of the lag phase, there is a strong predominance of glycoproteins and transmembrane proteins. This study identifies yeast deletion strains with survival advantage on prolonged storage and amplifies our understanding of the genes critical for getting out of the lag phase.

NASA Discipline Cell Biology↗