Search NASA⌕ Search

SEARCH · Search NASA

Results for “Ordinary least square”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Two biased estimation techniques in linear regression: Application to aircraft

Several ways for detection and assessment of collinearity in measured data are discussed. Because data collinearity usually results in poor least squares estimates, two estimation techniques which can limit a damaging effect of collinearity are presented. These two techniques, the principal components regression and mixed estimation, belong to a class of biased estimation techniques. Detection and assessment of data collinearity and the two biased estimation techniques are demonstrated in two examples using flight test data from longitudinal maneuvers of an experimental aircraft. The eigensystem analysis and parameter variance decomposition appeared to be a promising tool for collinearity evaluation. The biased estimators had far better accuracy than the results from the ordinary least squares technique.

Klein, Vladislav↗

The role of social support on midwestern farmers’ willingness to grow perennial bioenergy crops

The lack of farmers' willingness to grow perennial bioenergy crops (PBCs) presents a critical barrier to the emergence of cellulosic biofuel production. The willingness relies on a complex network of economic, environmental, and social drivers, among which the influence of social factors (e.g., the influence of neighborhood, community, and communication) is less understood. This study addresses this knowledge gap via a survey analysis of midwestern farmers. The survey data are analyzed through ordinary least square regression and structural equation model, which together investigate the individual and interactive impacts of multiple factors on farmers' decisions to adopt PBCs. Based on a farm-scale analysis, six statistically significant predictors of farmer willingness to grow PBCs are identified: perception of PBCs' environment benefits, education level, willingness to take risks, familiarity with PBCs, portion of peers already growing PBCs, and support of biorefineries locating in the local community. Among these, the latter three predictors are social support variables. It is found that familiarity with the crops is the most significant predictor of willingness; familiarity is also an important intermediate variable that mediates the influence of many other predictors. In addition, peer adoption can both directly and indirectly affect willingness via its influence on familiarity. Furthermore, these findings suggest that it is a pressing need to improve farmers’ knowledge of PBCs to promote the adoption of such crops.

09 BIOMASS FUELS↗

Evaluating the role of green infrastructure features in post-disaster recovery – Case Study of Beaumont, Texas after tropical storm imelda

While green infrastructure (GI) can provide multiple environmental benefits, its role in post-disaster economic and social recovery remains relatively underexplored. This article investigates whether different characteristics of GI, such as size, shape, connectivity, and amenities, affect the resilience of local businesses following Tropical Storm Imelda in Beaumont, Texas. The study utilizes SafeGraph mobility data to analyze foot traffic patterns to local businesses before, during, and after the disaster. FRAGSTATS indices measure GI characteristics (e.g., area, shape index, fractal dimension, proximity) while park features such as sports facilities, playgrounds, water features, and accessibility are cataloged through manual observation. Ordinary Least Squares regression models assess the relationship between park characteristics and post-recovery business performance, controlling for demographic variables including income, race, and poverty levels. Results indicate that certain GI attributes significantly enhance business recovery. Points of interest within walking distance (0.5 miles) of parks demonstrated better post-recovery status compared to those beyond this range. Specifically, parks with larger areas (p < 0.01) and more complex shapes measured by fractal dimension index (p < 0.01) had the strongest positive impact on surrounding businesses' recovery. Interestingly, playgrounds showed a negative correlation with recovery (p < 0.05), likely due to flood damage rendering them unusable during the immediate recovery period. Social vulnerability factors, including higher poverty rates and minority populations, negatively affected recovery outcomes despite park proximity.

Economic resilience↗

Geographical Insights into Suicide Mortality Through Spatial Machine Learning

Suicide mortality is a leading cause of death in the United States, with an upward trend that emphasizes its significance as a public health issue. Previous research has employed global models like ordinary least squares (OLS) regression and local models such as geographically weighted regression (GWR). While local models are useful for analyzing spatial variations in suicide mortality, they share limitations with traditional global models, particularly about their inability to handle multi-collinearity and non-linear relationships. Machine learning approaches, like random forests (RF), can address some of these limitations but often fail to account for spatial variability. This gap highlights the need for spatial ML models specifically designed to tackle suicide mortality. This research seeks to fill this void by using a geographically weighted random forest model (GWRF) to examine the associations between county-level suicide mortality in the U.S. from 2010 to 2020 and various social and environmental determinants of health. A key aspect of our methodology is disciplined feature selection, which reduces the pool of explanatory variables by about 90%. This refinement enhances the explanatory power of both global (R2 improved from 0.59 to 0.67) and local (R2 improved from 0.64 to 0.67) RF models while reducing their run times. An analysis of the importance scores for these selected features reveals that the drivers of suicide mortality vary by context. Thus, to effectively address regional disparities and inform targeted public health interventions, a holistic approach that incorporates multiple county-level characteristics is essential.

Lebakula, Viswadeep [ORNL] (ORCID:0000000152935914↗

What Shapes Transportation Charging Infrastructure Availability? Evidence from Tennessee

This study examines how community, travel, and freight characteristics relate to public charging infrastructure availability across Tennessee ZIP codes. We link Alternative Fuels Data Center station locations with traffic, socioeconomic, demographic, commuting, and freight employment data to build a ZIP code-level dataset. Ordinary least squares regression captures variation in chargers per 10,000 residents (R2=0.311). Quantile regressions at the 25th, 50th, and 75th percentiles, with pseudo R2 values up to 0.099, show that the determinants of infrastructure availability differ across low-, medium-, and high-availability areas. Percent female, percent car commuters, average household size, and median age are negatively associated with charging availability across much of the distribution. Truck traffic is positively associated only in lower-availability ZIP codes, while vehicle miles traveled shifts from a negative association at the lower end of the distribution to a positive association at the upper end. The results provide insight into how public charging deployment aligns with community characteristics, mobility demand, and freight activity across Tennessee. Future work can distinguish charger types and power levels, incorporate land-use and temporal rollout patterns, and examine how charging infrastructure needs differ across urban and rural contexts.

Calderón, Oriana [University of Tennessee, Knoxvil↗

Data-driven modeling to enhance municipal water demand estimates in response to dynamic climate conditions

Altered precipitation and temperature patterns from a changing climate will affect supply, demand, and overall municipal water system operations throughout the arid western U.S. While supply forecasts leverage hydrological models to connect climate influences with surface water availability, demand forecasts typically estimate water use independent of climate and other externalities. Stemming from an increased focus on seasonal water demand management, we use the Salt Lake City, Utah municipal water system as a test bed to assess model accuracy versus complexity trade-offs between simple climate-independent econometric-based models and complex climate-sensitive data-driven models to average to extreme wet and dry climate conditions—representative of a new climate normal. Here, the climate-independent model displayed low performance during extreme dry conditions with predictions exceeding 90% and 40% of the observed monthly and seasonal volumetric demands, respectively, which we attribute to insufficient model complexity. The climate-sensitive models displayed greater accuracy in all conditions, with an ordinary least squares model demonstrating a measurable reduction in prediction bias (3.4% vs. -27.3%) and RMSE (74.0 lpcd vs. 294 lpcd) compared to the climate-independent model. The climate-sensitive workflow increased model accuracy and characterized climate-demand interactions, demonstrating a novel tool to enhance water system management.

54 ENVIRONMENTAL SCIENCES↗

Data for The Role of Social Support on Midwestern Farmers’ Willingness to Grow Perennial Bioenergy Crops

The lack of farmers’ willingness to grow perennial bioenergy crops (PBCs) presents a critical barrier to the emergence of cellulosic biofuel production. The willingness relies on a complex network of economic, environmental, and social drivers, among which the influence of social factors (e.g., the influence of neighborhood, community, and communication) is less understood. This study addresses this knowledge gap via a survey analysis of midwestern farmers. The survey data are analyzed through ordinary least square regression and structural equation model, which together investigate the individual and interactive impacts of multiple factors on farmers’ decisions to adopt PBCs. Based on a farm-scale analysis, six statistically significant predictors of farmer willingness to grow PBCs are identified: perception of PBCs’ environment benefits, education level, willingness to take risks, familiarity with PBCs, portion of peers already growing PBCs, and support of biorefineries locating in the local community. Among these, the latter three predictors are social support variables. It is found that familiarity with the crops is the most significant predictor of willingness; familiarity is also an important intermediate variable that mediates the influence of many other predictors. In addition, peer adoption can both directly and indirectly affect willingness via its influence on familiarity. These findings suggest that it is a pressing need to improve farmers’ knowledge of PBCs to promote the adoption of such crops.

Economics↗

Mountain Basin Controls on the Snow-to-Streamflow Signal: An AIC-Weighted Multiple Linear Regression Framework

A regression-based analysis quantifies how basin characteristics modulate the snow-to-streamflow signal. First, we use the ERA5-Land reanalysis gridded product (European Centre for Medium Range Weather Forecasts reanalysis 5 -Land component) for 4,655 hydrologic unit code - 10 (HUC10) mountain basins across the western United States (US) for water years 1987–2024. Linear regressions are performed for peak snow water equivalent (SWE) and annual streamflow for each mountain basin. Models use ordinary least squares in Python’s statsmodels package. After which, an Akaike Information Criterion (AIC)–weighted ensemble multiple linear regression (MLR) framework with 47 watershed traits is used to predict the linear regression coefficient of determination (r-squared) defining the ability of peak SWE to predict annual streamflow across all mountain basin. Predictor sets are constrained to avoid multicollinearity by excluding models with variance inflation factors (VIF) greater than 5. Mountain basin traits included in the MLR include seasonal climate, topography, vegetation type and structure, and bedrock geology. Accepted models are considered if their AIC is within 2.0 of the model with the minimum AIC, or best model. To compare predictor influence across acceptable models, we computed standardized regression coefficients. To evaluate structural redundancy among models, we constructed binary inclusion vectors for each acceptable model, denoting whether a predictor was present (1) or absent (0). Core predictor variables are defined as occurring in at least 67% of the acceptable models. For this regional analysis, only one model was found acceptable, with higher snow-to-streamflow translation (higher r-squared) occurring in colder mountain basins with higher relative winter precipitation, more snow accumulation and a lower fraction of annual precipitation that falls in the spring and summer. The second component of the data package uses previously published, high-resolution output from an integrated hydrological model of the East River watershed using the U.S. Geological Survey Groundwater and Surface water Flow model (GSFLOW, doi:10.15485/1998576). East River MLR expands upon the approach described above to explore the response of five streamflow metrics—annual streamflow, runoff efficiency, 7-day minimum flow, low-flow duration, and non-perennial stream fraction to snow system indicators including peak SWE, snow-covered area, snow disappearance date, and the fraction of basin area characterized by low-to-no snow, as well as seasonal precipitation and temperature, and annual hydrologic variables representing soil moisture, evapotranspiration (ET), the partitioning of incoming precipitation to evapotranspiration (ET/P), groundwater storage, and groundwater inflow to streams. MLR was done on all water years (P0: 1987-2024) and for each period as determined in the split analysis using pooled regression techniques (P1: 1987-2011 and P2: 2012-2024) to evaluate shifting predictor variable emphasis on streamflow generation. Results indicate that since 2012, peak SWE has lost statistical strength in its prediction of annual streamflow and runoff efficiency, and the indirect influence of spring temperature has emerged as critically important. Low-flow metrics remain largely influenced by soil moisture, vegetation water use and groundwater inflows with summer precipitation becoming a direct influence on minimum summer flow. Together, these data and Python-based analysis tools provide a framework for identifying the key watershed characteristics that control how streamflow responds to snow from year to year. The package also helps quantify uncertainty in statistical models and assess how snow–streamflow relationships vary across regions and over time. This dataset contains comma-separated values files (.csv), text files (.txt), python code files (.py), figure files (.png), and shapefiles (.cpg, .dbf, .prj, .sbn, .sbx, .shp, .xml). Further details on file contents and MLR execution can be found in the readme file and the FLMD files. Work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Maximum Likelihood Estimation: Some Basics

The maximum likelihood estimation is a general estimation procedure. It is often compared to estimation procedures like the ordinary least squares regression or generalized method of moments, to name a few. We discuss some basics about the maximum likelihood estimation, its advantages and disadvantages, and provide an example application to a gamma distribution function.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Maximum Likelihood Estimation: Some Basics

The maximum likelihood estimation is a general estimation procedure. It is often compared to estimation procedures like the ordinary least squares regression or generalized method of moments, to name a few. We discuss some basics about the maximum likelihood estimation, its advantages and disadvantages, and provide an example application to a gamma distribution function.

97 MATHEMATICS AND COMPUTING↗

Standardising the “Gregory method” for calculating equilibrium climate sensitivity

The equilibrium climate sensitivity (ECS) – the equilibrium global mean temperature response to a doubling of atmospheric CO 2 – is a high-profile metric for quantifying the Earth system's response to human-induced climate change. A widely applied approach to estimating the ECS is the “Gregory method” (Gregory et al., 2004), which uses an ordinary least squares (OLS) regression between the net radiative flux, N, and surface air temperature anomalies, ΔT, from a 150 year experiment in which atmospheric CO 2 concentrations are quadrupled. The ECS is determined by extrapolating the linear fit to N=0, i.e. the ΔT-intercept, indicating the point at which the system is back in equilibrium. This method has been used to compare ECS estimates across the CMIP5 and CMIP6 ensembles and will likely be a key diagnostic for CMIP7. Despite its widespread application, there is little consistency or transparency between studies in how the climate model data is processed prior to the regression, leading to potential discrepancies in ECS estimates. We identify 32 alternative data processing pathways, varying by differences in global mean weighting, net radiative flux variable, anomaly calculation method, and linear regression fit. Using 44 CMIP6 models, we systematically assess the impact of these choices on ECS estimates and calculate uncertainty ranges using two bootstrap approaches. While the inter-model ECS range is insensitive to the data processing pathway, individual outlier models exhibit notable differences. Approximating a model's native grid cell area (if irregular) with cosine of the latitude can decrease the ECS by 11 %, the choice of N-variable can change the ECS by 6 %, and some anomaly calculation methods can introduce spurious temporal correlations in the processed data. Beyond data processing choices, we also evaluate an alternative linear regression method – total least squares (TLS) – which has a more statistically robust basis than OLS. However, for consistency with previous literature, and given TLS may reduce the ECS compared to OLS (by up to 24 %), thereby making a known bias in the Gregory method worse, we do not feel there is sufficient clarity to recommend a transition to TLS in all cases. To improve reproducibility and comparability in future studies, we recommend a standardised Gregory method: weighting the global mean by cell area, using the top of the atmosphere (as opposed to the top of model) N-variable, and calculating anomalies by first applying a rolling average to the preindustrial control timeseries then subtracting from the raw CO 2 quadrupling experiment. This approach accounts for model drift while reducing noise in the data to best meet the pre-conditions of the linear regression. While CMIP6 results of the multi-model mean ECS appear insensitive to these processing choices, similar assumptions may not hold for CMIP7, underscoring the need for standardised data preparation in future climate sensitivity assessments.

Geosciences↗

Aerodynamic parameters of an advanced fighter aircraft estimated from flight data. Preliminary results

Preliminary estimates of aerodynamic parameters of an advanced fighter aircraft were obtained from flight data of different values of the angle of attack from 8 to 54 deg. The data were analyzed by a stepwise regression with the ordinary least squares technique. The estimated stability and control derivatives are plotted against the angle of attack and compared with wind tunnel measurement and previous flight results. Also included is the data compatibility check of measured data. The effect of various input forms on the estimates is demonstrated in two examples using simulated data.

Klein, Vladislav↗

Aerodynamic Parameters of High Performance Aircraft Estimated from Wind Tunnel and Flight Test Data

A concept of system identification applied to high performance aircraft is introduced followed by a discussion on the identification methodology. Special emphasis is given to model postulation using time invariant and time dependent aerodynamic parameters, model structure determination and parameter estimation using ordinary least squares an mixed estimation methods, At the same time problems of data collinearity detection and its assessment are discussed. These parts of methodology are demonstrated in examples using flight data of the X-29A and X-31A aircraft. In the third example wind tunnel oscillatory data of the F-16XL model are used. A strong dependence of these data on frequency led to the development of models with unsteady aerodynamic terms in the form of indicial functions. The paper is completed by concluding remarks.

Klein, Vladislav↗

Aerodynamic Parameters of High Performance Aircraft Estimated from Wind Tunnel and Flight Test Data

A concept of system identification applied to high performance aircraft is introduced followed by a discussion on the identification methodology. Special emphasis is given to model postulation using time invariant and time dependent aerodynamic parameters, model structure determination and parameter estimation using ordinary least squares and mixed estimation methods. At the same time problems of data collinearity detection and its assessment are discussed. These parts of methodology are demonstrated in examples using flight data of the X-29A and X-31A aircraft. In the third example wind tunnel oscillatory data of the F-16XL model are used. A strong dependence of these data on frequency led to the development of models with unsteady aerodynamic terms in the form of indicial functions. The paper is completed by concluding remarks.

Klein, Vladislav↗

Baseline Observations of Hemispheric Sea Ice with the Nimbus 7 Scanning Multichannel Microwave Radiometer

The Scanning Multichannel Microwave Radiometer (SMMR) on board the NASA Nimbus 7 satellite was designed to obtain data for sea surface temperatures (SSTs), near-surface wind speeds, sea ice coverage and type, rainfall rates over the oceans, cloud water content, snow water equivalent, and soil moisture. In this paper, I shall emphasize the sea ice observations and mention briefly some important SST observations. A prime factor contributing to the importance of SMMR sea ice observations lies in their successful integration into a long-term time series, presently being extended by observations from the series of Special Sensor Microwave/Imager (SSMI) on board the DOD/DMSP F8, Fl1, and F12 satellites. This currently constitutes a 19-year data set. Almost half of this was provided by the SMMR. Unfortunately, the 4-year data set produced earlier by the single-channel Electrically Scanned Microwave Radiometer (ESMR) was not successfully integrated into the SMMR/SSMI data set. This resulted primarily from the lack of an overlap period to provide intersensor adjustment, but also because of the large difference between the algorithms to produce ice concentrations and large temporal gaps in the ESMR data. The lack of overlap between the SeaSat and Nimbus 7 SMMR data sets was an important consideration for also excluding the SeatSat one, but the spatial gaps especially in the Southern Hemisphere daily SeaSat observations was another. The sea ice observations will continue into the future by means of the Advanced Microwave Scanning Radiometer (AMSR) on board the ADEOS II and EOS satellites due to be launched in mid- and late-2000, respectively. Analysis of the sea ice data has been carried out by a number of different techniques. Long-term trends have been examined by means of ordinary least squares and band-limited regression. Oscillations in the data have been examined by band-limited Fourier analysis. Here, I shall present results from a novel combination of Principal Component analysis and the recently-developed Empirical Mode Decomposition (EMD). In this method, the data are first separated into spatial and temporal parts, and then the temporal parts of the first few PCs are broken into intrinsic modes by the EMD method.

Gloersen, Per↗

An Airline-Based Multilevel Analysis of Airfare Elasticity for Passenger Demand

Price elasticity of passenger demand for a specific airline is estimated. The main drivers affecting passenger demand for air transportation are identified. First, an Ordinary Least Squares regression analysis is performed. Then, a multilevel analysis-based methodology to investigate the pattern of variation of price elasticity of demand among the various routes of the airline under study is proposed. The experienced daily passenger demands on each fare-class are grouped for each considered route. 9 routes were studied for the months of February and May in years from 1999 to 2002, and two fare-classes were defined (business and economy). The analysis has revealed that the airfare elasticity of passenger demand significantly varies among the different routes of the airline.

Castelli, Lorenzo↗

Classes of Split-Plot Response Surface Designs for Equivalent Estimation

When planning an experimental investigation, we are frequently faced with factors that are difficult or time consuming to manipulate, thereby making complete randomization impractical. A split-plot structure differentiates between the experimental units associated with these hard-to-change factors and others that are relatively easy-to-change and provides an efficient strategy that integrates the restrictions imposed by the experimental apparatus. Several industrial and scientific examples are presented to illustrate design considerations encountered in the restricted randomization context. In this paper, we propose classes of split-plot response designs that provide an intuitive and natural extension from the completely randomized context. For these designs, the ordinary least squares estimates of the model are equivalent to the generalized least squares estimates. This property provides best linear unbiased estimators and simplifies model estimation. The design conditions that allow for equivalent estimation are presented enabling design construction strategies to transform completely randomized Box-Behnken, equiradial, and small composite designs into a split-plot structure.

Parker, Peter A.↗

Measurement System Characterization in the Presence of Measurement Errors

In the calibration of a measurement system, data are collected in order to estimate a mathematical model between one or more factors of interest and a response. Ordinary least squares is a method employed to estimate the regression coefficients in the model. The method assumes that the factors are known without error; yet, it is implicitly known that the factors contain some uncertainty. In the literature, this uncertainty is known as measurement error. The measurement error affects both the estimates of the model coefficients and the prediction, or residual, errors. There are some methods, such as orthogonal least squares, that are employed in situations where measurement errors exist, but these methods do not directly incorporate the magnitude of the measurement errors. This research proposes a new method, known as modified least squares, that combines the principles of least squares with knowledge about the measurement errors. This knowledge is expressed in terms of the variance ratio - the ratio of response error variance to measurement error variance.

Commo, Sean A.↗