Search NASA⌕ Search

SEARCH · Search NASA

Results for “Multivariate regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Large-Scale Inference of Multivariate Regression for Heavy-Tailed and Asymmetric Data

Large-scale multivariate regression is a fundamental statistical tool with a wide range of applications. Here, this study considers the problem of simultaneously testing a large number of general linear hypotheses, encompassing covariate-effect analysis, analysis of variance, and model comparisons. The challenge that accompanies a large number of tests is the ubiquitous presence of heavy-tailed and/or highly skewed measurement noise, which is the main reason for the failure of conventional least squares-based methods. For large-scale multivariate regression, we develop a set of robust inference methods to explore data features such as heavy tailedness and skewness, which are not visible to least squares methods. The new testing procedure is based on the data-adaptive Huber regression and a new covariance estimator of regression estimates. Under mild conditions, we show that our methods produce consistent estimates of the false discovery proportion. Extensive numerical experiments and an empirical study on quantitative linguistics demonstrate the advantage of the proposed method over many state-of-the-art methods when the data are generated from heavy-tailed and/or skewed distributions.

97 MATHEMATICS AND COMPUTING↗

Estimating Sparse Direct Effects in Multivariate Regression With the Spike-and-Slab LASSO

The multivariate regression interpretation of the Gaussian chain graph model simultaneously parametrizes (i) the direct effects of p predictors on q outcomes and (ii) the residual partial covariances between pairs of outcomes. We introduce a new method for fitting sparse versions of these models with spike-and-slab LASSO (SSL) priors. We develop an Expectation Conditional Maximization algorithm to obtain sparse estimates of the p × q matrix of direct effects and the q × q residual precision matrix. Our algorithm iteratively solves a sequence of penalized maximum likelihood problems with self-adaptive penalties that gradually filter out negligible regression coefficients and partial covariances. Because it adaptively penalizes individual model parameters, our method is seen to outperform fixed-penalty competitors on simulated data. We establish the posterior contraction rate for our model, buttressing our method’s excellent empirical performance with strong theoretical guarantees. Using our method, we estimated the direct effects of diet and residence type on the composition of the gut microbiome of elderly adults.

EM algorithm↗

Linear Multivariable Regression Models for Prediction of Eddy Dissipation Rate from Available Meteorological Data

Linear multivariable regression models for predicting day and night Eddy Dissipation Rate (EDR) from available meteorological data sources are defined and validated. Model definition is based on a combination of 1997-2000 Dallas/Fort Worth (DFW) data sources, EDR from Aircraft Vortex Spacing System (AVOSS) deployment data, and regression variables primarily from corresponding Automated Surface Observation System (ASOS) data. Model validation is accomplished through EDR predictions on a similar combination of 1994-1995 Memphis (MEM) AVOSS and ASOS data. Model forms include an intercept plus a single term of fixed optimal power for each of these regression variables; 30-minute forward averaged mean and variance of near-surface wind speed and temperature, variance of wind direction, and a discrete cloud cover metric. Distinct day and night models, regressing on EDR and the natural log of EDR respectively, yield best performance and avoid model discontinuity over day/night data boundaries.

MCKissick, Burnell T.↗

Fast and Non-Destructive Determination of Water Content in Ionic Liquids at Varying Temperatures by Raman Spectroscopy and Multivariate Regression Analysis

Imidazolium acetate ionic liquids (ILs) have been utilized as promising solvents in many applications that involve varying water content and temperature. These experimental variables affect the anion-cation intermolecular interactions, which in turn influence the performance of the ILs in these applications. Here, this paper shows Raman spectroscopy can be used as an operando method to measure water content in IL solvents when simultaneous temperature changes may occur. The Raman spectra of 1-alkyl-3-methylimidazolium acetate ILs (alkyl chain length n = 2, 4, 6, 8) with varying water content (from 0.028 to 0.899 water mole fraction) and temperature (from 78.1 K to 423.1 K) were measured. Increasing the water content or decreasing the temperature of the tested ILs weakens the anion-cation intermolecular interactions. The water content of these ILs can be quantified even in conditions when the temperature is changing using Raman spectroscopy combined with multivariate regression analysis, including principal component regression (PCR), partial-least-squares regression (PLSR), and artificial neural networks (ANNs). The ANN model combined with partial-least-squares (PLS) achieves the highest prediction accuracy of water content in ILs at varying temperatures (RMSECV = 0.017, R 2 CV = 99.1%, RMSEP = 0.019, R 2 P = 98.8%, RPD = 8.93). Raman spectroscopy provides a potential fast non-destructive operando method to monitor the water content of ILs even in applications when the temperature may be simultaneously altered; this information can lead to the optimized use of these ILs in many applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Uncertainty quantification in multivariable regression for material property prediction with Bayesian neural networks

With the increased use of data-driven approaches and machine learning-based methods in material science, the importance of reliable uncertainty quantification (UQ) of the predicted variables for informed decision-making cannot be overstated. UQ in material property prediction poses unique challenges, including multi-scale and multi-physics nature of materials, intricate interactions between numerous factors, limited availability of large curated datasets, etc. In this work, we introduce a physics-informed Bayesian Neural Networks (BNNs) approach for UQ, which integrates knowledge from governing laws in materials to guide the models toward physically consistent predictions. To evaluate the approach, we present case studies for predicting the creep rupture life of steel alloys. Experimental validation with three datasets of creep tests demonstrates that this method produces point predictions and uncertainty estimations that are competitive or exceed the performance of conventional UQ methods such as Gaussian Process Regression. Additionally, we evaluate the suitability of employing UQ in an active learning scenario and report competitive performance. The most promising framework for creep life prediction is BNNs based on Markov Chain Monte Carlo approximation of the posterior distribution of network parameters, as it provided more reliable results in comparison to BNNs based on variational inference approximation or related NNs with probabilistic outputs.

36 MATERIALS SCIENCE↗

Supervised machine learning-based multivariate regression of parallel closures for a high-collisionality deuterium-carbon plasma

Many plasmas of interest in laboratory experiments and space consist of multiple ion species. In tokamak edge plasmas, for instance, ionized impurities expelled from the vessel wall influence plasma transport. When describing multi-species plasmas using fluid equations, we need accurate closure relations to close the set of fluid equations. In this study, we introduce the development of fitting formulas for parallel closures using supervised machine learning, in conjunction with the recent closure theory, considering multi-ion collisions and arbitrary ion temperatures. We apply this approach to a high-collisionality deuterium-carbon plasma and demonstrate its effectiveness. As a result, the machine learning-based method for developing practical and accurate closures can be extended to a wider range of plasmas.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Comparing Designed Training Sets to Optimize Multivariate Regression Models for Pr, Nd, and Nitric Acid Using Spectrophotometry

Chemometric regression models were developed for the quantification of praseodymium (Pr, 0–1000 µg/mL), neodymium (Nd, 0–1000 µg/mL), and nitric acid (HNO 3 , 0.1–5 M) using spectrophotometry. Designed calibration sets were composed of 20 samples each: 10 model points and 10 lack-of-fit (LOF) points. The D-optimal designs effectively minimized the number of samples required to build models, and each design resulted in similar prediction performance, suggesting that statistical design of experiments can provide a reliable framework for selecting training set samples in three-variable systems. Partial least squares regression (PLSR) models were validated against a one-factor-at-a-time validation set composed of 125 samples (three variables, five levels). The top PLS-1 models resulted in average percent root mean square error of prediction error values of 3.5%, 1.7%, and 1.2% for Pr(III), Nd(III), and HNO 3 , respectively. Power set augmentations of the model and LOF samples were investigated to optimize the number of training set samples. PLSR models built using just required model points (10) had similar predictive capabilities as models including the LOF points (20) but with fewer samples. The number of validation samples was also varied systematically to learn how many samples are needed to validate regression models. This work addresses long-standing questions in the field of chemometrics to help make this approach amenable to the near-real-time quantification of hazardous species in remote settings.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multivariate regression modelling for gender prediction using volatile organic compounds from hand odor profiles via HS-SPME-GC-MS

The efficacy of using human volatile organic compounds (VOCs) as a form of forensic evidence has been well demonstrated with canines for crime scene response, suspect identification, and location checking. Although the use of human scent evidence in the field is well established, the laboratory evaluation of human VOC profiles has been limited. This study used Headspace-Solid Phase Microextraction-Gas Chromatography-Mass Spectrometry (HS-SPME-GC-MS) to analyze human hand odor samples collected from 60 individuals (30 Females and 30 Males). The human volatiles collected from the palm surfaces of each subject were interpreted for classification and prediction of gender. The volatile organic compound (VOC) signatures from subjects’ hand odor profiles were evaluated with supervised dimensional reduction techniques: Partial Least Squares-Discriminant Analysis (PLS-DA), Orthogonal-Projections to Latent Structures Discriminant Analysis (OPLS-DA), and Linear Discriminant Analysis (LDA). The PLS-DA 2D model demonstrated clustering amongst male and female subjects. The addition of a third component to the PLS-DA model revealed clustering and minimal separation of male and female subjects in the 3D PLS-DA model. The OPLS-DA model displayed discrimination and clustering amongst gender groups with leave one out cross validation (LOOCV) and 95% confidence regions surrounding clustered groups without overlap. The LDA had a 96.67% accuracy rate for female and male subjects. The culminating knowledge establishes a working model for the prediction of donor class characteristics using human scent hand odor profiles.

59 BASIC BIOLOGICAL SCIENCES↗

A Predictive Model for Lake Chad Total Surface Water Area Using Remotely Sensed and Modeled Hydrological and Meteorological Parameters and Multivariate Regression Analysis

Lake Chad is an endorheic lake in the Sahel region of Africa at the southern edge of the Sahara Desert. The lake, which is well known for its dramatic decrease in surface area during the 1970s and 1980s, experiences an annual flood resulting in a maximum total surface water area generally during February or March, though sometimes earlier or later. People along the shores of Lake Chad make their living fishing, farming, and raising livestock and have a vested interest in knowing when and how extensive the annual flooding will be, particularly those practicing recession farming in which the fertile ground of previously flooded area is used for planting crops. In this study, the authors investigate the relationship between lake and basin parameters, including rainfall, basin evapotranspiration, lake evapotranspiration, lake elevation, total surface water area, and the previous year’s total surface water area, and develop equations for each dry season month (except November) linking total surface water area to the other parameters. The resulting equations allow the user to estimate the December average monthly total surface water area of the lake in late November, and to make the estimates for January to May in early December. Based on the results of a Leave One Out Cross Validation analysis, the equations for lake area are estimated to have an average absolute error ranging from 5.3 percent (for February estimates) to 7.6 percent (for May estimates).

Lake Chad↗

Estuarine Sediment Deposition during Wetland Restoration: A GIS and Remote Sensing Modeling Approach

Restoration of the industrial salt flats in the San Francisco Bay, California is an ongoing wetland rehabilitation project. Remote sensing maps of suspended sediment concentration, and other GIS predictor variables were used to model sediment deposition within these recently restored ponds. Suspended sediment concentrations were calibrated to reflectance values from Landsat TM 5 and ASTER using three statistical techniques -- linear regression, multivariate regression, and an Artificial Neural Network (ANN), to map suspended sediment concentrations. Multivariate and ANN regressions using ASTER proved to be the most accurate methods, yielding r2 values of 0.88 and 0.87, respectively. Predictor variables such as sediment grain size and tidal frequency were used in the Marsh Sedimentation (MARSED) model for predicting deposition rates for three years. MARSED results for a fully restored pond show a root mean square deviation (RMSD) of 66.8 mm (<1) between modeled and field observations. This model was further applied to a pond breached in November 2010 and indicated that the recently breached pond will reach equilibrium levels after 60 months of tidal inundation.

Newcomer, Michelle↗

Assessing heterogeneity of patient and health system delay among TB in a population with internal migrants in China

Backgrounds The diagnostic delay of tuberculosis (TB) contributes to further transmission and impedes the implementation of the End TB Strategy. Therefore, we aimed to describe the characteristics of patient delay, health system delay, and total delay among TB patients in Shanghai, identify areas at high risk for delay, and explore the potential factors of long delay at individual and spatial levels. Method The study included TB patients among migrants and residents in Shanghai between January 2010 and December 2018. Patient and health system delays exceeding 14 days and total delays exceeding 28 days were defined as long delays. Time trends of long delays were evaluated by Joinpoint regression. Multivariable logistic regression analysis was employed to analyze influencing factors of long delays. Spatial analysis of delays was conducted using ArcGIS, and the hierarchical Bayesian spatial model was utilized to explore associated spatial factors. Results Overall, 61,050 TB patients were notified during the study period. Median patient, health system, and total delays were 12 days (IQR: 3–26), 9 days (IQR: 4–18), and 27 days (IQR: 15–43), respectively. Migrants, females, older adults, symptomatic visits to TB-designated facilities, and pathogen-positive were associated with longer patient delays, while pathogen-negative, active case findings and symptomatic visits to non-TB-designated facilities were associated with long health system delays (LHD). Spatial analysis revealed Chongming Island was a hotspot for patient delay, while western areas of Shanghai, with a high proportion of internal migrants and industrial parks, were at high risk for LHD. The application of rapid molecular diagnostic methods was associated with reduced health system delays. Conclusion Despite a relatively shorter diagnostic delay of TB than in the other regions in China, there was vital social-demographic and spatial heterogeneity in the occurrence of long delays in Shanghai. While the active case finding and rapid molecular diagnosis reduced the delay, novel targeted interventions are still required to address the challenges of TB diagnosis among both migrants and residents in this urban setting.

Sun, Ruoyao↗

G/SPLINES: A hybrid of Friedman's Multivariate Adaptive Regression Splines (MARS) algorithm with Holland's genetic algorithm

G/SPLINES are a hybrid of Friedman's Multivariable Adaptive Regression Splines (MARS) algorithm with Holland's Genetic Algorithm. In this hybrid, the incremental search is replaced by a genetic search. The G/SPLINE algorithm exhibits performance comparable to that of the MARS algorithm, requires fewer least squares computations, and allows significantly larger problems to be considered.

Rogers, David↗

Data and code from: Multivariate bayesian regression model for predicting disposed ash composition at U.S. coal fired power stations

This dataset contains the code and data files needed for implementation of a Multivariate Bayesian Regression model, described in Jin et al. (2025), for the historical prediction of the chemical composition of disposed coal ash at U.S. coal fired power plants as a function of annualized coal purchase data. The integrated coal supply data file (CoalSupplyDataset.csv) represents a compilation of monthly fuel purchase records for the period 1973-2022 at major U.S. power stations. These records were obtained from the U.S. Energy Information Administration. The CSV file also contains, for each coal purchase record, the coal region of the mine as defined by the U.S. Geological Survey. Data entry errors and data gaps in the EIA records were corrected as described in Jin et al. This CSV file represents the integrated coal supply data after corrections were made. The model structure and fitting parameters are encoded in pickle file format (Bayesian.pkl). The model was developed with the coal supply data and coal ash composition data, apportioned according to the Stratified Shuffle Split for training and testing subsets. The model was built using Python and the PyMC library. Reference Publication: Jin, Z.; Huang, J.; Hower, J.C.; Hsu-Kim, H.(2025). Predictive Assessment of the Chemical Composition of Coal Ash in Reserve at U.S. Disposal Sites. Environmental Science & Technology.

Coal ash composition↗

Multiclass Classification Using Bayesian Multivariate Adaptive Regression Splines

We present a new Bayesian model for the problem of multiclass classification. In this model, the probabilities of class membership of a given observation are determined by the mean of a latent Gaussian distribution. The mean functions of this latent distribution consist of combinations of highly flexible basis functions of the inputs: multivariate adaptive regression splines (MARS), first developed for multiple regression. We use reversible jump Markov chain Monte Carlo to make inference on the classification model, including the number of basis functions. We compare the probabilistic classification performance of our proposed approach to existing methods on simulated and benchmark data, and compare uncertainty estimates on simulated data. Our proposed method compares favorably with existing Bayesian and frequentist multiclass classification methods in out-of-sample probabilistic classification, and uncertainty estimation of these probabilistic classifications. We examine the fit of the proposed method to a data set of hurricane storm surge levels near Delaware Bay, US, and conclude that sea level rise is a key contributor to damage delivered by storm surge.

97 MATHEMATICS AND COMPUTING↗

Procedures for using signals from one sensor as substitutes for signals of another

Long-term monitoring of surface conditions may require a transfer from using data from one satellite sensor to data from a different sensor having different spectral characteristics. Two general procedures for spectral signal substitution are described in this paper, a principal-components procedure and a complete multivariate regression procedure. They are evaluated through a simulation study of five satellite sensors (MSS, TM, AVHRR, CZCS, and HRV). For illustration, they are compared to another recently described procedure for relating AVHRR and MSS signals. The multivariate regression procedure is shown to be best. TM can accurately emulate the other sensors, but they, on the other hand, have difficulty in accurately emulating its shortwave infrared bands (TM5 and TM7).

Suits, G.↗

A diagnostic analysis of the VVP single-doppler retrieval technique

A diagnostic analysis of the VVP (volume velocity processing) retrieval method is presented, with emphasis on understanding the technique as a linear, multivariate regression. Similarities and differences to the velocity-azimuth display and extended velocity-azimuth display retrieval techniques are discussed, using this framework. Conventional regression diagnostics are then employed to quantitatively determine situations in which the VVP technique is likely to fail. An algorithm for preparation and analysis of a robust VVP retrieval is developed and applied to synthetic and actual datasets with high temporal and spatial resolution. A fundamental (but quantifiable) limitation to some forms of VVP analysis is inadequate sampling dispersion in the n space of the multivariate regression, manifest as a collinearity between the basis functions of some fitted parameters. Such collinearity may be present either in the definition of these basis functions or in their realization in a given sampling configuration. This nonorthogonality may cause numerical instability, variance inflation (decrease in robustness), and increased sensitivity to bias from neglected wind components. It is shown that these effects prevent the application of VVP to small azimuthal sectors of data. The behavior of the VVP regression is further diagnosed over a wide range of sampling constraints, and reasonable sector limits are established.

Boccippio, Dennis J.↗