Search NASA⌕ Search

SEARCH · Search NASA

Results for “regression models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Predicting Initial Trans-Membrane Pressure for Optimized Operations in UF Unit Using Random Forest

With the growing scarcity of freshwater, innovative process design mechanisms like Reverse Osmosis (RO) are increasingly gaining attention among water treatment utilities to address the rising demand. Ensuring reliable water production necessitates efficient resource utilization, minimizing downtime in (ultra-filtration) UF systems. Recent advancements in machine learning (ML) have enabled the development of accurate data-driven models for Model Predictive Control (MPC), often requiring minimal prior knowledge of underlying physical processes. In this study, we present predictive regression models based on Random Forest (RF) and Auto-Regressive (AR) approaches to forecast the initial Trans-Membrane Pressure (TMP) for each filtration cycle in data generated by Direct Potable Reuse (DPR) systems. The proposed RF-based model demonstrates superior performance compared to baseline methods, including historical mean, Last Observation Carried Forward (LOCF), and naïve AR models, across various forecasting horizons in terms of root mean square error (RMSE) metric. To evaluate how different classes of process variables contribute to TMP dynamics over time, we examine the feature importance of independent covariates across multiple forecast horizons. This analysis provides insight into the temporal relevance of operational and sensor-derived features, guiding control and monitoring strategies. Additionally, the impact of hyperparameter tuning on TMP prediction performance is studied for both direct and recursive RF modelling approaches across increasing forecast horizons. Accurate prediction of initial TMP is critical for optimizing RO operations, as it enables the development of robust modelling frameworks by accurately estimating membrane fouling trends, thereby enhancing process efficiency and long-term reliability. The demonstrated efficacy of the RF-based approach highlights its potential as a tool for real-time decision-making in water treatment systems, paving the way for advanced process optimization and sustainable water resource management.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Predicting Initial Trans-Membrane Pressure for Optimized Operations in UF Unit Using Random Forest

With the growing scarcity of freshwater, innovative process design mechanisms like Ultra-filtration(UF) units are increasingly gaining attention among water treatment utilities to address the rising demand. Ensuring reliable water production necessitates efficient resource utilization, minimizing downtime in UF systems. Recent advancements in machine learning (ML) have enabled the development of accurate data-driven models for Model Predictive Control (MPC), often requiring minimal prior knowledge of underlying physical processes. In this study, we present predictive regression models based on Random Forest (RF) and Auto-Regressive (AR) approaches to forecast the initial Trans-Membrane Pressure (TMP) for each filtration cycle in data generated by Direct Potable Reuse (DPR) systems. The proposed RF-based model demonstrates superior performance compared to baseline methods, including historical mean, Last Observation Carried Forward (LOCF), and naïve AR models, across various forecasting horizons in terms of root mean square (RMSE) metric. Accurate prediction of initial TMP is critical for optimizing CCRO operations, as it enables the development of robust modelling frameworks that enhance process efficiency and reliability. The demonstrated efficacy of the RF-based approach highlights its potential as a tool for real-time decision-making in water treatment systems, paving the way for advanced process optimization and sustainable water resource management.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Estimation of Kalman filter gain from output residuals

This paper presents a procedure for estimating the Kalman filter gain from output residuals. The system state space model is assumed to be known, but the process and noise covariance are unknown. The proposed procedure consists of three basic steps. First, the output residuals are computed from the given model and a given set of input-output data. Second, a linear regression model for this part of the response is computed by a least squares solution. Third, the Kalman filter gain is then estimated from the coefficients of this model. Numerical results using experimental data are presented to illustrate the validity of the developed procedure.

Juang, Jer-Nan↗

Hypersonic Wind Tunnel Calibration Using the Modern Design of Experiments

A calibration of a hypersonic wind tunnel has been conducted using formal experiment design techniques and response surface modeling. Data from a compact, highly efficient experiment was used to create a regression model of the pitot pressure as a function of the facility operating conditions as well as the longitudinal location within the test section. The new calibration utilized far fewer design points than prior experiments, but covered a wider range of the facility s operating envelope while revealing interactions between factors not captured in previous calibrations. A series of points chosen randomly within the design space was used to verify the accuracy of the response model. The development of the experiment design is discussed along with tactics used in the execution of the experiment to defend against systematic variation in the results. Trends in the data are illustrated, and comparisons are made to earlier findings.

Rhode, Matthew N.↗

Argentina Food Security & Agriculture: Crop Monitoring and Forecasting for Argentina using NASA Satellite Observations

Early harvest information helps guide agricultural commodity assessments in Argentina, providing valuable planning information to identify potentially food-insecure regions, anticipate transportation and storage demands, predict price fluctuations, and project commodity trends. However, crop yield estimates are currently subjective, based on interviews with qualified informants (i.e., farmers, agribusiness actors). In partnership with the Buenos Aires Grain Exchange, we leveraged Terra Moderate Resolution Imaging Spectroradiometer (MODIS), Soil Moisture Active Passive (SMAP), and Integrated Multi-satellitE Retrievals for Global Precipitation Measurement (GPM IMERG) NASA Earth observations to develop a Google Earth Engine (GEE) toolset to monitor vegetation growth. The first component of the toolset produces spatial and temporal maps of temperature, precipitation, soil moisture, and the Normalized Difference Vegetation Index (NDVI), allowing users to visualize the influence of the region’s climate and weather. Next, we developed an autoregressive model to predict NDVI several months in advance. Lastly, we created a linear regression model of crop yield and NDVI for soybeans, corn, and wheat, and input the forecasted NDVI to generate a predicted crop yield output. The NDVI forecasting model produced accurate predictions at two, four, and six months when examining the most recent growing season. In the crop yield model, soybeans exhibited moderately strong correlation, wheat had consistent weak correlation, and corn varied from weak to strong correlation depending on zone. This information is vital for vegetation growth monitoring by identifying areas of high growth and allocating resources to areas of lower growth to efficiently maximize crop yields.

DEVELOP Tech Paper↗

An empirical model for ocean radar backscatter and its application in inversion routine to eliminate wind speed and direction effects

Several regression models were tested to explain the wind direction dependence of the 1975 JONSWAP (Joint North Sea Wave Project) scatterometer data. The models consider the radar backscatter as a harmonic function of wind direction. The constant term accounts for the major effect of wind speed and the sinusoidal terms for the effects of direction. The fundamental accounts for the difference in upwind and downwind returns, while the second harmonic explains the upwind-crosswind difference. It is shown that a second harmonic model appears to adequately explain the angular variation. A simple inversion technique, which uses two orthogonal scattering measurements, is also described which eliminates the effect of wind speed and direction. Vertical polarization was shown to be more effective in determining both wind speed and direction than horizontal polarization.

Dome, G. J.↗

Aircraft Anomaly Detection Using Performance Models Trained on Fleet Data

This paper describes an application of data mining technology called Distributed Fleet Monitoring (DFM) to Flight Operational Quality Assurance (FOQA) data collected from a fleet of commercial aircraft. DFM transforms the data into aircraft performance models, flight-to-flight trends, and individual flight anomalies by fitting a multi-level regression model to the data. The model represents aircraft flight performance and takes into account fixed effects: flight-to-flight and vehicle-to-vehicle variability. The regression parameters include aerodynamic coefficients and other aircraft performance parameters that are usually identified by aircraft manufacturers in flight tests. Using DFM, the multi-terabyte FOQA data set with half-million flights was processed in a few hours. The anomalies found include wrong values of competed variables, (e.g., aircraft weight), sensor failures and baises, failures, biases, and trends in flight actuators. These anomalies were missed by the existing airline monitoring of FOQA data exceedances.

Gorinevsky, Dimitry↗

Modeling approaches to estimate community annoyance due to sonic booms using data from repeated surveys

In an on-going project for NASA, we are developing a modeling approach to analyze the relationship between the noise level of sonic booms from an experimental supersonic plane and the level of annoyance measured through community response surveys. The goal of the project is to obtain a quantitative relationship between noise level and annoyance that is representative for the affected population. Particular modeling challenges include multiple annoyance measurements per survey respondent and very few occurrences of annoyance overall. To address these challenges, we propose a two-stage model for the presence of high annoyance, with the first stage modeling the probability that a respondent is ever highly annoyed and the second stage a multilevel logistic regression model for high annoyance based on noise level and demographic characteristics. We use a variation of Multilevel Regression and Poststratification (Gelman and Little, 1997) to obtain an overall representative noise-annoyance curve for the population. The approach is applied to data from a NASA pilot study.

Robyn Ferg↗

Modeling Approaches to Estimate Community Annoyance Due to Sonic Booms Using Data from Repeated Surveys

In an on-going project for NASA, we are developing a modeling approach to analyze the relationship between the noise level of sonic booms from an experimental supersonic plane and the level of annoyance measured through community response surveys. The goal of the project is to obtain a quantitative relationship between noise level and annoyance that is representative for the affected population. Particular modeling challenges include multiple annoyance measurements per survey respondent and very few occurrences of annoyance overall. To address these challenges, we propose a two-stage model for the presence of high annoyance, with the first stage modeling the probability that a respondent is ever highly annoyed and the second stage a multilevel logistic regression model for high annoyance based on noise level and demographic characteristics. We use a variation of Multilevel Regression and Poststratification (Gelman and Little, 1997) to obtain an overall representative noise-annoyance curve for the population. The approach is applied to data from a NASA pilot study.

Robyn Ferg↗

A data-driven approach to real-time vertical position estimation for NSTX-U vertical stability control

In this paper, a database of 77 996 plasma equilibrium reconstructions from 727 discharges during the initial operation of the NSTX-U spherical tokamak is analyzed to develop a statistically robust model of the plasma vertical position for real-time control. A variety of regression models are developed and tested, ranging in complexity from linear models to deep neural networks, and including input signals ranging from the four pairs of flux loops used historically on NSTX-U up to the full set of 389 real-time signals available to the plasma control system. A linear model based on 140 real-time magnetics signals is found to offer excellent accuracy, with a coefficient of determination R 2 = 0.906. The robustness of this model to limited training data, new operating scenarios, and signal errors is tested, and a procedure is demonstrated to tune the model parameters to optimize its robustness. A time-dependent plasma equilibrium solver, TokaMaker, is used to simulate vertical stability control in NSTX-U, demonstrating that it should be possible to iteratively tune the parameters of a linear vertical position model to stabilize both positive and negative triangularity plasmas in future experiments.

magnetic diagnostics↗

Heuristic Area Cost Estimation for Observational Coverage Schedulers

This paper presents a comparison of heuris- tics used to estimate the amount of time it would take for a spacecraft to image an area using Boustrophedon decomposition (Choset and Pignon 1998). Machine learning tech- niques are used to characterize algorithmic performance of coverage algorithms. It is shown that an ordinary least-squares linear model is among the most accurate in a set of constant and linear order regression models both in terms of memory consumption and schedule duration. These are demonstrated using the ASPEN planning system (Fukunaga et al. 1997) on the Eagle Eye domain.

Knight, Russell↗

MSFC solar activity predictions for satellite orbital lifetime estimation

The procedure to predict solar activity indexes for use in upper atmosphere density models is given together with an example of the performance. The prediction procedure employs a least square linear regression model to generate the predicted smoothed vinculum R sub 13 and geomagnetic vinculum A sub p(13) values. Linear regression equations are then employed to compute corresponding vinculum F sub 10.7(13) solar flux values from the predicted vinculum R sub 13 values. The output is issued principally for satellite orbital lifetime estimations.

Fuler, H. C.↗

Online LIBS–ML Framework for Dynamic Characterization of Heterogeneous Waste-Derived Gasification Feedstocks

LIBS−ML framework for real time feedstock characterization during continuous conveyor transport Heterogeneous waste derived feedstocks (e.g., waste coal, biomass and blends) introduce rapid variability in heating value and ash chemistry that affect gasifier operation, yet conventional laboratory characterization techniques are too slow to support proactive control. To address this gap, this study reports on an online, in situ, dynamic characterization framework that couple’s laser-induced breakdown spectroscopy (LIBS) with leakage safe machine learning (ML) regression to deliver real time, decision quality predictions of gasifier relevant properties. A controlled sample matrix spanning two different waste coals, two different biomasses, and engineered blends under two particle size conditions were constructed and benchmarked using standardized laboratory analyses for proximate/ultimate properties and ash composition. LIBS spectra were acquired dynamically as material flowed on a conveyor belt, using high energy 1064 nm laser ablation and shot averaging to improve repeatability and precision. Supervised regression models (multi layer perceptron (MLP) /artificial neural network (ANN), random forest (RF), and support vector regression (SVR)) and an optimized weighted ensemble were trained on emission line feature sets using nested cross validation with Bayesian hyperparameter tuning and validated against an independent hold out set. The proposed LIBS−ML workflow achieves near laboratory predictive fidelity across parametric targets (including higher heating value (HHV), ash content, fixed carbon, sulfur, major ash forming oxides, and initial deformation temperature (IDT)), with the weighted ensemble providing a robust default predictor under dynamic measurement conditions. These results demonstrate a practical pathway for real time feedstock characterization that can enable feedforward adjustments and more resilient gasifier operation for variable quality waste derived fuels.

Biomass↗

Evaluation of spatial, radiometric and spectral Thematic Mapper performance for coastal studies

The effect different wetland plant canopies have upon observed reflectance in Thematic Mapper bands is studied. The three major vegetation canopy types (broadleaf, gramineous and leafless) produce unique spectral responses for a similar quantity of live biomass. The spectral biomass estimate of a broadleaf canopy is most similar to the harvest biomass estimate when a broadleaf canopy radiance model is used. All major wetland vegetation species can be identified through TM imagery. Simple regression models are developed equating the vegetation index and the infrared index with biomass. The spectral radiance index largely agreed with harvest biomass estimates.

Klemas, V.↗

CDF and PDF Comparison Between Humacao, Puerto Rico and Florida

The knowledge of the atmospherics phenomenon is an important part in the communication system. The principal factor that contributes to the attenuation in a Ka band communication system is the rain attenuation. We have four years of tropical region observations. The data in the tropical region was taken in Humacao, Puerto Rico. Previous data had been collected at various climate regions such as desserts, template area and sub-tropical regions. Figure 1 shows the ITU-R rain zone map for North America. Rain rates are important to the rain attenuation prediction models. The models that predict attenuation generally are of two different kinds. The first one is the regression models. By using a data set these models provide an idea of the observed attenuation and rain rates distribution in the present, past and future. The second kinds of models are physical models which use the probability density functions (PDF).

Gonzalez-Rodriguez, Rosana↗

Comparison of Likelihood Methods for Generalized Linear Mixed Models with Application to Quiet Supersonic Flights 2018 Data

Repeated measurement will be a feature of the survey data collected during the Quesst missionX-59 community response tests (CRT). Since each participant will report his or her categorical level of annoyance in response to multiple events, the responses from any single individual may be correlated with one another. Several models within the class of generalized linear mixed models (GLMM) are pertinent to the analysis of correlated categorical outcomes; the random intercept logistic regression model is one example. Both Bayesian and frequentist methods for fitting these models are available, with frequentist methods relying on some form of approximation (of either an integral or the integrand) that appears in the marginal likelihood function. Given several anticipated similarities of the X-59 CRT data to data collected during a past risk reduction, Quiet Supersonic Flights 2018 (QSF18), this short note is intended to create awareness. It documents an instance in which a reported population average dose-response relationship derived from QSF18 single event data was distorted by the integral approximation applied in likelihood-based methods. We review some of the available literature on the topic, compare the outputs of several different computational approaches implemented in available statistical software, and present simple corrective actions that may be useful during the Quesst mission.

dose-response model↗

Multi-Axis Identifiability Using Single-Surface Parameter Estimation Maneuvers on the X-48B Blended Wing Body

The problem of parameter estimation on hybrid-wing-body type aircraft is complicated by the fact that many design candidates for such aircraft involve a large number of aero- dynamic control effectors that act in coplanar motion. This fact adds to the complexity already present in the parameter estimation problem for any aircraft with a closed-loop control system. Decorrelation of system inputs must be performed in order to ascertain individual surface derivatives with any sort of mathematical confidence. Non-standard control surface configurations, such as clamshell surfaces and drag-rudder modes, further complicate the modeling task. In this paper, asymmetric, single-surface maneuvers are used to excite multiple axes of aircraft motion simultaneously. Time history reconstructions of the moment coefficients computed by the solved regression models are then compared to each other in order to assess relative model accuracy. The reduced flight-test time required for inner surface parameter estimation using multi-axis methods was found to come at the cost of slightly reduced accuracy and statistical confidence for linear regression methods. Since the multi-axis maneuvers captured parameter estimates similar to both longitudinal and lateral-directional maneuvers combined, the number of test points required for the inner, aileron-like surfaces could in theory have been reduced by 50%. While trends were similar, however, individual parameters as estimated by a multi-axis model were typically different by an average absolute difference of roughly 15-20%, with decreased statistical significance, than those estimated by a single-axis model. The multi-axis model exhibited an increase in overall fit error of roughly 1-5% for the linear regression estimates with respect to the single-axis model, when applied to flight data designed for each, respectively.

Ratnayake, Nalin A.↗

The impact of kidney function on Alzheimer’s disease blood biomarkers: implications for predicting amyloid-β positivity

Impaired kidney function has a potential confounding effect on blood biomarker levels, including biomarkers for Alzheimer’s disease (AD). Given the imminent use of certain blood biomarkers in the routine diagnostic work-up of patients with suspected AD, knowledge on the potential impact of comorbidities on the utility of blood biomarkers is important. We aimed to evaluate the association between kidney function, assessed through estimated glomerular filtration rate (eGFR) calculated from plasma creatinine and AD blood biomarkers, as well as their influence over predicting Aβ-positivity. We included 242 participants from the Translational Biomarkers in Aging and Dementia (TRIAD) cohort, comprising cognitively unimpaired individuals (CU; n = 124), mild cognitive impairment (MCI; n = 58), AD dementia (n = 34), and non-AD dementia (n = 26) patients all characterized by [ 18 F] AZD-4694. Plasma samples were analyzed for Aβ42, Aβ40, glial fibrillary acidic protein (GFAP), neurofilament light chain (NfL), tau phosphorylated at threonine 181 (p-tau181), 217 (p-tau217), 231 (p-tau231) and N-terminal containing tau fragments (NTA-tau) using Simoa technology. Kidney function was assessed by eGFR in mL/min/1.73 m 2 , based on plasma creatinine levels, age, and sex. Participants were also stratified according to their eGFR-indexed stages of chronic kidney disease (CKD). We evaluated the association between eGFR and blood biomarker levels with linear models and assessed whether eGFR provided added predictive value to determine Aβ-positivity with logistic regression models. Biomarker concentrations were highest in individuals with CKD stage 3, followed by stages 2 and 1, but differences were only significant for NfL, Aβ42, and Aβ40 (not Aβ42/Aβ40). All investigated biomarkers showed significant associations with eGFR except plasma NTA-tau, with stronger relationships observed for Aβ40 and NfL. However, after adjusting for either age, sex or Aβ-PET SUVr, the association with eGFR was no longer significant for all biomarkers except Aβ40, Aβ42, NfL, and GFAP. When evaluating whether accounting for kidney function could lead to improved prediction of Aβ-positivity, we observed no improvements in model fit (Akaike Information Criterion, AIC) or in discriminative performance (AUC) by adding eGFR to a base model including each plasma biomarker, age, and sex. While covariates like age and sex improved model fit, eGFR contributed minimally, and there were no significant differences in clinical discrimination based on AUC values. We found that kidney function seems to be associated with AD blood biomarker concentrations. However, these associations did not remain significant after adjusting for age and sex, except for Aβ40, Aβ42, NfL, and GFAP. While covariates such as age and sex improved prediction of Aβ-positivity, including eGFR in the models did not lead to improved prediction for any biomarker. Our findings indicate that renal function, within the normal to mild impairment range, does not seem to have a clinically relevant impact when using highly accurate blood biomarkers, such as p-tau217, in a biomarker-supported diagnosis.

60 APPLIED LIFE SCIENCES↗