Search NASA⌕ Search

SEARCH · Search NASA

Results for “logistic regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Logistic Regression in Clinical Studies

• A logistic regression model is used when the outcome of interest is binary. The term “logistic” refers to the underlying “logit” (log odds) function that is used to model the binary outcome. • Odds ratios are produced from a logistic regression model and have a useful interpretation. • Tips, tricks and concepts used to fit logistic regression models are similar to those used in linear regression models. • Modeling building that is knowledge-based rather than automatic is preferred in most applications of logistic regression. • A logistic regression model that is overparameterized (ie, too many variables for too few events) can result in odds ratios that are implausibly large and confidence intervals that are wide and uninterpretable. These types of “overfitted” models should be avoided. • Logistic regression models can be fit using most standard statistical software.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

A Provably Accurate Randomized Sampling Algorithm for Logistic Regression

In statistics and machine learning, logistic regression is a widely-used supervised learning technique primarily employed for binary classification tasks. When the number of observations greatly exceeds the number of predictor variables, we present a simple, randomized sampling-based algorithm for logistic regression problem that guarantees high-quality approximations to both the estimated probabilities and the overall discrepancy of the model. Our analysis builds upon two simple structural conditions that boil down to randomized matrix multiplication, a fundamental and well-understood primitive of randomized numerical linear algebra. We analyze the properties of estimated probabilities of logistic regression when leverage scores are used to sample observations, and prove that accurate approximations can be achieved with a sample whose size is much smaller than the total number of observations. To further validate our theoretical findings, we conduct comprehensive empirical evaluations. Overall, our work sheds light on the potential of using randomized sampling approaches to efficiently approximate the estimated probabilities in logistic regression, offering a practical and computationally efficient solution for large-scale datasets.

Chowdhury, Agniva↗

Probabilistic Power Consumption Modeling for Commercial Buildings Using Logistic Regression Markov Chain

The total energy consumed by buildings takes up to 40% of U.S. energy use, in which a large portion is contributed by commercial buildings. Building performance optimization is desirable but requires accurate building models with uncertainties taken into account. This paper proposes a novel probabilistic modeling method using Logistic Regression Markov Chain (LRMC). The LRMC model enhances the performance of traditional Markov Chain (MC) models by adopting time-variant transition matrices calibrated using logistic regression with exogenous inputs. Compared with existing building models, the proposed model produces accurate multi-step modeling results with full probability distribution. The proposed probabilistic building model is tested using actual commercial building measurements and modeling performance is evaluated with two probabilisitc metrics. The results show that the LRMC model has higher accuracy than traditional MC model and Logistic Regression (LR) model in that it yields lower error scores under both evaluation metrics.

Building modeling↗

Dynamic logistic regression and variable selection: Forecasting and contextualizing civil unrest

Civil unrest can range from peaceful protest to violent furor, and researchers are working to monitor, forecast, and assess such events to allocate resources better. Twitter has become a real-time data source for forecasting civil unrest because millions of people use the platform as a social outlet. Additionally, daily word counts are used as model features, and predictive terms contextualize the reasons for the protest. To forecast civil unrest and infer the reasons for the protest, we consider the problem of Bayesian variable selection for the dynamic logistic regression model and propose using penalized credible regions to select parameters of the updated state vector. This method avoids the need for shrinkage priors, is scalable to high-dimensional dynamic data, and allows the importance of variables to vary in time as new information becomes available. A substantial improvement in both precision and F1-score using this approach is demonstrated through simulation. Finally, we apply the proposed model fitting and variable selection methodology to the problem of forecasting civil unrest in Latin America. Our dynamic logistic regression approach shows improved accuracy compared to the static approach currently used in event prediction and feature selection.

97 MATHEMATICS AND COMPUTING↗

Predicting fatigue from heart rate signatures using functional logistic regression

Physical fatigue can have adverse effects on humans in extreme environments. Therefore, being able to predict fatigue using easy to measure metrics such as heart rate (HR) signatures has potential to have an impact in real-life scenarios. We apply a functional logistic regression model that uses HR signatures to predict physical fatigue, where physical fatigue is defined in a data-driven manner. Data were collected using commercially available wearable devices on 47 participants hiking the 20.7-mile Grand Canyon rim-to-rim trail in a single day. Fitted model provides good predictions and interpretable parameters for real-life application.

60 APPLIED LIFE SCIENCES↗

classLog: Logistic regression for the classification of genetic sequences

Introduction Sequencing and phylogenetic classification have become a common task in human and animal diagnostic laboratories. It is routine to sequence pathogens to identify genetic variations of diagnostic significance and to use these data in realtime genomic contact tracing and surveillance. Under this paradigm, unprecedented volumes of data are generated that require rapid analysis to provide meaningful inference. Methods We present a machine learning logistic regression pipeline that can assign classifications to genetic sequence data. The pipeline implements an intuitive and customizable approach to developing a trained prediction model that runs in linear time complexity, generating accurate output rapidly, even with incomplete data. Our approach was benchmarked against porcine respiratory and reproductive syndrome virus (PRRSv) and swine H1 influenza A virus (IAV) datasets. Trained classifiers were tested against sequences and simulated datasets that artificially degraded sequence quality at 0, 10, 20, 30, and 40%. Results When applied to a poor-quality sequence data, the classifier achieved between >85% to 95% accuracy for the PRRSv and the swine H1 IAV HA dataset and this increased to near perfect accuracy when using the full dataset. The model also identifies amino acid positions used to determine genetic clade identity through a feature selection ranking within the model. These positions can be mapped onto a maximum-likelihood phylogenetic tree, allowing for the inference of clade defining mutations. Discussion Our approach is implemented as a python package with code available at https://github.com/flu-crew/classLog .

Zeller, Michael A.↗

Assessing heterogeneity of patient and health system delay among TB in a population with internal migrants in China

Backgrounds The diagnostic delay of tuberculosis (TB) contributes to further transmission and impedes the implementation of the End TB Strategy. Therefore, we aimed to describe the characteristics of patient delay, health system delay, and total delay among TB patients in Shanghai, identify areas at high risk for delay, and explore the potential factors of long delay at individual and spatial levels. Method The study included TB patients among migrants and residents in Shanghai between January 2010 and December 2018. Patient and health system delays exceeding 14 days and total delays exceeding 28 days were defined as long delays. Time trends of long delays were evaluated by Joinpoint regression. Multivariable logistic regression analysis was employed to analyze influencing factors of long delays. Spatial analysis of delays was conducted using ArcGIS, and the hierarchical Bayesian spatial model was utilized to explore associated spatial factors. Results Overall, 61,050 TB patients were notified during the study period. Median patient, health system, and total delays were 12 days (IQR: 3–26), 9 days (IQR: 4–18), and 27 days (IQR: 15–43), respectively. Migrants, females, older adults, symptomatic visits to TB-designated facilities, and pathogen-positive were associated with longer patient delays, while pathogen-negative, active case findings and symptomatic visits to non-TB-designated facilities were associated with long health system delays (LHD). Spatial analysis revealed Chongming Island was a hotspot for patient delay, while western areas of Shanghai, with a high proportion of internal migrants and industrial parks, were at high risk for LHD. The application of rapid molecular diagnostic methods was associated with reduced health system delays. Conclusion Despite a relatively shorter diagnostic delay of TB than in the other regions in China, there was vital social-demographic and spatial heterogeneity in the occurrence of long delays in Shanghai. While the active case finding and rapid molecular diagnosis reduced the delay, novel targeted interventions are still required to address the challenges of TB diagnosis among both migrants and residents in this urban setting.

Sun, Ruoyao↗

Pregnancy outcome after first trimester exposure to domperidone—An observational cohort study

Abstract Aim To assess the teratogenic risk of domperidone by comparing the incidence of major malformation with domperidone to a control. Methods Pregnancy outcome data were obtained for women at two Japanese facilities that provide counseling on drug use during pregnancy between April 1988 and December 2017. The incidence of major malformation was calculated among infants born to women taking domperidone ( n = 519), nonteratogenic drugs (control, n = 1673), or metoclopramide (reference, n = 241) during the first trimester of pregnancy. Using the control group as reference, the crude odds ratio (OR) of the incidence of major malformation in the domperidone and metoclopramide groups was calculated using univariable logistic regression analysis. Adjusted OR was also calculated using multivariable logistic regression analysis adjusted for various other factors. Results The incidence of major malformation was 2.9% (14/485, 95% confidence interval [CI]: 1.6–4.8) in the domperidone group, 1.7% (27/1554, 95%CI: 1.1–2.5) in the control group, and 3.6% (8/224, 95%CI: 1.6–6.9) in the metoclopramide group. The adjusted multivariable logistic regression analysis showed no significant difference in incidence between the control and domperidone groups (adjusted OR: 1.86 [95%CI: 0.73–4.70], p = 0.191) or between the control and metoclopramide groups (adjusted OR: 2.20 [95%CI: 0.69–6.98], p = 0.183). Conclusions This observational cohort study showed that domperidone exposure during the first trimester was not associated with increased risk of major malformation in infants. These results may help alleviate the anxiety of patients who took domperidone during pregnancy.

Hishinuma, Kayoko↗

When less is more: How increasing the complexity of machine learning strategies for geothermal energy assessments may not lead toward better estimates

Previous moderate- and high-temperature geothermal resource assessments of the western United States utilized data-driven methods and expert decisions to estimate resource favorability. Although expert decisions can add confidence to the modeling process by ensuring reasonable models are employed, expert decisions also introduce human and, thereby, model bias. This bias can present a source of error that reduces the predictive performance of the models and confidence in the resulting resource estimates. Our study aims to develop robust data-driven methods with the goals of reducing bias and improving predictive ability. We present and compare nine favorability maps for geothermal resources in the western United States using data from the U.S. Geological Survey's 2008 geothermal resource assessment. Two favorability maps are created using the expert decision-dependent methods from the 2008 assessment (i.e., weight-of-evidence and logistic regression). With the same data, we then create six different favorability maps using logistic regression (without underlying expert decisions), XGBoost, and support-vector machines paired with two training strategies. The training strategies are customized to address the inherent challenges of applying machine learning to the geothermal training data, which have no negative examples and severe class imbalance. We also create another favorability map using an artificial neural network. We demonstrate that modern machine learning approaches can improve upon systems built with expert decisions. We also find that XGBoost, a non-linear algorithm, produces greater agreement with the 2008 results than linear logistic regression without expert decisions, because the expert decisions in the 2008 assessment rendered the otherwise linear approaches non-linear despite the fact that the 2008 assessment used only linear methods. The F1 scores for all approaches appear low (F1 score < 0.10), do not improve with increasing model complexity, and, therefore, indicate the fundamental limitations of the input features (i.e., training data). Until improved feature data are incorporated into the assessment process, simple non-linear algorithms (e.g., XGBoost) perform equally well or better than more complex methods (e.g., artificial neural networks) and remain easier to interpret.

15 GEOTHERMAL ENERGY↗

Novel Application of Machine Learning Techniques for Rapid Source Apportionment of Aerosol Mass Spectrometer Datasets

In this work, we apply machine learning approaches sparse multinomial logistic regression to classify aerosol mass spectrometer (AMS) unit mass resolution (UMR) data followed by an ensemble regression technique for source apportionment of organic aerosols (OA). The classifier was trained on 60 well characterized laboratory and positive matrix factorization (PMF) deconvolved reference spectra to identify eight OA types. These include four laboratory-derived secondary organic aerosol (SOA) spectra, which include isoprene photooxidation SOA, isoprene epoxydiols (IEPOX) SOA, a monoterpene SOA type that includes a-pinene and ß-pinene SOA, and aromatic SOA from oxidation of naphthalene and m-xylene precursors, as well as PMF deconvolved spectra for three primary organic aerosol (POA) types, namely, hydrocarbon-like organic aerosol (HOA), biomass burning organic aerosol (BBOA), and cooking OA (COA), and a more oxidized oxygenated OA type (MO-OOA). A 5-fold cross-validation strategy, repeated 10 times, was used to assess the classifier’s performance. The classifier had high classification accuracy for COA, aromatic SOA, and isoprene SOA spectra but incorrectly classified ~9% by number of MO-OOA spectra as BBOA, 12% of BBOA spectra as HOA (and vice versa), and 18% of IEPOX-SOA spectra as aromatic SOA. Next, an ensemble regression model was trained on an artificially generated dataset consisting of mixtures of different OA types to assess its ability to predict fractional mass abundances from classification probabilities of various OA species obtained from the multinomial logistic regression classifier trained on the reference spectra. Ultimately, the proposed approach was applied for source apportionment of aircraft-based AMS measurements of OA UMR spectra during the HI-SCALE field campaign. On two representative days (May 6th and 18th, 2016), the algorithm determined that ~50-60% of OA by mass was MO-OOA, which represented a highly aged organic aerosol mixture from different sources. On both days, BBOA was determined to contribute less than 10% to OA by mass. However, on May 18th, the aromatic SOA fraction was higher compared to that on May 6th. The proposed approach is capable of rapidly analyzing AMS data in real time, making it suitable for applications where rapid source apportionment of AMS OA spectra is desirable.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Knowledge of lactation amenorrhea method among postpartum women in Ethiopia: a facility-based cross-sectional study

While the importance of knowledge about contraceptives in improving their utilization and thereby reducing the risk of unintended pregnancies is well documented, there are limited studies documented about the Lactational Amenorrhea Method (LAM). Thus, understanding the knowledge of postpartum mothers about LAM is essential for designing tailored interventions. This study assessed the level of knowledge about LAM and its associated factors among postpartum mothers in Ethiopia. A facility-based cross-sectional study was conducted among 3148 randomly selected postpartum participants. The study utilized multistage sampling approach in hospitals located across five regions and one city administration in Ethiopia. Data were collected using face-to-face interviews at discharge. A participant was categorized as having knowledge of LAM if she correctly answered the three LAM criteria: amenorrhea, the first 6 months, and exclusive breast feeding. A binary logistic regression model was used to identify factors associated with knowledge of LAM. Variables with p < 0.25 in the binary logistic regression were included in the multiple logistic regression. Then, associations were described using the adjusted odds ratio (AOR) along with the 95% confidence interval (CI), and statistical significance was declared at p < 0.05. Only four in 10 participants (40.6%; 95% CI 38.9–42.3) had knowledge of LAM. Participants who attended college or above educational level (AOR = 2.1, 95% CI 1.5–2.8), those with parity of two (AOR = 2.3; 95% CI 1.6–3.6) or more than two (AOR = 2.4; 95% CI 1.5–4.0), those who expressed a desire for further fertility (AOR = 1.3; 95% CI 1.1–1.5), individuals who received counselling on LAM (AOR = 3.0; 95% CI 2.6–3.7), and those who gave birth in hospital (AOR = 2.6; 95% CI 1.4–2.6) had higher odds of knowledge about LAM, compared to their counter parts. In contrary, participants resided far away from health facilities had 30% lower odd of knowledge about LAM compared to those resided near the health facilities (AOR = 0.70; 95% CI 0.6–0.8). The proportion of participants who had knowledge of LAM was low. Strengthening counseling about LAM during antenatal care and delivery with due attention to women with limited access to health facilities should be considered for increasing their level of knowledge on LAM.

60 APPLIED LIFE SCIENCES↗

Predicting nutrition and environmental factors associated with female reproductive disorders using a knowledge graph and random forests

Female reproductive disorders (FRDs) are common health conditions that may present with significant symptoms. Diet and environment are potential areas for FRD interventions. We utilized a knowledge graph (KG) method to predict factors associated with common FRDs (for example, endometriosis, ovarian cyst, and uterine fibroids). We harmonized survey data from the Personalized Environment and Genes Study (PEGS) on internal and external environmental exposures and health conditions with biomedical ontology content. We merged the harmonized data and ontologies with supplemental nutrient and agricultural chemical data to create a KG. We analyzed the KG by embedding edges and applying a random forest for edge prediction to identify variables potentially associated with FRDs. We also conducted logistic regression analysis for comparison. Across 9765 PEGS respondents, the KG analysis resulted in 8535 significant or suggestive predicted links between FRDs and chemicals, phenotypes, and diseases. Amongst these links, 32 were exact matches when compared with the logistic regression results, including comorbidities, medications, foods, and occupational exposures. Mechanistic underpinnings of predicted links documented in the literature may support some of our findings. Our KG methods are useful for predicting possible associations in large, survey-based datasets with added information on directionality and magnitude of effect from logistic regression. These results should not be construed as causal but can support hypothesis generation. This investigation enabled the generation of hypotheses on a variety of potential links between FRDs and exposures. Future investigations should prospectively evaluate the variables hypothesized to impact FRDs.

60 APPLIED LIFE SCIENCES↗

Machine Learning Emulation of Spatial Deposition from a Multi-Physics Ensemble of Weather and Atmospheric Transport Models

In the event of an accidental or intentional hazardous material release in the atmosphere, researchers often run physics-based atmospheric transport and dispersion models to predict the extent and variation of the contaminant spread. These predictions are imperfect due to propagated uncertainty from atmospheric model physics (or parameterizations) and weather data initial conditions. Ensembles of simulations can be used to estimate uncertainty, but running large ensembles is often very time consuming and resource intensive, even using large supercomputers. In this paper, we present a machine-learning-based method which can be used to quickly emulate spatial deposition patterns from a multi-physics ensemble of dispersion simulations. We use a hybrid linear and logistic regression method that can predict deposition in more than 100,000 grid cells with as few as fifty training examples. Logistic regression provides probabilistic predictions of the presence or absence of hazardous materials, while linear regression predicts the quantity of hazardous materials. The coefficients of the linear regressions also open avenues of exploration regarding interpretability—the presented model can be used to find which physics schemes are most important over different spatial areas. A single regression prediction is on the order of 10,000 times faster than running a weather and dispersion simulation. However, considering the number of weather and dispersion simulations needed to train the regressions, the speed-up achieved when considering the whole ensemble is about 24 times. Ultimately, this work will allow atmospheric researchers to produce potential contamination scenarios with uncertainty estimates faster than previously possible, aiding public servants and first responders.

97 MATHEMATICS AND COMPUTING↗

MAFLD identifies patients with significant hepatic fibrosis better than NAFLD

Abstract Background & Aims Diagnostic criteria for metabolic associated fatty liver disease (MAFLD) have been proposed, but not validated. We aimed to compare the diagnostic accuracy of the MAFLD definition vs the existing NAFLD criteria to identify patients with significant fibrosis and to characterize the impact of mild alcohol intake. Methods We enrolled 765 Japanese patients with fatty liver (median age 54 years). MAFLD and NAFLD were diagnosed in 79.6% and 70.7% of patients respectively. Significant fibrosis was defined by FIB‐4 index ≥1.3 and liver stiffness ≥6.6 kPa using shear wave elastography. Mild alcohol intake was defined as <20 g/day. Factors associated with significant fibrosis were analysed by logistic regression and decision‐tree analyses. Results Liver stiffness was higher in MAFLD compared to NAFLD (7.7 vs 6.8 kPa, P = .0010). In logistic regression, MAFLD (OR 4.401; 95% CI 2.144‐10.629; P < .0001), alcohol intake (OR 1.761; 95% CI 1.081‐2.853; P = .0234), and NAFLD (OR 1.721; 95%CI 1.009‐2.951; P = .0463) were independently associated with significant fibrosis. By decision‐tree analysis, MAFLD, but not NAFLD or alcohol consumption was the initial classifier for significant fibrosis. The sensitivity for detecting significant fibrosis was higher for MAFLD than NAFLD (93.9% vs 73.0%). In patients with MAFLD, even mild alcohol intake was associated with an increase in the prevalence of significant fibrosis (25.0% vs 15.5%; P = .0181). Conclusions The MAFLD definition better identifies a group with fatty liver and significant fibrosis evaluated by non‐invasive tests. Moreover, in patients with MAFLD, even mild alcohol consumption is associated with worsening of hepatic fibrosis measures.

Yamamura, Sakura↗

Risk Ratio and Risk Difference Estimation in Case-cohort Studies

Background: In case-cohort studies with binary outcomes, ordinary logistic regression analyses have been widely used because of their computational simplicity. However, the resultant odds ratio estimates cannot be interpreted as relative risk measures unless the event rate is low. The risk ratio and risk difference are more favorable outcome measures that are directly interpreted as effect measures without the rare disease assumption. Methods: We provide pseudo-Poisson and pseudo-normal linear regression methods for estimating risk ratios and risk differences in analyses of case-cohort studies. These multivariate regression models are fitted by weighting the inverses of sampling probabilities. Also, the precisions of the risk ratio and risk difference estimators can be improved using auxiliary variable information, specifically by adapting the calibrated or estimated weights, which are readily measured on all samples from the whole cohort. Finally, we provide computational code in R (R Foundation for Statistical Computing, Vienna, Austria) that can easily perform these methods. Results: Through numerical analyses of artificially simulated data and the National Wilms Tumor Study data, accurate risk ratio and risk difference estimates were obtained using the pseudo-Poisson and pseudo-normal linear regression methods. Also, using the auxiliary variable information from the whole cohort, precisions of these estimators were markedly improved. Conclusion: The ordinary logistic regression analyses may provide uninterpretable effect measure estimates, and the risk ratio and risk difference estimation methods are effective alternative approaches for case-cohort studies. These methods are especially recommended under situations in which the event rate is not low.

60 APPLIED LIFE SCIENCES↗