Search NASA⌕ Search

SEARCH · Search NASA

Results for “generalized linear model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Analysis of overlapping count data

Counts of a specific characteristic were obtained within regions defined on an object that was manufactured in a proprietary setting. The count regions were altered during production and resulted in misaligned or overlapping count data. A closed-formula maximum likelihood estimator (MLE) of the new region means is derived using all of the available count data and an independent Poisson model. The MLE is shown to be preferable to estimators constructed using generalized linear models for the overlapping data setting. This closed-form estimator extends to over-dispersed overlapping count data as the quasi-MLE and also performs well with correlated overlapping count data. Standard errors for the estimator are approximated and are validated with a simulation study. Additionally, the methods are extended to overlapping multinomial data. Illustrative examples of the methods are provided throughout the paper and are reproducible with the supplemental R code. Additionally, proofs of the paper’s results are also included in the supplemental material.

97 MATHEMATICS AND COMPUTING↗

Estimating soil N 2 O emissions induced by organic and inorganic fertilizer inputs using a Tier-2, regression-based meta-analytic approach for U.S. agricultural lands

Consistent methods are essential for generating country and region-specific estimates of greenhouse gas (GHG) emissions used for reporting and policymaking. The estimates of direct N 2 O emissions from U.S. agricultural soils have primarily relied on the use of emission factors (EFs, Tier-1) and process-based models (Tier-3). However, Tier-1 estimates are relatively crude while Tier-3 calculations can be costly. This work addressed this gap by developing a Tier-2, regression-based approach by leveraging a meta-database containing 1883 field N 2 O observations together with environmental and management covariates from 139 studies. Our results estimated higher monthly soil N 2 O emissions (N 2 O m , kg N/ha) during the growing season (0.38) than the fallow period (0.15), highlighting the importance of considering measurement periods when utilizing meta-databases for analyzing N 2 O drivers. Significantly different N 2 O m were found for tillage practices (conventional > no-till: 0.42 > 0.27), fertilizer type (liquid > solid manure: 0.55 > 0.32), and soil texture (fine > coarse: 0.36 > 0.22). The comparisons of the influence of crop type and rotation, water management, and soil order on N 2 O emissions are complicated by regional data availability and interactions among different factors. Additionally, the finding that N 2 O emissions reported based on area (N 2 O m ), N input rate (EF), or yield can alter treatment rankings underscores the need to establish transparent criteria for rewarding or discouraging regionally-based management practices using N 2 O metrics. Finally, we show how General Linear Models (GLMs) can be used to estimate country and regional Tier-2 N 2 O m using a suite of covariates. Our GLMs identified tillage, water management, N input type and rate, soil properties, and elevation as the most influential covariates for the conterminous U.S. The limited accuracy of regional-scale GLMs, however, suggests the need to further improve the quality and availability of GHG and covariate data through concerted efforts in data collection.

54 ENVIRONMENTAL SCIENCES↗

Calibration of cloud and aerosol related parameters for solar irradiance forecasts in WRF-solar

Model parameters are a major source of uncertainty in numerical weather prediction. Recently, the Weather Research and Forecasting model with Solar extensions (WRF-Solar) has been upgraded by enhancing the treatment of sub-grid scale cloud and aerosols with augmentations of a sub-grid scale cloud scheme (CLD3) and an upgraded aerosol-aware Thompson-Eidhammer scheme (TE14). However, the value of model parameters associated with these parameterizations are assigned based on limited measurements or theoretical calculations. Calibrating the most sensitive parameters has the potential to improve solar irradiance predictions. Here, we adopted a multiobjective surrogate-based optimization (SBO) framework to calibrate nine parameters used in CLD3 and TE14 that lead to the largest sensitivity in simulated irradiance. The normalized mean-absolute-error (NMAE) of global horizontal irradiance (GHI) and direct normal irradiance (DNI) are minimized by calibrating WRF-Solar over two regions including the Southern Great Plains (SGP) and Central California, in order to focus on parameter calibration under cloudy conditions with different aerosol loading. The results show that generalized linear model (GLM)-based surrogate models approximate physical models well, particularly when the third order and three-way interaction terms are considered. The SBO framework efficiently searches the parameter space for optimal solutions with less computational costs than directly calibrating the physical model. We first calibrate CLD3 parameters over the less-polluted SGP region. Optimized CLD3 parameters alone result in NMAE reduction by 14% for the site-mean and up to 33% for individual cases over the SGP region. With further calibration of TE14 parameters over the Central California during active fire periods, the optimized parameters lead to over 20% reductions of NMAE. Our investigation reveals, however, that optimizing TE14 has a limited impact on irradiance simulations under less-polluted conditions in the SGP.

14 SOLAR ENERGY↗

Soil Origin and Plant Genotype Modulate Switchgrass Aboveground Productivity and Root Microbiome Assembly

Switchgrass (Panicum virgatum) is a model perennial grass for bioenergy production that can be productive in agricultural lands that are not suitable for food production. There is growing interest in whether its associated microbiome may be adaptive in low- or no-input cultivation systems. However, the relative impact of plant genotype and soil factors on plant microbiome and biomass are a challenge to decouple. To address this, a common garden greenhouse experiment was carried out using six common switchgrass genotypes, which were each grown in four different marginal soils collected from long-term bioenergy research sites in Michigan and Wisconsin. We characterized the fungal and bacterial root communities with high-throughput amplicon sequencing of the ITS and 16S rDNA markers, and collected phenological plant traits during plant growth, as well as soil chemical traits. At harvest, we measured the total plant aerial dry biomass. Significant differences in richness and Shannon diversity across soils but not between plant genotypes were found. Generalized linear models showed an interaction between soil and genotype for fungal richness but not for bacterial richness. Community structure was also strongly shaped by soil origin and soil origin × plant genotype interactions. Overall, plant genotype effects were significant but low. Random Forest models indicate that important factors impacting switchgrass biomass included NO 3 – , Ca 2+ , PO 4 3– , and microbial biodiversity. We identified 54 fungal and 52 bacterial predictors of plant aerial biomass, which included several operational taxonomic units belonging to Glomeraceae and Rhizobiaceae, fungal and bacterial lineages that are involved in provisioning nutrients to plants.

plant biomass↗

Chemical mixture exposure patterns and obesity among U.S. adults in NHANES 2005–2012

The effect of chemical exposure on obesity has raised great concerns. Real-world chemical exposure always imposes mixture impacts, however their exposure patterns and the corresponding associations with obesity have not been fully evaluated. To discover obesity-related mixed chemical exposure patterns in the general U.S. population. Sparse Decompositional Regression (SDR), a model adapted from sparse representation learning technique, was developed to identify exposure patterns of chemical mixtures with exclusion (non-targeted model) and inclusion (targeted model) of health outcomes. We assessed the relationships between the identified chemical mixture patterns and obesity-related indexes. We also conducted a comprehensive evaluation of this SDR model by comparing to the existing models, including generalized linear regression model (GLM), principal component analysis (PCA), and Bayesian kernel machine regression (BKMR). Eight core exposure patterns were identified using the non-targeted SDR model. Patterns of high levels of MEP, high levels of naphthalene metabolites (ΣOH-Nap), and a pattern of high exposure levels of MCOP, MCNP, and MCPP were positively associated with obesity. Patterns of high levels of BP3, and a pattern of higher mixed levels of MPB, PPB, and MEP were found to have negative associations. Associations were strengthened using the targeted SDR model. In the single chemical analysis by GLM, BP3, MBP, PPB, MCOP, and MCNP showed significant associations with obesity or body indexes. The SDR model exceeded the performance of PCA in pattern identification. Both SDR and BKMR identified a positive contribution of ΣOH-Nap and MCOP, as well as a negative contribution of BP3 and PPB to obesity. Our study identified five core exposure patterns of chemical mixtures significantly associated with obesity using the newly developed SDR model. The SDR model could open a new avenue for assessing health effects of environmental mixture contaminants.

54 ENVIRONMENTAL SCIENCES↗

Neighborhood sociome factors and pediatric asthma exacerbations: Protective role of tree crown density and importance of pharmacy access in Chicago's south side

Abstract Background Pediatric asthma exacerbations remain a critical public health concern, particularly in historically underserved urban settings. Objective This study investigates sociome factors—the social context of disease—associated with asthma exacerbations among children living in Chicago's South Side, leveraging clinical and publicly available generalizable census tract‐level datasets from agencies including ChiVes, the City of Chicago Data Portal, EPA, Census Bureau, HUD, NOAA, and more. The aim is to uncover novel hypotheses for potential new interventions. Methods A generalized linear model assessed associations with the outcome of asthma exacerbations while accounting for clustering at the patient level. Predictors included all variables from the Sociome Data Commons, including social, environmental, behavioral, economic, housing, and school variables. Results Predictors of decreased risk included patient age (+4.8 years, −22%), tree crown density (+6% coverage, −17%), parks per acre (+0.41, −8%), and labor market engagement (+0.8 points, −9%). Conversely, predictors of increased risk included increased distance to the nearest pharmacy (+0.28 miles, +12%), limited English skills (+2.3%, +10%), higher inequality (+0.08 points, +8%), and visits in the Spring (+11%) and Fall (+20%). Conclusion The results suggest that tree crown density, a novel finding in the context of asthma exacerbations, may play a protective role. Limited access to health care facilities such as pharmacies continues to complicate care. Clinical Implications These findings provide hypotheses for future interventions for long‐standing asthma disparities.

Allergy↗

A Simulator for Neyer Tests of Explosives

Explosives and explosive devices such as detonators are typically tested by applying a range of stimuli such as voltage or mechanical shock, and recording binary “detonated/did not detonate” responses. These are analyzed using maximum likelihood or generalized linear models to provide estimates of quantities such as the all-fire and no-fire points. Given that the true threshold for detonation is unknown a priori , sequential design methods are typically used to optimize the set of test points. One popular method, implemented in commercial software, is Neyer’s algorithm. To support simulation and experimental design, we have developed code in the R programming language to duplicate the functions of the Neyer software. We provide code for the simulator along with a description and examples of usage.

42 ENGINEERING↗

Seroprevalence and risk factors for brucellosis amongst livestock and humans in a multi-herd ranch system in Kagera, Tanzania

Background Brucellosis remains a significant health and economic challenge for livestock and humans globally. Despite its public health implications, the factors driving the endemic persistence ofBrucellaat the human-livestock interface in Tanzania remain poorly elucidated. This study aimed to identify the seroprevalence ofBrucellainfection in livestock and humans within a ranching system and determine associated risk factors for disease endemicity. Methods A cross-sectional sero-epidemiological study was conducted in 2023 in Tanzania’s Karagwe District, involving 725 livestock (cattle, goats, sheep) from 10 herds and 112 humans from associated camps. Seroprevalence was assessed using competitive ELISA while epidemiological data were collected via questionnaires. Generalized Linear Models and Contrast Analysis were used to identify risk factors for infection. Results Overall seroprevalence was 34% in livestock and 41% in humans. Goats exhibited the highest prevalence (69.2%), while cattle had the lowest (22.6%). Mixed-species herds (Odds Ratio, OR = 2.96, CI [1.90–4.60]) and small ruminants-only herds (OR = 6.54, CI [3.65–11.72]) showed a significantly higher risk of seropositivity compared to cattle-only herds. Older cattle (OR = 5.23, CI [2.70–10.10]) and lactating females (OR = 2.87, CI [1.78–4.63]) represented significant risks for brucellosis in livestock. In humans, close contact with animals (OR = 7.20, CI [1.97–36.31]) and handling animals during parturition or aborted fetuses (OR = 2.37, CI [1.01–5.58]) were significant risk factors. Notably, no spatial association was found in seroprevalence between herds and nearby human communities. Conclusion The lack of spatial correlation between livestock and human seroprevalence suggests complex transmission dynamics, potentially involving endemic circulation in livestock and human infections from multiple sources of exposure to livestock. This study highlights the need for comprehensive zoonotic risk education and targeted intervention strategies. Further research is crucial to elucidate transmission pathways and improveBrucellainfection control. This includes developing robust methods for identifying infective species and implementing effective strategies to mitigateBrucellainfection in endemic regions.

Public, Environmental & Occupational Health↗

Sensitivity of Optical Satellites to Estimate Windthrow Tree-Mortality in a Central Amazon Forest

Windthrow (i.e., trees broken and uprooted by wind) is a major natural disturbance in Amazon forests. Images from medium-resolution optical satellites combined with extensive field data have allowed researchers to assess patterns of windthrow tree-mortality and to monitor forest recovery over decades of succession in different regions. Although satellites with high spatial-resolution have become available in the last decade, they have not yet been employed for the quantification of windthrow tree-mortality. Here, we address how increasing the spatial resolution of satellites affects plot-to-landscape estimates of windthrow tree-mortality. We combined forest inventory data with Landsat 8 (30 m pixel), Sentinel 2 (10 m), and WorldView 2 (2 m) imagery over an old-growth forest in the Central Amazon that was disturbed by a single windthrow event in November 2015. Remote sensing estimates of windthrow tree-mortality were produced from Spectral Mixture Analysis and evaluated with forest inventory data (i.e., ground true) by using Generalized Linear Models. Field measured windthrow tree-mortality (3 transects and 30 subplots) crossing the entire disturbance gradient was 26.9 ± 11.1% (mean ± 95% CI). Although the three satellites produced reliable and statistically similar estimates (from 26.5% to 30.3%, p < 0.001), Landsat 8 had the most accurate results and efficiently captured field-observed variations in windthrow tree-mortality across the entire gradient of disturbance (Sentinel 2 and WorldView 2 produced the second and third best results, respectively). As expected, mean-associated uncertainties decreased systematically with increasing spatial resolution (i.e., from Landsat 8 to Sentinel 2 and WorldView 2). However, the overall quality of model fits showed the opposite pattern. We suggest that this reflects the influence of a relatively minor disturbance, such as defoliation and crown damage, and the fast growth of natural regeneration, which were not measured in the field nor can be captured by coarser resolution imagery. Our results validate the reliability of Landsat imagery for assessing plot-to-landscape patterns of windthrow tree-mortality in dense and heterogeneous tropical forests. Satellites with high spatial resolution can improve estimates of windthrow severity by allowing the quantification of crown damage and mortality of lower canopy and understory trees. However, this requires the validation of remote sensing metrics using field data at compatible scales.

54 ENVIRONMENTAL SCIENCES↗

Feasibility of Adding Twitter Data to Aid Drought Depiction: Case Study in Colorado

The use of social media, such as Twitter, has changed the information landscape for citizens’ participation in crisis response and recovery activities. Given that drought progression is slow and also spatially extensive, an interesting set of questions arise, such as how the usage of Twitter by a large population may change during the development of a major drought alongside how the changing usage facilitates drought detection. For this reason, contemporary analysis of how social media data, in conjunction with meteorological records, was conducted towards improvement in the detection of drought and its progression. The research utilized machine learning techniques applied over satellite-derived drought conditions in Colorado. Three different machine learning techniques were examined: the generalized linear model, support vector machines and deep learning, each applied to test the integration of Twitter data with meteorological records as a predictor of drought development. It is found that the integration of data resources is viable given that the Twitter-based model outperformed the control run which did not include social media input. Eight of the ten models tested showed quantifiable improvements in the performance over the control run model, suggesting that the Twitter-based model was superior in predicting drought severity. Future work lies in expanding this method to depict drought in the western U.S.

54 ENVIRONMENTAL SCIENCES↗

Racoons ( Procyon Lotor ) Show Higher Trypanosoma Cruzi Detection Rates than Virginia Opossums ( Didelphis Virginiana ) in South Carolina, USA

Chagas disease, a significant public health concern in the Americas, is caused by a protozoan parasite, Trypanosoma cruzi. The life cycle of T. cruzi involves kissing bugs (Triatoma spp.) functioning as vectors and mammalian species serving as hosts. Raccoons (Procyon lotor) and opossums (Didelphis virginiana) have been identified as important reservoir species in the life cycle of T. cruzi, but prevalence in both species in the southeastern US is currently understudied. We quantified T. cruzi prevalence in these two key reservoir species across our study area in South Carolina, US, and identified factors that may influence parasite detection. We collected whole blood from 183 raccoons and 126 opossums and used PCR to detect the presence of T. cruzi. We then used generalized linear models with parasite detection status as a binary response variable and predictor variables of land cover, distance to water, sex, season, and species. Our analysis indicated that raccoons experienced significantly higher parasite detection rates than Virginia opossums, with T. cruzi prevalence found to be 26.5% (95% confidence interval [CI], 20.0–33.8) in raccoons and 10.5% (95% CI, 5.51–17.5) in opossums. Altogether, our results concur with previous studies, in that T. cruzi is established in reservoir host populations in natural areas of the southeastern US.

60 APPLIED LIFE SCIENCES↗

Generalized Tensor-on-Tensor Regression (GToTR)

SAND2026-23069O Generalized Tensor-on-Tensor Regression (GToTR) is a Python-based tool for conducting generalized tensor-on-tensor regression. It provides Canonical Polyadic (CP)-based generalized tensor regression models, support for generalized linear model-like families and links, alternating-optimization model fitting methods, and a standard statistics software interface. The tool supports tensor-valued responses and covariates using the open-source Python Tensor Toolbox (pyttb) software package. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Dunlavy, Daniel [Sandia National Lab. (SNL-CA), Li↗

Interpretable Machine Learning for Characterizing Electric Vehicle Charging Behavior: Insights from Real-World Data

As electric vehicle (EV) adoption rises globally, concerns about the impact on aging electrical grids grow, particularly regarding the charging behavior of EV drivers. This study analyzes real-world driving and charging data from Ford battery electric vehicles (BEVs) collected between 2018 and 2019 to develop interpretable models that characterize charging behavior and quantify influencing factors. Prior research has relied on assumptions regarding driver behavior, often overlooking actual charging patterns. By employing generalized linear mixed models (GLMMs), this work offers insights into how various elements, such as next trip distance and state of charge (SOC), influence charging decisions. The dataset comprises over three million park-trip pairs from 1,997 vehicles, revealing that features related to driving behavior significantly dictate charging behavior, while infrastructure and regional factors have lesser impacts. The findings suggest that existing simulation models may oversimplify EV charging behavior assumptions. This work utilizes real-world EV driving and charging data to train interpretable models that describe charging behavior and quantify the factors most associated with how drivers use charging infrastructure. This research underscores the need for interpretable, data-driven methodologies to inform future EV infrastructure planning and grid management.

29 - ENERGY PLANNING, POLICY AND ECONOMY↗

High-resolution (30-m) urban land cover projections for Los Angeles California Urban Area: 2010 to 2100 under SSP5

These data represent simulations of future land use and land cover (LULCC) for Los Angeles urban area (U.S. Census Bureau defined area) as raster tiff images at a 30-m pixel resolution and at decadal time steps from 2010 to 2100. LULCC classes in this product follow the National Land Cover Dataset (NLCD) classification. NLCD 21-24 correspond to open developed, low developed, medium developed, and high developed urban land classes, respectively. Only urban land cover classes (NLCD class 21, 22, 23, and 24) are dynamic over time; however, all NLCD classes are included in the final product. Therefore, NLCD classes that do not convert to an urban class will be similar to year 2000. The products were developed using a hybridized statistical and cellular automata approach. Linear mixed models (LMMs) were used to estimate future urban land budgets based on 1-km urban land fraction projections from Gao and Pesaresi (2021), whereas separate generalized linear mixed models (GLMMs) were used to estimate shifts in urban land intensities based on retrospective shifts in NLCD urban class intensities over a 20- year period. Based on urban land allocations from the statistical models, a cellular-automata and downscaling routine was used to simulate dynamic urban land expansion at a 30-m resolution based on suitability criteria. Scenarios of future urban landcover change projections include variant solutions for the Shared Socioeconomic Pathway 5 (SSP5) based on different population assumptions, different land use intensification assumptions, variable land zoning constraints, and iterative adjustments to correct for over allocation of urban expansion across decadal time periods from 2010 to 2100. This results in 320 raster products.

Land↗

High-resolution (30-m) urban land cover projections for Los Angeles California Urban Area: 2010 to 2100 under SSP3 and SSP5 [Updated simulations based on population-driven urban intensity transitions]

These data (v3) are updated from previous versions (1 and 2) in that they include consider the effects of population on transitions in urban land intensity. This leads to more reasonable differences in urban land projections under variant SSPs. For the present dataset, both SSP3 and SSP5 are provided. These data represent simulations of future land use and land cover (LULCC) for Los Angeles urban area (U.S. Census Bureau defined area) as raster tiff images at a 30-m pixel resolution and at decadal time steps from 2010 to 2100. LULCC classes in this product follow the National Land Cover Dataset (NLCD) classification. NLCD 21-24 correspond to open developed, low developed, medium developed, and high developed urban land classes, respectively. Only urban land cover classes (NLCD class 21, 22, 23, and 24) are dynamic over time; however, all NLCD classes are included in the final product. Therefore, NLCD classes that do not convert to an urban class will be similar to year 2000. The products were developed using a hybridized statistical and cellular automata approach. Linear mixed models (LMMs) were used to estimate future urban land budgets based on 1-km urban land fraction projections from Gao and Pesaresi (2021), whereas separate generalized linear mixed models (GLMMs) were used to estimate shifts in urban land intensities based on retrospective shifts in NLCD urban class intensities over a 20- year period. Based on urban land allocations from the statistical models, a cellular-automata and downscaling routine was used to simulate dynamic urban land expansion at a 30-m resolution based on suitability criteria. Scenarios of future urban landcover change projections include variant solutions for the Shared Socioeconomic Pathway 5 (SSP5) and SSP 3 based on different population assumptions, different land use intensification assumptions, variable land zoning constraints, and iterative adjustments to correct for over allocation of urban expansion across decadal time periods from 2010 to 2100. This results in 320 raster products.

Land↗

Predicting fatigue from heart rate signatures using functional logistic regression

Physical fatigue can have adverse effects on humans in extreme environments. Therefore, being able to predict fatigue using easy to measure metrics such as heart rate (HR) signatures has potential to have an impact in real-life scenarios. We apply a functional logistic regression model that uses HR signatures to predict physical fatigue, where physical fatigue is defined in a data-driven manner. Data were collected using commercially available wearable devices on 47 participants hiking the 20.7-mile Grand Canyon rim-to-rim trail in a single day. Fitted model provides good predictions and interpretable parameters for real-life application.

60 APPLIED LIFE SCIENCES↗

Personal and environmental predictors of polycyclic aromatic hydrocarbon exposure identified through repeated silicone wristband sampling

This study integrates quantitative data on personal exposure to polycyclic aromatic hydrocarbons (PAHs) in 162 silicone wristbands with demographics, behavioral information, and housing characteristics to explore contributions to residential exposure in a superfund-adjacent community over the course of a year. Forty-six residents completed questionnaires and wore silicone wristbands as personal passive samplers for seven consecutive days on up to four separate occasions in alternating months between November 2022 and June 2023. It was hypothesized that individual behaviors and housing characteristics are sources of dependence and correlation between personal PAH exposures. 50 PAHs were detected at least once, 17 of which were alkylated PAHs. Exposure to PAHs of similar molecular weight was often correlated, notably between naphthalenes (2-rings) and higher molecular weight PAHs (3 or more rings). Generalized linear mixed models identified flooring type, participant age, and sampling month as important predictors of increased PAH exposure, and flooring type, and use of wood stoves or heavy machinery as predictors of increased naphthalene exposure relative to higher molecular weight PAHs. Individual chemical models based on concentration data and detection frequencies corroborated these findings across multiple PAHs. We demonstrate that personal exposure is not static and the degree of variability in personal exposure is individual. Hence, identification of influential exposure factors through repeated measures of chemical exposure and characterization of variability in personal exposure as performed in this study, is important in the development of exposure mitigation strategies.

Bonner, Emily↗