Search NASA⌕ Search

Engineering topics

Streich, Jared

Publications and source records attributed to Streich, Jared.

Using iterative random forest to find geospatial environmental and Sociodemographic predictors of suicide attempts

Despite a recent global decrease in suicide rates, death by suicide has increased in the United States. It is therefore imperative to identify the risk factors associated with suicide attempts to combat this growing epidemic. In this study, we aim to identify potential risk factors of suicide attempt using geospatial features in an Artificial intelligence framework. We use iterative Random Forest, an explainable artificial intelligence method, to predict suicide attempts using data from the Million Veteran Program. This cohort incorporated 405,540 patients with 391,409 controls and 14,131 attempts. Our predictive model incorporates multiple climatic features at ZIP-code-level geospatial resolution. We additionally consider demographic features from the American Community Survey as well as the number of firearms and alcohol vendors per 10,000 people to assess the contributions of proximal environment, access to means, and restraint decrease to suicide attempts. In total 1,784 features were included in the predictive model. Our results show that geographic areas with higher concentrations of married males living with spouses are predictive of lower rates of suicide attempts, whereas geographic areas where males are more likely to live alone and to rent housing are predictive of higher rates of suicide attempts. We also identified climatic features that were associated with suicide attempt risk by age group. Additionally, we observed that firearms and alcohol vendors were associated with increased risk for suicide attempts irrespective of the age group examined, but that their effects were small in comparison to the top features. Taken together, our findings highlight the importance of social determinants and environmental factors in understanding suicide risk among veterans.

60 APPLIED LIFE SCIENCES↗

Few-Shot Learning Enables Population-Scale Analysis of Leaf Traits in Populus trichocarpa

Plant phenotyping is typically a time-consuming and expensive endeavor, requiring large groups of researchers to meticulously measure biologically relevant plant traits, and is the main bottleneck in understanding plant adaptation and the genetic architecture underlying complex traits at population scale. In this work, we address these challenges by leveraging few-shot learning with convolutional neural networks to segment the leaf body and visible venation of 2,906 Populus trichocarpa leaf images obtained in the field. In contrast to previous methods, our approach (a) does not require experimental or image preprocessing, (b) uses the raw RGB images at full resolution, and (c) requires very few samples for training (e.g., just 8 images for vein segmentation). Traits relating to leaf morphology and vein topology are extracted from the resulting segmentations using traditional open-source image-processing tools, validated using real-world physical measurements, and used to conduct a genome-wide association study to identify genes controlling the traits. In this way, the current work is designed to provide the plant phenotyping community with (a) methods for fast and accurate image-based feature extraction that require minimal training data and (b) a new population-scale dataset, including 68 different leaf phenotypes, for domain scientists and machine learning researchers. All of the few-shot learning code, data, and results are made publicly available.

59 BASIC BIOLOGICAL SCIENCES↗

Longitudinal Effects on Plant Species Involved in Agriculture and Pandemic Emergence Undergoing Changes in Abiotic Stress

In this work we identify changes in high-resolution zones across the globe linked by environmental similarity that have implications for agriculture, bioenergy, and zoonosis. We refine exhaustive vector comparison methods with improved similarity metrics as well as provide multiple methods of amalgamation across 744 months of climatic data. The results of the vector comparison are captured as networks which are analyzed using static and longitudinal comparison methods to reveal locations around the globe experiencing dramatic changes in abiotic stress. Specifically we (i) incorporate updated similarity scores and provide a comparison between similarity metrics, (ii) implement a new feature for resource optimization, (iii) compare an agglomerative view to a longitudinal view, (iv) compare across 2-way and 3-way vector comparisons, (v) implement a new form of analysis, and (vi) demonstrate biological applications and discuss implications across a diverse set of species distributions by detecting changes that affect their habitats. Species of interest are related to agriculture (e.g., coffee, wine, chocolate), bioenergy (e.g., poplar, switchgrass, pennycress), as well as those living in zones of concern for zoonotic spillover that may lead to pandemics (e.g., eucalyptus, flying foxes).

Cashman, Mikaela↗

Supporting information for Few-Shot Learning Enables Population-Scale Analysis of Leaf Traits in Populus trichocarpa

In this work, we use few-shot learning to segment the body and vein architecture of P. trichocarpa leaves from high-resolution scans obtained in the UC Davis common garden. Leaf and vein segmentation are formulated as separate tasks, in which convolutional neural networks (CNNs) are used to iteratively expand partial segmentations until reaching stopping criteria. Our leaf and vein segmentation approaches use just 50 and 8 manually traced images for training, respectively, and are applied to a set of 2,634 top and bottom leaf scans. We show that both methods achieve high segmentation accuracy, in some cases exceeding even human-level segmentation. The leaf and vein segmentations are subsequently used to extract 68 morphological traits using traditional open-source image processing tools, which are validated using real-world physical measurements. For a biological perspective, we perform a genome-wide association study using the vein density trait to discover novel genetic architectures associated with multiple physiological processes relating to leaf development and function. In addition to sharing all of the few-shot learning code (see https://github.com/jlager/few-shot-leaf-segmentation), we are releasing all images, manual segmentations, model predictions, 68 extracted leaf phenotypes, and a new set of SNPs called against the v4 P. trichocarpa genome for 1,419 genotypes. The data folder includes all images, ground truth segmentations, predicted segmentations, and extracted leaf traits. All images encode the sample ID in the file name by indicating the treatment, block, row, position, and leaf side, respectively. For example, the file, C_1_1_2_bot.jpeg, indicates the control treatment, block 1, row 1, position 2, and the bottom side of the leaf. Tabulated results include position IDs as well as the corresponding genotype IDs. The images folder includes the 2,906 high-resolution leaf scans taken in the field. The leaf_masks folder includes 50 ground truth segmentations used for training the leaf tracing algorithm. The leaf_preds folder includes the 2,906 predicted segmentations from the leaf tracing algorithm. The vein_masks folder includes 8 ground truth segmentations used for training the vein growing algorithm. The vein_preds folder includes the 1,453 predicted segmentations from the vein growing algorithm. The vein_probs folder includes the 1,453 predicted probability maps from the vein growing algorithm before thresholding. The genomes folder includes the set of SNPs called against the v4 P. trichocarpa genome for 1,419 genotypes with a README file detailing the steps taken. The results folder includes: (i) raw values of the 68 predicted leaf traits in digital_traits.tsv, (ii) manually measured values of petiole length and width in manual_traits.tsv, (iii) thin plate spline (TPS) adjusted values of the vein density trait in vein_density_tps_adj.tsv, (iv) best linear unbiased prediction (BLUP) adjusted values of the vein density trait in vein_density_blups.tsv, and (v) GWAS results for the vein density trait, including chromosome positions and corresponding P values, in gwas_results.csv.

09 BIOMASS FUELS↗

Climatic clustering and longitudinal analysis with impacts on food, bioenergy, and pandemics

Predicted growth in world population will put unparalleled stress on the need for sustainable energy and global food production, as well as increase the likelihood of future pandemics. In this work, we identify high-resolution environmental zones in the context of a changing climate and predict longitudinal processes relevant to these challenges. We do this using exhaustive vector comparison methods that measure the climatic similarity between all locations on earth at high geospatial resolution relative to global-scale analyses. The results are captured as networks, in which edges between geolocations are defined if their historical climate similarities exceed a threshold. We apply Markov clustering and our novel Correlation of Correlations method to the resulting climatic networks, which provides unprecedented agglomerative and longitudinal views of climatic relationships across the globe. The methods performed here resulted in the fastest (9.37x10 18 operations/sec) and one of the largest (168.7x10 21 operations) scientific computations ever performed, with more than 100 quadrillion edges considered for a single climatic network. Our climatic analysis reveals areas of the world experiencing rapid environmental changes, which can have important implications for global carbon fluxes and zoonotic spillover events. Correlation and network analyses of this kind are widely applicable across computational and predictive biology domains, including systems biology, ecology, carbon cycles, biogeochemistry, and zoonosis research.

59 BASIC BIOLOGICAL SCIENCES↗

Supporting data for climatic clustering and longitudinal analysis with impacts on food, bioenergy, and pandemics

This data supports the conclusions found in climatic clustering and longitudinal analysis with impacts on food, bioenergy, and pandemics. Included here are (i) the binarized geolocation vectors used for exhaustive vector comparisons, (ii) the resulting climatic networks, (iii) the results of applying Markov clustering to the climatic networks, and (iv) the results of applying Correlation-of-Correlations (cor-cor) to the climatic networks. The set of binarized geolocation vectors that are used as inputs for the Combinatorial Metrics library (CoMet) are of the form comet-UUUUUxVVVVV-XXXX-YYYY.shuffled.tped where UUUUU is the number of vectors, VVVVV is the length of each vector, XXXX is the starting year, and YYYY is the ending year. Each line corresponds to a geolocation vector of binary elements A (i.e., 0) and T (i.e., 1). The set of climatic networks that are used for downstream network analysis are of the form network-U-way-XXXX-YYYY.parsed.txt where U is the order of the comparison (2-way or 3-way), XXXX is the starting year, and YYYY is the ending year. Each line corresponds to an edge linking two geolocations (defined by latitude and longitude) with its corresponding edge weight (i.e., DUO score). The set of cluster results are of the form clusters-U-way-XXXX-YYYY-thresh-VVVV-inflation-WWW.clustered.txt where U is the order of the comparison (2-way or 3-way), XXXX is the starting year, YYYY is the ending year, VVVV is the similarity threshold, and WWW is the Markov clustering inflation rate. Each line corresponds to a single cluster and is composed of a number of corresponding geolocations (defined by latitude and longitude). The set of cor-cor results are of the form corcor-U-way-XXXX-YYYY.cumulative.txt where U is the order of the comparison (2-way or 3-way), XXXX is the starting year, and YYYY is the ending year. Each line corresponds to a single geolocation with it's corresponding cor-cor value.

54 ENVIRONMENTAL SCIENCES↗

Quinoa Phenotyping Methodologies: An International Consensus

Quinoa is a crop originating in the Andes but grown more widely and with the genetic potential for significant further expansion. Due to the phenotypic plasticity of quinoa, varieties need to be assessed across years and multiple locations. To improve comparability among field trials across the globe and to facilitate collaborations, components of the trials need to be kept consistent, including the type and methods of data collected. Here, an internationally open-access framework for phenotyping a wide range of quinoa features is proposed to facilitate the systematic agronomic, physiological and genetic characterization of quinoa for crop adaptation and improvement. Mature plant phenotyping is a central aspect of this paper, including detailed descriptions and the provision of phenotyping cards to facilitate consistency in data collection. High-throughput methods for multi-temporal phenotyping based on remote sensing technologies are described. Tools for higher-throughput post-harvest phenotyping of seeds are presented. A guideline for approaching quinoa field trials including the collection of environmental data and designing layouts with statistical robustness is suggested. To move towards developing resources for quinoa in line with major cereal crops, a database was created. The Quinoa Germinate Platform will serve as a central repository of data for quinoa researchers globally.

59 BASIC BIOLOGICAL SCIENCES↗

Potentially adaptive SARS-CoV-2 mutations discovered with novel spatiotemporal and explainable AI models

Abstract Background A mechanistic understanding of the spread of SARS-CoV-2 and diligent tracking of ongoing mutagenesis are of key importance to plan robust strategies for confining its transmission. Large numbers of available sequences and their dates of transmission provide an unprecedented opportunity to analyze evolutionary adaptation in novel ways. Addition of high-resolution structural information can reveal the functional basis of these processes at the molecular level. Integrated systems biology-directed analyses of these data layers afford valuable insights to build a global understanding of the COVID-19 pandemic. Results Here we identify globally distributed haplotypes from 15,789 SARS-CoV-2 genomes and model their success based on their duration, dispersal, and frequency in the host population. Our models identify mutations that are likely compensatory adaptive changes that allowed for rapid expansion of the virus. Functional predictions from structural analyses indicate that, contrary to previous reports, the Asp 614 Gly mutation in the spike glycoprotein (S) likely reduced transmission and the subsequent Pro 323 Leu mutation in the RNA-dependent RNA polymerase led to the precipitous spread of the virus. Our model also suggests that two mutations in the nsp13 helicase allowed for the adaptation of the virus to the Pacific Northwest of the USA. Finally, our explainable artificial intelligence algorithm identified a mutational hotspot in the sequence of S that also displays a signature of positive selection and may have implications for tissue or cell-specific expression of the virus. Conclusions These results provide valuable insights for the development of drugs and surveillance strategies to combat the current and future pandemics.

59 BASIC BIOLOGICAL SCIENCES↗

Can exascale computing and explainable artificial intelligence applied to plant biology deliver on the United Nations sustainable development goals?

Human population growth and accelerated climate change necessitate agricultural improvements using designer crop ideotypes (idealized plants that can grow in niche environments). Diverse and highly skilled research groups must integrate efforts to bridge the gaps needed to achieve international goals toward sustainable agriculture. Given the scale of global agricultural needs and the breadth of multiple types of omics data needed to optimize these efforts, explainable artificial intelligence (AI with a decipherable decision making process that provides a meaningful explanation to humans) and exascale computing (computers that can perform 1018 floating-point operations per second, or exaflops) are crucial. Accurate phenotyping and daily-resolution climatype associations are equally important for refining ideotype production to specific environments at various levels of granularity. In this article, we review advances toward tackling technological hurdles to solve multiple United Nations Sustainable Development Goals and discuss a vision to overcome gaps between research and policy.

59 BASIC BIOLOGICAL SCIENCES↗