Search NASA⌕ Search

Engineering topics

Gaziano, J. Michael

Publications and source records attributed to Gaziano, J. Michael.

Diversity and scale: Genetic architecture of 2068 traits in the VA Million Veteran Program

One of the justifiable criticisms of human genetic studies is the underrepresentation of participants from diverse populations. Lack of inclusion must be addressed at-scale to identify causal disease factors and understand the genetic causes of health disparities. We present genome-wide associations for 2068 traits from 635,969 participants in the Department of Veterans Affairs Million Veteran Program, a longitudinal study of diverse United States Veterans. Systematic analysis revealed 13,672 genomic risk loci; 1608 were only significant after including non-European populations. Fine-mapping identified causal variants at 6318 signals across 613 traits. One-third (n = 2069) were identified in participants from non-European populations. This reveals a broadly similar genetic architecture across populations, highlights genetic insights gained from underrepresented groups, and presents an extensive atlas of genetic associations.

59 BASIC BIOLOGICAL SCIENCES↗

Type 1 Diabetes Genetic Risk in 109,954 Veterans With Adult-Onset Diabetes: The Million Veteran Program (MVP)

OBJECTIVE To characterize high type 1 diabetes (T1D) genetic risk in a population where type 2 diabetes (T2D) predominates. RESEARCH DESIGN AND METHODS Characteristics typically associated with T1D were assessed in 109,594 Million Veteran Program participants with adult-onset diabetes, 2011–2021, who had T1D genetic risk scores (GRS) defined as low (0 to <45%), medium (45 to <90%), high (90 to <95%), or highest (≥95%). RESULTS T1D characteristics increased progressively with higher genetic risk (P < 0.001 for trend). A GRS ≥90% was more common with diabetes diagnoses before age 40 years, but 95% of those participants were diagnosed at age ≥40 years, and their characteristics resembled those of individuals with T2D in mean age (64.3 years) and BMI (32.3 kg/m2). Compared with the low-risk group, the highest-risk group was more likely to have diabetic ketoacidosis (low GRS 0.9% vs. highest GRS 3.7%), hypoglycemia prompting emergency visits (3.7% vs. 5.8%), outpatient plasma glucose <50 mg/dL (7.5% vs. 13.4%), a shorter median time to start insulin (3.5 vs. 1.4 years), use of a T1D diagnostic code (16.3% vs. 28.1%), low C-peptide levels if tested (1.8% vs. 32.4%), and glutamic acid decarboxylase antibodies (6.9% vs. 45.2%), all P < 0.001. CONCLUSIONS Characteristics associated with T1D were increased with higher genetic risk, and especially with the top 10% of risk. However, the age and BMI of those participants resemble those of people with T2D, and a substantial proportion did not have diagnostic testing or use of T1D diagnostic codes. T1D genetic screening could be used to aid identification of adult-onset T1D in settings in which T2D predominates.

Yang, Peter K. (ORCID:0000000193796981)↗

A multi-ancestry GWAS of Fuchs corneal dystrophy highlights the contributions of laminins, collagen, and endothelial cell regulation

Fuchs endothelial corneal dystrophy (FECD) is a leading indication for corneal transplantation, but its molecular etiology remains poorly understood. We performed genome-wide association studies (GWAS) of FECD in the Million Veteran Program followed by multi-ancestry meta-analysis with the previous largest FECD GWAS, for a total of 3970 cases and 333,794 controls. We confirm the previous four loci, and identify eight novel loci: SSBP3, THSD7A, LAMB1, PIDD1, RORA, HS3ST3B1, LAMA5, and COL18A1. We further confirm the TCF4 locus in GWAS for admixed African and Hispanic/Latino ancestries and show an enrichment of European-ancestry haplotypes at TCF4 in FECD cases. Among the novel associations are low frequency missense variants in laminin genes LAMA5 and LAMB1 which, together with previously reported LAMC1, form laminin-511 (LM511). AlphaFold 2 protein modeling, validated through homology, suggests that mutations at LAMA5 and LAMB1 may destabilize LM511 by altering inter-domain interactions or extracellular matrix binding. Finally, phenome-wide association scans and colocalization analyses suggest that the TCF4 CTG18.1 trinucleotide repeat expansion leads to dysregulation of ion transport in the corneal endothelium and has pleiotropic effects on renal function.

59 BASIC BIOLOGICAL SCIENCES↗

Centralized Interactive Phenomics Resource: an integrated online phenomics knowledgebase for health data users

Development of clinical phenotypes from electronic health records (EHRs) can be resource intensive. Several phenotype libraries have been created to facilitate reuse of definitions. However, these platforms vary in target audience and utility. Here, we describe the development of the Centralized Interactive Phenomics Resource (CIPHER) knowledgebase, a comprehensive public-facing phenotype library, which aims to facilitate clinical and health services research. The platform was designed to collect and catalog EHR-based computable phenotype algorithms from any healthcare system, scale metadata management, facilitate phenotype discovery, and allow for integration of tools and user workflows. Phenomics experts were engaged in the development and testing of the site. The knowledgebase stores phenotype metadata using the CIPHER standard, and definitions are accessible through complex searching. Phenotypes are contributed to the knowledgebase via webform, allowing metadata validation. Data visualization tools linking to the knowledgebase enhance user interaction with content and accelerate phenotype development. The CIPHER knowledgebase was developed in the largest healthcare system in the United States and piloted with external partners. The design of the CIPHER website supports a variety of front-end tools and features to facilitate phenotype development and reuse. Health data users are encouraged to contribute their algorithms to the knowledgebase for wider dissemination to the research community, and to use the platform as a springboard for phenotyping. CIPHER is a public resource for all health data users available at https://phenomics.va.ornl.gov/ which facilitates phenotype reuse, development, and dissemination of phenotyping knowledge.

60 APPLIED LIFE SCIENCES↗

Multimodal representation learning for predicting molecule–disease relations

Motivation: Predicting molecule–disease indications and side effects is important for drug development and pharmacovigilance. Comprehensively mining molecule–molecule, molecule–disease and disease–disease semantic dependencies can potentially improve prediction performance. Methods: We introduce a Multi-Modal REpresentation Mapping Approach to Predicting molecular-disease relations (M2REMAP) by incorporating clinical semantics learned from electronic health records (EHR) of 12.6 million patients. Specifically, M2REMAP first learns a multimodal molecule representation that synthesizes chemical property and clinical semantic information by mapping molecule chemicals via a deep neural network onto the clinical semantic embedding space shared by drugs, diseases and other common clinical concepts. To infer molecule–disease relations, M2REMAP combines multimodal molecule representation and disease semantic embedding to jointly infer indications and side effects. Results: We extensively evaluate M2REMAP on molecule indications, side effects and interactions. Results show that incorporating EHR embeddings improves performance significantly, for example, attaining an improvement over the baseline models by 23.6% in PRC-AUC on indications and 23.9% on side effects. Further, M2REMAP overcomes the limitation of existing methods and effectively predicts drugs for novel diseases and emerging pathogens. Availability and implementation: The code is available at https://github.com/celehs/M2REMAP, and prediction results are provided at https://shiny.parse-health.org/drugs-diseases-dev/.

59 BASIC BIOLOGICAL SCIENCES↗