Search NASA⌕ Search

SEARCH · Search NASA

Results for “Phenotypes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Phenotypic Evaluation of Saccharum spp. Genotypes during the Plant-Cane Crop for Biomass Production in Northcentral Mississippi

Saccharum is relatively new to 33° N latitude. S. spontaneum readily hybridizes with commercial sugarcane (Saccharum spp.) and lends cold tolerance and greater yield to the hybrid progeny, called energycane. Since 2007, there have been numerous new hybrid and backcross energycane genotypes developed but there is a paucity of information about them. Twenty energycane genotypes were tested in the first season of growth from cane propagules (plant cane; PC) against Ho 02-113 (a control) for two site-years in northcentral Mississippi. Grand (exponential) growth continued into October. The prevailing paradigm is that tonnage is what matters. Except for percentage cellulose, all factors tested (dry matter yield, extractable juice volume, °Brix, theoretical ethanol from fermentation, theoretical ethanol from cellulose, and total theoretical ethanol) were greater from the second site-location compared to the first. Dry matter yield (DMY) and total theoretical ethanol yield (TTEY) were moderately correlated. Over the two years of this test only Ho 14-9213 exceeded in mean DMY of Ho 02-113. Sixteen of the 19 test genotypes in this test equaled or exceeded the mean TTEY of Ho 02-113.

60 APPLIED LIFE SCIENCES↗

Data from: Biomass yield, yield components and growing season phenotypic measurements of Miscanthus

For sustainable biomass production of Miscanthus × giganteus (hereafter miscanthus), understanding the impact of stand age and nitrogen (N) fertilization on biomass yield is crucial. This study investigated the effects of varying N fertilization rates (0, 56, 112, and 168 kg N ha-1) on yield components (tiller height, density, and weight) and their correlations with end-of-season biomass yield in miscanthus. We also explored end-of-season biomass yield prediction using in-season traits (canopy height, leaf area index (LAI), and leaf chlorophyll content (LCC)). The study was conducted at two sites in Illinois: a previously unfertilized 10-year-old miscanthus research stand at Urbana and a 16-year-old commercial stand at Pesotum with a history of annual 56N application. Results from 2018-2021 in Urbana and 2020-2021 in Pesotum showed increased biomass yields with N fertilization, varying by rate, year, and location. Biomass yield in Pesotum peaked at 56N, while in Urbana, it increased significantly at 112 kg N ha-1. Biomass yield was strongly correlated with tiller height and weight measured at Urbana across N rates. Morphological traits measured every 2-3 weeks during the 2020 and 2021 growing seasons showed that canopy height was the strongest single predictor of miscanthus biomass yield, followed by LCC. Mid-August to September measurements of these traits were the best predictors of biomass yield. Multiple regressions involving the canopy height and LCC further improved yield predictions. We conclude that while N enhances biomass yields of aging miscanthus, the optimum rate depends on the site, environmental conditions, and management.

Aging↗

Lost and Found: Rediscovering Microbiome-Associated Phenotypes that Reshape Agricultural Sustainability

Overview Code and data repository for NIL Manuscript. Documentation includes sequence processing examples and data analysis. Supplemental sequence processing and R statistical analysis for publication, which compares the microbiome of teosinte-B73 Near Isogenic Lines. Sample Data Amplicon sequence data for 16S rRNA genes, the fungal ITS2 region, and nitrogen-cycling functional genes are available through the NCBI Sequence Read Archive (SRA) under accession number PRJNA1042643(https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1042643). Raw metabolomic data are available on Metabolomics Workbench, Project ID: PR002654. This study is available at the NIH Common Fund's National Metabolomics Data Repository (NMDR) website, the Metabolomics Workbench, https://www.metabolomicsworkbench.org where it has been assigned Study ID ST004211. The data can be accessed directly via its Project DOI: http://dx.doi.org/10.21228/M8KV8T.

Near Isogeneic Lines↗

Centralized Interactive Phenomics Resource: an integrated online phenomics knowledgebase for health data users

Development of clinical phenotypes from electronic health records (EHRs) can be resource intensive. Several phenotype libraries have been created to facilitate reuse of definitions. However, these platforms vary in target audience and utility. Here, we describe the development of the Centralized Interactive Phenomics Resource (CIPHER) knowledgebase, a comprehensive public-facing phenotype library, which aims to facilitate clinical and health services research. The platform was designed to collect and catalog EHR-based computable phenotype algorithms from any healthcare system, scale metadata management, facilitate phenotype discovery, and allow for integration of tools and user workflows. Phenomics experts were engaged in the development and testing of the site. The knowledgebase stores phenotype metadata using the CIPHER standard, and definitions are accessible through complex searching. Phenotypes are contributed to the knowledgebase via webform, allowing metadata validation. Data visualization tools linking to the knowledgebase enhance user interaction with content and accelerate phenotype development. The CIPHER knowledgebase was developed in the largest healthcare system in the United States and piloted with external partners. The design of the CIPHER website supports a variety of front-end tools and features to facilitate phenotype development and reuse. Health data users are encouraged to contribute their algorithms to the knowledgebase for wider dissemination to the research community, and to use the platform as a springboard for phenotyping. CIPHER is a public resource for all health data users available at https://phenomics.va.ornl.gov/ which facilitates phenotype reuse, development, and dissemination of phenotyping knowledge.

60 APPLIED LIFE SCIENCES↗

Phenome‐to‐genome insights for evaluating root system architecture in field studies of maize

Abstract Understanding the genetic basis of root system architecture (RSA) in crops requires innovative approaches that enable both high‐throughput and precise phenotyping in field conditions. In this study, we evaluated multiple phenotyping and analytical frameworks for quantifying RSA in mature, field‐grown maize in three field experiments. We used forward and reverse genetic approaches to evaluate >1700 maize root crowns, including a diversity panel, a biparental mapping population, and maize mutant and wild‐type alleles at two known RSA genes,DEEPER ROOTING 1(DRO1) andRootless1(Rt1). We show the utility of increasing the dimensionality of traditional two‐dimensional (2D) techniques, referred to as the “2D multi‐view” method, to improve the capture of whole root system information for mapping genetic variation influencing RSA. Comparison of univariate and multivariate genome‐wide association study (GWAS) approaches revealed that multivariate traits were effective at dissecting complex RSA phenotypes and identifying pleiotropic quantitative trait loci (QTLs). Overall, three‐dimensional (3D) root models generated from X‐ray computed tomography and digital phenotyping captured a larger proportion of RSA trait variations compared to other methods of root phenotyping, as evidenced by both genome‐wide and single‐gene analyses. Among the individual root traits, root pulling force emerged as a highly heritable estimate of RSA that identified the largest number of shared QTLs with 3D phenotypes. Our study shows that integrating complementary phenotyping technologies helps to provide a more comprehensive understanding of the genetic architecture of RSA in field‐grown maize.

Genetics & Heredity↗

Data for "Genetics of flooding tolerance in an F2 Miscanthus sacchariflorus ssp. lutarioriparius × M. sinensis population"

This dataset contains all data and supplementary materials from "Genetics of flooding tolerance in an F2 Miscanthus sacchariflorus ssp. lutarioriparius × M. sinensis population". 1. The dataset S1 table contains the raw phenotypic data collected during the experiment. 2. The dataset S2 table contains the LSmean values for the 24 traits studied. 3. The dataset S3 table contains the TASSEL GBSv2 map, marker information, and genotype data used for mapping. 4. The dataset S4 table contains information on candidate genes found in each of the QTL intervals. 5. The dataset S5 table contains the GO annotations and KEGG enrichment analyses for those candidate genes. 6. The dataset S6 table contains information on the sequences used to classify AP2 ERF transcription factors. 7. The dataset S7 table contains information on AP2 ERF orthologs between Miscanthus and rice based on synteny. 8. Supplementary file 1 contains the ANOVA results using the raw phenotypic data collected from protocol "A". 9. Supplementary file 2 contains the ANOVA results using the raw phenotypic data collected from protocol "B". 10. Supplementary file 3 contains notes on the comparison of SNP calling methods. 11. Supplementary file 4 is a script for analyzing candidate genes found in QTL intervals.

Miscanthus, flood, partial submergence, complete s↗

Five amino acid mismatches in the zinc-finger domains of Cellulose Synthase 5 and Cellulose Synthase 6 cooperatively modulate their functional properties by controlling homodimerization in Arabidopsis

Cellulose synthase 5 (CESA5) and CESA6 are known to share substantial functional overlap. In the zinc-finger domain (ZN) of CESA5, there are five amino acid (AA) mismatches when compared to CESA6. These mismatches in CESA5 were replaced with their CESA6 counterparts one by one until all were replaced, generating nine engineered CESA5s. Each N-terminal enhanced yellow fluorescent protein-tagged engineered CESA5 was introduced to prc1-1, a cesa6 null mutant, and resulting mutants were subjected to phenotypic analyses. For this work, we found that five single AA-replaced CESA5 proteins partially rescue the prc1-1 mutant phenotypes to different extents. Multi-AA replaced CESA5s further rescued the mutant phenotypes in an additive manner, culminating in full recovery by CESA5 G43R + S49T+S54P+S80A+Y88F . Investigations in cellulose content, cellulose synthase complex (CSC) motility, and cellulose microfibril organization in the same mutants support the results of the phenotypic analyses. Bimolecular fluorescence complementation assays demonstrated that the level of homodimerization in every engineered CESA5 is substantially higher than CESA5. The mean fluorescence intensity of CSCs carrying each engineered CESA5 fluctuates with the degree to which the prc1-1 mutant phenotypes are rescued by introducing a corresponding engineered CESA5. Taken together, these five AA mismatches in the ZNs of CESA5 and CESA6 cooperatively modulate the functional properties of these CESAs by controlling their homodimerization capacity, which in turn imposes proportional changes on the incorporation of these CESAs into CSCs.

59 BASIC BIOLOGICAL SCIENCES↗

CarbStor: Development, Analysis and Modification of Carbon Storing Model Soil Communities

Soil microbial communities carry out a number of key processes including plant growth promotion, bioremediation and cycling of nutrients. Carbon cycling is among the most important of these nutrients that are metabolized and processed by the soil microbial community. Many of the carbon inputs are converted to alternative organic forms of carbon that can be used by plants or act as biomass for microbial growth. However, inorganic forms of carbon can also be produced by soil microbial communities including calcium carbonate (CaCO 3 ). Production of calcium carbonate is beneficial for the ecosystem in several ways: it can stabilize soils and improve soil health, especially denser soils with high clay content, it can act as a method of bioremediation, it can serve as an alternative carbon source for plants and it can be a way to store carbon in soil in a stable, inorganic manner for the long term. While the chemistry surrounding individual species carrying out this process is well known what is lacking is an understanding of how species interact in a community to drive carbonate production. As all microbial species in soil exist in a community setting gaining this knowledge is critical to our predicting and controlling this microbial phenotype to greatly improve soil health. The CarbStor project is focused on developing, analyzing and modifying defined microbial soil consortia that express phenotypes at both the species and community level to convert carbon into recalcitrant stable sources such as precipitated carbonate or microbial necromass. To take full advantage of the soil community for this process we will need to fill several key knowledge gaps (KG), three of which are the focus of CarbStor. KG1: Whether and to what degree microbial communities can be developed that produce precipitated carbon via microbial metabolism. KG2: What interspecies interactions drive the individual member phenotypes in defined communities that lead to carbon precipitation. KG3: How can these interactions be modified to enhance carbon sequestration beyond what native communities are capable of. We hypothesize that in a carbon sequestering community only a subset of species will express phenotypes related to carbon storage processes. We also hypothesize that these phenotypes are expressed as a result of interactions with other species in the community that are not involved in carbon storage processes and that these interactions can be harnessed to enhance community carbon sequestration.

54 ENVIRONMENTAL SCIENCES↗

Stochastic journeys of cell progenies through compartments and the role of self-renewal, symmetric and asymmetric division

Abstract Division and differentiation events by which cell populations with specific functions are generated often take place as part of a developmental programme, which can be represented by a sequence of compartments. A compartment is the set of cells with common characteristics; sharing, for instance, a spatial location or a phenotype. Differentiation events are transitions from one compartment to the next. Cells may also die or divide. We consider three different types of division events: (i) where both daughter cells inherit the mother’s phenotype (self-renewal), (ii) where only one of the daughters changes phenotype (asymmetric division), and (iii) where both daughters change phenotype (symmetric division). The self-renewal probability in each compartment determines whether the progeny of a single cell, moving through the sequence of compartments, is finite or grows without bound. We analyse the progeny stochastic dynamics with probability generating functions. In the case of self-renewal, by following one of the daughters after any division event, we may construct lifelines containing only one cell at any time. We analyse the number of divisions along such lines, and the compartment where lines terminate with a death event. Analysis and numerical simulations are applied to a five-compartment model of the gradual differentiation of hematopoietic stem cells and to a model of thymocyte development: from pre-double positive to single positive (SP) cells with a bifurcation to either SP4 or SP8 in the last compartment of the sequence.

Science & Technology - Other Topics↗

Cryo-EM confirms a common fibril fold in the heart of four patients with ATTRwt amyloidosis

ATTR amyloidosis results from the conversion of transthyretin into amyloid fibrils that deposit in tissues causing organ failure and death. This conversion is facilitated by mutations in ATTRv amyloidosis, or aging in ATTRwt amyloidosis. ATTRv amyloidosis exhibits extreme phenotypic variability, whereas ATTRwt amyloidosis presentation is consistent and predictable. Previously, we found unique structural variabilities in cardiac amyloid fibrils from polyneuropathic ATTRv-I84S patients. In contrast, cardiac fibrils from five genotypically different patients with cardiomyopathy or mixed phenotypes are structurally homogeneous. To understand fibril structure’s impact on phenotype, it is necessary to study the fibrils from multiple patients sharing genotype and phenotype. Here we show the cryo-electron microscopy structures of fibrils extracted from four cardiomyopathic ATTRwt amyloidosis patients. Our study confirms that they share identical conformations with minimal structural variability, consistent with their homogenous clinical presentation. Our study contributes to the understanding of ATTR amyloidosis biopathology and calls for further studies.

59 BASIC BIOLOGICAL SCIENCES↗

Host-specific adaptation in Fusarium oxysporum correlates with distinct accessory chromosome content in human and plant pathogenic strains

ABSTRACT Fusarium oxysporumis a cross-kingdom pathogen. While some strains cause disseminated fusariosis and blinding corneal infections in humans, others are responsible for devastating vascular wilt diseases in plants. To better understand the distinct adaptations ofF. oxysporumto animal or plant hosts, we conducted a comparative phenotypic and genetic analysis of two strains: MRL8996 (isolated from a keratitis patient) and Fol4287 (isolated from a wilted tomato [Solanum lycopersicum]). Infection of mouse corneas and tomato plants revealed that, while both strains cause symptoms in both hosts, MRL8996 caused more severe corneal disease in mice, whereas Fol4287 induced more pronounced wilting symptoms in tomato plants.In vitroassays using abiotic stress treatments revealed that the human pathogen MRL8996 was better adapted to elevated temperatures, whereas the plant pathogen Fol4287 was more tolerant to osmotic and cell wall stresses. Both strains displayed broad resistance to antifungal treatment, with MRL8996 exhibiting the paradoxical effect of increased tolerance to higher concentrations of the antifungal caspofungin. We identified a set of accessory chromosomes (ACs) that encode genes with different functions and have distinct transposon profiles between MRL8996 and Fol4287. Interestingly, ACs from both genomes also encode proteins with shared functions, such as chromatin remodeling and post-translational protein modifications. Our phenotypic assays and comparative genomics analyses lay the foundation for future studies correlating genotypes with phenotype and for developing targeted antifungals for agricultural and clinical uses. IMPORTANCE Fusarium oxysporumis a cross-kingdom fungal pathogen that infects both plants and animals. In addition to causing many devastating wilt diseases, this group of organisms was recently recognized by the World Health Organization as a high-priority threat to human health. Climate change has increased the risk ofFusariuminfections, asFusariumstrains are highly adaptable to changing environments. Deciphering fungal adaptation mechanisms is crucial to developing appropriate control strategies. We performed a comparative analysis ofFusariumstrains using an animal (mouse) and plant (tomato) host andin vitroconditions that mimic abiotic stress. We also performed comparative genomics analyses to highlight the genetic differences between human and plant pathogens and correlate their phenotypic and genotypic variations. We uncovered important functional hubs shared by plant and human pathogens, such as chromatin modification, transcriptional regulation, and signal transduction, which could be used to identify novel antifungal targets.

Microbiology↗

Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1

KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer

Joachimiak, Marcin [Lawrence Berkeley National Lab↗

Data for Promoter Deletion in the Soybean Compact Mutant Leads to Overexpression of a Gene with Homology to the C20-Gibberellin 2-Oxidase Family

Height is a critical component of plant architecture, significantly affecting crop yield. The genetic basis of this trait in soybean remains unclear. In this study, we report the characterization of the Compact mutant of soybean, which has short internodes. The candidate gene was mapped to chromosome 17, and the interval containing the causative mutation was further delineated using biparental mapping. Whole-genome sequencing of the mutant revealed an 8.7 kb deletion in the promoter of the Glyma.17g145200 gene, which encodes a member of the class III gibberellin (GA) 2-oxidases. The mutation has a dominant effect, likely via increased expression of the GA 2-oxidase transcript observed in green tissue, as a result of the deletion in the promoter of Glyma.17g145200. We further demonstrate that levels of GA precursors are altered in the Compact mutant, supporting a role in GA metabolism, and that the mutant phenotype can be rescued with exogenous GA3. We also determined that overexpression of Glyma.17g145200 in Arabidopsis results in dwarfed plants. Thus, gain of promoter activity in the Compact mutant leads to a short internode phenotype in soybean through altered metabolism of gibberellin precursors. These results provide an example of how structural variation can control an important crop trait and a role for Glyma.17g145200 in soybean architecture, with potential implications for increasing crop yield.

Biomass Analytics↗

Data for Creating Yellow Seed Camelina sativa with Enhanced Oil Accumulation by CRISPR-Mediated Disruption of Transparent Testa 8

Camelina ( Camelina sativa L.), a hexaploid member of the Brassicaceae family, is an emerging oilseed crop being developed to meet the increasing demand for plant oils as biofuel feedstocks. In other Brassicas, high oil content can be associated with a yellow seed phenotype, which is unknown for camelina. We sought to create yellow seed camelina using CRISPR/Cas9 technology to disrupt its Transparent Testa 8 (TT8) transcription factor genes and to evaluate the resulting seed phenotype. We identified three TT8 genes, one in each of the three camelina subgenomes, and obtained independent CsTT8 lines containing frameshift edits. Disruption of TT8 caused seed coat colour to change from brown to yellow reflecting their reduced flavonoid accumulation of up to 44%, and the loss of a well-organized seed coat mucilage layer. Transcriptomic analysis of CsTT8-edited seeds revealed significantly increased expression of the lipid-related transcription factors LEC1, LEC2, FUS3, and WRI1 and their downstream fatty acid synthesis-related targets. These changes caused metabolic remodelling with increased fatty acid synthesis rates and corresponding increases in total fatty acid (TFA) accumulation from 32.4% to as high as 38.0% of seed weight, and TAG yield by more than 21% without significant changes in starch or protein levels compared to parental line. These data highlight the effectiveness of CRISPR in creating novel enhanced-oil germplasm in camelina. The resulting lines may directly contribute to future net-zero carbon energy production or be combined with other traits to produce desired lipid-derived bioproducts at high yields.

Biofuels↗

Diverse signatures of convergent evolution in cactus-associated yeasts

Many distantly related organisms have convergently evolved traits and lifestyles that enable them to live in similar ecological environments. However, the extent of phenotypic convergence evolving through the same or distinct genetic trajectories remains an open question. Here, we leverage a comprehensive dataset of genomic and phenotypic data from 1,049 yeast species in the subphylum Saccharomycotina (Kingdom Fungi, Phylum Ascomycota) to explore signatures of convergent evolution in cactophilic yeasts, ecological specialists associated with cacti. We inferred that the ecological association of yeasts with cacti arose independently approximately 17 times. Using a machine learning–based approach, we further found that cactophily can be predicted with 76% accuracy from both functional genomic and phenotypic data. The most informative feature for predicting cactophily was thermotolerance, which we found to be likely associated with altered evolutionary rates of genes impacting the cell envelope in several cactophilic lineages. We also identified horizontal gene transfer and duplication events of plant cell wall–degrading enzymes in distantly related cactophilic clades, suggesting that putatively adaptive traits evolved independently through disparate molecular mechanisms. Notably, we found that multiple cactophilic species and their close relatives have been reported as emerging human opportunistic pathogens, suggesting that the cactophilic lifestyle—and perhaps more generally lifestyles favoring thermotolerance—might preadapt yeasts to cause human disease. This work underscores the potential of a multifaceted approach involving high-throughput genomic and phenotypic data to shed light onto ecological adaptation and highlights how convergent evolution to wild environments could facilitate the transition to human pathogenicity.

59 BASIC BIOLOGICAL SCIENCES↗

Root genetics in the field to understand drought adaptation and carbon sequestration (Final Scientific/Technical Report)

For all crop plants, roots play a critical role in growth. Roots anchor the plants, and are the primary site of nutrient and water uptake. Roots are also the main source of C to soil in the form of root tissues and exudates, and thus greatly influence SOM stocks. To perform these functions, primary roots extend into soil, producing a network of branching roots of characteristic form, known as its root system architecture (RSA). RSA varies among species, and among varieties within a species that are adapted to different environments. Root traits are major targets for the second green revolution because of their potential to improve crop productivity, increase drought tolerance and nutrient acquisition, and increase C capture of soil. Improving the quality of roots in maize will be particularly valuable, since this crop is planted on over 92 million acres annually in the US. The future sustainability of agricultural systems relies on their ability to enhance soil organic matter (SOM) storage and reduce GHG emissions, while maintaining or enhancing productivity. This program had two components, Sensors and Models. For the first component, we designed and built a high-throughput phenotyping platform for root pulling of maize plants. This eliminated the physical labor of manually pulling up plants and reduced the number of personnel required down to one. The standardized pulling mechanism allowed recording force curves during the pulling process, providing additional information. We validated that the maximum force for pulling the root system was well-correlated with the root system mass and provided root crowns for further RSA analysis. These root crowns identified significant correlations with 2D root area and root depth, along with 3D root volume, total root length and number of root tips. We then used this system for field-based studies in maize on the genetics of root system architecture and its relation to nitrogen-use efficiency (NUE), including using lines relevant to the Corteva breeding program. Varieties were also evaluated at Corteva sites in the cornbelt and Danforth farm in Missouri, to establish responses across sites. From these studies we have identified genetic loci associated with root traits and created mutant lines for these loci and correlations of root traits with NUE. For the Models component, we worked to incorporate root and soil characteristics into the MEMS 2.0 soil and ecosystem biogeochemical model. Existing soil C models, such as Century, are unable to represent specific root trait interactions with the soil environment and therefore to accurately forecast the potential C sequestration benefits of root breeding under different climatic and soil type conditions. We have developed the MEMS 2.0 ecosystem biogeochemical model to improve quantification of farm-scale soil carbon and greenhouse gas emissions. The new knowledge and large datasets produced by this project will be used to develop and drive an innovative model capable of forecasting the impacts on soil C stocks and nutrient dynamics. An innovation was to use the empirical data from the field studies (in 1, above) to model genetic variation in nitrogen use efficiencies and soil C input. Our work demonstrated that maize root-derived C rapidly replaces existing soil C and after 3 years of continuous maize, up to 20% of soil organic C in the topsoil (0-15cm) and 3% in the subsoil (15-30cm) was contributed by maize. However, this contribution did not entirely represent a net increase. Root C contribution to soil was affected by maize genetics. We have analyzed soils derived from the CSU field trials for C and N stocks, in the different soil physical fractions represented by the MEMS model, using both physical fractionation with elemental analyses, and Fourier transformed infrared spectroscopy. Data will be used to link crop nitrogen use efficiencies with soil C sequestration and provide data to bridge the field trials with the model development, for verification of model predictions. The project had a number of successful outcomes: we have used the new phenotyping platform to identify new genetic loci that can enhance root phenotypes; we have partnered with multiple maize seed companies phenotype varieties in their breeding programs; we have developed the MEMS model that can help inform industry on the potential for carbon sequestration in the agricultural sector, and which is now available at the CSU Soil Carbon Solutions Center for use.

59 BASIC BIOLOGICAL SCIENCES↗

Consistent performance of large language models in rare disease diagnosis across ten languages and 4917 cases

Background Large language models (LLMs) are increasingly used medicine for diverse applications including differential diagnostic support. The training data used to create LLMs such as the Generative Pretrained Transformer (GPT) predominantly consist of English-language texts, but LLMs could be used across the globe to support diagnostics if language barriers could be overcome. Initial pilot studies on the utility of LLMs for differential diagnosis in languages other than English have shown promise, but a large-scale assessment on the relative performance of these models in a variety of European and non-European languages on a comprehensive corpus of challenging rare-disease cases is lacking. Methods We created 4917 clinical vignettes using structured data captured with Human Phenotype Ontology (HPO) terms with the Global Alliance for Genomics and Health (GA4GH) Phenopacket Schema. These clinical vignettes span a total of 360 distinct genetic diseases with 2525 associated phenotypic features. We used translations of the Human Phenotype Ontology together with language-specific templates to generate prompts in English, Chinese, Czech, Dutch, French, German, Italian, Japanese, Spanish, and Turkish. We applied GPT-4o, version gpt-4o-2024-08-06, and the medically fine-tuned Meditron3-70B to the task of delivering a ranked differential diagnosis using a zero-shot prompt. An ontology-based approach with the Mondo disease ontology was used to map synonyms and to map disease subtypes to clinical diagnoses in order to automate evaluation of LLM responses. Findings For English, GPT-4o placed the correct diagnosis at the first rank 19.9% and within the top-3 ranks 27.0% of the time. In comparison, for the nine non-English languages tested here the correct diagnosis was placed at rank 1 between 16.9% and 20.6%, within top-3 between 25.4% and 28.6% of cases. The Meditron3 model placed the correct diagnosis within the first 3 ranks for 20.9% of cases in English and between 19.9% and 24.0% for the other nine languages. Interpretation The differential diagnostic performance of LLMs across a comprehensive corpus of rare-disease cases was largely consistent across the ten languages tested. This suggests that the utility of LLMs in clinical settings may extend to non-English clinical settings.

Artificial intelligence↗

APPL Hyperspectral_Imaging_Dataset_for_Heritability_Analysis_in_Populus_trichocarpa

This dataset contains hyperspectral imaging data collected at the Advanced Plant Phenotyping Laboratory (APPL) at Oak Ridge National Laboratory. Natural variants of Populus trichocarpa were imaged using a high-throughput hyperspectral phenotyping pipeline to quantify spectral reflectance traits for downstream quantitative genetics analyses. The dataset includes hyperspectral image files and derived reflectance data products suitable for extracting spectral features across the measured wavelength range (e.g., VNIR and/or SWIR, depending on instrument configuration), along with associated sample metadata (e.g., genotype identifiers, experimental design factors, and imaging run identifiers). These data were generated to support analyses of broad-sense heritability of hyperspectral traits and their relationships with biochemical phenotypes (including lignin traits from Py-MBMS).

APPL↗