Search NASA⌕ Search

SEARCH · Search NASA

Results for “Genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (June to September 2017)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2017 in June (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) in an active meander (Meander C) of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (15-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (50-88 cm depth below surface). Sediments were homogenized from the ~10 cm cores for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0151851. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 405 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (May to September 2018)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2018 in May (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) near the Meander C/Pumphouse floodplain sites of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (30-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (70-80 cm depth below surface). Sediments were homogenized from the 10 cm core for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0141020. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 478 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES↗

High Throughput Genome Releaser

In this study, we present the development of a High Throughput Genome Releaser, an innovative device addressing common challenges in screening PCR. This genome DNA releaser is designed for rapid, cost-effective, and efficient DNA extraction, optimized for subsequent PCR reactions. Our experimentation with various synthetic materials led us to select a particular type of plastic that mirrors the properties of glass cover slides, providing a smooth surface and effective compression capabilities. We engineered a 96-well device equipped with a 96-well plate and a top rod, operable both manually and automatically, which is compatible with widely used liquid-handling robot decks. This compatibility enhances ease of use in high-throughput PCR setups. Additionally, we developed software to support its automatic functions. The genome releaser facilitates the extraction of PCR-amplifiable genomic DNA from 96 samples within minutes, eliminates the need for extraction buffers, and is adaptable to a wide range of microorganisms and cells. This versatility could significantly advance biomanufacturing processes.

42 ENGINEERING↗

A standards perspective on genomic data reusability and reproducibility

Genomic and metagenomic sequence data provides an unprecedented ability to re-examine findings, offering a transformative potential for advancing research, developing computational tools, enhancing clinical applications, and fostering scientific collaboration. However, effective and ethical reuse of genomics data is hampered by numerous technical and social challenges. The International Microbiome and Multi’Omics Standards Alliance (IMMSA, https://www.microbialstandards.org/) and the Genomic Standards Consortium (GSC, https://gensc.org) hosted a 5-part seminar series “A Year of Data Reuse” in 2024 to explore challenges and opportunities of data reuse and reproducibility across disparate domains of the genomic sciences. Addressing these challenges will require a multifaceted approach, including common metadata reporting, clear communication, standardized protocols, improved data management infrastructure, ethical guidelines, and collaborative policies that prioritize transparency and accessibility. We offer strategies to enable responsible and technically feasible data reuse, recognition of data reproducibility challenges, and emphasizing the importance of cross-disciplinary efforts in the pursuit of open science and data-driven innovation.

59 BASIC BIOLOGICAL SCIENCES↗

Genes encoding calmodulin-binding proteins in the Arabidopsis genome

Analysis of the recently completed Arabidopsis genome sequence indicates that approximately 31% of the predicted genes could not be assigned to functional categories, as they do not show any sequence similarity with proteins of known function from other organisms. Calmodulin (CaM), a ubiquitous and multifunctional Ca(2+) sensor, interacts with a wide variety of cellular proteins and modulates their activity/function in regulating diverse cellular processes. However, the primary amino acid sequence of the CaM-binding domain in different CaM-binding proteins (CBPs) is not conserved. One way to identify most of the CBPs in the Arabidopsis genome is by protein-protein interaction-based screening of expression libraries with CaM. Here, using a mixture of radiolabeled CaM isoforms from Arabidopsis, we screened several expression libraries prepared from flower meristem, seedlings, or tissues treated with hormones, an elicitor, or a pathogen. Sequence analysis of 77 positive clones that interact with CaM in a Ca(2+)-dependent manner revealed 20 CBPs, including 14 previously unknown CBPs. In addition, by searching the Arabidopsis genome sequence with the newly identified and known plant or animal CBPs, we identified a total of 27 CBPs. Among these, 16 CBPs are represented by families with 2-20 members in each family. Gene expression analysis revealed that CBPs and CBP paralogs are expressed differentially. Our data suggest that Arabidopsis has a large number of CBPs including several plant-specific ones. Although CaM is highly conserved between plants and animals, only a few CBPs are common to both plants and animals. Analysis of Arabidopsis CBPs revealed the presence of a variety of interesting domains. Our analyses identified several hypothetical proteins in the Arabidopsis genome as CaM targets, suggesting their involvement in Ca(2+)-mediated signaling networks.

NASA Discipline Plant Biology↗

Evolution, language and analogy in functional genomics

Almost a century ago, Wittgenstein pointed out that theory in science is intricately connected to language. This connection is not a frequent topic in the genomics literature. But a case can be made that functional genomics is today hindered by the paradoxes that Wittgenstein identified. If this is true, until these paradoxes are recognized and addressed, functional genomics will continue to be limited in its ability to extrapolate information from genomic sequences.

Evolution, Molecular↗

Identification of transcribed sequences in Arabidopsis thaliana by using high-resolution genome tiling arrays

Using a maskless photolithography method, we produced DNA oligonucleotide microarrays with probe sequences tiled throughout the genome of the plant Arabidopsis thaliana. RNA expression was determined for the complete nuclear, mitochondrial, and chloroplast genomes by tiling 5 million 36-mer probes. These probes were hybridized to labeled mRNA isolated from liquid grown T87 cells, an undifferentiated Arabidopsis cell culture line. Transcripts were detected from at least 60% of the nearly 26,330 annotated genes, which included 151 predicted genes that were not identified previously by a similar genome-wide hybridization study on four different cell lines. In comparison with previously published results with 25-mer tiling arrays produced by chromium masking-based photolithography technique, 36-mer oligonucleotide probes were found to be more useful in identifying intron-exon boundaries. Using two-dimensional HPLC tandem mass spectrometry, a small-scale proteomic analysis was performed with the same cells. A large amount of strongly hybridizing RNA was found in regions "antisense" to known genes. Similarity of antisense activities between the 25-mer and 36-mer data sets suggests that it is a reproducible and inherent property of the experiments. Transcription activities were also detected for many of the intergenic regions and the small RNAs, including tRNA, small nuclear RNA, small nucleolar RNA, and microRNA. Expression of tRNAs correlates with genome-wide amino acid usage.

Arabidopsis/genetics↗

A genomic view of Earth’s biomes

Microorganisms are essential to all life on Earth through critical roles in key biological processes and diverse interactions with other organisms that shape ecosystems, drive biogeochemical cycles and influence both human health and environmental health. High-throughput sequencing from environmental samples has revolutionized the understanding of microbial diversity and functions. With vast amounts of genomes now available across Earth’s biomes, these data provide a blueprint of microbial life that can be harnessed for a more holistic understanding of microbiome structure and function across the various ecosystems on Earth. Here we review the application of genome-centric approaches, including recent advances in single-cell sequencing and functional profiling, to survey microbial and viral diversity. Furthermore, we highlight some of the most impactful evolutionary and functional discoveries, explore the spatial diversity and temporal dynamics of microorganisms across diverse environments, and discuss genome-enabled insights into host-associated microorganisms.

Ecology↗

Soybean genomics research community strategic plan: A vision for 2024–2028

Abstract This strategic plan summarizes the major accomplishments achieved in the last quinquennial by the soybean [Glycine max(L.) Merr.] genetics and genomics research community and outlines key priorities for the next 5 years (2024–2028). This work is the result of deliberations among over 50 soybean researchers during a 2‐day workshop in St Louis, MO, USA, at the end of 2022. The plan is divided into seven traditional areas/disciplines: Breeding, Biotic Interactions, Physiology and Abiotic Stress, Functional Genomics, Biotechnology, Genomic Resources and Datasets, and Computational Resources. One additional section was added, Training the Next Generation of Soybean Researchers, when it was identified as a pressing issue during the workshop. This installment of the soybean genomics strategic plan provides a snapshot of recent progress while looking at future goals that will improve resources and enable innovation among the community of basic and applied soybean researchers. We hope that this work will inform our community and increase support for soybean research.

Genetics & Heredity↗

Rapid DNA unwinding accelerates genome editing by engineered CRISPR-Cas9

Thermostable clustered regularly interspaced short palindromic repeats (CRISPR) and CRISPR-associated (Cas9) enzymes could improve genome-editing efficiency and delivery due to extended protein lifetimes. However, initial experimentation demonstrated Geobacillus stearothermophilus Cas9 (GeoCas9) to be virtually inactive when used in cultured human cells. Laboratory-evolved variants of GeoCas9 overcome this natural limitation by acquiring mutations in the wedge (WED) domain that produce >100-fold-higher genome-editing levels. Cryoelectron microscopy (cryo-EM) structures of the wild-type and improved GeoCas9 (iGeoCas9) enzymes reveal extended contacts between the WED domain of iGeoCas9 and DNA substrates. Biochemical analysis shows that iGeoCas9 accelerates DNA unwinding to capture substrates under the magnesium-restricted conditions typical of mammalian but not bacterial cells. These findings enabled rational engineering of other Cas9 orthologs to enhance genome-editing levels, pointing to a general strategy for editing enzyme improvement. Together, these results uncover a new role for the Cas9 WED domain in DNA unwinding and demonstrate how accelerated target unwinding dramatically improves Cas9-induced genome-editing activity.

59 BASIC BIOLOGICAL SCIENCES↗

The Elements of Life, Photosynthesis and Genomics

I am a Professor of Biochemistry, Biophysics and Structural Biology and Plant and Microbial Biology at the University of California in Berkeley. I was born and raised in India, emigrated to the United States to attend university, earning a B.S. in Molecular Biology and a Ph.D. in Biochemistry at the University of Wisconsin in Madison. Following post-doctoral studies with Lawrence Bogorad at Harvard University where I became interested in genetic control of trace element quotas, I joined the department of Chemistry and Biochemistry at UCLA. One of the first to appreciate essential trace metals as potential regulators of gene expression, I articulated the details of the nutritional Cu regulon in Chlamydomonas. In parallel, I used genetic approaches to discover the genes governing missing steps in tetrapyrrole metabolism, including the attachment of heme to apocytochromes in the thylakoid lumen and the factors catalyzing the formation of ring V in chlorophyll. After biochemistry and classical genetics, I embraced genomics, taking a leadership role on the Joint Genome Institute’s efforts on the Chlamydomonas genome and more recently, contributing to high quality assemblies of several genomes in the green algal radiation, and large transcriptomic and proteomic datasets — focusing on the diel metabolic cycle in synchronized cultures and acclimation to key environmental and nutritional stressors — that are well-used and appreciated by the community. Finally, a new venture in Berkeley is the promotion of Auxenochlorella protothecoides as the true “green yeast” and as a platform for engineering algae to produce useful bioproducts.

59 BASIC BIOLOGICAL SCIENCES↗

Structure of an RNA G-quadruplex from the West Nile virus genome

Potential G-quadruplex sites have been identified in the genomes of DNA and RNA viruses and proposed as regulatory elements. The genus Orthoflavivirus contains arthropod-transmitted, positive-sense, single-stranded RNA viruses that cause significant human disease globally. Computational studies have identified multiple potential G-quadruplex sites that are conserved across members of this genus. Subsequent biophysical studies established that some G-quadruplexes predicted in Zika and tickborne encephalitis virus genomes can form and known quadruplex binders reduced viral yields from cells infected with these viruses. The susceptibility of RNA to degradation and the variability of loop regions have made structure determination challenging. Despite these difficulties, we report a high-resolution structure of the NS5-B quadruplex from the West Nile virus genome. Analysis reveals two stacked tetrads that are further stabilized by a stacked triad and transient noncanonical base pairing. This structure expands the landscape of solved RNA quadruplex structures and demonstrates the diversity and complexity of biological quadruplexes. We anticipate that the availability of this structure will assist in solving further viral RNA quadruplexes and provides a model for a conserved antiviral target in Orthoflavivirus genomes.

60 APPLIED LIFE SCIENCES↗

Genome-resolved biogeography of Phaeocystales, cosmopolitan bloom-forming algae

Phaeocystales, comprising the genus Phaeocystis and an uncharacterized sister lineage, are nanoplanktonic haptophytes widespread in the global ocean. Several species form mucilaginous colonies and influence key biogeochemical cycles, yet their underlying diversity and ecological strategies remain underexplored. Here, we present new genomic data from 13 strains, including three high-quality reference genomes (N50 > 30 kbp), and integrate previous metagenome-assembled genomes to resolve a robust phylogeny. Divergence timing of P. antarctica aligns with Miocene cooling and Southern Ocean isolation. Genomic traits reveal metabolic flexibility, including mixotrophic nitrogen acquisition in temperate waters and gene expansions linked to polar nutrient adaptation. Concordantly, transcriptomic comparisons between temperate and polar Phaeocystis suggest Southern Ocean populations experience iron and B12 limitation. We also identify signatures of horizontal gene transfer and endogenous giant virus/virophage insertions. Together, these findings highlight Phaeocystales as an ecologically versatile and geographically widespread lineage shaped by evolutionary innovation and adaptation to contrasting environmental stressors.

Füssy, Zoltán↗

Genome resources for three modern cotton lines guide future breeding efforts

Cotton ( Gossypium hirsutum L.) is the key renewable fibre crop worldwide, yet its yield and fibre quality show high variability due to genotype-specific traits and complex interactions among cultivars, management practices and environmental factors. Modern breeding practices may limit future yield gains due to a narrow founding gene pool. Precision breeding and biotechnological approaches offer potential solutions, contingent on accurate cultivar-specific data. Here we address this need by generating high-quality reference genomes for three modern cotton cultivars (‘UGA230’, ‘UA48’ and ‘CSX8308’) and updating the ‘TM-1’ cotton genetic standard reference. Despite hypothesized genetic uniformity, considerable sequence and structural variation was observed among the four genomes, which overlap with ancient and ongoing genomic introgressions from ‘Pima’ cotton, gene regulatory mechanisms and phenotypic trait divergence. Differentially expressed genes across fibre development correlate with fibre production, potentially contributing to the distinctive fibre quality traits observed in modern cotton cultivars. These genomes and comparative analyses provide a valuable foundation for future genetic endeavours to enhance global cotton yield and sustainability.

59 BASIC BIOLOGICAL SCIENCES↗

Identifying genomic data use with the Data Citation Explorer

Increases in sequencing capacity, combined with rapid accumulation of publications and associated data resources, have increased the complexity of maintaining associations between literature and genomic data. As the volume of literature and data have exceeded the capacity of manual curation, automated approaches to maintaining and confirming associations among these resources have become necessary. Here we present the Data Citation Explorer (DCE), which discovers literature incorporating genomic data that was not formally cited. This service provides advantages over manual curation methods including consistent resource coverage, metadata enrichment, documentation of new use cases, and identification of conflicting metadata. The service reduces labor costs associated with manual review, improves the quality of genome metadata maintained by the U.S. Department of Energy Joint Genome Institute (JGI), and increases the number of known publications that incorporate its data products. The DCE facilitates an understanding of JGI impact, improves credit attribution for data generators, and can encourage data sharing by allowing scientists to see how reuse amplifies the impact of their original studies.

59 BASIC BIOLOGICAL SCIENCES↗

GenomeDepot: data management system for microbial comparative genomics

Summary GenomeDepot is an open-source web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of websites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, Basic Local Alignment Search Tool (BLAST) search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools. Availability and implementation GenomeDepot is open source and distributed under the GNU General Public License via GitHub (https://github.com/aekazakov/genome-depot). GenomeDepot is implemented in Python and was tested in Ubuntu Linux. Full installation instructions and documentation are available at https://aekazakov.github.io/genome-depot/. GenomeDepot demo server is freely accessible at https://iseq.lbl.gov/demogd/.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

The genome of the polyextremophilic yeast, Naganishia friedmannii, reveals adaptations involved in stress response pathways, carbohydrate metabolism expansion, and a limited DNA repair repertoire

Here we report the draft genome sequence of Naganishia friedmannii (formerly Cryptococcus friedmannii) isolate, a Basidiomycota yeast commonly found in some of the most extreme environments of the Earth's cryosphere. We isolated N. friedmannii strain Llullensis from soils at 6000 m above sea level on Volcán Llullaillaco, Argentina. The genome was 22.2 Mb with 6251 identified protein coding genes. Proteins known to be associated with thermal, osmotic, and radiation stress were identified in the genome. Comparative analysis with seven other Naganishia genomes revealed unique features underlying its polyextremophilic lifestyle. Naganishia friedmannii showed an expansion of genes involved in breaking down plant-derived carbohydrates, supporting the hypothesis that it survives at high elevations by metabolizing wind-deposited organic matter. Surprisingly, many genes involved in cell-cycle checkpoints and DNA repair were missing, as in several other Naganishia species. This extensive loss may be adaptive in extreme environments prone to abiotic stress, where a high mutation rate could generate advantageous traits, and reduced cell-cycle control may allow for faster reproduction that would be advantageous for rapid growth during brief periods of soil wetting following rare snow events.

Vimercati, Lara↗

Genomic factors shaping codon usage across the Saccharomycotina subphylum

Codon usage bias, or the unequal use of synonymous codons, is observed across genes, genomes, and between species. It has been implicated in many cellular functions, such as translation dynamics and transcript stability, but can also be shaped by neutral forces. We characterized codon usage across 1,154 strains from 1,051 species from the fungal subphylum Saccharomycotina to gain insight into the biases, molecular mechanisms, evolution, and genomic features contributing to codon usage patterns. We found a general preference for A/T-ending codons and correlations between codon usage bias, GC content, and tRNA-ome size. Codon usage bias is distinct between the 12 orders to such a degree that yeasts can be classified with an accuracy >90% using a machine learning algorithm. We also characterized the degree to which codon usage bias is impacted by translational selection. We found it was influenced by a combination of features, including the number of coding sequences, BUSCO count, and genome length. Our analysis also revealed an extreme bias in codon usage in the Saccharomycodales associated with a lack of predicted arginine tRNAs that decode CGN codons, leaving only the AGN codons to encode arginine. Analysis of Saccharomycodales gene expression, tRNA sequences, and codon evolution suggests that avoidance of the CGN codons is associated with a decline in arginine tRNA function. Consistent with previous findings, codon usage bias within the Saccharomycotina is shaped by genomic features and GC bias. However, we find cases of extreme codon usage preference and avoidance along yeast lineages, suggesting additional forces may be shaping the evolution of specific codons.

59 BASIC BIOLOGICAL SCIENCES↗