Search NASA⌕ Search

SEARCH · Search NASA

Results for “functional genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Data for A Fluorescence-Based Transient Expression Assay for the Analysis of Upstream Open Reading Frames in Plants

Scripts for the manuscript "A fluorescence-based transient expression assay for the analysis of upstream open reading frames in plant" by Haas et al. Upstream open reading frames (uORFs) are regulatory elements present in the 5′ leaders of mRNA that can significantly impact downstream gene expression in eukaryotes. In crop engineering, editing of uORFs can provide an avenue to upregulate expression of native genes without the need to add persistent transgenic copies. Even with genome- wide methods to identify translated uORFs such as ribosome profiling, their functional characterization depends on validation through reporter gene assays and mutagenesis studies. Current screening methods for plants use luciferases or protoplasts to measure differential gene expression between wild- type and mutated transcript leaders, which requires tissue processing and/or substrate addition. Here, we present a time- and cost- efficient alternative to investigate transcript leaders by co- expression of two fluorescent proteins in Nicotiana benthamiana leaf tissue and test our assay on genes involved in photoprotection, editing of which could provide a pathway to increase CO2 assimilation during sun–shade transitions.

Gene Editing↗

UnigeneFinder: An Automated Pipeline for Gene Calling From Transcriptome Assemblies Without a Reference Genome

ABSTRACT For most species, transcriptome data are much more readily available than genome data. Without a reference genome, gene calling is cumbersome and inaccurate because of the high degree of redundancy in de novo transcriptome assemblies. To simplify and increase the accuracy of de novo transcriptome assembly in the absence of a reference genome, we developed UnigeneFinder. Combining several clustering methods, UnigeneFinder substantially reduces the redundancy typical of raw transcriptome assemblies. This pipeline offers an effective solution to the problem of inflated transcript numbers, achieving a closer representation of the actual underlying genome. UnigeneFinder performs comparably or better, compared with existing tools, on plant species with varying genome complexities. UnigeneFinder is the only available transcriptome redundancy solution that fully automates the generation of primary transcript, coding region, and protein sequences, analogous to those available for high‐quality reference genomes. These features, coupled with the pipeline’s cross‐platform implementation, focus on automation, and an accessible, user‐friendly interface, make UnigeneFinder a useful tool for many downstream sequence‐based analyses in nonmodel organisms lacking a reference genome, including differential gene expression analysis, accurate ortholog identification, functional enrichments, and evolutionary analyses. UnigeneFinder also runs efficiently both on high‐performance computing (HPC) systems and personal computers, further reducing barriers to use.

Xue, Bo [Plant Resilience Institute Michigan State↗

Tropical intertidal microbiome response to the 2024 Marine Honour oil spill

Marine fuel oil (MFO) spills in tropical coastal environments are under-characterized despite increasing risk from maritime activities. Microbial and geochemical responses to the June 2024 Marine Honour MFO spill on Singapore's intertidal sediments were analyzed in real time over 185 days. Using metagenomics and hydrocarbon profiling, microbial community shifts and hydrocarbon degradation were quantified across visibly oiled (high-impact) and clean (low-impact) sites. Microbiomes at all sites adapted rapidly to the spill through increased diversity and abundance of genes encoding alkane and aromatic compound degradation, detoxification, and biosurfactant production. The dominant hydrocarbon-degrading bacteria differed markedly from those reported in other crude oil spills and in regions with different climates. Oil deposition intensity strongly influenced microbial succession and hydrocarbon-degrading gene profiles, and this reflected early toxicity constraints in heavily oiled areas. The persistence of hydrocarbon degradation genes beyond hydrocarbon detection in sediments suggested long-term functional priming may occur. The study provides novel genome-resolved insight into the microbial response to MFO pollution, advances understanding of marine environmental biodegradation, and provides urgently needed baseline data for oil spill response strategies in Southeast Asia and beyond.

Coastal pollution↗

Secure biosystems design in Saccharomyces cerevisiae establishes effective biocontainment strategies and mechanisms of escape

The widespread application of recombinant DNA and synthetic biology approaches for microbial metabolic engineering pursuits has motivated the development of biocontainment strategies, targeting safe and secure deployment of genetically modified microorganisms (GMMs). However, the design rules and mechanistic drivers governing biocontainment efficacy, as well as impacts of biocontainment upon microbial fitness, remain to be comprehensively evaluated, hindering predictive design and application of these strategies. We have developed a platform for high-resolution analysis of a transactivated kill switch in laboratory and industrial strains of Saccharomyces cerevisiae to assess modes of biocontainment escape and establish design rules for development of kill switch systems in diverse microbes. A camphor-regulated, RelE toxin system was systematically deployed to assess the impacts of differential kill switch copy number and ploidy in laboratory vs industrial strains. CRISPR-mediated integration of the biocontainment system at various loci revealed rapid escape events driven, in part, by mutations to both the Cam-transactivator (cam-TA) and RelE toxin. Genetic engineering enabled recapitulation of escape phenotypes, confirming mechanisms of escape and establishing structure-function relationships in the cam-TA system. Interestingly, genomic resequencing of escape mutants also revealed a series of off-target mutations, implicating additional modes of kill switch escape. Multi-copy integration of the kill switch system mitigated these effects by orders of magnitude, without compromising the biosynthetic capacity of the microbes, but proved insufficient to establish sustained biocontainment. The resultant data define a series of key design rules for next-generation biocontainment strategies and add to a growing foundational knowledge base targeting establishment of secure biosystems designs.

59 BASIC BIOLOGICAL SCIENCES↗

Assessment of the Effect of Deleting the African Swine Fever Virus Gene R298L on Virus Replication and Virulence of the Georgia2010 Isolate

African swine fever (ASF) is a lethal disease of domestic pigs that is currently challenging swine production in large areas of Eurasia. The causative agent, ASF virus (ASFV), is a large, double-stranded and structurally complex virus. The ASFV genome encodes for more than 160 proteins; however, the functions of most of these proteins are still in the process of being characterized. The ASF gene R298L, which has previously been characterized as able to encode a functional serine protein kinase, is expressed late in the virus infection cycle and may be part of the virus particle. There is no description of the importance of the R298L gene in basic virus functions such as replication or virulence in the natural host. Based on its evolution, it is proposed that there are four different phenotypes of R298L of ASFV in nature, which may have potential implications for R298L functionality. We report here that a recombinant virus lacking the R298L gene in the Georgia 2010 isolate, ASFV-G-∆R298L, does not exhibit significant changes in its replication in primary cultures of swine macrophages. In addition, when experimentally inoculated in pigs, ASFV-G-∆R298L induced a fatal form of the disease similar to that caused by the parental virulent ASFV-G. Therefore, deletion of R298L does not significantly affect virus replication and virulence in domestic pigs of the ASFV Georgia 2010 isolate.

Virology↗

Deletion of the African Swine Fever Virus Gen I196L in the Georgia2010 Isolate Genome Does Not Affect Virus Replication or Virulence in Domestic Pigs

African swine fever (ASF) is a lethal disease of domestic pigs that is currently challenging swine production in large areas of Eurasia and the Caribbean. The causative agent, ASF virus (ASFV), is a large, double-stranded, and structurally complex virus. The ASFV genome encodes for more than 160 proteins; however, the functions of most of them are still in the process of being characterized. Recently, ASFV gene I196L has been reported as being critically involved in disease production in domestic pigs. We report here that a recombinant virus derived from the Georgia 2010 isolate (ASFV-G) lacking the I196L gene, ASFV-G-∆I196L, had the same ability to replicate in primary cultures of swine macrophage and, when experimentally inoculated in pigs, produced a fatal form of the disease similar to that caused by the parental virulent ASFV-G. Therefore, deletion of the I196L gene does not significantly affect virus replication and virulence in domestic pigs of the ASFV Georgia 2010 isolate.

Virology↗

Genetic dissection of cardiac growth control pathways

Cardiac muscle cells exhibit two related but distinct modes of growth that are highly regulated during development and disease. Cardiac myocytes rapidly proliferate during fetal life but exit the cell cycle irreversibly soon after birth, following which the predominant form of growth shifts from hyperplastic to hypertrophic. Much research has focused on identifying the candidate mitogens, hypertrophic agonists, and signaling pathways that mediate these processes in isolated cells. What drives the proliferative growth of embryonic myocardium in vivo and the mechanisms by which adult cardiac myocytes hypertrophy in vivo are less clear. Efforts to answer these questions have benefited from rapid progress made in techniques to manipulate the murine genome. Complementary technologies for gain- and loss-of-function now permit a mutational analysis of these growth control pathways in vivo in the intact heart. These studies have confirmed the importance of suspected pathways, have implicated unexpected pathways as well, and have led to new paradigms for the control of cardiac growth.

NASA Program Biomedical Research and Countermeasur↗

Analysis of the myosins encoded in the recently completed Arabidopsis thaliana genome sequence

BACKGROUND: Three types of molecular motors play an important role in the organization, dynamics and transport processes associated with the cytoskeleton. The myosin family of molecular motors move cargo on actin filaments, whereas kinesin and dynein motors move cargo along microtubules. These motors have been highly characterized in non-plant systems and information is becoming available about plant motors. The actin cytoskeleton in plants has been shown to be involved in processes such as transportation, signaling, cell division, cytoplasmic streaming and morphogenesis. The role of myosin in these processes has been established in a few cases but many questions remain to be answered about the number, types and roles of myosins in plants. RESULTS: Using the motor domain of an Arabidopsis myosin we identified 17 myosin sequences in the Arabidopsis genome. Phylogenetic analysis of the Arabidopsis myosins with non-plant and plant myosins revealed that all the Arabidopsis myosins and other plant myosins fall into two groups - class VIII and class XI. These groups contain exclusively plant or algal myosins with no animal or fungal myosins. Exon/intron data suggest that the myosins are highly conserved and that some may be a result of gene duplication. CONCLUSIONS: Plant myosins are unlike myosins from any other organisms except algae. As a percentage of the total gene number, the number of myosins is small overall in Arabidopsis compared with the other sequenced eukaryotic genomes. There are, however, a large number of class XI myosins. The function of each myosin has yet to be determined.

NASA Discipline Plant Biology↗

Metagenome-assembled genomes from topsoils collected during NEON campaign in East River, CO (06/14/2018-06/28/2018)

The Watershed Function Science Focus Area (WF SFA) at Lawrence Berkeley National Lab is working to build a mechanistic understanding of the distribution and dynamics of biogeochemical processes in mountainous watersheds and their response to perturbation. In June 2018, the NEON (National Ecological Observatory Network) Airborne Observatory Platform (AOP) performed a taskable airborne imaging campaign to collect visible to shortwave infrared (VSWIR) imaging spectroscopy and LiDAR data across 330 km2 in the Upper East River at Crested Butte, CO. We conducted a parallel ground sampling campaign to sample vegetation traits, as well as soil physical, chemical, and microbiological characteristics. We collected these samples from 438 sites across 12 locations spanning much of the elevation, topographic, and geologic variability across the study area. A subset of 250 samples were used for soil metagenomics which is presented here. In addition, at each site, vegetation samples were collected to measure species-specific leaf water content and leaf mass area, foliar elemental composition and foliar CN stable isotope ratios. Soil samples were collected to measure soil physical properties which include bulk density and soil texture analysis. A suite of soil chemical properties was measured from the samples collected at each site, including pH, organic matter, concentrations exchangeable cations, total elemental composition, and the concentrations of extractable N pools (e.g. total free amino acids, ammonium, nitrate, dissolved organic N, and total dissolved N). Additionally, we have measured soil microbial biomass CN stoichiometry. Here, we present 1982 metagenome-assembled genomes (MAGs) for the bacterial and archaeal community from topsoil collected from during NEON 2018 campaign. All metagenomes were sequenced at JGI (Joint Genome Institute) (GOLD Study ID: Gs0149986). Metagenomes were assembled using JGI Metagenome Workflow (10.1128/mSystems.00804-20). The dataset includes (1) zip files for 1982 MAG fasta files (neon_genomes1-5.tar.gz, split into 5 tarballs to keep tarballs under 0.5 GB), (2) neon_Gs0149986_samples_soilproperties_metagenomes.csv: the sample information together with the accession numbers for the underlying metagenomes and the associated soil physical and chemical measurements in NMDC (National Microbiome Data Collaborative) compliant format, (3) neon_Gs0149986.kml: location bounding box file for the sampled locations, (4) samples.csv: sample metadata file used to register Internationall Generic Sample Numbers (IGSNs), (5) flmd.csv: file level metadata file, and (6) dd.csv: data dictionary file. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

CAZyme domain architectures suggest fine-scale functional differentiation among anaerobic fungi and bacteria during lignocellulose conversion to volatile fatty acids

Anaerobic fermentation with microbial communities (microbiomes) is an emerging platform for conversion of lignocellulosic biomass to biofuels and bioproducts. The process relies on diverse anaerobic microbes that interact to deconstruct and convert lignocellulosic biomass into a range of products, such as volatile fatty acids (VFAs), which can be achieved by arresting methanogenesis during fermentation. However, defining the distinct functional roles played by various fungi and bacteria during anaerobic biodegradation remains poorly understood. Here, we performed parallel enrichment experiments from cow faeces, goat faeces, and anaerobic digester sludge, selecting for fungal or bacterial dominated communities that convert sorghum biomass into VFAs. Subsequently we reconstructed metabolic networks across these enrichments based on recovered bacterial metagenome-assembled genomes (MAGs) and fungal isolate genomes and profiled their metabolic activity using metatranscriptomics to identify potential functional niches. Our findings implicate diverse bacteria affiliated with the Bacteroidales and Lachnospiraceae in the direct conversion of lignocellulosic biomass to propionate and butyrate, respectively, whereas Neocallimastix-dominated fungal enrichments converted lignocellulose to lactate, acetate and formate. Analysis of carbohydrate-active enzymes (CAZymes) revealed fine-scale differences between microbes that expressed unique multi-functional enzymes linking two or more CAZymes together with distinct carbohydrate binding motifs, implicating lignocellulose structure as a key driver of selection and niche differentiation. Most of these multi-functional enzymes localized complementary degradation functions together, likely conferring synergistic degradation effects within and between microbiome members. We anticipate that these findings will help inform efforts to develop synthetic microbiomes with tailored functionality for low-cost conversion of lignocellulosic biomass to fuels and bio-based chemicals.

Lawson, Christopher E [University of Toronto;]↗

Exploring life’s hidden majority: microbial dark matter symposium highlights

The Microbial Dark Matter Symposium held on August 28–29, 2025, in Laguna Beach, Orange County, CA, convened a multidisciplinary group of scientists to address the vast unknowns in microbial life—from uncultured taxa and uncharacterized proteins to elusive viruses and spacefaring microbes. Set against a scenic coastal backdrop, the symposium highlighted advances in single-cell genomics, proximity ligation sequencing, and artificial intelligence-ready bioinformatics, while also probing the limits of microbial persistence, metabolism, and ecological distribution. Sessions explored microbial dark matter from multiple dimensions: cultivability, where new strategies are enabling recovery of elusive microbes; functional ambiguity, where metagenomic dark zones are illuminated by computational annotation; and genomic representation, where single-cell methods bridge gaps left by shotgun community sequencing. Researchers shared breakthroughs in identifying atmospheric microbiomes, “dark oxygen” production in groundwater ecosystems, and microbial survival on the International Space Station. The symposium emphasized integration of methods, disciplines, and ecosystems, advancing a collective push to illuminate the microbial dark matter on Earth and beyond. By highlighting emerging tools, pressing questions, and cross-domain insights, the symposium underscored the need for collaborative, open, and adaptive approaches to study the microbial unknown. The meeting marks a pivotal moment in microbiology, where cultivating knowledge of the uncultivated promises transformative understanding of life, everywhere.

Podar, Mircea [ORNL] (ORCID:0000000327760205)↗

Human RNome Project draft human RNome sequence of GM12878, B-cell line, obtained by mass-spectrometry sequencing, long-read sequencing and short-read sequencing.

Here we report the first draft of the human RNome sequence, a reference map of RNA chemical modifications in a human B-cell line. RNA carries a diverse repertoire of chemical modifications that regulate gene expression, cellular function, and responses to physiological and pathological cues. Yet, unlike the genome, no reference map of RNA modifications is available for any human cell. To generate this resource, the Human RNome Project Consortium analyzed a shared RNA preparation from the well-characterized GM12878 B-cell line using short-read sequencing, long-read direct RNA sequencing, and mass spectrometry, generating more than 7.1 billion sequencing reads spanning approximately 1.2 trillion nucleotides. The resulting maps of the human RNome reveal that RNA modifications are organized according to function, transcript architecture, and cellular identity. Modifications concentrate at functional centers of ribosomal and transfer RNAs, follow the canonical topology of N6-methyladenosine in coding transcripts, and form coordinated hotspots in immune regulatory genes. This first reference human RNome provides a foundation for understanding how RNA chemistry shapes cellular identity, human disease, and the development of RNA-based therapeutics.

59 BASIC BIOLOGICAL SCIENCES↗

Identification of functional non-coding variants associated with orofacial cleft

Oral facial cleft (OFC) comprises cleft lip with or without cleft palate (CL/P) or cleft palate only. Genome wide association studies (GWAS) of isolated OFC have identified common single nucleotide polymorphisms (SNPs) in many genomic loci where the presumed effector gene (for example, IRF6 in the 1q32 locus) is expressed in embryonic oral epithelium. To identify candidates for functional SNPs at eight such loci we conduct a massively parallel reporter assay in a fetal oral epithelial cell line, revealing SNPs with allele-specific effects on enhancer activity. We filter these SNPs against chromatin-mark evidence of enhancers and test a subset in traditional reporter assays, which support the candidacy of SNPs at loci containing FOXE1, IRF6, MAFB, TFAP2A, and TP63. For two SNPs near IRF6 and one near FOXE1, we engineer the genome of induced pluripotent stem cells, differentiate the cells into embryonic oral epithelium, and discover allele-specific effects on the levels of effector gene expression, and, in two cases, the binding affinity of transcription factors FOXE1 or ETS2. Conditional analyses of GWAS data suggest the two functional SNPs near IRF6 account for the majority of risk for CL/P at this locus. This study connects genetic variation associated with OFC to mechanisms of pathogenesis.

Kumari, Priyanka↗

Depth-dependent Metagenome-Assembled Genomes of Agricultural Soils under Managed Aquifer Recharge

Abstract Managed Aquifer Recharge (MAR) systems, which intentionally replenish groundwater aquifers with excess water, are critical for addressing water scarcity exacerbated by demographic shifts and climate variability. To date, little is known about the functional diversity of the soil microbiome at different soil depth inhabiting agricultural soils used for MAR. Knowing the functional diversity is pivotal in regulating nutrient cycling and maintaining soil health. Metagenomics, particularly Metagenome-Assembled Genomes (MAGs), provide a powerful tool to explore the diversity of uncultivated soil microbes, facilitating in-depth investigations into microbial functions. In a field experiment conducted in a California vineyard, we sequenced soil DNA before and after water application of MAR. Through this process, we assembled 146 medium and 14 high-quality MAGs, uncovering a wide array of archaeal and bacterial taxa across different soil depths. These findings advance our understanding of the microbial ecology and functional diversity of soils used for MAR, contributing to the development of more informed and sustainable land management strategies.

Science & Technology - Other Topics↗

Probing interspecies metabolic interactions within a synthetic binary microbiome using genome-scale modeling

Metabolic interactions within a microbial community play a key role in determining the structure, function, and composition of the community. However, due to the complexity and intractability of natural microbiomes, limited knowledge is available on interspecies interactions within a community. In this work, using a binary synthetic microbiome, a methanotroph-photoautotroph (M-P) coculture, as the model system, we examined different genome-scale metabolic modeling (GEM) approaches to gain a better understanding of the metabolic interactions within the coculture, how they contribute to the enhanced growth observed in the coculture, and how they evolve over time. Using batch growth data of the model M-P coculture, we compared three GEM approaches for microbial communities. Two of the methods are existing approaches: SteadyCom, a steady state GEM, and dynamic flux balance analysis (DFBA) Lab, a dynamic GEM. We also proposed an improved dynamic GEM approach, DynamiCom, for the M-P coculture. SteadyCom can predict the metabolic interactions within the coculture but not their dynamic evolutions; DFBA Lab can predict the dynamics of the coculture but cannot identify interspecies interactions. DynamiCom was able to identify the cross-fed metabolite within the coculture, as well as predict the evolution of the interspecies interactions over time. A new dynamic GEM approach, DynamiCom, was developed for a model M-P coculture. Constrained by the predictions from a validated kinetic model, DynamiCom consistently predicted the top metabolites being exchanged in the M-P coculture, as well as the establishment of the mutualistic N-exchange between the methanotroph and cyanobacteria. The interspecies interactions and their dynamic evolution predicted by DynamiCom are supported by ample evidence in the literature on methanotroph, cyanobacteria, and other cyanobacteria-heterotroph cocultures.

59 BASIC BIOLOGICAL SCIENCES↗

Computer Modeling of Protocellular Functions: Peptide Insertion in Membranes

Lipid vesicles became the precursors to protocells by acquiring the capabilities needed to survive and reproduce. These include transport of ions, nutrients and waste products across cell walls and capture of energy and its conversion into a chemically usable form. In modem organisms these functions are carried out by membrane-bound proteins (about 30% of the genome codes for this kind of proteins). A number of properties of alpha-helical peptides suggest that their associations are excellent candidates for protobiological precursors of proteins. In particular, some simple a-helical peptides can aggregate spontaneously and form functional channels. This process can be described conceptually by a three-step thermodynamic cycle: 1 - folding of helices at the water-membrane interface, 2 - helix insertion into the lipid bilayer and 3 - specific interactions of these helices that result in functional tertiary structures. Although a crucial step, helix insertion has not been adequately studied because of the insolubility and aggregation of hydrophobic peptides. In this work, we use computer simulation methods (Molecular Dynamics) to characterize the energetics of helix insertion and we discuss its importance in an evolutionary context. Specifically, helices could self-assemble only if their interactions were sufficiently strong to compensate the unfavorable Free Energy of insertion of individual helices into membranes, providing a selection mechanism for protobiological evolution.

Rodriquez-Gomez, D.↗

Omics Research on the International Space Station

The International Space Station (ISS) is an orbiting laboratory whose goals include advancing science and technology research. Completion of ISS assembly ushered a new era focused on utilization, encompassing multiple disciplines such as Biology and Biotechnology, Physical Sciences, Technology Development and Demonstration, Human Research, Earth and Space Sciences, and Educational Activities. The research complement planned for upcoming ISS Expeditions 45&46 includes several investigations in the new field of omics, which aims to collectively characterize sets of biomolecules (e.g., genomic, epigenomic, transcriptomic, proteomic, and metabolomic products) that translate into organismic structure and function. For example, Multi‐Omics is a JAXA investigation that analyzes human microbial metabolic cross‐talk in the space ecosystem by evaluating data from immune dysregulation biomarkers, metabolic profiles, and microbiota composition. The NASA OsteoOmics investigation studies gravitational regulation of osteoblast genomics and metabolism. Tissue Regeneration uses pan‐omics approaches with cells cultured in bioreactors to characterize factors involved in mammalian bone tissue regeneration in microgravity. Rodent Research‐3 includes an experiment that implements pan‐omics to evaluate therapeutically significant molecular circuits, markers, and biomaterials associated with microgravity wound healing and tissue regeneration in bone defective rodents. The JAXA Mouse Epigenetics investigation examines molecular alterations in organ specific gene expression patterns and epigenetic modifications, and analyzes murine germ cell development during long term spaceflight. Lastly, Twins Study ("Differential effects of homozygous twin astronauts associated with differences in exposure to spaceflight factors"), NASA's first foray into human omics research, applies integrated analyses to assess biomolecular responses to physical, physiological, and environmental stressors associated with spaceflight.

Love, John↗

Leaky ribosomal scanning enables tunable translation of bicistronic ORFs in green algae

Advances in sequencing technology have unveiled examples of nucleus-encoded polycistrons, once considered rare. Exclusively polycistronic transcripts are prevalent in green algae, although the mechanism by which multiple polypeptides are translated from a single transcript is unknown. Here, we used bioinformatic and in vivo mutational analyses to evaluate competing mechanistic models for translation of bicistronic mRNAs in green algae. High-confidence manually curated datasets of bicistronic loci from two divergent green algae, Chlamydomonas reinhardtii and Auxenochlorella protothecoides, revealed a preference for weak Kozak-like sequences for ORF 1 and an underrepresentation of potential initiation codons before the ORF 2 start codon, which are suitable conditions for leaky ribosome scanning to allow ORF 2 translation. We used mutational analysis in A. protothecoides to test the mechanism. In vivo manipulation of the ORF 1 Kozak-like sequence and start codon altered reporter expression at ORF 2, with a weaker Kozak-like sequence enhancing expression and a stronger one diminishing it. A synthetic bicistronic dual reporter demonstrated inversely adjustable activity of green fluorescent protein expressed from ORF 1 and luciferase from ORF 2, depending on the strength of the ORF 1 Kozak-like sequence. Our findings demonstrate that translation of multiple ORFs in green algal bicistronic transcripts is consistent with episodic leaky scanning of ORF 1 to allow translation at ORF 2. This work has implications for the potential functionality of upstream open reading frames (uORFs) found across eukaryotic genomes and for transgene expression in synthetic biology applications.

59 BASIC BIOLOGICAL SCIENCES↗