Search NASA⌕ Search

SEARCH · Search NASA

Results for “Bioinformatics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Developing A Hybrid Spacesuit Simulator as A Research Tool for Assessing Extravehicular Activity Relevant Workload

Conducting human tests in a pressurized spacesuit is limited by availability, cost, and manpower; however, pressurized spacesuits are not always needed depending on the objectives of testing, including the development and testing of new informatics capabilities. The Human Physiology, Performance, Protection & Operations Laboratory (H-3PO) at NASA is developing a Hybrid Spacesuit Simulator (HS3) to support testing and characterization of human performance during analog planetary exploration extravehicular activities (EVAs). The goal of HS3 is to create a low-cost, modular, and unpressurized spacesuit simulator as a research tool that provides relevant physical and cognitive workload approximations with EVA-like immersion. HS3 consists of a soft outer suit, thermal control, gloves, boots, helmet, and integrated bioinformatics and communications. Baseline HS3 assessments were performed during 3-hour EVA simulations in two different subjects (DEMO1 and DEMO2) that included traverses at variable resistances and geological sampling activities. Liquid cooling garment (LCG) temperature, mean skin temperature, heart rate, motion capture, and metabolic rate were collected during each 3-hour simulated EVA. During DEMO1 and DEMO2, baseline metabolic rates at rest were 836 ± 327 BTU/hr and 869 ± 207 BTU/hr and increased to 2124 ± 548 BTU/hr and 2269 ± 559 BTU/hr, respectively, during 500m traverse. Average inlet LCG temperatures were 29.57 ± 6.62 °C and 25.63 ± 6.48 °C for DEMO1 and DEMO2 with increased outlet LCG temperatures of 33.53 ± 6.62 °C and 29.21 ± 4.79 °C, respectively. Overall, HS3 will enable future studies to characterize EVA tasks, human performance, and test future EVA capabilities in analog test environments without the need for pressurized suited environments.

Suit simulator↗

NASA GeneLab Multi-study Visualization Portal

NASA GeneLab has helped advance the field of Space Biology by providing a public repository where researchers can store, share, analyze and visualize the results of space flight related omics experiments. The GeneLab data visualization portal allows any user, regardless of bioinformatics knowledge or access to computational resources, to interact with the experimental data, draw their own conclusions, and gain insights about the effects of space on living systems. These tools help democratize scientific research and foster the NASA Open Science initiative. The new multi-study feature of the GeneLab visualization platform allows users to mine study metadata from RNA sequencing (RNA-seq) experiments to identify samples of interest by filtering datasets based on organism, tissue, assay technology type, and/or factor. Once samples are selected from multiple datasets, users can combine and normalize the sample data, then utilize the visualization displays, including Principal Component Analysis (PCA) plots, to assess sample distributions. Finally, users can perform differential gene expression analysis on the combined data and visualize the results through PCA plots, Volcano plots, Pair plots, Heatmap, Ideogram and Gene Set Enrichment Analysis. All user-generated results and visualizations will be available for download. Here, we present a biological study using samples from multiple GeneLab RNA-seq datasets and analyzed using the multi-study visualization platform to demonstrate inter- and intra-study variability, as well as commonly differentially expressed genes between spaceflight and ground control conditions across datasets. This new feature opens a wide range of possibilities and opportunities for further development including combining other assay technology types and integration with batch effect correction techniques and machine learning applications. Overall, this tool allows users to increase the statistical power of individual experiments, validate hypothesis, identify patterns, and opens the door to new and exciting research.

space biology↗

Earth Science Data Processing With Nextflow

Earth science data processing tasks present many challenges. These tasks often process large input datasets and require scores of CPU-hours to generate results. All but the simplest tasks will be decomposed into a series of computational or data manipulation steps, also known as a scientific workflow. In order to reduce the burden of orchestrating and running the dependent processing steps, a workflow execution engine is required. This poster describes the lessons learned by the CLARREO Pathfinder (CPF) team while developing multiple scientific workflows and utilizing the open-source Nextflow engine to execute them in a cloud computing environment. The Nextflow engine is designed with the following stated goals: first, the engine does not dictate how individual steps in the task are implemented (i.e. it is language and interface agnostic); second, the engine supports easy configuration and modularity at the workflow level so that others can easily execute our workflows to reproduce results; lastly, the engine eases development by transparently scaling execution from local to remote environments. Nextflow was developed for the bioinformatics domain but is a good fit for other scientific workflows where the overall task is well-described by a dataflow diagram. The CPF team has developed Nextflow pipelines (i.e. scientific workflows) to simulate CLARREO radiance, generate large look-up tables for inter-calibration algorithms, and generate L4 intercalibration data products. These pipelines consume from single-digits to hundreds of thousands of CPU-hours. In the development and evolution of these pipelines we have discovered many design patterns, pitfalls, and solutions to common problems. Our goal is to demonstrate important aspects of how to design, implement, run, and ultimately share Nextflow pipelines in the domain of Earth science.

Aron D Bartle↗

Does Collection Time Bias the Ecology of Cleanroom Air Samples?

Microbial monitoring of astromaterials collections has taken on increased importance with the return of biologically sensitive samples from the asteroids Ryugu and Bennu and the initiation of the Mars Sample Return Program. Terrestrial bacteria and fungi can alter the mineralogy and organic composition of our collections causing irreversible contamination of pristine samples and increasing the risk of false positives for life detection measurements. NASA has conducted routine microbial monitoring of its existing collections since 20181. Initial monitoring focused on surface samples collected with foam swabs. Although, airborne microbiology is often decoupled from surface microbiology in the built environment2 culture-based air sampling techniques like impactors were not compliant with existing contamination control requirements. Bringing organic rich media, gelatin or liquids into curation cleanrooms presents an unacceptable risk to pristine samples. In 2022 NASA purchased a materials complaint air sampler and began collecting air samples from the cleanrooms in addition to surface samples3. The new instrument uses an electret filter to collect samples that are suitable for cultivating organisms or for direct DNA sequencing. Preliminary DNA sequencing results appeared to indicate that longer sampling times biased the microbial community in favor of hearty, spore-forming bacteria3. We present the results of a study comparing overnight sampling (17 hours) to short (1 hour) sampling of unoccupied curation cleanrooms. The results will help us optimize our monitoring protocols and develop a more detailed inventory of the ecology of astromaterials curation cleanrooms. Methods: We analyzed 72 paired air samples from six different cleanrooms including the meteorite processing lab (ISO 7 equivalent, 16 samples), the lunar lab (ISO 6 equivalent, 10 samples), the stardust lab (ISO 5 equivalent 14 samples), the OSIRIS-REx lab (ISO 5 equivalent, 12 samples), the Hayabusa2 lab (ISO 5 equivalent, 14 samples), and the Genesis lab (ISO 4 equivalent, 6 samples). All the samples were collected with an InnovaPrep Bobcat air sampler operating at a sampling rate of 200 L/min. The sampler operates for 5 minutes out of every 20 minute period. Half of the samples were collected by filtering 3,000L (15 min. of active sampling) of air across an electret filter for one hour. The rest of the samples were collected by filtering approximately 51,000 L air across the filter overnight (~17 hours, 255 min. of active sampling). Cells were eluted from the filter using 6-7 ml of pressurized 0.15% tween 20 in PBS (phosphate buffered saline). This liquid was used to cultivate bacteria according to previously published methods1,4,5 and for DNA extraction and next generation sequencing. DNA was extracted with a Qiagen MagAttract PowerMicrobiome kit6. To identify bacteria and archaea, the 16S rRNA gene was amplified using Earth Microbiome primers for the V4 region 7. The amplified DNA was sequenced on an Illumina MiSeq using a V3 reagent kit. The resulting sequences were processed using DADA2 and QIIME2 as implemented on the EDGE bioinformatics platform8–10. Results: Only two of the 72 samples had no amplifiable DNA. Amplified DNA concentrations ranged from 2.67 – 0.272 ng/µl. The median concentration of amplified DNA for the 1 hour samples was 0.770 ± 0.368 ng/µl. The median concentration of amplified DNA for the overnight samples was 0.877 ± 0.434 ng/µl. On average the overnight samples had slightly more sequences (58,960 vs. 59,456) and ASV’s (amplicon sequence variants) (60 vs 64.5) than the one hour samples, but these differences are not statistically significant. The most abundant ASV in every sample mapped to the genus Cupravidus. ASV’s mapping to the genuses Bacillus, Schlegelella, Thermus, and Staphylococcus were also common. Discussion and Future Work: Alpha diversity statistics like Shannon Entropy and Faith Phylogenetic Diversity are used to describe the diversity of organisms in a single sample. If a longer sampling time was biasing the data, we would expect to see a change in these diversity statistics vs. sample time. However, we did not observe this in our data. The median Shannon entropy was slightly higher for the overnight samples (3.773 vs 3.611) as was the Faith Phylogenetic Diversity (4.042 vs 3.596), but both values were within a standard deviation of each other for the two sampling times (Fig. 1). It is unlikely, that the longer sampling time is introducing bias into our data. We do observe a significant decrease in diversity when comparing the air samples by lab. The Genesis lab (ISO 4 equivalent) has a lower median number of ASV’s (45.5) than the other labs (62). Median values for Shannon Entropy (3.717 vs. 3.430) and Faith Phylogenetic Diversity (3.796 vs. 3.548) are also lower for Genesis, but those values are with one standard deviation of each other for the different sampling times. This is consistent with previous culture-based results suggesting that the environment in cleanrooms tends to select for a core group of organisms capable of surviving under dry, low nutrient, conditions. The presence of the ASV’s mapping to Cupravidus and Thermus in our sequencing blanks and controls suggests that several of the most common organisms in our samples represent contaminants from the reagents used to perform the DNA extractions and sequencing. Further work is needed to identify these contaminants, remove them from our data and recalculate the diversity statistics. This is a systematic error. Therefore, we do not expect removing the sequencing contaminants to change our conclusions. Longer air sample collection times appear to result in slightly higher diversity and do not bias the results towards “hardy” bacteria like spore-formers. Based on these preliminary results we conclude that sampling at least 3,000 liters of air is sufficient to capture the microbial diversity of cleanrooms, and that air samples can also be collected overnight without negatively impacting diversity. These results allow us to be flexible when designing microbial monitoring plans so that they do not interfere with routine lab activity. References: 1. Regberg, A. B. et al. 49th Lunar and Planetary Science Conference (2018). 2. The United States Pharmacopeial Convention. USP General Chapter <1116> (2013). 3. Regberg, A. B., et al. 54th Lunar and Planetary Science Conference (2023). 4. Regberg, A. B. et al. 53rd Lunar and Planetary Science Conference ( 2022). 5. Davis, R. E.,et al. 50th Lunar and Planetary Science Conference (2019). 6. Qiagen. MagAttract® PowerMicrobiome® DNA/RNA EP Kit Handbook. (2018). 7. Walters, W. et al. mSystems 1, (2015). 8. Callahan, B. J. et al. Nat. Methods 13, 581–583 (2016). 9. Hall, M. & Beiko, R. G. Microbiome Analysis: Methods and Protocols113–129 (Springer, 2018). 10. Philipson, C. et al. Bio-Protoc. 7, e2622 (2017).

A. B. Regberg↗

Beyond Fair: Engagement, Data Usability, and Open Community Productivity through the NASA Open Science Data Repository

The FAIR principle (findable, accessible, interoperable, and reusable) governs the storage and sharing of NASA space biology and health data[1]. These guiding principles maximize reuse of data and the reproducibility of scientific findings. The NASA Open Science Data Repository (OSDR; an expansion of NASA GeneLab) was built on the FAIR principles and houses over 500 studies and close to 1000 datasets from decades of space life sciences experiments. OSDR embodies the FAIR principles through data governance that includes mediated, embargoed, and fully open access data. The FAIR data governance principles were recently proposed to be expanded to encompass a FAIREST framework for assessing research data repositories (FAIR + Engagement, Social connections, and Trust)[2]. FAIREST emphasizes the importance of data repositories engaging with the scientific community and gaining the trust of researchers regarding data quality. Trust also refers to the TRUST principles developed for assessment of digital repositories: Transparency, Responsibility, User Focus, Sustainability, Technology[3]. We present the “Open Science for Life in Space” Analysis Working Groups (AWGs) as evidence regarding the power of engagement, social connections, and trust which has enhanced OSDR’s capabilities and productivity. AWG members engage in two main activities. One, members provide feedback on OSDR scientific standards for data ingestion, curation, and reuse (study, subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability). Two, AWG members collaborate to mine-reuse OSDR data to conduct scientific analysis. With nearly 800 active members, the AWGs have resulted in 32 publications re-using OSDR data and contributed many papers in two major special issues in Cell (2020) and Nature (2024). AWGs also serve as networking groups, facilitate social connections between researchers at all levels of experience, and also have a social online ‘Forum’ used to keep members informed on projects and opportunities. This community-centric, productive, and trustworthy data culture has resulted in a broader effect with international space agencies, academics, and the commercial space sector wanting to submit their data to OSDR. Ten studies of Inspiration 4 data were recently publicly released by OSDR, as were some JAXA human data. Coming up soon in OSDR are data submissions from the European Space Agency, Virgin Galactic PIs, and SpaceX Polaris Dawn. A major benefit of OSDR is the array of standardized and uniformly formatted data (which was developed through AWG member consensus), from which visualization tools, analysis tools, and machine learning models can be built or trained. This talk will cover the Multi-Study Visualization Tool, the Environmental Data Application, RadLab, and a UCSF-NSF funded knowledge graph biomedical health discovery tool ‘SPOKE’ currently being integrated with OSDR. OSDR also provides training programs in bioinformatics and machine learning to improve the scientific community’s awareness of data availability and to boost their ability to perform data analysis. The increasing engagement of the scientific community and the public with technologies powered by artificial intelligence (AI) heightens the need for data analysis to be transparent. The AI for Life in Space initiative leverages the data products provided in OSDR to train AI models, with an emphasis on explainable and trustworthy AI, which would not be possible without FAIR data and metadata. Overall, here we will demonstrate the importance for NASA life sciences data repositories to adhere to the FAIREST framework, by providing examples and success stories from different aspects of OSDR.

data↗

Benchmarking Computational Tools for Calling SNPs and Indels in Complex Microbial Populations

The NASA BioNutrients missions seek to understand the suitability of microorganisms for bioproduction during space flight. One topic of interest is the stability of microbial genomes during long-term ambient storage and subsequent rehydration and growth. To address these questions, samples from 8 species were flown to ISS for 5 years of desiccated storage at ambient temperature (Stasis Packs) and 2 species were packaged along with powdered media inside a bioreactor system to allow hydration and growth in microgravity (Production Packs). For both systems, Whole Genome Sequencing (WGS) of the DNA extracted from the returned samples and paired ground controls will be conducted to identify changes in genome stability due to time, storage conditions and growth in space. Across the technical replicates, ground controls, 10 timepoints, and multiple experimental conditions, ~300 samples have been selected for initial analysis with WGS sequencing to 100x coverage. A flexible and resource efficient mutation calling pipeline is needed to process this large dataset and allow for comparisons between species. Many bioinformatics tools for calling Indels and Single Nucleotide Variants (SNVs) are designed for use with pure isolates, where true variations from the reference genome are expected to dominate the reads aligning to the location of mutation. In contrast, DNA from the Stasis Pack (SP) samples was collected directly after recovery from desiccated storage and the Production Pack (PP) samples were collected after fermentation. In this context, reads with mutations are expected to be less frequent than reads that align with the reference genome, as each sample will include multiple lines of cells. Thus, BioNutrients samples are expected to be similar to samples from cancer cell or “pooled” sequencing approaches. In preparation for the analysis of the BioNutrients samples, we have tested three mutation calling tools (GATK for Microbes, BreSeq and DiscoSNP) designed for complex samples. A challenge of validating mutation identification pipelines is a lack of “Ground Truth” datasets, especially for complex samples. To compare these three tools, we sought to identify mutations in pre-existing WGS data collected from populations of Chlamydomonas reinhardtii that were exposed to UV mutagenesis and growth in LEO as part of the Space Algae-1 mission. Here we present a summary of these tools against the analysis originally conducted using the CRISP tool. Critical metrics are compared such as runtime, the number of SNPs, the number and size of Indels, and patterns of transversion and transitions identified by each tool are reported. By sharing these benchmarking results collected in support of the BioNutrients mission, we aim to guide others seeking to identify SNVs in similarly complex microbial samples.

Biology↗

Does the International Space Station Leak DNA? Preliminary Results from the ISS External Microorganisms Payload

Existing crewed spacecraft like the ISS (International Space Station) leak by design. The ISS routinely releases gas to maintain life support systems and when astronauts exit the station to perform space walks. The chemical component of this leakage is well characterized, but the biological components are not. The ISS is not subject to planetary protection requirements, but planned missions to Mars will use similar systems and will be subject to planetary protection requirements. If detectable microorganisms are escaping through vents and or airlocks we may need to redesign our crewed habitats to minimize this type of contamination. To test the hypothesis that microorganisms from inside ISS are detectable on exterior surfaces an astronaut used the ISS External Microorganisms sampling kit (Rucker et al. 2018) to sample exterior surfaces of the ISS during an EVA (Extra Vehicular Activity) in January of 2025. These samples were returned to Earth for DNA extraction and sequencing. We successfully, extracted and sequenced bacterial, fungal and viral DNA from these samples that was not present in the negative controls. These results should help NASA refine the planetary protection requirements for crewed missions. Methods: The samples were collected using sterile, DNA free, buccal swabs (23 mm. diameter) housed in custom canisters. Each canister uses a 0.2 μm Teflon filter to maintain sterility as the caddy, holding 8 swabs moves in and out of vacuum. The astronaut sampled the: 1) airlock vestibule, 2) airlock thermal cover, 3) a gap in the micrometeorite shielding near the airlock, 4) a handrail near the airlock, 5) the Carbon Dioxide Removal Assembly vent, and 6) the Vacuum Exhaust System vent. The seventh swab was exposed to vacuum during the EVA without touching it to a surface. The eighth swab, a negative control, was not opened until the caddy returned to Earth. DNA was extracted from the swabs using a QIamp UCP Pathogen kit and prepared for sequencing on an Aviti (Element Biosciences) sequencer (Arslan et al. 2024). The resulting sequences were analyzed using the EDGE Bioinformatics platform (Li et al. 2017). The sequences were analyzed individually using tools like BLAST, GOTTCHA2, Kraken2, and PanGIA. The data were also assembled into metagenome assembled genomes) using tools like CONCOCT, MaxBin2 and MetaBAT2. Results: We successfully extracted and sequenced bacterial, archaeal, fungal and viral DNA from all seven samples. The handrail swab had the lowest number of reads (768,651) and the airlock thermal cover had the highest number of reads (8,819,230). These samples contain DNA from human associated bacteria (e.g. Crynebacterium riegelii ), fungi (.e.g. Penicillium rubens ), and viruses (e.g Alphapapillomavirus ). Conclusion: Preliminary interpretation suggest that the airlock and the space suits themselves are the largest sources of contaminant DNA. Most if not all of the DNA is from organisms known to be present inside the ISS. Vents attached to life support systems may be a lesser source of biological contamination. Further analysis should help NASA address planetary protection knowledge gaps for crewed missions.

Aaron B Regberg↗

RCSB protein data Bank: Next‐generation advanced search for exploration of experimental structures and computed structure models

Abstract The Protein Data Bank (PDB), established in 1971, is the primary global, open‐access archive for experimentally determined 3D macromolecular structures (proteins, RNA, DNA). The research‐focused RCSB.org web‐portal provides access to these data alongside more than one million machine‐learning‐predicted structure models, greatly expanding the available structural landscape. Rapid growth of both experimental and computational structures has increased the need for powerful yet accessible search tools that serve a broad and diverse scientific community. Herein, we describe a redesigned RCSB Protein Data Bank RCSB.org Advanced Search capability that supports intuitive discovery of 3D structures through a unified interface. This interface integrates annotation‐, sequence‐, and 3D structure‐based searches, embeds an interactive 3D viewer, and incorporates curated biological knowledge, such as catalytic site definitions from Mechanism and Catalytic Site Atlas and ligand‐guided structural motifs, for constructing geometry‐driven queries. A new Chemical Search tool allows definition of chemical queries via an integrated drawing tool or standard identifiers, seamlessly combining them with annotation filters. By allowing query definition directly within spatial and chemical contexts, these search interfaces reduce the need for detailed knowledge of residue numbering, chain identifiers, or external cheminformatics software. This capability enables efficient exploration of structures, chemical diversity, and structure–function relationships across all life domains. The redesigned interfaces can be accessed directly at rcsb.org/search/advanced for Advanced Search and rcsb.org/search/chemical for Chemical Search.

Rose, Yana [Research Collaboratory for Structural ↗

The Unified Phenotype Ontology : a framework for cross-species integrative phenomics

Phenotypic data are critical for understanding biological mechanisms and consequences of genomic variation, and are pivotal for clinical use cases such as disease diagnostics and treatment development. For over a century, vast quantities of phenotype data have been collected in many different contexts covering a variety of organisms. The emerging field of phenomics focuses on integrating and interpreting these data to inform biological hypotheses. A major impediment in phenomics is the wide range of distinct and disconnected approaches to recording the observable characteristics of an organism. Phenotype data are collected and curated using free text, single terms or combinations of terms, using multiple vocabularies, terminologies, or ontologies. Integrating these heterogeneous and often siloed data enables the application of biological knowledge both within and across species. Existing integration efforts are typically limited to mappings between pairs of terminologies; a generic knowledge representation that captures the full range of cross-species phenomics data is much needed. We have developed the Unified Phenotype Ontology (uPheno) framework, a community effort to provide an integration layer over domain-specific phenotype ontologies, as a single, unified, logical representation. uPheno comprises (1) a system for consistent computational definition of phenotype terms using ontology design patterns, maintained as a community library; (2) a hierarchical vocabulary of species-neutral phenotype terms under which their species-specific counterparts are grouped; and (3) mapping tables between species-specific ontologies. This harmonized representation supports use cases such as cross-species integration of genotype-phenotype associations from different organisms and cross-species informed variant prioritization.

59 BASIC BIOLOGICAL SCIENCES↗

MIBiG 4.0: advancing biosynthetic gene cluster curation through global collaboration

Specialized or secondary metabolites are small molecules of biological origin, often showing potent biological activities with applications in agriculture, engineering and medicine. Usually, the biosynthesis of these natural products is governed by sets of co-regulated and physically clustered genes known as biosynthetic gene clusters (BGCs). To share information about BGCs in a standardized and machine-readable way, the Minimum Information about a Biosynthetic Gene cluster (MIBiG) data standard and repository was initiated in 2015. Since its conception, MIBiG has been regularly updated to expand data coverage and remain up to date with innovations in natural product research. Here, we describe MIBiG version 4.0, an extensive update to the data repository and the underlying data standard. In a massive community annotation effort, 267 contributors performed 8304 edits, creating 557 new entries and modifying 590 existing entries, resulting in a new total of 3059 curated entries in MIBiG. Particular attention was paid to ensuring high data quality, with automated data validation using a newly developed custom submission portal prototype, paired with a novel peer-reviewing model. MIBiG 4.0 also takes steps towards a rolling release model and a broader involvement of the scientific community. MIBiG 4.0 is accessible online at https://mibig.secondarymetabolites.org/.

59 BASIC BIOLOGICAL SCIENCES↗

Mondo: integrating disease terminology across communities

Precision medicine aims to enhance diagnosis, treatment, and prognosis by integrating multimodal data at the point of care. However, challenges arise due to the vast number of diseases, differing methods of classification, and conflicting terminological coding systems and practices used to represent molecular definitions of disease. This lack of interoperability artificially constrains the potential for diagnosis, clinical decision support, care outcome analysis, as well as data linkage across research domains to support the development or repurposing of therapeutics. There is a clear and pressing need for a unified system for managing disease entities⁠—including identifiers, synonyms, and definitions. To address these issues, we created the Mondo disease ontology—a community-driven, open-source, unified disease classification system that harmonizes diverse terminologies into a consistent, computable framework. Mondo integrates key medical and biomedical terminologies, including Online Mendelian Inheritance in Man (OMIM), Orphanet, Medical Subject Headings (MeSH), National Cancer Institute Thesaurus (NCIt), and more, to provide a comprehensive and accurate representation of disease concepts with fully provenanced and attributed links back to the sources. Mondo can be used as the handle for curation of gene–disease associations utilized in diagnostic applications, research applications such as computational phenotyping, and in clinical coding systems in clinical decision support by pointing the clinician to the numerous knowledge resources linked to the Mondo identifier. Mondo's community-centric approach, stewarded by the Monarch Initiative's expertise in ontologies, ensures that the ontology remains adaptable to the evolving needs of biomedical research and clinical communities, as well as the knowledge providers.

biomedical informatics↗

MolViewSpec: a Mol* extension for describing and sharing molecular visualizations

Data visualization is a pivotal component of a structural biologist’s arsenal. The Mol* Viewer makes molecular visualizations available to broader audiences via most web browsers. While Mol* provides a wide range of functionality, it has a steep learning curve and is only available via a JavaScript interface. To enhance the accessibility and usability of web-based molecular visualization, we introduce MolViewSpec (molstar.org/mol-view-spec), a standardized approach for defining molecular visualizations that decouples the definition of complex molecular scenes from their rendering. Scene definition can include references to commonly used structural, volumetric, and annotation data formats together with a description of how the data should be visualized and paired with optional annotations specifying colors, labels, measurements, and custom 3D geometries. Developed as an open standard, this solution paves the way for broader interoperability and support across different programming languages and molecular viewers, enabling more streamlined, standardized, and reproducible visual molecular analyses. MolViewSpec is freely available as a Mol* extension and a standalone Python package.

Midlik, Adam [European Bioinformatics Institute (U↗

Simultaneous Overexpression of FERULOYL‐CoA 6′‐HYDROXYLASE 1 and COUMARIN SYNTHASE Leads to Coumarin‐Enriched Lignin and Improved Saccharification in Greenhouse‐ and Field‐Grown Poplar

ABSTRACT The urgent need for renewable resources has increased the interest in woody biomass to manufacture bio‐based products. However, lignin recalcitrance limits the enzymatic conversion of wood into fermentable sugars, posing a major challenge for biomass deconstruction. To address this problem, we aimed at incorporating the coumarin scopoletin into the lignin polymer of poplar ( Populus tremula × P . alba ) by expressing FERULOYL‐CoA 6′‐HYDROXYLASE 1 ( F6′H1 ) and COUMARIN SYNTHASE ( COSY ) in lignifying cells. Three constructs were evaluated: two bicistronic constructs, SCOP1 ( COSY followed by F6′H1 ) and SCOP2 ( F6′H1 followed by COSY ), and one monocistronic, SCOP3 (only F6′H1 ). SCOP1 poplars produced most free scopoletin without altering overall lignin, cellulose or hemicellulose content. SCOP2 poplars were overall less efficient in scopoletin production and most of these lines showed a severe biomass yield penalty, whereas SCOP3 caused plant lethality. NMR and metabolic analyses confirmed that scopoletin cross‐coupled with G and S monomers during lignification in SCOP1 lines. In addition to scopoletin, the detection of benzodioxane structures revealed the incorporation of dihydroxycoumarins. Overall coumarin incorporation in lignin amounted up to 2.3%. After alkaline pretreatment, wood from greenhouse‐grown SCOP1 poplars released up to 29% more glucose compared to the wild type upon limited saccharification. Field‐testing of three SCOP1 lines showed a 6 to 11% increase in saccharification efficiency, with the line containing the lowest scopoletin levels maintaining normal growth. These results demonstrate that engineering lignin composition in poplar can improve saccharification, and emphasize the importance of construct design, translational research and field validation.

alternative lignin monomers↗

nf-core/proteinfamilies: a scalable pipeline for the generation of protein families

The growth of metagenomics-derived amino acid sequence data has transformed our understanding of protein function, microbial diversity, and evolutionary relationships. However, the vast majority of these proteins remain functionally uncharacterized. Grouping the millions of such uncharacterized sequences with the few experimentally characterized ones allows the transfer of annotations, while the inspection of conserved residues with multiple sequence alignments can provide clues to function, even in the absence of existing functional information. To address the challenges associated with this data surge and the need to group sequences, we present a scalable, open-source, parametrizable Nextflow pipeline (nf-core/proteinfamilies) that generates nascent protein families or assigns new proteins to existing families. The computational benchmarks demonstrated that resource usage scales approximately linearly with input size, and the biological benchmarks showed that the generated protein families closely resemble manually curated families in widely used databases.

Nextflow↗

Oleaginous Yeast Biology Elucidated With Comparative Transcriptomics

ABSTRACT Extremophilic yeasts have favorable metabolic and tolerance traits for biomanufacturing‐ like lipid biosynthesis, flavinogenesis, and halotolerance – yet the connection between these favorable phenotypes and strain genotype is not well understood. To this end, this study compares the phenotypes and gene expression patterns of biotechnologically relevant yeasts Yarrowia lipolytica , Debaryomyces hansenii , and Debaryomyces subglobosus grown under nitrogen starvation, iron starvation, and salt stress. To analyze the large data set across species and conditions, two approaches were used: a “network‐first” approach where a generalized metabolic network serves as a scaffold for mapping genes and a “cluster‐first” approach where unsupervised machine learning co‐expression analysis clusters genes. Both approaches provide insight into strain behavior. The network‐first approach corroborates that Yarrowia upregulates lipid biosynthesis during nitrogen starvation and provides new evidence that riboflavin overproduction in Debaryomyces yeasts is overflow metabolism that is routed to flavin cofactor production under salt stress. The cluster‐first approach does not rely on annotation; therefore, the coexpression analysis can identify known and novel genes involved in stress responses, mainly transcription factors and transporters. Therefore, this work links the genotype to the phenotype of biotechnologically relevant yeasts and demonstrates the utility of complementary computational approaches to gain insight from transcriptomics data across species and conditions.

Weintraub, Sarah J. [Department of Bioinformatics ↗

WorkflowHub: a registry for computational workflows

The rising popularity of computational workflows is driven by the need for repetitive and scalable data processing, sharing of processing know-how, and transparent methods. As both combined records of analysis and descriptions of processing steps, workflows should be reproducible, reusable, adaptable, and available. Workflow sharing presents opportunities to reduce unnecessary reinvention, promote reuse, increase access to best practice analyses for non-experts, and increase productivity. In reality, workflows are scattered and difficult to find, in part due to the diversity of available workflow engines and ecosystems, and because workflow sharing is not yet part of research practice. WorkflowHub provides a unified registry for all computational workflows that links to community repositories, and supports both the workflow lifecycle and making workflows findable, accessible, interoperable, and reusable (FAIR). By interoperating with diverse platforms, services, and external registries, WorkflowHub adds value by supporting workflow sharing, explicitly assigning credit, enhancing FAIRness, and promoting workflows as scholarly artefacts. The registry has a global reach, with hundreds of research organisations involved, and more than 800 workflows registered.

97 MATHEMATICS AND COMPUTING↗

Genomic factors limiting the diversity of Saccharomycotina plant pathogens

The Saccharomycotina fungi have evolved to inhabit a vast diversity of habitats over their 400-million-year evolution. There are, however, only a few known fungal pathogens of plants in this subphylum, primarily belonging to the genera Eremothecium and Geotrichum. We compared the genomes of 12 plant-pathogenic Saccharomycotina strains to 360 plant-associated strains to identify features unique to the phytopathogens. Characterization of the oxylipin synthesis genes, a compound believed to be involved in Eremothecium pathogenicity, did not reveal any differences in gene presence within or between the plant-pathogenic and plant-associated strains. A reverse-ecological approach, however, revealed that plant pathogens lack several metabolic enzymes known to assist other phytopathogens in overcoming plant defenses. This includes L-rhamnose metabolism, formamidase and nitrilase genes. This result suggests that the Saccharomycotina plant pathogens are limited to infecting ripening fruits as they are without the necessary enzymes to degrade common phytohormones and secondary metabolites produced by plants.

Saccharomycotina, fungi, phytopathogen, reverse ec↗

PDB-IHM: A System for Deposition, Curation, Validation, and Dissemination of Integrative Structures

Structures of many large biomolecular assemblies are now being determined using integrative approaches. In these approaches, information derived from multiple experimental and computational methods is combined to compute three-dimensional structures of multi-protein complexes and other macromolecular machines. A standalone prototype data resource for integrative structures called PDB-Dev was built, based on recommendations of the Integrative and Hybrid Methods (IHM) Task Force of the Worldwide Protein Data Bank (wwPDB). This effort included developing data standards and software tools for collecting, curating, validating, visualizing, archiving, and disseminating integrative structures that span diverse spatiotemporal scales and conformational states. Mechanisms have been created to validate integrative structures based on the experimental data underpinning them. Building upon this foundational framework, PDB-Dev has been further expanded to handle large dynamic macromolecular systems and integrative structures that combine, for example, experimental restraints with atomic coordinates computed by machine learning algorithms. Data standards and supporting tools have also been extended to capture information about biomolecular dynamics, such as conformational transitions and related kinetic data derived from biophysical methods. Recently, PDB-Dev was unified with the PDB archive and rebranded as PDB-IHM (pdb-ihm.org), further promoting FAIR (Findable, Accessible, Interoperable, and Reusable) principles of data stewardship for integrative structural biology.

IHMCIF↗