Search NASA⌕ Search

SEARCH · Search NASA

Results for “manual curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

VISTA Enhancer browser: an updated database of tissue-specific developmental enhancers

Regulatory elements (enhancers) are major drivers of gene expression in mammals and harbor many genetic variants associated with human diseases. Here, we present an updated VISTA Enhancer Browser (https://enhancer.lbl.gov), a database of transgenic enhancer assays conducted in developing mouse embryos in vivo. Since the original publication in 2007, the database grew nearly 20-fold from 250 to over 4500 experiments and currently harbors over 23 500 images. The updated database provides structured information on experiments conducted at different stages of embryonic development, including enhancer activities of human pathogenic and synthetic variants and sequences derived from a variety of species. In addition to manually curated results of thousands of individual experiments, the new database also features hundreds of manually curated comparisons between alleles. The VISTA Enhancer Browser provides a crucial resource for study of human genetic variation, gene regulation and developmental biology.

59 BASIC BIOLOGICAL SCIENCES↗

Identifying genomic data use with the Data Citation Explorer

Increases in sequencing capacity, combined with rapid accumulation of publications and associated data resources, have increased the complexity of maintaining associations between literature and genomic data. As the volume of literature and data have exceeded the capacity of manual curation, automated approaches to maintaining and confirming associations among these resources have become necessary. Here we present the Data Citation Explorer (DCE), which discovers literature incorporating genomic data that was not formally cited. This service provides advantages over manual curation methods including consistent resource coverage, metadata enrichment, documentation of new use cases, and identification of conflicting metadata. The service reduces labor costs associated with manual review, improves the quality of genome metadata maintained by the U.S. Department of Energy Joint Genome Institute (JGI), and increases the number of known publications that incorporate its data products. The DCE facilitates an understanding of JGI impact, improves credit attribution for data generators, and can encourage data sharing by allowing scientists to see how reuse amplifies the impact of their original studies.

59 BASIC BIOLOGICAL SCIENCES↗

Data Citation Explorer (DCE) v1.0

Increases in sequencing capacity, combined with rapid accumulation of publications and associated data resources, have increased the complexity of maintaining associations between literature and genomic data. As the volume of literature and data have exceeded the capacity of manual curation, automated approaches to maintaining and confirming associations among these resources have become necessary. Here we present the Data Citation Explorer (DCE), which discovers literature incorporating genomic data whether or not provenance was clearly indicated. This service provides advantages over manual curation methods including consistent resource coverage, metadata enrichment, documentation of new use cases, and identification of conflicting metadata. The service reduces labor costs associated with manual review, improves the quality of genome metadata maintained by the U.S. Department of Energy Joint Genome Institute (JGI), and increases the number of known publications that incorporate its data products. The DCE facilitates an understanding of JGI impact, improves credit attribution for data generators, and can encourage data sharing by allowing scientists to see how reuse amplifies the impact of their original studies.

Parker, Charles↗

Integrative Modeling and Analysis of Fungal Central Carbon Metabolism

Over a thousand fungal genomes have been sequenced, yet manually curated genome-scale metabolic models (GEMs) are available for only a limited number of species. Moreover, these models have often been developed independently, leading to inconsistencies in namespaces, compartment definitions, and pathway representations that hinder comparative analysis, the systematic reuse of prior curation efforts, and the integration of consolidated metabolic knowledge. Here, we present the Consolidated Fungal Core Metabolism Model (CFCMM), constructed by integrating thirteen published fungal models spanning Ascomycota, Mucoromycota, and both Crabtree-positive and Crabtree-negative yeasts. We harmonized metabolites and reactions into a non-redundant shared ModelSEED ontological space, standardized compartmentalization, and refined gene–protein–reaction (GPR) rules. Using pathway-level visualization and systematic gap detection, we further improved the integrated network through literature-guided curation to correct stoichiometry, stereospecificity, and pathway architecture. Orthologous protein family reconstruction and functional annotation workflows were used to validate and inform GPR associations, with particular emphasis on ambiguous enzyme superfamilies and membrane-associated components. Using the resulting CFCMM, we built high-quality central carbon core models for each fungus and performed flux balance analysis to quantify ATP-yield variation under aerobic and anaerobic conditions, explicitly evaluating scenarios driven by differences in electron transport chain (ETC) composition. Simulations reproduced the expected fermentative yield of approximately 2 mmol ATP per mmol glucose under anaerobic conditions and separated the thirteen fungi into two bioenergetic groups under aerobic respiration based on Complex I status, with predicted yields of approximately 30 versus 22 mmol ATP per mmol glucose. Forcing flux through the alternative oxidase bypass further reduced ATP yields to approximately 12 and 4 mmol ATP per mmol glucose in Complex I-containing and Complex I-lacking fungi, respectively. Collectively, this work provides a manually curated, ModelSEED-consistent, and extensible fungal core metabolic template, deployed in DOE KBase as a resource for automated reconstruction of central carbon core models from any sequenced fungal genome. In addition, the CFCMM provides modular components for developing GEMs with more accurate energy predictions and enables robust comparative analyses of fungal bioenergetics and core metabolic diversity

59 BASIC BIOLOGICAL SCIENCES↗

Knowledge-matching based computational framework for genome-scale metabolic model refinement

Genome-scale metabolic models (GEMs) are mathematically structured knowledge base reconstructed from annotated genome of different organisms. With the advancement of next-generation sequencing technology, many organisms have had their genomes sequenced. However, obtaining a high-quality GEM is highly time-consuming, even with the introduction of several genome-scale reconstruction tools that offer automated draft network generation and gap filling. It has been recognized that the iterative process of manual curation and refinement is the limiting step of GEM development, and how to expedite the GEM refinement is still an open question. As cellular metabolism is a complex system with very high degree of freedom and redundancy, the principles and techniques developed in process systems engineering can be adapted to expedite GEM refinement. In this paper we present a knowledge-matching based computation framework for GEM refinement, and demonstrate the effectiveness of the proposed solution using the refinement of a GEM for Clostridium tyrobutyricum.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES): a method for populating knowledge bases using zero-shot learning

Abstract Motivation Creating knowledge bases and ontologies is a time consuming task that relies on manual curation. AI/NLP approaches can assist expert curators in populating these knowledge bases, but current approaches rely on extensive training data, and are not able to populate arbitrarily complex nested knowledge schemas. Results Here we present Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES), a Knowledge Extraction approach that relies on the ability of Large Language Models (LLMs) to perform zero-shot learning and general-purpose query answering from flexible prompts and return information conforming to a specified schema. Given a detailed, user-defined knowledge schema and an input text, SPIRES recursively performs prompt interrogation against an LLM to obtain a set of responses matching the provided schema. SPIRES uses existing ontologies and vocabularies to provide identifiers for matched elements. We present examples of applying SPIRES in different domains, including extraction of food recipes, multi-species cellular signaling pathways, disease treatments, multi-step drug mechanisms, and chemical to disease relationships. Current SPIRES accuracy is comparable to the mid-range of existing Relation Extraction methods, but greatly surpasses an LLM’s native capability of grounding entities with unique identifiers. SPIRES has the advantage of easy customization, flexibility, and, crucially, the ability to perform new tasks in the absence of any new training data. This method supports a general strategy of leveraging the language interpreting capabilities of LLMs to assemble knowledge bases, assisting manual knowledge curation and acquisition while supporting validation with publicly-available databases and ontologies external to the LLM. Availability and implementation SPIRES is available as part of the open source OntoGPT package: https://github.com/monarch-initiative/ontogpt.

59 BASIC BIOLOGICAL SCIENCES↗

Meta2DB: Curated Shotgun Metagenomic Feature Sets and Metadata for Health State Prediction

Meta2DB is a curated metagenomic and metadata database that provides structurally consistent microbiome taxonomy feature count tables for 13 897 samples across 84 studies, 23 disease states, and 34 geographical locations. All samples were uniformly processed using a streamlined metagenomic classification pipeline that employs a unique and comprehensive reference database indexed to contain all sequences across all kingdoms of life that were present in the NCBI Nucleotide (nt) database retrieved on 4 January 2023. This pipeline leverages high-performance computing (HPC) resources at Lawrence Livermore National Laboratory and was used to process 50TB of publicly available raw metagenomic sequence data. Extensive metadata curation was carried out through a combination of manual curation and automated parsing, producing a consistent inter-study metadata table specifically structured to facilitate training of ML models for prediction of human health.

Kok, C [Lawrence Livermore National Laboratory (LL↗

Plant sulfate transporter protein sequences for phylogenetic analysis

Sulfur is an essential macronutrient that supports plant growth, development, and responses to environmental stress. Sulfate is the predominant inorganic form of sulfur in soils, and its uptake by roots and translocation to shoots are facilitated by the sulfate transporter (SULTR) family of proteins. Although the first plant SULTR gene was identified nearly three decades ago, several subfamily members, particularly those in the expansive and angiosperm-specific SULTR3 group, remain poorly characterized. To support comprehensive phylogenetic and sequence-based analyses, we compiled a curated dataset of 262 SULTR protein sequences from 22 plant species spanning the evolutionary breadth of land plants. This collection includes representatives from two basal lineages, two early-divergent angiosperms, six monocots, and ten dicots. All sequences were extracted from genome assemblies available in Phytozome v13 (Joint Genome Institute) and manually curated, with cross-referencing to additional databases such as NCBI when needed. This dataset provides a valuable resource for reconstructing the evolutionary history of the SULTR family, with particular emphasis on the diversification of SULTR3 transporters in flowering plants. This resource may also support functional annotation, comparative genomics, and structural modeling of sulfate transport proteins.

CBI↗

Algorithmic Classification of Raman Spectra Biosignatures: Improving Life Detection Confidence

“Agnostic” biosignatures – indicators of life (or the absence of life), independent of a particular biochemistry – are increasingly considered a high standard for life detection. The Ladder of Life Detection (2018) called for investigating how combinations of independent and different potential biosignatures affect confidence. To address this gap, statistical classification of elemental abundances, isotopic fractionation, and reflectance spectroscopy (VNIR) has been implemented. Raman spectroscopy, highly desirable due to its wide availability, has the potential to improve this predictive power. This work implemented biosignature classification algorithms on Raman data alone, in preparation for combination with the other data types. Raman spectroscopy data was collected from published databases and papers as part of a manually curated dataset of “indicative” and “non-indicative of life” samples. These currently include 61 non-indicative samples (meteorites, magnetite); 3 indicative living samples (bacteria); 20 indicative non-living samples (chalk, bone); and 12 indicative mixed (with non-indicative material) samples (soil, microbial mats). Laboratory work is ongoing to characterize additional samples, particularly a greater breadth of mixed systems. Spectra were interpolated, filtered with the Savitzsky-Golay filter, and de-noised. For a preliminary examination, agnostic features were manually extracted including mean intensity, number of peaks, and mean peak width. Different peak prominences and filtering polynomials were used to refine features. Classification algorithms were implemented: k-nearest neighbors (KNN), logistic regression (LR), linear support vector machines (SVM), random forest (RF), Gaussian naïve bayes (GNB). Lastly, Monte Carlo simulations on 1,000 50%-train-test-splits were used to validate classification performance and feature significance. The preliminary feature set achieved its highest AUC of 0.52 with LR, with no strongly discriminatory features. Work to improve feature extraction, such as through deep learning with back propagation, is planned. In future work, the Raman data will be combined with the other data types, and potentially new data types such as enantiomeric excess. This project was partially supported through the NASA Ames Project EXcellence (APEX) incubator program.

Astrobiology↗

Climate Data Initiative: A Geocuration Effort to Support Climate Resilience

Curation is traditionally defined as the process of collecting and organizing information around a common subject matter or a topic of interest and typically occurs in museums, art galleries, and libraries. The task of organizing data around specific topics or themes is a vibrant and growing effort in the biological sciences but to date this effort has not been actively pursued in the Earth sciences. In this paper, we introduce the concept of geocuration and define it as the act of searching, selecting, and synthesizing Earth science data/metadata and information from across disciplines and repositories into a single, cohesive, and useful collection. We present the Climate Data Initiative (CDI) project as a prototypical example. The CDI project is a systematic effort to manually curate and share openly available climate data from various federal agencies. CDI is a broad multi-agency effort of the U.S. government and seeks to leverage the extensive existing federal climate-relevant data to stimulate innovation and private-sector entrepreneurship to support national climate-change preparedness. We describe the geocuration process used in the CDI project, lessons learned, and suggestions to improve similar geocuration efforts in the future.

Metada↗

Climate Data Initiative: A Geocuration Effort to Support Climate Resilience

Curation is traditionally defined as the process of collecting and organizing information around a common subject matter or a topic of interest and typically occurs in museums, art galleries, and libraries. The task of organizing data around specific topics or themes is a vibrant and growing effort in the biological sciences but to date this effort has not been actively pursued in the Earth sciences. In this paper, we introduce the concept of geocuration and define it as the act of searching, selecting, and synthesizing Earth science data/metadata and information from across disciplines and repositories into a single, cohesive, and useful compendium We present the Climate Data Initiative (CDI) project as an exemplar example. The CDI project is a systematic effort to manually curate and share openly available climate data from various federal agencies. CDI is a broad multi-agency effort of the U.S. government and seeks to leverage the extensive existing federal climate-relevant data to stimulate innovation and private-sector entrepreneurship to support national climate-change preparedness. We describe the geocuration process used in CDI project, lessons learned, and suggestions to improve similar geocuration efforts in the future.

virtual collections↗

Geocuration Lessons Learned from the Climate Data Initiative Project

Curation is traditionally defined as the process of collecting and organizing information around a common subject matter or a topic of interest and typically occurs in museums, art galleries, and libraries. The task of organizing data around specific topics or themes is a vibrant and growing effort in the biological sciences but to date this effort has not been actively pursued in the Earth sciences. This presentation will introduce the concept of geocuration, which we define it as the act of searching, selecting, and synthesizing Earth science data/metadata and information from across disciplines and repositories into a single, cohesive, and useful compendium. We also present the Climate Data Initiative (CDI) project as an prototypical example. The CDI project is a systematic effort to manually curate and share openly available climate data from various federal agencies. CDI is a broad multi-agency effort of the U.S. government and seeks to leverage the extensive existing federal climate-relevant data to stimulate innovation and private-sector entrepreneurship to support national climate change preparedness. The geocuration process used in the CDI project, key lessons learned, and suggestions to improve similar geocuration efforts in the future will be part of this presentation.

climate↗

Protein–Protein Interaction Networks Derived from Classical and Machine Learning-Based Natural Language Processing Tools

The study of protein-protein interactions (PPIs) provides insight into various biological mechanisms, including the binding of antibodies to antigens, enzymes to inhibitors or promoters, and receptors to ligands. Recent studies of PPIs have led to significant biological breakthroughs. For example, the study of PPIs involved in the human:SARS-CoV-2 viral infection mechanism aided in the development of the SARS-CoV-2 vaccines. Though several databases exist for the manual curation of PPI networks, text mining methods have been routinely demonstrated as useful alternatives for newly studied or understudied species where databases are incomplete. Here, the relationship extraction (RE) performance of several open-source classical text processing, machine learning (ML)-based natural language processing (NLP), and large language model (LLM)-based NLP tools were compared. Overall, our results indicated that networks derived from classical methods tend to have high true positive rates at the expense of having overconnected-networks, ML-based NLP methods have lower true positive rates but networks with the closest structures to the target network, and LLM-based NLP methods tend to exist in-between the two other approaches, with variable performances. Finally, the selection of a specific NLP approach should be tied to the needs of a study and text availability, as models varied in performance due to the amount of text provided.

59 BASIC BIOLOGICAL SCIENCES↗

A Climatology and Life‐Cycle Characteristics of Atmospheric Fronts and Their Associated Precipitation

Abstract Atmospheric fronts are one of the main sources of mid‐latitude variability. We employ a novel method for identifying and tracking fronts and frontal precipitation. Thermal and dynamical variables are used to identify fronts as areal objects in space, which are tracked in time using the open‐source TempestExtremes software package. Precipitation objects are co‐located to identify frontal precipitation. The method is subjected to validation and sensitivity tests using manually curated data from the National Weather Service. Climatologies of fronts and frontal precipitation are computed from reanalysis and observations; fronts are present upwards of 14% of the time in the storm tracks, and represent the majority (up to 90%) of total and extreme precipitation. Novel aspects of the method are showcased through the lifetime characteristics of fronts across North America. Three sets of warm and cold fronts were discovered, and their duration, distance‐traveled, and translation velocity are examined. Plain Language Summary Mid‐latitude low‐pressure systems and weather fronts are important for our day‐to‐day experience of weather events, particularly in the mid‐latitudes. This work makes use of standardized atmospheric data and creates a method of automatically tracking these important atmospheric features and their precipitation to quantify their relative role in global precipitation. Weather fronts are persistent in the mid‐latitudes and are associated with the majority of precipitation–particularly the most intense precipitation. Trajectories of fronts over North America are categorized to create a set of archetypal fronts that occur in that region. The differences between these types of fronts are characterized. Key Points An automated, efficient, and skillful frontal detection algorithm is developed and validated Fronts contribute a larger fraction of extreme precipitation than all precipitation in mid‐latitude storm tracks Fronts across North America have substantial variation in characteristics depending on their origin location

extratropical cyclone↗

A rule-free workflow for the automated generation of databases from scientific literature

Abstract In recent times, transformer networks have achieved state-of-the-art performance in a wide range of natural language processing tasks. Here we present a workflow based on the fine-tuning of BERT models for different downstream tasks, which results in the automated extraction of structured information from unstructured natural language in scientific literature. Contrary to existing methods for the automated extraction of structured compound-property relations from similar sources, our workflow does not rely on the definition of intricate grammar rules. Hence, it can be adapted to a new task without requiring extensive implementation efforts and knowledge. We test our data-extraction workflow by automatically generating a database for Curie temperatures and one for band gaps. These are then compared with manually curated datasets and with those obtained with a state-of-the-art rule-based method. Furthermore, in order to showcase the practical utility of the automatically extracted data in a material-design workflow, we employ them to construct machine-learning models to predict Curie temperatures and band gaps. In general, we find that, although more noisy, automatically extracted datasets can grow fast in volume and that such volume partially compensates for the inaccuracy in downstream tasks.

36 MATERIALS SCIENCE↗

Global characterization of biosynthetic gene clusters in non-model eukaryotes using domain architectures

Abstract The majority of pharmaceuticals are derived from natural products, bioactive compounds naturally synthesized by organisms to provide evolutionary advantages. Although the rich evolutionary history of eukaryotic algal species implicates a high potential for natural product-based drug discovery, it remains largely untouched. This study investigates 2762 putative biosynthetic gene clusters (BGCs) from 212 eukaryotic algal genomes. To analyze a vast set of structurally diverse BGCs, we employed comparative analysis based on the vectorization of biosynthetic domains, referred to as biosynthetic domain architecture (BDA). By characterizing core biosynthetic machineries through BDA, we identified key BDAs of modular BGCs in diverse eukaryotes and introduced 16 candidate modular BGCs with similar BDAs to previously validated BGCs. This study provides a global characterization of eukaryotic algal BGCs, offering an alternative to laborious manual curation for BGC prioritization.

59 BASIC BIOLOGICAL SCIENCES↗

Ontologizing health systems data at scale: making translational discovery a reality

Common data models solve many challenges of standardizing electronic health record (EHR) data but are unable to semantically integrate all of the resources needed for deep phenotyping. Open Biological and Biomedical Ontology (OBO) Foundry ontologies provide computable representations of biological knowledge and enable the integration of heterogeneous data. However, mapping EHR data to OBO ontologies requires significant manual curation and domain expertise. We introduce OMOP2OBO, an algorithm for mapping Observational Medical Outcomes Partnership (OMOP) vocabularies to OBO ontologies. Using OMOP2OBO, we produced mappings for 92,367 conditions, 8611 drug ingredients, and 10,673 measurement results, which covered 68–99% of concepts used in clinical practice when examined across 24 hospitals. When used to phenotype rare disease patients, the mappings helped systematically identify undiagnosed patients who might benefit from genetic testing. By aligning OMOP vocabularies to OBO ontologies our algorithm presents new opportunities to advance EHR-based deep phenotyping.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Organ-delimited gene regulatory networks provide high accuracy in candidate transcription factor selection across diverse processes

Organ-specific gene expression datasets that include hundreds to thousands of experiments allow the reconstruction of organ-level gene regulatory networks (GRNs). However, creating such datasets is greatly hampered by the requirements of extensive and tedious manual curation. Here, we trained a supervised classification model that can accurately classify the organ-of-origin for a plant transcriptome. This K-Nearest Neighbor-based multiclass classifier was used to create organ-specific gene expression datasets for the leaf, root, shoot, flower, and seed in Arabidopsis thaliana . A GRN inference approach was used to determine the: i. influential transcription factors (TFs) in each organ and, ii. most influential TFs for specific biological processes in that organ. These genome-wide, organ-delimited GRNs (OD-GRNs), recalled many known regulators of organ development and processes operating in those organs. Importantly, many previously unknown TF regulators were uncovered as potential regulators of these processes. As a proof-of-concept, we focused on experimentally validating the predicted TF regulators of lipid biosynthesis in seeds, an important food and biofuel trait. Of the top 20 predicted TFs, eight are known regulators of seed oil content, e.g., WRI1, LEC1, FUS3. Importantly, we validated our prediction of MybS2, TGA4, SPL12, AGL18, and DiV2 as regulators of seed lipid biosynthesis. We elucidated the molecular mechanism of MybS2 and show that it induces purple acid phosphatase family genes and lipid synthesis genes to enhance seed lipid content. This general approach has the potential to be extended to any species with sufficiently large gene expression datasets to find unique regulators of any trait-of-interest.

09 BIOMASS FUELS↗