Search NASASearch

SEARCH · Search NASA

Results for “Curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Knowledge-guided learning with curated prior genetic biomarkers for robust model interpretation

Abstract Motivation Knowledge-guided learning offers effective and robust model training strategies in data-scarce settings by incorporating established domain knowledge, thereby enhancing generalization, robustness, and interpretability. By contrast, conventional deep learning approaches rely purely on data-driven learning, which can limit robust model interpretability, particularly in high-dimensional settings with limited size samples. In computational biology, knowledge-guided learning has primarily leveraged network- and structural-based knowledge, leading to biologically interpretable representations and enhanced predictive performance compared to conventional approaches. However, curated biomarkers, one of the most accessible forms of biological knowledge, remain largely unexplored within knowledge-guided paradigms. Results In this study, we propose a model-agnostic training paradigm, Biomarker-driven Explainable Prior-guided Learning (BioExPL), that can be applied to any neural networks that incorporates curated prior knowledge. BioExPL enforces neural networks to reflect curated biomarker priors in their latent representations through a novel knowledge-alignment loss. BioExPL consistently demonstrated significantly improved predictive performance and enhanced model interpretability with minimized computational overhead in simulation studies and intensive experiments on multiple cancer datasets. BioExPL not only integrates prior curated knowledge into the model but also accurately identifies unknown associated signals additionally. BioExPL is model-agnostic and domain-independent, enabling its integration into diverse neural network architectures. Availability and implementation The open-source is publicly available at: https://github.com/datax-lab/BioExPL.

Baek, Beomsu [Department of Computer Science, Univ

Atomistic Simulation of Glasses and Amorphous Materials: Challenges and Opportunities for the Next Decade

Atomistic simulations have become indispensable tools for understanding glass structure, dynamics, and properties, yet persistent challenges limit their predictive power. This perspective examines three interconnected issues, namely glass formation procedures, interatomic potential development, and machine learning applications, which emerged from the 5th International Workshop on Challenges of Atomistic Simulations of Glasses and Amorphous Materials. We identify convergent community priorities for (i) standardized validation protocols, (ii) curated benchmark datasets with complete metadata, and (iii) open repositories for glasses. A systematic was forward is provided by a hierarchical validation framework for assessing the structural fidelity, property prediction, and behavioral realism of simulation techniques. Looking ahead, transformative advances are promised by the fusion of classical techniques with machine learning based approaches, for instance, by integrating swap Monte Carlo with machine-learning (ML) potentials, leveraging foundation models through transfer learning, and finetuning ML potentials with experimental data. Progress depends on the community committing to validated models, reproducible protocols, and sustained data sharing.

Krishnan, N. M. Anoop

Large language model-driven database for thermoelectric materials

Thermoelectric materials have the ability to convert waste heat into electricity, offering a valuable solution for energy harvesting. However, their widespread use is hindered by low conversion efficiency, the reliance on expensive rare earth elements, and the environmental and regulatory concerns associated with lead-based materials. A fast and cost-effective way to identify highly efficient thermoelectric materials is through data-driven methods. These approaches rely on robust and comprehensive datasets to train models. Although there are several databases on thermoelectric materials, there is still a need to collect and integrate experimental data from peer-reviewed research articles to capture diverse compositions and properties of materials. Here, in this work, we developed a comprehensive database of 7,123 thermoelectric compounds, containing key information such as chemical composition, structural detail, seebeck coefficient, electrical and thermal conductivity, power factor, and figure of merit (ZT). We used the GPTArticleExtractor workflow, powered by large language models (LLM), to extract and curate data automatically from the scientific literature published in Elsevier journals. This process enabled the creation of a structured database that addresses the challenges of manual data collection. The open access database could stimulate data-driven research and advance thermoelectric material analysis and discovery.

Database

Full ribosomal operon sequencing of anaerobic gut fungi (phylum Neocallimastigomycota ): insights on its markers and phylogenetic resolution

The phylogenetic affiliations of anaerobic gut fungi (Neocallimastigomycota) are typically evaluated using single-gene markers. However, this approach often fails to resolve relationships between closely related lineages. To address this issue and identify alternative markers, we created a curated database comprising the complete ribosomal operon sequences of 156 isolates, representing 20 of the 22 recognized genera and two new genus-level clades. Using long-read sequencing, we obtained ~9 kbp operon sequences and developed a robust analysis pipeline. Incorporating both coding genes and non-coding regions (excluding IGS1) improved phylogenetic resolution. This phylogenetic approach successfully resolved the Cyllamyces and Caecomyces clades (hard-to-distinguish genetically), as well as seven analysed Piromyces species. We also scanned the operon for markers that are suitable for short-read sequencing platforms, with the aim of enhancing biodiversity and phylogenetic studies. Notably, the ETS1 genetic region also enabled the distinction between these lineages, indicating its phylogenetic value within the ribosomal operon. The resulting database is a valuable resource for expanding and strengthening phylogenetic frameworks.

High-throughput sequencing

Application of Nuclear Technology to Art Identification Problems: First Annual Report

Throughout the centuries, as man has become increasingly affluent, his interest in antiquities and objects of art, and the price he has been willing to pay for them, has increased. Inevitably, one result has been to attract the unscrupulous to the production and sale of counterfeit paintings, sculptures, and other works of art. The skill of forgers varies but often is highly sophisticated. However, dealers, museum curators, and other experts have also become more skillful in the detection of forgeries. Scientific tools of increasing sensitivity and sophistication have gradually been applied to the examination of the materials of art and archaeology. Such tools frequently support and render more reliable the judgment of the experts. The quality of forgeries may improve further as their makers become acquainted with the new methods of detection and in turn learn to circumvent or confound the methods. Therefore, the development of still more advanced methods of examination has great utility in the art world. This is particularly true if the ultimate effect is to make circumvention of the methods of examination so costly that the economic incentives for producing forgeries may be substantially reduced, if not eliminated.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH