Search NASASearch

SEARCH · Search NASA

Results for “Sample curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

50 records · Page 3

Toward accelerating rare-earth metal extraction using equivariant neural networks

The separation of rare-earth metals, vital for numerous advanced technologies, is hampered by their similar chemical properties, making ligand discovery a significant challenge. Traditional experimental and quantum chemistry approaches for identifying effective ligands are often resource-intensive. We introduce a machine learning protocol based on an equivariant neural network, Allegro, for the rapid and accurate prediction of binding energies in rare-earth complexes. Key to this work is our newly curated dataset of rare-earth metal complexes—made publicly available to foster further research—systematically generated using the Architector program. This dataset distinctively features functionalized derivatives of proven rare-earth-chelating scaffolds, hydroxypyridinone (HOPO), catecholamide (CAM), and their thio-analogues, selected for their established efficacy in binding these elements. Trained on this valuable resource, our Allegro models demonstrate excellent performance, particularly when trained to directly predict DFT-level binding energies, yielding highly accurate results that closely correlate with theoretical calculations on a diverse test set. Furthermore, this strategy exhibited strong out-of-sample generalization, accurately predicting binding energies for an isomeric HOPO-derivative ligand not seen during training. By substantially reducing computational demands, this machine learning framework, alongside the provided dataset, represent powerful tools to accelerate the high-throughput screening and rational design of novel ligands for efficient rare-earth metal separation.

Gupta, Ankur K. [Lawrence Berkeley National Labora

Project 2.8: Biological Specimen Repository for the Mayak Project

Curation of the existing specimens, including Standard Operating Procedures: The Southern Urals Biophysics Institute (SUBI) will continue to manage the day-to-day operations of the biorepository, but will have to assume the financial costs of maintaining and purchasing equipment such as freezers; Standard Operating Procedures are already in place for all aspects of the operations, and the Georgetown University team will be available on a voluntary and ad hoc basis to answer any technical questions in the future. Acquisition and tracking of new specimens: SUBI will decide on the collection of new tissues and blood samples in the future, and on what scale, depending on available resources and research needs. The Georgetown University team is willing to provide advice on a voluntary and ad hoc basis. Procedures for receiving and approving specimen requests from users: SUBI can take over this function; Chris Loffredo and other qualified research scientists at Georgetown University would be willing to serve as Tissue Review Committee members on a voluntary and ad hoc basis. Biospecimens remaining at Georgetown University: pursuant to the original terms of the Biospecimen Transfer Agreement between SUBI and Georgetown University, all unused portions of specimens (FFPE blocks, slides, and frozen tissues) will be returned to SUBI at the conclusion of the approved scientific research for which they were transferred. It was anticipated that the return shipment of biospecimens would occur in the fall of 2023. However, at this time, due to international developments, there are no shipping companies who can provide such deliveries to Russia. For now, the biospecimens can remain at Georgetown University where they are stored at the Genomics and Epigenomics Shared Resource. There have not been any storage costs to date, but this could change in the future.

59 BASIC BIOLOGICAL SCIENCES

Identifying Anomalous DESI Galaxy Spectra with a Variational Autoencoder

The tens of millions of spectra being captured by the Dark Energy Spectroscopic Instrument (DESI) provide tremendous discovery potential. In this work we show how Machine Learning, in particular Variational Autoencoders (VAE), can detect anomalies in a sample of approximately 200,000 DESI spectra comprising galaxies, quasars and stars. We demonstrate that the VAE can compress the dimensionality of a spectrum by a factor of 100, while still retaining enough information to accurately reconstruct spectral features. We then detect anomalous spectra as those with high reconstruction error and those which are isolated in the VAE latent representation. The anomalies identified fall into two categories: spectra with artefacts and spectra with unique physical features. Awareness of the former can help to improve the DESI spectroscopic pipeline; whilst the latter can lead to the identification of new and unusual objects. To further curate the list of outliers, we use the Astronomaly package which employs Active Learning to provide personalised outlier recommendations for visual inspection. In this work we also explore the VAE latent space, finding that different object classes and subclasses are separated despite being unlabelled. We demonstrate the interpretability of this latent space by identifying tracks within it that correspond to various spectral characteristics. For example, we find tracks that correspond to increasing star formation and increase in broad emission lines along the Balmer series. In upcoming work we hope to apply the methods presented here to search for both systematics and astrophysically interesting objects in much larger datasets of DESI spectra.

Nicolaou, C. [University Coll. London] (ORCID:0000

A curated benchmark for cofolding models on kinase conformational states

Abstract Protein kinases are critical drug targets, requiring therapeutics that can modulate their active and inactive conformational states. While cofolding models can generate global folds directly from kinase sequences and ligand SMILES strings, these models have not yet been tested on their ability to recover ligand-induced-fit conformational states of the kinase proteins. Here, we introduce KinConfBench, a curated benchmark of 2225 high-quality human kinase chains to evaluate the ability of four state-of-the-art cofolding models—Boltz-2, Chai-1, Protenix, and RoseTTAFold-All-Atom—to recover both canonical and rare conformational states. We show that geometric success metrics of a ligand pose in the active site do not correlate strongly with the correct kinase conformational state, motivating a new set of dynamical benchmarks for assessing cofolding models. While all four cofolding models achieve ~60–80% prediction accuracy for kinase conformational classification, they exhibit severe mode collapse when performing multiple inferences, show negligible structural diversity in sampling induced-fit motions, and display a prevalent “apo-drift” in which most cofolding models predominantly predict the kinase to be in its ligand-free state. Our results highlight that capturing ligand-induced protein conformational diversity, not just geometric fit, is critical for next-generation structure-based drug discovery.

Sun, Kunyang

Two-dimensional heteronuclear single quantum coherence (HSQC) NMR spectra of lignin isolated from field grown transgenic poplar

Here we present a curated dataset of a series of two-dimensional heteronuclear single quantum coherence (HSQC) nuclear magnetic resonance (NMR) spectra of lignin isolated from a field grown transgenic poplar engineered with a monolignol 4-O-methyltransferase (MOMT4). The poplar was collected from a 2-year-old rotation trees within a three-year field trial experiment. The poplar was Soxhlet-extracted with toluene/ethanol and the extractives-free poplar was then ball-milled in a Retsch PM100 planetary ball mill using a porcelain jar with ceramic balls at 600 rpm for 2 h (in 5 min on and 5 min off cycles to avoid excessive sample heating). The ball-milled materials were then subjected to enzymatic hydrolysis for 48 h followed by centrifugation and washing with deionized water. The solid residue was extracted twice with 96:4 (v/v) 1,4-dioxane/water mixture at room temperature overnight. The extracts were combined, rotary evaporated, and freeze-dried to recover the lignin. The dry lignin samples were dissolved in deuterated dimethyl sulfoxide for NMR experiments. 13C–1H HSQC experiments were performed in a Bruker Avance III HD 500 MHz NMR spectrometer operating at a frequency of 125.12 MHz for the 13C nucleus using a standard Bruker pulse sequence (hsqcetgpsisp2.2) on a Prodigy platform cryoprobe. The NMR spectra were acquired under the following acquisition conditions: 220 ppm spectral width in F1 (13C) dimension with 256 data points and 12 ppm spectral width in F2 (1H) dimension with 1024 data points, a 90° pulse, a one bond C–H coupling constant of 145 Hz, a 1.0 s pulse delay, and 64 scans. All the data was processed using the Bruker’s TopSpin 3.6 software. The NMR spectra provides structural characteristics information about lignin in field grown transgenic MOMT4 poplar. Additional meta data is embedded in the raw spectra figures.

Lignin structure, HSQC, poplar, field trial, MOMT4

Phenotypically anchored transcriptomics across diverse agrichemicals reveals conserved pathways and unique gene expression signatures in zebrafish

Agrichemicals such as herbicides, fungicides, insecticides, and biocides are widely used in agriculture, yet some are associated with adverse effects in humans and the environment. While many of these chemicals have been extensively studied in vitro and are included in the EPA’s ToxCast program, comprehensive in vivo comparisons using RNA sequencing across structurally diverse agrichemicals, in a single screening platform, are lacking. In this study, we examined structurally diverse agrichemicals found in the U.S. Environmental Protection Agency’s (EPA) Toxcast Phase I and II library by statically exposing early life stage zebrafish at 6 h post fertilization (hpf) until 120 hpf at concentrations ranging from 0.25 to 100 µM. Morphological outcomes were assessed at 120 hpf across 10 endpoints, including yolk sac edema, craniofacial malformations, and axis abnormalities. Chemicals that produced robust concentration-response relationships were selected for transcriptomic profiling. For transcriptomic analysis, zebrafish were statically exposed to each chemical and sampled at 48 hpf, prior to the onset of morphological effects observed at 120 hpf. Differential expression analysis identified between 0 and 4,538 differentially expressed genes (DEGs) per chemical, with no clear correlation to morphological severity. Both DEG and co-expression network analyses revealed chemical-specific expression patterns that converged on shared biological pathways, including neurodevelopment and cytoskeletal organization. Key regulatory genes such as mylpfa and krt4 were identified within co-expression modules, suggesting their potential role in conserved toxicity mechanisms. Semantic similarity analysis of enriched gene ontology (GO) terms, when compared to existing datasets, highlighted gaps in the annotation of neurodevelopmental processes, indicating that some in vivo effects may not be fully captured by current curated resources. The results provide new insights into the modes of action of diverse agrichemicals and establish a framework for understanding how agrichemical structure relates to biological function in a vertebrate model.

agrichemical

Carbon-13 NMR spectra of lignin isolated from field grown transgenic poplar

Here we present a curated dataset of a series of 13C nuclear magnetic resonance (NMR) spectra of lignin isolated from transgenic monolignol 4-O-methyltransferase (MOMT4) engineered poplar. The transgenic poplar was collected from a 3-year field trial experiment. The poplar was Soxhlet-extracted with toluene/ethanol and the extractives-free poplar was then ball-milled in a Retsch PM100 planetary ball mill using a porcelain jar with ceramic balls at 600 rpm for 2 h. The ball-milled materials were then subjected to enzymatic hydrolysis for 48 h followed by centrifugation and washing with deionized water. The solid residue was extracted twice with 96:4 (v/v) 1,4-dioxane/water mixture at room temperature overnight. The extracts were combined, rotary evaporated, and freeze-dried to recover the lignin. The dry lignin samples were dissolved in deuterated dimethyl sulfoxide for NMR characterization. 13C experiments were performed in a Bruker Avance III HD 500 MHz NMR spectrometer operating at a frequency of 125.12 MHz for the 13C nucleus using a standard Bruker pulse sequence (zgpg) on a Prodigy platform cryoprobe. The NMR spectra were acquired under the following conditions: spectra width 229 ppm, 64k data points, 1s pulse delay, and 6k scans. All the data was processed using the Bruker’s TopSpin 3.6 software. Additional meta data is embedded in the raw spectra files.

13C NMR, lignin, poplar, field trial, MOMT4, CBI

Cellulose crystallinity index (CrI) of switchgrass measured by solid-state NMR

Here we present a curated dataset of switchgrass cellulose crystallinity index (CrI) measured using solid state nuclear magnetic resonance (NMR) spectroscopy. Seventy-two topline switchgrass lines grown in greenhouse were collected and ground to -20/+80 mesh. The switchgrass was then Soxhlet-extracted with toluene/ethanol for 24 h to remove extractives. The extractives-free switchgrass was holopulped by using peracetic acid at 5 g loading per g biomass and the solution consistency was adjusted to 5% with DI water. Holopulping was conducted at room temperature for 24 h with stirring. The obtained holocellulose was washed excessively with DI water and air-dried at room temperature for 24 h. The dried holocellulose was treated with hydrochloric acid (2.5 M) for 2 h to remove hemicellulose. The isolated cellulose was collected by filtration, rinsed with DI water, and used to analyze cellulose crystallinity by solid-state NMR. The NMR samples were prepared by packing the moisturized cellulose into 4-mm cylindrical Zirconia MAS rotors. Cross polarization/magic angle spinning (CP/MAS) NMR analysis of cellulose was carried out on a Bruker Avance III 400-MHz spectrometer operating at 100.59 MHz for 13C in a Bruker double-resonance MAS probe head at spinning speeds of 8 kHz. The CP/MAS experiments utilized a 5 ms (90°) proton pulse, 1.5 ms contact pulse, 4 s recycle delay and 4000 scans. The cellulose crystallinity index was determined from the areas of the cellulose crystalline C4 signal (δ 86-92 ppm) over the entire C4 regions (δ 79-92 ppm) in the NMR spectra.

Cellulose, crystallinity index, switchgrass, solid

Dark Energy Survey Year 6 Results: Photometric Dataset for Cosmology

We describe the photometric dataset assembled from the full 6 yr of observations by the Dark Energy Survey (DES) in support of static-sky cosmology analyses. DES Y6 Gold is a curated dataset derived from DES Data Release 2 (DR2) that incorporates improved measurement, photometric calibration, object classification and value-added information. Y6 Gold comprises nearly 5000 deg$^{2}$ of grizY imaging in the south Galactic cap and includes 669 million objects with a depth of i$_{AB}$ ∼ 23.4 mag at a signal-to-noise ratio ∼ 10 for extended objects and a top-of-the-atmosphere photometric uniformity <2 mmag. Y6 Gold augments DES DR2 with simultaneous fits to multiepoch photometry for more robust galaxy shapes, colors, and photometric redshift estimates. Y6 Gold features improved morphological star–galaxy classification with an efficiency of 98.6% and a contamination of 0.8% for galaxies with 17.5 < i$_{AB}$ < 22.5. Additionally, it includes per-object quality information, and accompanying maps of the footprint coverage, masked regions, imaging depth, survey conditions, and astrophysical foregrounds that are used for cosmology analyses. After quality selections, benchmark samples contain 448 million galaxies and 120 million stars. This publication is complemented by data access and documentation.

79 ASTRONOMY AND ASTROPHYSICS

Two-dimensional heteronuclear single quantum coherence (HSQC) NMR spectra of lignin isolated from Populus trichocarpa residues after CELF pretreatment and CBP fermentation

Here we present a curated dataset of a series of two-dimensional heteronuclear single quantum coherence (HSQC) nuclear magnetic resonance (NMR) spectra of lignin isolated from a woody energy crop (Populus trichocarpa) residues after co-solvent enhanced lignocellulosic fractionation (CELF) pretreatment and consolidated bioprocessing (CBP) process. The natural poplar variant GW-9947 from the Center for Bioenergy Innovation (CBI) was used. The poplar was knife milled and passed through a 1 mm sieve. The CELF pretreatment was performed in a Parr autoclave reactor with 7.5 wt % solids loading, 0.5 wt% H2SO4 as catalyst at 150°C with 15, 25 and 30 minutes, respectively. Tetrahydrofuran was added in a 1:1 mass ratio with water as the pretreatment solvent. The residues from CELF pretreatment were then subjected to CBP using the bacterium C. thermocellum DSM 1313. CBP fermentations were performed at 60 °C in a shaker at 50 grams/L solids loadings. Lignin was isolated from the pretreated samples after ball-milling in a porcelain jar with ceramic balls via Retsch PM 200 at 580 rpm for 2.5 h followed by enzymatic hydrolysis in acetate buffer (pH 4.8, 50 °C) for 48 h. The lignin samples were characterized using 13C–1H HSQC experiments which were performed in a Bruker Avance III HD 500 MHz NMR spectrometer operating at a frequency of 125.12 MHz for the 13C nucleus. A standard Bruker pulse sequence was used on a Prodigy platform cryoprobe. The dry lignin samples were dissolved in deuterated dimethylsulfoxide for HSQC experiments. The spectra were acquired under the following acquisition conditions: 210 ppm spectral width in F1 (13C) dimension with 256 data points and 11 ppm spectral width in F2 (1H) dimension with 1024 data points, a 90° pulse, a one bond C–H coupling constant of 145 Hz, a 1.0 s pulse delay, and 64 scans. All the data was processed using the TopSpin 3.6 software (Bruker BioSpin). The NMR spectra provides structural characteristics information about lignin remaining in solids after CELF (150 °C with 15, 25 and 30 minutes) process and C. thermocellum CBP.

Lignin structure, HSQC, poplar, CELF, CBP, CBI

Heteronuclear single quantum coherence (HSQC) NMR spectra of lignin isolated from switchgrass residues after fermentation with milling

Here we present a curated dataset of two-dimensional heteronuclear single quantum coherence (HSQC) nuclear magnetic resonance (NMR) spectra of lignin isolated from a herbaceous energy crop (Panicum virgatum L.). The lowland variant “Timber” switchgrass from Ernst seeds was used. The switchgrass was knife milled and passed through a 2 mm sieve prior to consolidated bioprocessing (CBP) process. Switchgrass was suspended in Milli-Q water and autoclaved for 90 min on liquid cycles. The residues after autoclaving were then subjected to CBP using coculture of Clostridium thermocellum (C. thermocellum DSM 1313 (LL1004)) and Thermoanaerobacterium thermosaccarolyticum ( T. thermosaccharolyticum HG-8 ATCC 31960 (LL1244)). Once-fermented and twice-fermented (FF) switchgrass were subjected to ball and disc milling in bioreactors at 55 °C and 60, 48 grams/L solids loadings for primary and secondary fermentation respectively. When fermentations were completed, the residual solids were rinsed with milli-Q water. Lignin was isolated from the pretreated residues after ball-milling in a porcelain jar with ceramic balls via Retsch PM 200 at 600 rpm for 2 h followed by enzymatic hydrolysis in acetate buffer (pH 4.8, 50 °C) for 48 h. The dry lignin samples were dissolved in deuterated dimethyl sulfoxide (d6) and characterized using 13C–1H HSQC in a Bruker Avance III HD 500-MHz NMR spectrometer. A standard Bruker pulse sequence (hsqcetgpsisp.2) was used on a Prodigy platform cryoprobe. The spectra were acquired with the following acquisition conditions: 230 ppm spectral width in F1 (13C) dimension with 256 data points and 12 ppm spectral width in F2 (1H) dimension with 2048 data points, a 90° pulse, with a C–H coupling constant of 145 Hz, a 1.0 s pulse delay, and 64 scans. Spectra were processed using the Bruker TopSpin 3.6 software.

Lignin, HSQC, Switchgrass, CBP , Ball mill, Disc m

HSQC spectra of lignin isolated from poplar roots

Here we present a curated dataset of two-dimensional heteronuclear single quantum coherence (HSQC) nuclear magnetic resonance (NMR) spectra of lignin isolated from roots of a greenhouse grown natural population of an energy crop poplar (Populus trichocarpa). Dormant cuttings of field-grown poplar were grown in 6-liter pots in a peat-based media containing bark, perlite, vermiculite, dolomite lime and a wetting agent in an environmentally controlled greenhouse. Temperatures were between 21 and 23 °C, with supplemental lighting to support a 16-h day length using 1000-watt high-pressure sodium lights in greenhouse. Once established, all plants were cut-back, allowed to regrow and harvested at the same time following an eight-month long growth period. Plants were harvested and the belowground roots were washed off soils, blotted, dried in an oven at 70 °C for 3 days, and Wiley milled (mesh size 20). The roots were Soxhlet-extracted with toluene/ethanol for 24 h to remove extractives. The extracted roots were ball-milled in a Retsch PM100 planetary ball mill using a porcelain jar with ceramic balls at 600 rpm for 2 h (in 5 min on and 5 min off cycles to avoid excessive sample heating). The ball-milled materials were then subjected to enzymatic hydrolysis for 48 h followed by centrifugation and washing with deionized water. The solid residue was extracted twice with 96% (v/v) 1,4-dioxane/water mixture at room temperature overnight. The extracts were combined, rotary evaporated, and freeze-dried to recover lignin. The dry lignin samples were dissolved in deuterated dimethyl sulfoxide (d6) and transferred into a 5 mm tube. 13C–1H HSQC experiments were performed in a Bruker Avance III HD 500 MHz NMR spectrometer operating at a frequency of 125.12 MHz for the 13C nucleus using a standard Bruker pulse sequence on a Prodigy platform cryoprobe. The NMR spectra were acquired under the following acquisition conditions: 220 ppm spectral width in F1 (13C) dimension with 256 data points and 12 ppm spectral width in F2 (1H) dimension with 1024 data points, a 90° pulse, a one bond C–H coupling constant of 145 Hz, a 1.0 s pulse delay, and 64 scans. Spectra were processed using the Bruker TopSpin software. Additional meta data is embedded in the raw spectra figures.

HSQC, lignin, poplar, roots, CBI

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES

Mondo: integrating disease terminology across communities

Precision medicine aims to enhance diagnosis, treatment, and prognosis by integrating multimodal data at the point of care. However, challenges arise due to the vast number of diseases, differing methods of classification, and conflicting terminological coding systems and practices used to represent molecular definitions of disease. This lack of interoperability artificially constrains the potential for diagnosis, clinical decision support, care outcome analysis, as well as data linkage across research domains to support the development or repurposing of therapeutics. There is a clear and pressing need for a unified system for managing disease entities⁠—including identifiers, synonyms, and definitions. To address these issues, we created the Mondo disease ontology—a community-driven, open-source, unified disease classification system that harmonizes diverse terminologies into a consistent, computable framework. Mondo integrates key medical and biomedical terminologies, including Online Mendelian Inheritance in Man (OMIM), Orphanet, Medical Subject Headings (MeSH), National Cancer Institute Thesaurus (NCIt), and more, to provide a comprehensive and accurate representation of disease concepts with fully provenanced and attributed links back to the sources. Mondo can be used as the handle for curation of gene–disease associations utilized in diagnostic applications, research applications such as computational phenotyping, and in clinical coding systems in clinical decision support by pointing the clinician to the numerous knowledge resources linked to the Mondo identifier. Mondo's community-centric approach, stewarded by the Monarch Initiative's expertise in ontologies, ensures that the ontology remains adaptable to the evolving needs of biomedical research and clinical communities, as well as the knowledge providers.

biomedical informatics