Search NASA⌕ Search

SEARCH · Search NASA

Results for “Sample curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

414 records · Page 23

Identifying Anomalous DESI Galaxy Spectra with a Variational Autoencoder

The tens of millions of spectra being captured by the Dark Energy Spectroscopic Instrument (DESI) provide tremendous discovery potential. In this work we show how Machine Learning, in particular Variational Autoencoders (VAE), can detect anomalies in a sample of approximately 200,000 DESI spectra comprising galaxies, quasars and stars. We demonstrate that the VAE can compress the dimensionality of a spectrum by a factor of 100, while still retaining enough information to accurately reconstruct spectral features. We then detect anomalous spectra as those with high reconstruction error and those which are isolated in the VAE latent representation. The anomalies identified fall into two categories: spectra with artefacts and spectra with unique physical features. Awareness of the former can help to improve the DESI spectroscopic pipeline; whilst the latter can lead to the identification of new and unusual objects. To further curate the list of outliers, we use the Astronomaly package which employs Active Learning to provide personalised outlier recommendations for visual inspection. In this work we also explore the VAE latent space, finding that different object classes and subclasses are separated despite being unlabelled. We demonstrate the interpretability of this latent space by identifying tracks within it that correspond to various spectral characteristics. For example, we find tracks that correspond to increasing star formation and increase in broad emission lines along the Balmer series. In upcoming work we hope to apply the methods presented here to search for both systematics and astrophysically interesting objects in much larger datasets of DESI spectra.

Nicolaou, C. [University Coll. London] (ORCID:0000↗

A curated benchmark for cofolding models on kinase conformational states

Abstract Protein kinases are critical drug targets, requiring therapeutics that can modulate their active and inactive conformational states. While cofolding models can generate global folds directly from kinase sequences and ligand SMILES strings, these models have not yet been tested on their ability to recover ligand-induced-fit conformational states of the kinase proteins. Here, we introduce KinConfBench, a curated benchmark of 2225 high-quality human kinase chains to evaluate the ability of four state-of-the-art cofolding models—Boltz-2, Chai-1, Protenix, and RoseTTAFold-All-Atom—to recover both canonical and rare conformational states. We show that geometric success metrics of a ligand pose in the active site do not correlate strongly with the correct kinase conformational state, motivating a new set of dynamical benchmarks for assessing cofolding models. While all four cofolding models achieve ~60–80% prediction accuracy for kinase conformational classification, they exhibit severe mode collapse when performing multiple inferences, show negligible structural diversity in sampling induced-fit motions, and display a prevalent “apo-drift” in which most cofolding models predominantly predict the kinase to be in its ligand-free state. Our results highlight that capturing ligand-induced protein conformational diversity, not just geometric fit, is critical for next-generation structure-based drug discovery.

Sun, Kunyang↗

In-Lab Rapid Analytical Detection of Lunar Volatiles By Universal Gas Analyzer With Comparison to GC-MS System

Introduction: The curation of permanently shadowed regions (PSRs) [1] on the lunar surface is centered around studies based upon the observed volatiles from the LCROSS mission [2]. The rapid detection of important volatile gases and vapors present in planetary bodies and Astromaterials by a standalone analytical device is an area of intense research interest in our group and Planetary Exploration & Astromaterials Research Laboratory (PEARL) facility and this work is relevant to the future preparation of viable lunar simulants for testing curation efforts down the road. The groundbreaking results obtained from the LCROSS Mission [2] open the requirements for the direct detection of volatiles present in regolith materials collected from the lunar surface. The mass spectrometry of volatile chemicals is a general technique that utilizes a set of instruments that creates charged ions from a gaseous chemical species and measures the intensities vs. mass-to-charge ratio (m/z) [3]. In this context, we discuss in-lab experimental results and procedures for rapid qualitative analysis of main LCROSS volatiles (water, H2S, NH3, CO2, and CH3OH) by a Universal Gas Analyzer (UGA) instrument. Additionally, the instrument performance was evaluated by measuring the isotopic abundance ratio of atmospheric Ar-40 to Ar-36 present in room air since, argon is a relevant gas in planetary studies as it can provide an insight and reference point to isotope studies [4]. Additional, cross comparisons were attempted and made between the two instruments to develop a robust analytical technique by comparing mass spectral data for H2S headspace samples with a Trace-1310/ ISQ 7000 (ThermoFisher Scientific.) GC-MS system. Background: The benchtop UGA System is equipped with an SRS UGA 300 quadrupole mass spectrometer designed and built by Stanford Research Systems [5]. This system can be configured for several types of gaseous chemical analysis. The inlet line continuously samples gases at low flow rates (several milliliters per minute) through a capillary limiting the intake pressure making the instrument ideal for online analysis of select gases and/or room atmosphere. Moreover, in our current UGA system, a change in composition at the inlet can be detected in about 200 milliseconds and a complete spectrum is acquired (for a range of 1-100 amu) in under 45 sec with masses measured at rates up to 25 msec per point [5]. This system provides a quick upstream analytical data that we can then compare to results obtained by our GC-MS system. Sample Preparation: Small volume (2-4 mL) of analyte sample was taken in a 10 mL glass vial and sealed with a crimped cap and purged with pure Ar or N2 gas to displace air from the top. The headspace sample was scanned by the UGA instrument at analog, histogram, and pressure vs. time modes. The isotopic abundance ratio for 40Ar-to-36Ar was estimated by measuring partial pressure vs. time scans and setting the mass at 40 and 36 respectively. Results and Discussions: In this work, we have investigated the applicability of the UGA system by qualitative analysis of a series of LCROSS volatiles measured individually. Fig. 1 demonstrates a set of vertically offset spectra for the partial pressures measured as a function of mass-to-charge (m/z) ratios. The average acquisition time for each spectrum was less than a minute suggesting that the UGA system is ideal for quick analysis of geochemical volatiles. For the cross-comparison, we analyzed an H2S headspace sample by a Trace-1310/ISQ-7000 system and compared mass spectral data with previously measured UGA histogram scan data (Fig. 2). In both cases, major peak positions are the same, however, the intensities of fragment ions ([1H132S]+ and [32S]+) are higher for UGA suggesting that the fragment ionization process is stronger in UGA compared to that of GC-MS. To investigate how the integrated area under each chromatogram varies with the headspace sample volume, a set of five H2S headspace samples with increasing volumes was analyzed by the GC-MS system (Fig. 3, inset). A small volume (e.g., 200 to 1000 µL) of H2S/H2O vapor was withdrawn from a 20 mL stock sample vial containing ~5 mL of 0.4% H2S in water by a gas-tight syringe and added to another 20 mL vial filled with argon and analyzed by the GC-MS system. Finally, the UGA detector sensitivity was evaluated by calculating the atmospheric 40Ar-to-36Ar isotopic abundance ratio in room air by running a partial pressure vs. time scan with setting the atomic mass at 40, and 36. Fig. 4(a) shows a ~25 min duration “P vs. time” scan for 40Ar (plot for 36Ar is not shown). The partial pressure values (~100 points) were corrected by subtracting the corresponding background pressure value for 37Ar and utilized to calculate 40Ar-to-36Ar isotopic abundance ratios as shown by Fig. 4b. The average isotopic abundance ratio is ~306 with a 2*STDEV ~13. This abundance ratio is significantly close to the previously reported value of 298.56 [6] and the ratio obtained by our GC-MS system (303 for a UHP grade Ar sample). Conclusions: Our study strongly evidenced that the benchtop UGA system is a valuable analytical tool for the detection of major LCROSS volatiles. The rapid scanning capability, the inexpensiveness of the whole system, and impressive detection sensitivity prove its worthiness as an essential device for advanced geochemical applications. Moreover, cross comparisons with the GC-MS provide important bridges into advanced curatorial efforts into the future. References: [1] Bickel, V.T., et al. (2021) Nat Commun 12, 5607. [2] Colaprete, A., et al. (2010) Science, 330, 463-468. [3] Glavin, D. P. et al. (2012) 2012 IEEE Aerospace Conference, 1-11. [4] Willett, C. D., et al. (2022) Geochimica et Cosmochimica Acta 329, 119-134. [5] Operation Manual and Programming Reference. (2018) Universal gas Analyzers, Stanford Research Systems. [6] Lee, J. Y., et al. (2006) Geochimica et Cosmochimica Acta 70, 4507–4512. Notes: (4 figures are attached with text as shown by the attached file)

Curation↗

Two-dimensional heteronuclear single quantum coherence (HSQC) NMR spectra of lignin isolated from field grown transgenic poplar

Here we present a curated dataset of a series of two-dimensional heteronuclear single quantum coherence (HSQC) nuclear magnetic resonance (NMR) spectra of lignin isolated from a field grown transgenic poplar engineered with a monolignol 4-O-methyltransferase (MOMT4). The poplar was collected from a 2-year-old rotation trees within a three-year field trial experiment. The poplar was Soxhlet-extracted with toluene/ethanol and the extractives-free poplar was then ball-milled in a Retsch PM100 planetary ball mill using a porcelain jar with ceramic balls at 600 rpm for 2 h (in 5 min on and 5 min off cycles to avoid excessive sample heating). The ball-milled materials were then subjected to enzymatic hydrolysis for 48 h followed by centrifugation and washing with deionized water. The solid residue was extracted twice with 96:4 (v/v) 1,4-dioxane/water mixture at room temperature overnight. The extracts were combined, rotary evaporated, and freeze-dried to recover the lignin. The dry lignin samples were dissolved in deuterated dimethyl sulfoxide for NMR experiments. 13C–1H HSQC experiments were performed in a Bruker Avance III HD 500 MHz NMR spectrometer operating at a frequency of 125.12 MHz for the 13C nucleus using a standard Bruker pulse sequence (hsqcetgpsisp2.2) on a Prodigy platform cryoprobe. The NMR spectra were acquired under the following acquisition conditions: 220 ppm spectral width in F1 (13C) dimension with 256 data points and 12 ppm spectral width in F2 (1H) dimension with 1024 data points, a 90° pulse, a one bond C–H coupling constant of 145 Hz, a 1.0 s pulse delay, and 64 scans. All the data was processed using the Bruker’s TopSpin 3.6 software. The NMR spectra provides structural characteristics information about lignin in field grown transgenic MOMT4 poplar. Additional meta data is embedded in the raw spectra figures.

Lignin structure, HSQC, poplar, field trial, MOMT4↗

Phenotypically anchored transcriptomics across diverse agrichemicals reveals conserved pathways and unique gene expression signatures in zebrafish

Agrichemicals such as herbicides, fungicides, insecticides, and biocides are widely used in agriculture, yet some are associated with adverse effects in humans and the environment. While many of these chemicals have been extensively studied in vitro and are included in the EPA’s ToxCast program, comprehensive in vivo comparisons using RNA sequencing across structurally diverse agrichemicals, in a single screening platform, are lacking. In this study, we examined structurally diverse agrichemicals found in the U.S. Environmental Protection Agency’s (EPA) Toxcast Phase I and II library by statically exposing early life stage zebrafish at 6 h post fertilization (hpf) until 120 hpf at concentrations ranging from 0.25 to 100 µM. Morphological outcomes were assessed at 120 hpf across 10 endpoints, including yolk sac edema, craniofacial malformations, and axis abnormalities. Chemicals that produced robust concentration-response relationships were selected for transcriptomic profiling. For transcriptomic analysis, zebrafish were statically exposed to each chemical and sampled at 48 hpf, prior to the onset of morphological effects observed at 120 hpf. Differential expression analysis identified between 0 and 4,538 differentially expressed genes (DEGs) per chemical, with no clear correlation to morphological severity. Both DEG and co-expression network analyses revealed chemical-specific expression patterns that converged on shared biological pathways, including neurodevelopment and cytoskeletal organization. Key regulatory genes such as mylpfa and krt4 were identified within co-expression modules, suggesting their potential role in conserved toxicity mechanisms. Semantic similarity analysis of enriched gene ontology (GO) terms, when compared to existing datasets, highlighted gaps in the annotation of neurodevelopmental processes, indicating that some in vivo effects may not be fully captured by current curated resources. The results provide new insights into the modes of action of diverse agrichemicals and establish a framework for understanding how agrichemical structure relates to biological function in a vertebrate model.

agrichemical↗

Carbon-13 NMR spectra of lignin isolated from field grown transgenic poplar

Here we present a curated dataset of a series of 13C nuclear magnetic resonance (NMR) spectra of lignin isolated from transgenic monolignol 4-O-methyltransferase (MOMT4) engineered poplar. The transgenic poplar was collected from a 3-year field trial experiment. The poplar was Soxhlet-extracted with toluene/ethanol and the extractives-free poplar was then ball-milled in a Retsch PM100 planetary ball mill using a porcelain jar with ceramic balls at 600 rpm for 2 h. The ball-milled materials were then subjected to enzymatic hydrolysis for 48 h followed by centrifugation and washing with deionized water. The solid residue was extracted twice with 96:4 (v/v) 1,4-dioxane/water mixture at room temperature overnight. The extracts were combined, rotary evaporated, and freeze-dried to recover the lignin. The dry lignin samples were dissolved in deuterated dimethyl sulfoxide for NMR characterization. 13C experiments were performed in a Bruker Avance III HD 500 MHz NMR spectrometer operating at a frequency of 125.12 MHz for the 13C nucleus using a standard Bruker pulse sequence (zgpg) on a Prodigy platform cryoprobe. The NMR spectra were acquired under the following conditions: spectra width 229 ppm, 64k data points, 1s pulse delay, and 6k scans. All the data was processed using the Bruker’s TopSpin 3.6 software. Additional meta data is embedded in the raw spectra files.

13C NMR, lignin, poplar, field trial, MOMT4, CBI↗

Revisiting the Origin of Macromolecular Carbon (MMC) in Lunar Basalts 15556 & 10044

Volatile elements influence the geo-chemical evolution of planetary bodies and they are in magmas at every stage, from melting within planetary interiors to eruption at the surface. Analyses of lunar mare basalts supported the hypothesis that lunar mag-mas were depleted in volatiles (H-C-F-Cl-S), relative to their terrestrial analogs [1]. Nevertheless, several early studies of samples returned during the Apollo program proposed that the mare basalt eruptions, including the “fire fountain” eruptions, were propelled by the oxidation of magmatic graphite to CO (and/or CO2) gas [2, 3]. Seminal studies during the 1970’s measured the bulk concentration and isotopic compositions of C from Apollo 11 samples, and identified several carbonaceous compounds, including: (a) gaseous (CO, CO2, and traces of CH4), (b) metallic carbide, and (c) potentially elemental carbon [4-6]. These studies reported a relatively broad range of C contents (~100-400 μg/g) and isotopic values (δC13 = -30 to +20), and suggested that these heterogeneities can be explained by contribution from multiple factors, including: (a) indigenous carbon, (b) solar wind implantation, (c) bombardment and/or meteorite impact, and (d) terrestrial contamination [5,6]. However, unequivocal observations of magmatic graphite in lunar basalts have never been made. Macromolecular carbon (MMC)—graphitic carbon varying from nearly amorphous to highly crystalline varieties—was identified as inclusions hosted by igneous pyroxenes from Martian meteorites and were attributed to being indigenous to Mars [8]. The authors carefully considered the textural and mineralogical relationship of the MMC phases, and concluded that the subset of MMC located within and/or adjacent to cracks, or at a disrupted surface (e.g., cut) were most consistent with terrestrial contamination. The near absence of con-firmed instances of lunar magmatic MMC within the literature [9], combined with the wide range of isotopic values and bulk carbon contents measured in lunar bas-alts begs the question as to whether previously measured carbon is of an indigenous origin, or the result of terrestrial contamination. Using Raman spectroscopy, we have observed MMC in lunar basalts subjected to different forms of anthropogenic modification related to sample preparation including polished sections, sawn surfaces, and fractured surfaces adjacent to sawn sur-faces. We have observed MMC of unknown origin in all of these settings. Here we report the preliminary textural and spectroscopic characteristics of MMC hosted within the groundmass of Apollo 15 (15556) and Apollo 11 basalts (10044) as part of our ongoing investigation of the origin of these carbonaceous materials.

lunar↗

Cellulose crystallinity index (CrI) of switchgrass measured by solid-state NMR

Here we present a curated dataset of switchgrass cellulose crystallinity index (CrI) measured using solid state nuclear magnetic resonance (NMR) spectroscopy. Seventy-two topline switchgrass lines grown in greenhouse were collected and ground to -20/+80 mesh. The switchgrass was then Soxhlet-extracted with toluene/ethanol for 24 h to remove extractives. The extractives-free switchgrass was holopulped by using peracetic acid at 5 g loading per g biomass and the solution consistency was adjusted to 5% with DI water. Holopulping was conducted at room temperature for 24 h with stirring. The obtained holocellulose was washed excessively with DI water and air-dried at room temperature for 24 h. The dried holocellulose was treated with hydrochloric acid (2.5 M) for 2 h to remove hemicellulose. The isolated cellulose was collected by filtration, rinsed with DI water, and used to analyze cellulose crystallinity by solid-state NMR. The NMR samples were prepared by packing the moisturized cellulose into 4-mm cylindrical Zirconia MAS rotors. Cross polarization/magic angle spinning (CP/MAS) NMR analysis of cellulose was carried out on a Bruker Avance III 400-MHz spectrometer operating at 100.59 MHz for 13C in a Bruker double-resonance MAS probe head at spinning speeds of 8 kHz. The CP/MAS experiments utilized a 5 ms (90°) proton pulse, 1.5 ms contact pulse, 4 s recycle delay and 4000 scans. The cellulose crystallinity index was determined from the areas of the cellulose crystalline C4 signal (δ 86-92 ppm) over the entire C4 regions (δ 79-92 ppm) in the NMR spectra.

Cellulose, crystallinity index, switchgrass, solid↗

Dark Energy Survey Year 6 Results: Photometric Dataset for Cosmology

We describe the photometric dataset assembled from the full 6 yr of observations by the Dark Energy Survey (DES) in support of static-sky cosmology analyses. DES Y6 Gold is a curated dataset derived from DES Data Release 2 (DR2) that incorporates improved measurement, photometric calibration, object classification and value-added information. Y6 Gold comprises nearly 5000 deg$^{2}$ of grizY imaging in the south Galactic cap and includes 669 million objects with a depth of i$_{AB}$ ∼ 23.4 mag at a signal-to-noise ratio ∼ 10 for extended objects and a top-of-the-atmosphere photometric uniformity <2 mmag. Y6 Gold augments DES DR2 with simultaneous fits to multiepoch photometry for more robust galaxy shapes, colors, and photometric redshift estimates. Y6 Gold features improved morphological star–galaxy classification with an efficiency of 98.6% and a contamination of 0.8% for galaxies with 17.5 < i$_{AB}$ < 22.5. Additionally, it includes per-object quality information, and accompanying maps of the footprint coverage, masked regions, imaging depth, survey conditions, and astrophysical foregrounds that are used for cosmology analyses. After quality selections, benchmark samples contain 448 million galaxies and 120 million stars. This publication is complemented by data access and documentation.

79 ASTRONOMY AND ASTROPHYSICS↗

Two-dimensional heteronuclear single quantum coherence (HSQC) NMR spectra of lignin isolated from Populus trichocarpa residues after CELF pretreatment and CBP fermentation

Here we present a curated dataset of a series of two-dimensional heteronuclear single quantum coherence (HSQC) nuclear magnetic resonance (NMR) spectra of lignin isolated from a woody energy crop (Populus trichocarpa) residues after co-solvent enhanced lignocellulosic fractionation (CELF) pretreatment and consolidated bioprocessing (CBP) process. The natural poplar variant GW-9947 from the Center for Bioenergy Innovation (CBI) was used. The poplar was knife milled and passed through a 1 mm sieve. The CELF pretreatment was performed in a Parr autoclave reactor with 7.5 wt % solids loading, 0.5 wt% H2SO4 as catalyst at 150°C with 15, 25 and 30 minutes, respectively. Tetrahydrofuran was added in a 1:1 mass ratio with water as the pretreatment solvent. The residues from CELF pretreatment were then subjected to CBP using the bacterium C. thermocellum DSM 1313. CBP fermentations were performed at 60 °C in a shaker at 50 grams/L solids loadings. Lignin was isolated from the pretreated samples after ball-milling in a porcelain jar with ceramic balls via Retsch PM 200 at 580 rpm for 2.5 h followed by enzymatic hydrolysis in acetate buffer (pH 4.8, 50 °C) for 48 h. The lignin samples were characterized using 13C–1H HSQC experiments which were performed in a Bruker Avance III HD 500 MHz NMR spectrometer operating at a frequency of 125.12 MHz for the 13C nucleus. A standard Bruker pulse sequence was used on a Prodigy platform cryoprobe. The dry lignin samples were dissolved in deuterated dimethylsulfoxide for HSQC experiments. The spectra were acquired under the following acquisition conditions: 210 ppm spectral width in F1 (13C) dimension with 256 data points and 11 ppm spectral width in F2 (1H) dimension with 1024 data points, a 90° pulse, a one bond C–H coupling constant of 145 Hz, a 1.0 s pulse delay, and 64 scans. All the data was processed using the TopSpin 3.6 software (Bruker BioSpin). The NMR spectra provides structural characteristics information about lignin remaining in solids after CELF (150 °C with 15, 25 and 30 minutes) process and C. thermocellum CBP.

Lignin structure, HSQC, poplar, CELF, CBP, CBI↗

Statistical Classification of Biosignature Information using Multiple Instrument Observations

The accurate identification of biosignatures (indications of life) from data taken from remote or in situ planetary exploration is one of the most important challenges in astrobiology, the interdisciplinary field examining habitability and the potential for extraterrestrial life. This study employs machine learning algorithms to optimize the identification of biosignatures, with an emphasis on those which are agnostic to a specific biochemical basis. We exploit the wealth of terrestrial data available from biogenic and abiogenic systems to enhance efficient feature prioritization. Our dataset, pulled from public databases and laboratory recorded measurements, includes elemental abundance, isotopic fractionation, and VNIR/Raman spectra The data curation process included standardization for detection limits and ranges. Subsequent feature extraction yielded detailed inputs for machine learning, including combinations of elemental content, isotopic ratios, and parameters of spectral peaks and troughs. Feature significance was evaluated across diverse machine learning methodologies, such as k-nearest neighbors, logistic regression, Random Forest, support vector machines, and Gaussian Naïve Bayes, along with a combined voting classifier. We utilized Receiver Operating Characteristic Area Under the Curve (ROC AUC) across 2,000 50% test-train splits as a robust metric of model performance. Results revealed a promising ROC AUC of 0.853 for the combined voting classifier. Removing elemental abundance data notably reduced model accuracy (13% decrease in AUC), highlighting its critical role in biosignature detection. Several other individual data features exhibited significance within their respective data types, offering additional granularity. This research fortifies the relevance of machine learning to astrobiology, potentially enhancing life detection missions by allowing algorithmic prioritization of high-interest samples for further investigation. Future work will refine data standardization, expand the dataset to include more terrestrial systems, and incorporate convolutional neural networks for spectral feature extraction. The potential for public data sharing is also under exploration, reinforcing our commitment to collective scientific advancement.

Statistical↗

A Combined Study Investigating the Insoluble and Soluble Organic Compounds in Category 3 Carbonaceous Itokawa Particles Recovered by the Hayabusa Mission

At the 3rd International Announcement of Opportunity (AO), we have been approved for five Category 3 carbonaceous Itokawa particles (RA-QD02-0012, RA-QD02-0078, RB-CV-0029, RB-CV-0080 and RB-QD04-0052) recovered by the first Hayabusa mission of JAXA. In this investigation, we aim to provide a comprehensive study to characterize and account for the presence of carbon-bearing phases as suggested by the initial Scanning Electron Microscopy (SEM) analysis carried out by JAXA at the curation facility, and to describe the mineralogical components of the particles. The insoluble organic content of Itokawa particle has been investigated with the use of micro-Raman spectroscopy by Kitajima and co-workers [1]. The Raman spectra of Itokawa particles show broad G- and D-bands typical of low temperature material which offers an interesting contrast to the high metamorphic grade (LL4-6) of the Itokawa parent body. Amino acid analysis has been conducted by Naraoka et al. [2] to study the soluble organic component of Itokawa particles, but since it was a preliminary study and thus did not have the opportunity to target on Category 3 carbonaceous particles, only terrestrial contaminants were identified. The investigation will be carried out in the following order prioritized according to the progressive damage the analytical techniques can induce: (1) micro-Raman spectrometry, (2) two-step laser mass spectrometry (micro-L2MS), (3) ultra-high performance liquid chromatography with fluorescence detection and time-of-flight mass spectrometry (LC-FD/ToF-MS), and optimally if we can recover the particles after wet chemistry analysis, we will mount the samples and perform (4) electron beam microscopy (SEM, electron back-scattered diffraction [EBSD]) and (5) carbon X-ray absorption near edge structure spectroscopy (C-XANES). We will begin the analytical procedures upon receiving the samples in September/October. This work will provide us with an understanding of the variety and origins of the carbon-bearing phases present in primitive solar system bodies from a direct sample-returned mission, which is less likely hampered by risks of terrestrial contamination as compared to meteorite finds and falls.

Chan, Q. H. S.↗

Heteronuclear single quantum coherence (HSQC) NMR spectra of lignin isolated from switchgrass residues after fermentation with milling

Here we present a curated dataset of two-dimensional heteronuclear single quantum coherence (HSQC) nuclear magnetic resonance (NMR) spectra of lignin isolated from a herbaceous energy crop (Panicum virgatum L.). The lowland variant “Timber” switchgrass from Ernst seeds was used. The switchgrass was knife milled and passed through a 2 mm sieve prior to consolidated bioprocessing (CBP) process. Switchgrass was suspended in Milli-Q water and autoclaved for 90 min on liquid cycles. The residues after autoclaving were then subjected to CBP using coculture of Clostridium thermocellum (C. thermocellum DSM 1313 (LL1004)) and Thermoanaerobacterium thermosaccarolyticum ( T. thermosaccharolyticum HG-8 ATCC 31960 (LL1244)). Once-fermented and twice-fermented (FF) switchgrass were subjected to ball and disc milling in bioreactors at 55 °C and 60, 48 grams/L solids loadings for primary and secondary fermentation respectively. When fermentations were completed, the residual solids were rinsed with milli-Q water. Lignin was isolated from the pretreated residues after ball-milling in a porcelain jar with ceramic balls via Retsch PM 200 at 600 rpm for 2 h followed by enzymatic hydrolysis in acetate buffer (pH 4.8, 50 °C) for 48 h. The dry lignin samples were dissolved in deuterated dimethyl sulfoxide (d6) and characterized using 13C–1H HSQC in a Bruker Avance III HD 500-MHz NMR spectrometer. A standard Bruker pulse sequence (hsqcetgpsisp.2) was used on a Prodigy platform cryoprobe. The spectra were acquired with the following acquisition conditions: 230 ppm spectral width in F1 (13C) dimension with 256 data points and 12 ppm spectral width in F2 (1H) dimension with 2048 data points, a 90° pulse, with a C–H coupling constant of 145 Hz, a 1.0 s pulse delay, and 64 scans. Spectra were processed using the Bruker TopSpin 3.6 software.

Lignin, HSQC, Switchgrass, CBP , Ball mill, Disc m↗

Carbon and Oxygen Isotope Measurements of Ordinary Chondrite (OC) Meteorites from Antarctica Indicate Distinct Carbonate Species Using a Stepped Acid Extraction Procedure

The purpose of this study is to characterize the stable isotope values of terrestrial, secondary carbonate minerals from five Ordinary Chondrite (OC) meteorites collected in Antarctica. These samples were identified and requested from NASA based upon their size, alteration history, and collection proximity to known Martian meteorites. They are also assumed to be carbonate-free before falling to Earth. This research addresses two questions involving Mars carbonates: 1) characterize terrestrial, secondary carbonate isotope values to apply to Martian meteorites for isolating in-situ carbonates, and 2) increase understanding of carbonates formed in cold and arid environments with Antarctica as an analog for Mars. Two samples from each meteorite, each approximately 0.5 grams, were crushed and dissolved in pure phosphoric acid for 3 sequential reactions: a) R times 0 for 1 hour at 30 degrees Centigrade (fine calcite extraction), b) R times 1 for 18 hours at 30 degrees Centigrade (course calcite extraction), and c) R times 2 for 3 hours at 150 degrees Centigrade (siderite and/or magnesite extraction). CO (sub 2) was distilled by freezing with liquid nitrogen from each sample tube, then separated from organics and sulfides with a TRACE GC using a Restek HayeSep Q 80/100 6 foot 2 millimeter stainless column, and then analyzed on a Thermo MAT 253 Isotope Ratio Mass Spectrometer (IRMS) in Dual Inlet mode. This system was built at NASA/JSC over the past 3 years and proof-tested with known carbonate standards to develop procedures, assess yield, and quantify expected error bands. Two distinct species of carbonates are found: 1) calcite, and 2) non-calcite carbonate (future testing will attempt to differentiate siderite from magnesite). Preliminary results indicate the terrestrial carbonates are formed at approximately sigma (sup 13) C equal to plus 5 per mille, which is consistent with atmospheric CO (sub 2) sigma (sup 13) C equal to minus 7 per mille and fractionation of plus12 per mille based upon polar temperature of -20 degrees Centigrade. The oxygen values fractionate sigma (sup 18) O equal to minus 10-20 per mille lighter between the R times 0 and R times 1 reactions at 30 degrees Centigrade. The carbonate oxygen isotope measurements are consistently heavier than expected with meteoric water and temperatures from Antarctica, perhaps due to secondary carbonate formation during curation in Houston, TX.

Evans, Michael E.↗

HSQC spectra of lignin isolated from poplar roots

Here we present a curated dataset of two-dimensional heteronuclear single quantum coherence (HSQC) nuclear magnetic resonance (NMR) spectra of lignin isolated from roots of a greenhouse grown natural population of an energy crop poplar (Populus trichocarpa). Dormant cuttings of field-grown poplar were grown in 6-liter pots in a peat-based media containing bark, perlite, vermiculite, dolomite lime and a wetting agent in an environmentally controlled greenhouse. Temperatures were between 21 and 23 °C, with supplemental lighting to support a 16-h day length using 1000-watt high-pressure sodium lights in greenhouse. Once established, all plants were cut-back, allowed to regrow and harvested at the same time following an eight-month long growth period. Plants were harvested and the belowground roots were washed off soils, blotted, dried in an oven at 70 °C for 3 days, and Wiley milled (mesh size 20). The roots were Soxhlet-extracted with toluene/ethanol for 24 h to remove extractives. The extracted roots were ball-milled in a Retsch PM100 planetary ball mill using a porcelain jar with ceramic balls at 600 rpm for 2 h (in 5 min on and 5 min off cycles to avoid excessive sample heating). The ball-milled materials were then subjected to enzymatic hydrolysis for 48 h followed by centrifugation and washing with deionized water. The solid residue was extracted twice with 96% (v/v) 1,4-dioxane/water mixture at room temperature overnight. The extracts were combined, rotary evaporated, and freeze-dried to recover lignin. The dry lignin samples were dissolved in deuterated dimethyl sulfoxide (d6) and transferred into a 5 mm tube. 13C–1H HSQC experiments were performed in a Bruker Avance III HD 500 MHz NMR spectrometer operating at a frequency of 125.12 MHz for the 13C nucleus using a standard Bruker pulse sequence on a Prodigy platform cryoprobe. The NMR spectra were acquired under the following acquisition conditions: 220 ppm spectral width in F1 (13C) dimension with 256 data points and 12 ppm spectral width in F2 (1H) dimension with 1024 data points, a 90° pulse, a one bond C–H coupling constant of 145 Hz, a 1.0 s pulse delay, and 64 scans. Spectra were processed using the Bruker TopSpin software. Additional meta data is embedded in the raw spectra figures.

HSQC, lignin, poplar, roots, CBI↗

GeneLab: A Systems Biology Platform for Spaceflight Omics Data

NASA's mission includes expanding our understanding of biological systems to improve life on Earth and to enable long-duration human exploration of space. Resources to support large numbers of spaceflight investigations are limited. NASA's GeneLab project is maximizing the science output from these experiments by: (1) developing a unique public bioinformatics database that includes space bioscience relevant "omics" data (genomics, transcriptomics, proteomics, and metabolomics) and experimental metadata; (2) partnering with NASA-funded flight experiments through bio-sample sharing or sample augmentation to expedite omics data input to the GeneLab database; and (3) developing community-driven reference flight experiments. The first database, GeneLab Data System Version 1.0, went online in April 2015. V1.0 contains numerous flight datasets and has search and download capabilities. Version 2.0 will be released in 2016 and will link to analytic tools. In 2015 Genelab partnered with two Biological Research in Canisters experiments (BBRIC-19 and BRIC-20) which examine responses of Arabidopsis thaliana to spaceflight. GeneLab also partnered with Rodent Research-1 (RR1), the maiden flight to test the newly developed rodent habitat. GeneLab developed protocols for maxiumum yield of RNA, DNA and protein from precious RR-1 tissues harvested and preserved during the SpaceX-4 mission, as well as from tissues from mice that were frozen intact during spaceflight and later dissected. GeneLab is establishing partnerships with at least three planned flights for 2016. Organism-specific nationwide Science Definition Teams (SDTs) will define future GeneLab dedicated missions and ensure the broader scientific impact of the GeneLab missions. GeneLab ensures prompt release and open access to all high-throughput omics data from spaceflight and ground-based simulations of microgravity and radiation. Overall, GeneLab will facilitate the generation and query of parallel multi-omics data, and deep curation of metadata for integrative analysis, allowing researchers to uncover cellular networks as observed in systems biology platforms. Consequently, the scientific community will have access to a more complete picture of functional and regulatory networks responsive to the spaceflight environment.. Analysis of GeneLab data will contribute fundamental knowledge of how the space environment affects biological systems, and enable emerging terrestrial benefits resulting from mitigation strategies to prevent effects observed during exposure to space. As a result, open access to the data will foster new hypothesis-driven research for future spaceflight studies spanning basic science to translational science.

proteomics↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗

Mondo: integrating disease terminology across communities

Precision medicine aims to enhance diagnosis, treatment, and prognosis by integrating multimodal data at the point of care. However, challenges arise due to the vast number of diseases, differing methods of classification, and conflicting terminological coding systems and practices used to represent molecular definitions of disease. This lack of interoperability artificially constrains the potential for diagnosis, clinical decision support, care outcome analysis, as well as data linkage across research domains to support the development or repurposing of therapeutics. There is a clear and pressing need for a unified system for managing disease entities⁠—including identifiers, synonyms, and definitions. To address these issues, we created the Mondo disease ontology—a community-driven, open-source, unified disease classification system that harmonizes diverse terminologies into a consistent, computable framework. Mondo integrates key medical and biomedical terminologies, including Online Mendelian Inheritance in Man (OMIM), Orphanet, Medical Subject Headings (MeSH), National Cancer Institute Thesaurus (NCIt), and more, to provide a comprehensive and accurate representation of disease concepts with fully provenanced and attributed links back to the sources. Mondo can be used as the handle for curation of gene–disease associations utilized in diagnostic applications, research applications such as computational phenotyping, and in clinical coding systems in clinical decision support by pointing the clinician to the numerous knowledge resources linked to the Mondo identifier. Mondo's community-centric approach, stewarded by the Monarch Initiative's expertise in ontologies, ensures that the ontology remains adaptable to the evolving needs of biomedical research and clinical communities, as well as the knowledge providers.

biomedical informatics↗