Search NASA⌕ Search

SEARCH · Search NASA

Results for “Sequencing data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Automation of Laser Plasma Focused Ion Beam Microscopy for Next-Gen Energy Materials

Automation can revolutionize the use of ultrafast laser ablation and plasma-focused ion beam (PFIB) techniques for high-throughput, reproducible cross-sectioning and various sample preparation in materials characterization. As these methods become essential for analyzing complex energy materials and next-generation devices, efficient, standardized workflows are needed to minimize variability and enhance precision. This work highlights our advancements in developing automated processes for sample preparation that integrates machine learning, workflow optimization, and large-scale data acquisition to improve efficiency and scalability in applications such as electrolyzers, photovoltaic cells, and microelectronics. To streamline cross-sectioning and lamella fabrication, we have implemented fully automated workflows that standardize laser ablation and PFIB milling sequences. These workflows incorporate pre-programmed protocols for material removal, alignment, and thinning, reducing user intervention and ensuring consistency across different sample types. Machine learning algorithms further enhance automation by predicting optimal milling strategies and adapting parameters based on material properties and sectioning requirements. This approach significantly improves throughput while maintaining the structural integrity of prepared samples for high-resolution imaging and analysis, including transmission electron microscopy. Beyond sample preparation, our automation platform enables the acquisition of large, high-resolution datasets through serial sectioning, image alignment, and 3D reconstruction. These automated routines facilitate multi-scale characterization, capturing structural and compositional details from the nanoscale to the device level. By reducing variability and increasing efficiency, our automated approach enhances defect analysis, failure diagnostics, and process optimization, accelerating advancements in materials research and device engineering.

36 MATERIALS SCIENCE↗

Comprehensive SNP Data for 1,323 GWAS Population in Populus trichocarpa and Combined Annotation Files for P. trichocarpa v3.0 and v3.1

The VCF dataset includes genetic variations found in 1,323 Populus trichocarpa genotypes, providing valuable information for scientists studying plant genetics. Researchers have generated this dataset using whole-genome DNA short-read sequencing on the Illumina Genome Analyzer, HiSeq 2000, and HiSeq 2500 platforms. This sequencing effort ensured a minimum expected sequencing depth of 15×. The dataset comprises more than 9.7 million single nucleotide polymorphisms (SNPs) and indel variants. The combined annotation files are derived from P. trichocarpa v3.0 and v3.1. We merged these files to create a comprehensive annotation file used for GWAS analysis. In total, 38,830 genes overlapped between the two versions. For overlapping genes, we defined the start as the smaller and the end as the larger among the two versions to increase the likelihood of locating candidate genetic loci. Additionally, we included 2,505 unique genes from v3.0 and 4,120 unique genes from v3.1, resulting in a total of 45,455 genes in the updated annotation file.

09 BIOMASS FUELS↗

Genomes OnLine Database (GOLD) v.10: new features and updates

The Genomes OnLine Database (GOLD; https://gold.jgi.doe.gov/) at the Department of Energy Joint Genome Institute is a comprehensive online metadata repository designed to catalog and manage information related to (meta)genomic sequence projects. GOLD provides a centralized platform where researchers can access a wide array of metadata from its four organization levels namely Study, Organism/Biosample, Sequencing Project and Analysis Project. GOLD continues to serve as a valuable resource and has seen significant growth and expansion since its inception in 1997. With its expanded role as a collaborative platform, it not only actively imports data from other primary repositories like National Center for Biotechnology Information but also supports contributions from researchers worldwide. This collaborative approach has enriched the database with diverse datasets, creating a more integrated resource to enhance scientific insights. As genomic research becomes increasingly integral to various scientific disciplines, more researchers and institutions are turning to GOLD for their metadata needs. To meet this growing demand, GOLD has expanded by adding diverse metadata fields, intuitive features, advanced search capabilities and enhanced data visualization tools, making it easier for users to find and interpret relevant information. This manuscript provides an update and highlights the new features introduced over the last 2 years.

59 BASIC BIOLOGICAL SCIENCES↗

Pixel-Resolved Long-Context Learning for Turbulence at Exascale: Resolving Small-scale Eddies Toward the Viscous Limit

Turbulence plays a crucial role in multiphysics applications, including aerodynamics, fusion, and combustion. Accurately capturing turbulence's multiscale characteristics is essential for reliable predictions of multiphysics interactions, but remains a grand challenge even for exascale supercomputers and advanced deep learning models. The extreme-resolution data required to represent turbulence, ranging from billions to trillions of grid points, pose prohibitive computational costs for models based on architectures like vision transformers. To address this challenge, we introduce a multiscale hierarchical Turbulence Transformer that reduces sequence length from billions to a few millions and a novel RingX sequence parallelism approach that enables scalable long-context learning. We perform scaling and science runs on the Frontier supercomputer. Our approach demonstrates excellent performance up to 1.1 EFLOPS on 32,768 AMD GPUs, with a scaling efficiency of 94\%. To our knowledge, this is the first AI model for turbulence that can capture small-scale eddies down to the dissipative range in three-dimensional turbulence at high Reynolds numbers.

Yin, Junqi [ORNL] (ORCID:0000000338435520)↗

Transportability of exogenous microbial community correlates with interwell connectivity in deep aquifers

Subsurface resource engineering operations often utilize continuous injection of externally-sourced water into geological reservoirs for formation pressure maintenance, resource recovery or energy/waste storage. Such injected water generally contains naturally occurring microbes. Little is known, however, about how the injectate microbes transport through geological media as a community, how such transportability is affected by injector-producer connectivity, and whether such knowledge can be utilized for flowpath characterization. In this study, we analyzed daily-to-weekly timeseries microbial community data from the injected- and produced-fluids of a ten-month flow test at a deep, well-characterized engineered aquifer. We found that the injectate microbial community was distinct from the indigenous community at the amplicon sequence variant (ASV) level, and that the transportability of injectate community towards a given producer, quantified by an “nASV-Overlap” metric we propose, had strong and significant positive correlation with known injector-producer connectivities at our site. This suggests that the better the connectivity, the higher the probability for more injectate species to flow through the interwell region and arrive at a producer. Because interwell connectivity is an important yet usually unknown parameter in subsurface resource engineering, such correlation in turn points to nASV-Overlap as a useful indicator of interwell connectivity for aquifer characterization and long-term monitoring. Based on our findings, an nASV-Overlap-based microbial tracing approach was developed for characterizing and monitoring the relative connectivities across multiple producers with a given injector. A side-by-side comparison between the new nASV-Overlap approach and traditional artificial tracer methods is presented, and their respective strengths and limitations are discussed.

Deep biosphere↗

High-throughput micro-scale bandgap mapping for perovskite-inspired materials with complex composition space

Abstract To realize the full promise of high-throughput experimental workflows, the rate of sample synthesis must be matched by that of characterization. Of growing interest are contactless optical techniques that can rapidly measure material homogeneity and properties. Here, we present a hyperspectral imaging method to measure local optical bandgap distributions within samples, utilizing spatially-resolved reflectance spectra coupled with automated data analysis. We collect approximately one million optical bandgap data across the compositional space of Cs 3 (Bi x Sb 1-x ) 2 (Br y I 1-y ) 9 perovskite-inspired materials. Our results show non-monotonic bandgap variations (i.e., bandgap bowing) along six composition gradient sequences, in addition to identifying samples with multiple bandgaps in statistics. High-throughput transient absorption spectroscopy reveals that within these compositions, the depletion of the ground state carriers to excited states occurred at discrete energy levels with independent carrier dynamics, consistent with the bandgap observation and indicative of phase separation. This work demonstrates the potential for rapid optical measurements to assess material quality and homogeneity in a high-throughput experimental setting, supporting screening and recipe optimization of optoelectronic material candidates with desired carrier dynamics and optical properties.

Science & Technology - Other Topics↗

Automating the Analysis of Large Language Models Responses through Zero-Shot Question Answering

Recent advancements in Large Language Models (LLMs) have shown significant potential in various applications, yet their evaluation, particularly in zero-shot question answering scenarios, remains a challenging task. In this study, our objective was to explore precision metrics for Large Language Models (LLM) and design and implement a software pipeline to automatically evaluate LLMs' outputs under zero-shot question answering. Zero-shot question answering involves a model providing answers to questions about topics it hasn't seen during training. It leverages the principles of zero-shot learning by relying on semantic understanding and generalization from related knowledge. The data used was metadata from medical databases on congenital heart disease. We explored eleven LLM metrics and selected three for our evaluation: BLEU, BERTScore, and MoverScore. BLEU calculates a score based on the overlap of n-grams (contiguous sequences of n items, typically words) between the machine-generated translation and the reference translations. Higher BLEU scores indicate better correspondence between the machine-generated and human-generated translations. BERTScore is a metric used to evaluate the quality of machine-generated text by measuring the similarity of token embeddings produced by BERT (Bidirectional Encoder Representations from Transformers) between the generated text and reference text. MoverScore is a metric that quantifies the dissimilarity between the distributions of word embeddings from machine-generated text and reference text, emphasizing semantic similarity over exact token overlap. We also introduced HBKI, a composite metric summarizing these approaches. We tested five models —GPT-3, Llama-2, Gemini 1.5 Pro, Solar 10.7B, and Mixtral-8x7b. Our software pipeline, designed and implemented using Object-Oriented Programming principles, allows users to customize the selection and extraction of features for topics of interest in their own research. Our results show that MoverScore delivered the most precise evaluation of the LLM's outputs, while Mixtral-8x7b achieved the best overall performance in extracting metadata from the databases.

97 MATHEMATICS AND COMPUTING↗

Two-dimensional heteronuclear single quantum coherence (HSQC) NMR spectra of lignin isolated from Populus trichocarpa residues after CELF pretreatment and CBP fermentation

Here we present a curated dataset of a series of two-dimensional heteronuclear single quantum coherence (HSQC) nuclear magnetic resonance (NMR) spectra of lignin isolated from a woody energy crop (Populus trichocarpa) residues after co-solvent enhanced lignocellulosic fractionation (CELF) pretreatment and consolidated bioprocessing (CBP) process. The natural poplar variant GW-9947 from the Center for Bioenergy Innovation (CBI) was used. The poplar was knife milled and passed through a 1 mm sieve. The CELF pretreatment was performed in a Parr autoclave reactor with 7.5 wt % solids loading, 0.5 wt% H2SO4 as catalyst at 150°C with 15, 25 and 30 minutes, respectively. Tetrahydrofuran was added in a 1:1 mass ratio with water as the pretreatment solvent. The residues from CELF pretreatment were then subjected to CBP using the bacterium C. thermocellum DSM 1313. CBP fermentations were performed at 60 °C in a shaker at 50 grams/L solids loadings. Lignin was isolated from the pretreated samples after ball-milling in a porcelain jar with ceramic balls via Retsch PM 200 at 580 rpm for 2.5 h followed by enzymatic hydrolysis in acetate buffer (pH 4.8, 50 °C) for 48 h. The lignin samples were characterized using 13C–1H HSQC experiments which were performed in a Bruker Avance III HD 500 MHz NMR spectrometer operating at a frequency of 125.12 MHz for the 13C nucleus. A standard Bruker pulse sequence was used on a Prodigy platform cryoprobe. The dry lignin samples were dissolved in deuterated dimethylsulfoxide for HSQC experiments. The spectra were acquired under the following acquisition conditions: 210 ppm spectral width in F1 (13C) dimension with 256 data points and 11 ppm spectral width in F2 (1H) dimension with 1024 data points, a 90° pulse, a one bond C–H coupling constant of 145 Hz, a 1.0 s pulse delay, and 64 scans. All the data was processed using the TopSpin 3.6 software (Bruker BioSpin). The NMR spectra provides structural characteristics information about lignin remaining in solids after CELF (150 °C with 15, 25 and 30 minutes) process and C. thermocellum CBP.

Lignin structure, HSQC, poplar, CELF, CBP, CBI↗

Data for Comparison of Genotyping Assays for Detection of Targeted CRISPR/Cas Mutagenesis in Highly Polyploid Sugarcane

Sugarcane ( Saccharum spp.) is an important biofuel feedstock and a leading source of global table sugar. Saccharum hybrid cultivars are highly polyploid (2n = 100–130), containing large numbers of functionally redundant hom(e)ologs in their genomes. Genome editing with sequence-specific nucleases holds tremendous promise for sugarcane breeding. However, identification of plants with the desired level of co-editing within a pool of primary transformants can be difficult. While DNA sequencing provides direct evidence of targeted mutagenesis, it is cost-prohibitive as a primary screening method in sugarcane and most other methods of identifying mutant lines have not been optimized for use in highly polyploid species. In this study, non-sequencing methods of mutant screening, including capillary electrophoresis (CE), Cas9 RNP assay, and high-resolution melt analysis (HRMA), were compared to assess their potential for CRISPR/Cas9-mediated mutant screening in sugarcane. These assays were used to analyze sugarcane lines containing mutations at one or more of six sgRNA target sites. All three methods distinguished edited lines from wild type, with co-mutation frequencies ranging from 2% to 100%. Cas9 RNP assays were able to identify mutant sugarcane lines with as low as 3.2% co-mutation frequency, and samples could be scored based on undigested band intensity. CE was highlighted as the most comprehensive assay, delivering precise information on both mutagenesis frequency and indel size to a 1 bp resolution across all six targets. This represents an economical and comprehensive alternative to sequencing-based genotyping methods which could be applied in other polyploid species.

Genomics↗

Canted antiferromagnetism and spin reorientation in corner-shared single chain quasi-one-dimensional Ba 2 ⁢FeSe 3

Here, we report the canted antiferromagnetic (AFM) structure together with a spin reorientation in a single chain quasi-one-dimensional (Q-1D) iron chalcogenide Ba 2⁢ FeSe 3 . Ba 2 ⁢FeSe 3 crystallizes in Pnma (No. 62) orthorhombic structure with linear single iron chains consisting of corner-shared distorted FeSe 4 tetrahedra along the 𝑏 axis. Ba 2 ⁢FeSe 3 is a narrow-gap semiconductor and orders AFM below 60 K. Modeling of neutron powder diffraction data reveals a canted AFM ground state of magnetic space group 𝑃⁢𝑎⁢21/𝑐 (BNS No. 14.80) with commensurate propagation vector 𝐤 =(0, $\frac{1}{2}$, 0), where the Fe ion spins are AFM aligned with up-down-up-down (↑−↓−↑−↓) sequence along the Q-1D chain direction of the 𝑏 axis. In the magnetically ordered state, the canting of magnetic moments reorients from the 𝑎⁢𝑐 plane to the 𝑎⁢𝑏 plane below 30 K, with a 10° tilting angle toward the 𝑎 axis, and the magnetic moment does not induce a net moment in either orientation. The density functional theory results indicate that an ↑−↓−↑−↓ AFM state is stabilized along the chain direction. In this work, we elucidate the unique canted AFM of the iron chalcogenide and pave the way for searching exotic physics in Q-1D Ba 2⁢ FeSe 3 .

Gao, Fei [Univ. of Texas at Dallas, Richardson, TX↗

Estimating Sparse Direct Effects in Multivariate Regression With the Spike-and-Slab LASSO

The multivariate regression interpretation of the Gaussian chain graph model simultaneously parametrizes (i) the direct effects of p predictors on q outcomes and (ii) the residual partial covariances between pairs of outcomes. We introduce a new method for fitting sparse versions of these models with spike-and-slab LASSO (SSL) priors. We develop an Expectation Conditional Maximization algorithm to obtain sparse estimates of the p × q matrix of direct effects and the q × q residual precision matrix. Our algorithm iteratively solves a sequence of penalized maximum likelihood problems with self-adaptive penalties that gradually filter out negligible regression coefficients and partial covariances. Because it adaptively penalizes individual model parameters, our method is seen to outperform fixed-penalty competitors on simulated data. We establish the posterior contraction rate for our model, buttressing our method’s excellent empirical performance with strong theoretical guarantees. Using our method, we estimated the direct effects of diet and residence type on the composition of the gut microbiome of elderly adults.

EM algorithm↗

Functional characterization of glycosyltransferases in duckweed to enable predictive biology

Glycosyltransferases (GTs) catalyze the formation of glycosidic linkages to produce almost all complex carbohydrates. This project used a multi-disciplinary, high-throughput (HTP) biochemical and computational biology approach focused on duckweed as a model energy crop, to study carbohydrate metabolic processes. To achieve this, developed and carried out out high-throughput (HTP) functional characterization of plant glycosyltransferases (GTs) role of enzymatic microenvironments be assessed through a combined proteomic and computational biology approach, and the combined data was used to populate deep-learning frameworks to predict plant GT function. Functional validation achieved through this research is being used to assign gene function and study plant processes at the systems level to efficiently link the genome sequence with gene function. Together, the combined approaches used within this study provide a foundation for how computational prediction, in combination with high-throughput functional validation, can be used to study plant processes at the systems level and translate knowledge gained to efficiently link genome sequence with gene function in a species agnostic manner.

09 BIOMASS FUELS↗

Frosted Tracks

SAND2025-01893O Frosted Tracks is a software tool to group trajectories according to sequences of their behavior. The goal is to start with a very large number of trajectories and identify groups that exhibit similar behavior patterns. The application combines TICC and Metric DBSCAN clustering algorithms for behavioral segmentation and labeling of air/sea trajectory data. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Dalbey, Keith↗

Single-cell proteomics of Arabidopsis leaf mesophyll reveals dynamic protein responses to water-deficit stress

Background The application of single-cell omics tools to biological systems can provide unique insights into diverse cellular populations and their heterogeneous responses to internal and external perturbations. Thus far, most single-cell studies in plant systems have been limited to RNA-sequencing approaches, which only provide indirect readouts of cellular functions. Results Here, we present a single-cell proteomics workflow for plant cells that integrates tape-sandwich protoplasting, piezoelectric cell sorting, nanoPOTS sample preparation, and ion mobility-based MS data acquisition method for label-free single-cell proteomics analysis of Arabidopsis leaf mesophyll cells. From a single leaf protoplast, over 3,000 proteins were quantified with high precision. The workflow is demonstrated to identify stress associated changes in protein abundance by analyzing 117 protoplasts from well-watered and water-deficit stressed plants. Additionally, we describe a new approach for constructing covarying protein networks at the single-cell level and demonstrate how single-cell protein covariation analysis can reveal previously unrecognized protein functions while also capturing stress-induced changes in protein–protein dynamics. Conclusions The label-free scProteomic approach presented here represents a significant advance through the demonstration of a facile protoplast isolation method combined with deep and precise proteomic coverage of Arabidopsis leaf mesophyll cell types. We believe this study will serve as an informative reference to future plant scProteomic investigations.

Arabidopsis↗

A 1-year study on SARS-CoV-2 variant shifts in wastewater using dPCR: comparison with clinical and GISAID data

Wastewater testing can be used to monitor SARS-CoV-2 infections in communities. Data from PCR-based wastewater testing are usually available to public health authorities within 5–7 days after excreta and other body fluids enter the sewer. While PCR-based methods can accurately detect and quantify SARS-CoV-2, sequencing-based methods are usually required to distinguish between variants, delaying the results and adding cost to the process. We developed and assessed a novel, customizable digital PCR (dPCR)-based genotyping method for SARS-CoV-2 variant detection in wastewater, which is more cost-effective, faster, and more accessible than sequencing. This approach was applied to more than 1,400 wastewater samples

Wilton, Rose↗

Data for Phylogenetic diversity of light-dependent phosphorylation of Thr78 in Rubisco activase

Rubisco activase is an ATP-dependent chaperone that facilitates dissociation of inhibitory sugar phosphates from the catalytic sites of Rubisco during photosynthesis. In Arabidopsis, Rubisco activase is negatively regulated by dark-dependent phosphorylation of Thr78. The prevalence of Thr78 in Rubisco activase was investigated across sequences from 91 plant species, finding that 29 (∼32%) species shared a threonine in the same position. Analysis of seven C3 species with an antibody raised against a Thr78 phospho-peptide demonstrated that this position is phosphorylated in multiple genera. However, light-dependent dephosphorylation of Thr78 was observed only in Arabidopsis. Further, phosphorylation of Thr78 could not be detected in any of the four C4 grass species examined. The results suggest that despite conservation of Thr78 in Rubisco activase from a wide range of species, a regulatory role for phosphorylation at this site is more limited. This provides a case study for how variation in post-translational regulation can amplify functional divergence across the phylogeny of plants beyond what is explained by sequence variation in a metabolically important protein.

photosynthesis↗

Genome-scale model development and genomic sequencing of the oleaginous clade Lipomyces

The Lipomyces clade contains oleaginous yeast species with advantageous metabolic features for biochemical and biofuel production. Limited knowledge about the metabolic networks of the species and limited tools for genetic engineering have led to a relatively small amount of research on the microbes. Here, a genome-scale metabolic model (GSM) of Lipomyces starkeyi NRRL Y-11557 was built using orthologous protein mappings to model yeast species. Phenotypic growth assays were used to validate the GSM (66% accuracy) and indicated that NRRL Y-11557 utilized diverse carbohydrates but had more limited catabolism of organic acids. The final GSM contained 2,193 reactions, 1,909 metabolites, and 996 genes and was thus named iLst996. The model contained 96 of the annotated carbohydrate-active enzymes. iLst996 predicted a flux distribution in line with oleaginous yeast measurements and was utilized to predict theoretical lipid yields. Twenty-five other yeasts in the Lipomyces clade were then genome sequenced and annotated. Sixteen of the Lipomyces species had orthologs for more than 97% of the iLst996 genes, demonstrating the usefulness of iLst996 as a broad GSM for Lipomyces metabolism. Pathways that diverged from iLst996 mainly revolved around alternate carbon metabolism, with ortholog groups excluding NRRL Y-11557 annotated to be involved in transport, glycerolipid, and starch metabolism, among others. Overall, this study provides a useful modeling tool and data for analyzing and understanding Lipomyces species metabolism and will assist further engineering efforts in Lipomyces .

59 BASIC BIOLOGICAL SCIENCES↗

Discovery of FoTO1 and Taxol genes enables biosynthesis of baccatin III

Abstract Plants make complex and potent therapeutic molecules 1,2 , but sourcing these molecules from natural producers or through chemical synthesis is difficult, which limits their use in the clinic. A prominent example is the anti-cancer therapeutic paclitaxel (sold under the brand name Taxol), which is derived from yew trees (Taxusspecies) 3 . Identifying the full paclitaxel biosynthetic pathway would enable heterologous production of the drug, but this has yet to be achieved despite half a century of research 4 . WithinTaxus’ large, enzyme-rich genome 5 , we suspected that the paclitaxel pathway would be difficult to resolve using conventional RNA-sequencing and co-expression analyses. Here, to improve the resolution of transcriptional analysis for pathway identification, we developed a strategy we term multiplexed perturbation × single nuclei (mpXsn) to transcriptionally profile cell states spanning tissues, cell types, developmental stages and elicitation conditions. Our data show that paclitaxel biosynthetic genes segregate into distinct expression modules that suggest consecutive subpathways. These modules resolved seven new genes, allowing a de novo 17-gene biosynthesis and isolation of baccatin III, the industrial precursor to Taxol, inNicotiana benthamianaleaves, at levels comparable with the natural abundance inTaxusneedles. Notably, we found that a nuclear transport factor 2 (NTF2)-like protein, FoTO1, is crucial for promoting the formation of the desired product during the first oxidation, resolving a long-standing bottleneck in paclitaxel pathway reconstitution. Together with a new β-phenylalanine-CoA ligase, the eight genes discovered here enable the de novo biosynthesis of 3’-N-debenzoyl-2’-deoxypaclitaxel. More broadly, we establish a generalizable approach to efficiently scale the power of co-expression analysis to match the complexity of large, uncharacterized genomes, facilitating the discovery of high-value gene sets.

Science & Technology - Other Topics↗