Search NASA⌕ Search

SEARCH · Search NASA

Results for “Sequencing data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Bottom-Up Simulation, Reconstruction, and Quantification of Macromolecule Sequences from Experimental Polymerizations

Motivated by the canonical sequence–structure–function paradigm, tools to characterize chemical patterning in natural biomacromolecules, from proteins to nucleic acids, have grown exponentially in recent years. However, analogous strategies for synthetic macromolecules remain in nascent stages, complicated by sequence polydispersity and analytical limitations. To address this, we have developed a comprehensive and open-source Python package, PRISM (polymer rate insights and sequence modeling), an end-to-end workflow that provides a path from experimental kinetics measurements to quantitative and qualitative metrics for describing chemical patterning in stochastic polymers. First, a numerical integration strategy was constructed to simulate and fit experimental data from reversible addition–fragmentation chain transfer (RAFT) polymerization kinetics, enabling the facile estimation of relevant reactivity ratios. These ratios were then used in a mechanism-specific stochastic kinetic simulation strategy to simulate sequence ensembles corresponding to model systems spanning experimental copolymers, classes of statistical polymers (e.g., alternating, block, and gradient), and multiblock copolymers. Lastly, inspired by sequence homology metrics from bioinformatics, we introduce visualization strategies and quantitative metrics to facilitate comparisons of different sequence ensembles. As the sequence–structure–function paradigm becomes increasingly central in de novo design of synthetic macromolecules, this toolkit provides a first step toward accurate and representative sequence description and featurization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A modular and extensible CHARMM-compatible model for all-atom simulation of polypeptoids

Peptoids (N-substituted glycines) are a class of sequence-defined synthetic peptidomimetic polymers with applications including drug delivery, catalysis, and biomimicry. Classical molecular simulations have been used to predict and understand the conformational dynamics of single chains and their self-assembly into morphologies including sheets, tubes, spheres, and fibrils. The CGenFF-NTOID model based on the CHARMM General Force Field has demonstrated success in accurate all-atom molecular modeling of peptoid structure and thermodynamics. Extension of this force field to new peptoid side chains has historically required reparameterization of side chain bonded interactions against ab initio data. This fitting protocol improves the accuracy of the force field but is also burdensome and precludes modular extensibility of the model to arbitrary peptoid sequences. In this work, we develop and demonstrate a Modular Side Chain CGenFF-NTOID (MoSiC-CGenFF-NTOID) as an extension of CGenFF-NTOID employing a modular decomposition of the peptoid backbone and side chain parameterizations, wherein arbitrary side chains within the large family of substituted methyl groups (i.e., –CH 3 , –CH 2 R, –CHRR', and –CRR'R") are directly ported from CGenFF. We validate this approach against ab initio calculations and experimental data to develop a MoSiC-CGenFF-NTOID model for all 20 natural amino acid side chains along with 13 commonly used synthetic side chains and present an extensible paradigm to efficiently determine whether a novel side chain can be directly incorporated into the model or whether refitting of the CGenFF parameters is warranted. We make the model freely available to the community along with a tool to perform automated initial structure generation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Innovating the next generation of commercial smart building software

Nearly 30% of commercial building energy use is wasted due to equipment faults and HVAC controls problems. The result is increased emissions, compromised comfort and productivity, and less reliable coordination of building power needs with a clean grid. The energy impact alone represents $17 billion in potential savings. Today’s smart building software provides a robust solution to address these operational deficiencies. Energy management and information systems (EMIS) are saving up to 9% on average, with two-year paybacks. They are being incorporated into energy management processes, commissioning services, and utility programs. As effective as they are, two barriers prevent even deeper benefits; limited personnel to fix problems once they are identified, and the expense and time to manually implement changes in control systems. In partnership with the research community, the EMIS industry is developing new capabilities to overcome these barriers. Moving beyond siloed products for either fault detection and diagnostics, or optimal control, these new capabilities empower users to not only automatically identify faults, but also to push corrective action, and control improvements to their buildings. In this paper, several areas for enhancements are documented: ‘one-time’ correction of faults such as setpoints, schedules, and economizer lockouts; short-term active testing for automated proportional integral derivative (PID) loop tuning and functional testing; and continuous supervisory control for demand flexibility and year-round efficiency. Results are presented from a pair of partner implementations out of a dozen providers integrating these enhancements into their products, including field tests from across the country, and insights into operator acceptance and integration into operations and maintenance practices.

Casillas, Armando↗

Using DNA Affinity Purification sequencing (DAP-seq) to identify in vitro binding sites of transcription factors potentially involved in aromatic degradation

The genome-wide binding sites of 44 transcription factors from the aromatic metabolizing Alphaproteobacterium Novosphingobium aromaticivorans were identified using DNA Affinity Purification sequencing (DAP-seq). We report 32 of these transcription factors have at least one area of enrichment. These data will be valuable for better understanding of aromatic metabolism.

aromatic metabolism↗

RCSB protein data Bank: Next‐generation advanced search for exploration of experimental structures and computed structure models

Abstract The Protein Data Bank (PDB), established in 1971, is the primary global, open‐access archive for experimentally determined 3D macromolecular structures (proteins, RNA, DNA). The research‐focused RCSB.org web‐portal provides access to these data alongside more than one million machine‐learning‐predicted structure models, greatly expanding the available structural landscape. Rapid growth of both experimental and computational structures has increased the need for powerful yet accessible search tools that serve a broad and diverse scientific community. Herein, we describe a redesigned RCSB Protein Data Bank RCSB.org Advanced Search capability that supports intuitive discovery of 3D structures through a unified interface. This interface integrates annotation‐, sequence‐, and 3D structure‐based searches, embeds an interactive 3D viewer, and incorporates curated biological knowledge, such as catalytic site definitions from Mechanism and Catalytic Site Atlas and ligand‐guided structural motifs, for constructing geometry‐driven queries. A new Chemical Search tool allows definition of chemical queries via an integrated drawing tool or standard identifiers, seamlessly combining them with annotation filters. By allowing query definition directly within spatial and chemical contexts, these search interfaces reduce the need for detailed knowledge of residue numbering, chain identifiers, or external cheminformatics software. This capability enables efficient exploration of structures, chemical diversity, and structure–function relationships across all life domains. The redesigned interfaces can be accessed directly at rcsb.org/search/advanced for Advanced Search and rcsb.org/search/chemical for Chemical Search.

Rose, Yana [Research Collaboratory for Structural ↗

scPlantAnnotate: an accurate and robust transformer-based model for plant cell type annotation

Accurate cell type annotation remains a major bottleneck in plant single-cell RNA sequencing (scRNA-seq), where existing tools are often adapted from animal studies and perform sub-optimally on plant data. The lack of plant-specific computational frameworks limits the construction of plant cell atlases and downstream biological discovery. We develop and evaluate scPlantAnnotate, a Transformer-based reference annotation framework tailored for plant scRNA-seq data, and benchmark it against state-of-the-art deep learning and conventional methods across multiple plant species. Species-specific scPlantAnnotate models were trained using curated datasets from Arabidopsis thaliana, Zea mays, Oryza sativa, and Glycine max. We compared scPlantAnnotate with leading baselines under both standard random-split evaluation and a more stringent leave-one-dataset-out setting, which tests robustness to completely unseen datasets and tissue types. scPlantAnnotate consistently outperforms existing approaches across all four species under random-split evaluation. In the leave-one-dataset-out setting for A. thaliana, where performance drops markedly for all methods due to strong batch effects and dataset heterogeneity, scPlantAnnotate nonetheless achieves the highest Accuracy, Macro-F1, Balanced Accuracy, and Macro-AUROC on average and ranks first on most held-out datasets. These results demonstrate improved robustness to dataset shifts, a critical yet underexplored challenge in plant scRNA-seq analysis. A freely accessible web server enables users to annotate their own datasets using pretrained models. scPlantAnnotate provides a plant-specific, Transformer-based framework for single-cell annotation that delivers state-of-the-art performance and enhanced robustness to unseen datasets. By addressing limitations of existing tools and enabling scalable reference-based annotation, scPlantAnnotate supports the development of comprehensive plant cell atlases and facilitates broader use of single-cell genomics in plant biology.

Bioinformatics↗

Discovery of potent small-molecule inhibitors of lipoprotein(a) formation

Lipoprotein(a) (Lp(a)), an independent, causal cardiovascular risk factor, is a lipoprotein particle that is formed by the interaction of a low-density lipoprotein (LDL) particle and apolipoprotein(a) (apo(a)). Apo(a) first binds to lysine residues of apolipoprotein B-100 (apoB-100) on LDL through the Kringle IV (K IV ) 7 and 8 domains, before a disulfide bond forms between apo(a) and apoB-100 to create Lp(a). Here we show that the first step of Lp(a) formation can be inhibited through small-molecule interactions with apo(a) K IV 7–8. We identify compounds that bind to apo(a) K IV 7–8, and, through chemical optimization and further application of multivalency, we create compounds with subnanomolar potency that inhibit the formation of Lp(a). Oral doses of prototype compounds and a potent, multivalent disruptor, LY3473329 (muvalaplin), reduced the levels of Lp(a) in transgenic mice and in cynomolgus monkeys. Although multivalent molecules bind to the Kringle domains of rat plasminogen and reduce plasmin activity, species-selective differences in plasminogen sequences suggest that inhibitor molecules will reduce the levels of Lp(a), but not those of plasminogen, in humans. These data support the clinical development of LY3473329—which is already in phase 2 studies—as a potent and specific orally administered agent for reducing the levels of Lp(a).

59 BASIC BIOLOGICAL SCIENCES↗

Identification of candidate host-specificity genes in Exserohilum turcicum using comparative genomics and transcriptomics

Abstract Exserohilum turcicum causes northern corn leaf blight and sorghum leaf blight. While the same species cause disease in both crops, the strains are host-specific. Here, we report the sequence and de novo annotated assemblies of one sorghum- and one maize-specific E. turcicum strain. The strains were sequenced using the PacBio Sequel II system. The total genome length for both assemblies was between 44 and 45 Mb with N50 of ∼2.5 Mb. Ninety-eight percent of the Benchmarking Universal Single-Copy Orthologs (BUSCO) for both assemblies had complete status. The estimated number of genes was 11,762 and 12,029 in the sorghum- and maize-specific isolates, respectively. Funannotate, EffectorP, SignalP, and transcriptome data were used to create functional annotation of each genome. The whole-genome comparison identified ten large-scale inversions and three translocations between the maize- and sorghum-specific strains, along with homologous genes and gene duplications. RNA was sequenced from the maize- and sorghum-specific isolate 10 days post-inoculation in maize and sorghum and from axenic cultures. Gene expression data from planta and axenic growth experiments were compared for each strain. Candidate host-specificity genes were identified by combining results from whole-genome comparison, synteny analysis, gene annotations, and transcriptome data. Overall, this study identified several candidate host-specificity genes that provide insights into E. turcicum interaction with its hosts.

Krone, Mara J. (ORCID:0000000159006624)↗

Soil microbiome resilience to short-term (30 days, 90 days) and long-term (1000 days) drought

This dataset contains data used for the paper "Drought duration does not impact soil microbiome resilience". The Related References will be updated with a full citation when available. Increasing global droughts exert large but poorly understood effects on the microbial communities and ecology of soil. Microbial communities generally show resilience and return to pre-drought conditions when short-term droughted soils are rewet; soils exposed to long-term drought, however, often show a lag upon rewetting, after which microbial communities may or may not return to their pre-stressed conditions. Though short-term droughts have been widely studied, long-term drought manipulation experiments remain rare, especially those that compare microbial response to short-term and long-term drought in tandem. We conducted a 1000-day drought simulation in controlled laboratory conditions with soil cores collected from a tidal freshwater ecosystem in Washington state, USA, and subsequently exposed them to rewetting for two weeks. We also included short-term (30-day and 90-day) drought and rewet treatments to directly compare microbial community and organic matter responses across drought durations. We found distinct microbial taxa belonging to Firmicutes and Actinobacteria enriched after the 1000-day drought, but not after the short-term droughts. While we hypothesized that the microbial community would recover from a short-term drought after rewetting to resemble pre-drought conditions, our results revealed community dissimilarities between rewet and pre-drought conditions across all drought durations. These findings suggest unique microbial life history strategies within certain microbial phyla that make them successful colonizers during an extended drought period, and the influence of environmental and physiological context on microbial responses to rewetting. The 16SrRNA gene amplicon dataset contains processed DNA sequences in the form of an ASV table with raw unrarefied read counts and representative sequences in .fasta format as described in the ESS-DIVE amplicon sequence reporting format (https://ess-dive.gitbook.io/amplicon-sequencing-reporting-format/instructions). The Fourier Transform Ion Cyclotron Resonance Mass Spectrometry (FTICR-MS) dataset consists of processed files containing presence absence data of molecular formulae and molecular characterization of FTICR resolved peaks. The Nuclear Magnetic Resonance (NMR) dataset contains files relevant to NMR spectra and peaks. A sample key file and a sample metadata file is included for the FTICR/NMR and 16S dataset respectively.

1000-day drought↗

High-throughput spin-bath characterization of spin defects in semiconductors

Detailed knowledge of the local environments of spin defects in semiconductors, such as nitrogenvacancy (NV) centers in diamond or divacancies in silicon carbide, is crucial for optimizing control and entanglement protocols in quantum sensing and information applications. However, at present a direct experimental characterization of individual defect environments is not scalable, as conventional spin-bath measurements are time consuming and difficult to automate. Achieving high-throughput characterization requires short experiments to probe the spin bath. However, with fewer and noisier measurements, the inverse problem of recovering spin-bath properties from measured data becomes ill posed, with multiple spin baths having a high likelihood of yielding the same data. In this work, we present a set of computational tools to resolve the ill-posed inverse problem of recovering the atomic positions and hyperfine couplings of random nuclei surrounding spin defects from sparse, noisy experimental coherence data, which can be obtained in hours. Here, we use a trans-dimensional Bayesian approach that incorporates ab initio data to yield full posterior distributions over nuclear spin environments, enabling robust recovery from limited data. We also provide practical tools and guidelines to determine the limits of detectability for hyperfine couplings under specific dynamical decoupling sequences and sampling conditions. In addition, we demonstrate how the tools developed here, in combination with ab initio simulations of spin baths, can guide the design of efficient experimental protocols for application-specific high-throughput screening. To showcase the utility of our approach, we apply it to design fast dynamical decoupling experiments to characterize the spin baths often individual NV centers in diamond. While the primary focus is on accelerating spin-bath characterization of spin defects, this Bayesian approach also lays the foundation for digital-twin studies of spin defects, where a virtual model of the spin-defect system evolves in real time with ongoing experimental measurements. Together, the set of tools we designed and applied paves the way for scalable deployment of spin defects in semiconductors for quantum sensing and information applications.

Bayesian methods↗

Microbial vitamin biosynthesis links gut microbiota dynamics to chemotherapy toxicity

ABSTRACT Dose-limiting toxicities pose a major barrier to cancer treatment. While preclinical studies show that the gut microbiota influences and is influenced by anticancer drugs, data from patients paired with careful side effect monitoring remains limited. Here, we investigate capecitabine (CAP)-microbiome interactions through longitudinal metagenomic sequencing of stool from 56 advanced colorectal cancer patients. CAP significantly altered the gut microbiome, enriching for menaquinol (vitamin K2) biosynthesis genes. Transposon library screens, targeted gene deletions, and media supplementation revealed that menaquinol biosynthesis protectsEscherichia colifrom drug toxicity. Stool menaquinol gene and metabolite levels were associated with decreased peripheral sensory neuropathy. Machine learning models trained in this cohort predicted toxicities in an independent cohort. Taken together, these results suggest treatment-associated increases in microbial vitamin biosynthesis serve a chemoprotective role for bacterial and host cells. Further, our findings provide a foundation for in-depth mechanistic dissection, human intervention studies, and extension to other cancer treatments. IMPORTANCE Side effects are common during the treatment of cancer. The trillions of microbes found within the human gut are sensitive to anticancer drugs, but the effects of treatment-induced shifts in gut microbes for side effects remain poorly understood. We profiled gut microbes in colorectal cancer patients treated with capecitabine and carefully monitored side effects. We observed a marked expansion in genes for producing vitamin K2 (menaquinone). Vitamin K2 rescued gut bacterial growth and was associated with decreased side effects in patients. We then used information about gut microbes to develop a predictive model of drug toxicity that was validated in an independent cohort. These results suggest that treatment-associated increases in bacterial vitamin production protect both bacteria and host cells from drug toxicity, providing new opportunities for intervention and motivating the need to better understand how dietary intake and bacterial production of micronutrients like vitamin K2 influence cancer treatment outcomes.

Microbiology↗

cjohnson-LANL/GRL_Kilauea

The python routines are outlined in detail to perform the methods and results in the manuscript under review in the journal Geophysical Research Letters titled “Seismic features predict ground motions during repeating caldera collapse sequence” with LA-UR-23-33345. All routines are written in open source python and were applied to publicly available data sets. The codes formats the data into the appropriate structure required to train a boosted tree regression model. Other codes produce figure results.

Johnson, Christopher W↗

Carbon-13 NMR spectra of lignin isolated from field grown transgenic poplar

Here we present a curated dataset of a series of 13C nuclear magnetic resonance (NMR) spectra of lignin isolated from transgenic monolignol 4-O-methyltransferase (MOMT4) engineered poplar. The transgenic poplar was collected from a 3-year field trial experiment. The poplar was Soxhlet-extracted with toluene/ethanol and the extractives-free poplar was then ball-milled in a Retsch PM100 planetary ball mill using a porcelain jar with ceramic balls at 600 rpm for 2 h. The ball-milled materials were then subjected to enzymatic hydrolysis for 48 h followed by centrifugation and washing with deionized water. The solid residue was extracted twice with 96:4 (v/v) 1,4-dioxane/water mixture at room temperature overnight. The extracts were combined, rotary evaporated, and freeze-dried to recover the lignin. The dry lignin samples were dissolved in deuterated dimethyl sulfoxide for NMR characterization. 13C experiments were performed in a Bruker Avance III HD 500 MHz NMR spectrometer operating at a frequency of 125.12 MHz for the 13C nucleus using a standard Bruker pulse sequence (zgpg) on a Prodigy platform cryoprobe. The NMR spectra were acquired under the following conditions: spectra width 229 ppm, 64k data points, 1s pulse delay, and 6k scans. All the data was processed using the Bruker’s TopSpin 3.6 software. Additional meta data is embedded in the raw spectra files.

13C NMR, lignin, poplar, field trial, MOMT4, CBI↗

SITCOMTN-163: Source Selection for Abell 360 in LSSTComCam Data Preview 1

We cover the source selection done for the Abell 360 LSSTComCam cluster study focusing on color and photo-z cuts. Identification of cluster members via the red sequence and photo-z offers a science focused validation of photometry with LSSTComCam. We are able to cross-match with DESI spectroscopic redshifts to offer an independent validation suite on photo-z estimates. We make various color and photo-z cuts to generate N(z)s and shear profiles. We find that both cuts are quite robust to various choices made and can produce a shear profile for Abell 360. Both methods should be applicable to generating a consistent mass estimate for the cluster.

79 ASTRONOMY AND ASTROPHYSICS↗

Prediction of Specificity of α-Conotoxins to Subtypes of Human Nicotinic Acetylcholine Receptors with Semi-supervised Machine Learning

Conotoxins are a family of highly toxic neurotoxins composed of cysteine-rich peptides produced by marine cone snails. The most lethal cone snail species to humans is Conus geographus, with fatality rates of up to ∼65% from a single sting, which is caused mostly by the activity of α-conotoxins against human nicotinic acetylcholine receptors (nAChRs). While sequence-based machine learning (ML) classifiers have been trained to identify targets of conotoxins binding voltage-gated ion channels, no ML model has been built to predict the subtype-specific nAChR targets of α-conotoxins. Here, we trained an ML model in a semi-supervised manner to predict the specificity of α-conotoxin binding toward different human nAChR subtypes to overcome the challenge of limited data in subtype-specific nAChR targets of α-conotoxins and the issue that one α-conotoxin can bind multiple nAChR subtypes with high selectivity. We considered additional features of sequences of α-conotoxins in training our ML model, including the secondary structure propensities and electrostatic properties, which resulted in better prediction capability for the ML model. Notably, we identify that most α-conotoxins bind to α3β2, α1γδ, and α7 subtypes of human nAChRs. Our findings from this study provide a framework for predicting targets of various kinds of toxins.

59 BASIC BIOLOGICAL SCIENCES↗

DIVA/DeviceEditor v6.1.2

DIVA is an end-to-end DNA design and construction management platform that streamlines how researchers design, build, and receive sequence-verified DNA constructs. Through a web-based BioCAD interface (DeviceEditor), researchers independently design DNA constructs and submit them to a centralized queue with a single action. Designs progress transparently through standardized states which allow researchers to track status and access finished constructs via a central DNA repository. Submitted designs are reviewed by dedicated staff for feasibility and optimization, reducing costly failures and improving downstream execution. Automated DNA assembly software optimizes construction strategies by reusing existing parts where possible and sourcing synthetic DNA only when needed. Standardized, sequence-agnostic assembly methods enable many independent constructs to be built in parallel using lab automation, dramatically increasing throughput. High-throughput next-generation sequencing is used to verify construct accuracy, with flexible platforms selected based on task requirements. Throughout the process, detailed success and failure data are captured and analyzed, enabling continuous improvement of assembly protocols. Compared to traditional, manual DNA construction workflows, DIVA offers higher scalability, transparency, reproducibility, and data-driven optimization.

Plahar, Hector [Lawrence Berkeley National Laborat↗

Old Woman Creek Wetland Sediment and Electrochemical Sensor Microbial Community, 2023

We are developing a technique to monitor microbiological activities referred to as zero resistance ammetry, which entails the deployment of graphite electrodes in sediments. Measurement of current between electrodes of contrasting redox regimes and/or predominant terminal electron accepting processes can be used as an indicator of the extents of microbiological activity. We deployed an electrode array at depths of 2 mm, 4 mm, 76 mm, 78 mm, 152 mm, 154 mm, 227 mm, and 229 mm below the wetland sediment water interface in the Old Woman Creek National Estuarine Research Center, Huron, OH, USA (Lat. = 41.380833, Long. = -82.508889). A core was collected from adjacent sediment and subsamples were collected from depth intervals of 0 – 25 mm, 25 – 127 mm, 127 – 128 mm, and below 178 mm. To determine if the microbial communities attached to the electrodes were reflective of the adjacent sediment-associated microbial community, we conducted a 16S rRNA gene-based (V4 region) survey of these respective materials. This data package contains the results of these surveys, including metadata on the depths from which samples were collected (samples.csv), DNA extraction and sequencing information (OWC_DEPTH_AMPLICON_SEQUENCING_METADATA), sequence processing information (OWC_DEPTH_BIOINFORMATIC_METADATA.csv), an operational taxonomic unit (OTU) table (OWC_DEPTH_97OTUS_TABLE.csv), and nucleotide sequences of OTUs (OWC_DEPTH_97OTUS_SEQS.fasta). All files can be opened using a text-editing application. The fasta file is compatible with bioinformatics applications.

54 ENVIRONMENTAL SCIENCES↗

TRACE Input Modernization

This work presents a Tom’s Obvious Minimal Language (TOML)-based representation of input for the US Nuclear Regulatory Commission’s TRAC/RELAP Advanced Computational Engine (TRACE) thermal hydraulics code. Implemented using the Workbench Analysis Sequence Processor (WASP), the approach maps traditional TRACE input structures to a hierarchical format composed of named parameters, typed values, and native data collections. The resulting representation preserves TRACE’s existing modeling capabilities while providing a modern, structured interface for model development and management. WASP further extends TOML through a file import directive that supports modular model composition and reusable input organization. In addition, WASP provides extended array data entry convenience with various data repeat and interpolation capabilities. Examples of the new TOML syntax are provided for major TRACE input categories, including hydraulic components, heat structures, control systems, and trip logic. The TOML representation establishes a foundation for improved validation, tooling, automation, and model maintainability while remaining compatible with existing TRACE workflows. To facilitate migration to the TOML-based input format, the TRACE executable now supports conversion of native TRACE input into an intermediate JSON representation. A Python utility subsequently transforms the JSON data into an equivalent TOML model. Lastly, the TRACE executable now supports execution using TOML-formatted input.

Lefebvre, Robert A. [Oak Ridge National Laboratory↗