Search NASASearch

SEARCH · Search NASA

Results for “Sequence Alignment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

67 records · Page 4

High-throughput Single-Cell Proteomics and Transcriptomics from the Same Cells with a Nanoliter-Scale Spin-Transfer Approach

Single-cell multiomic platforms provide a comprehensive snapshot of cellular states and cell types by offering critical insights into the spatiotemporal regulation of biomolecular networks at a systems level, thereby defining the basis of multicellularity. Here, we introduce nanoSPINS, an advanced platform that enables high-throughput profiling and integrative analysis of the transcriptome and proteome from the same single cells using RNA sequencing and isobaric labeling LC-MS-based proteomics, respectively. NanoSPINS can efficiently transfer mRNA-containing droplets across two microarrays via a centrifugation-based approach, while proteins are retained on the initial platform. Benchmarking of nanoSPINS on two cell lines demonstrates its ability to generate global proteomic and transcriptomic profiles that align well with previously established methodologies/platforms. The incorporation of isobaric TMTpro labeling into this single-cell multiomics platform significantly enhances the throughput of single-cell proteomic analyses. Through the high-throughput quantification of the proteome and transcriptome, nanoSPINS not only facilitates the identification of molecular features at both mRNA and protein level but also provides larger sample sizes for improved statistical power in clustering and differential abundance. Given the broad applicability of single-cell multiomics in biological research and clinical settings, we believe nanoSPINS represents a powerful platform for the characterization of heterogeneous cell populations.

multi 'omics

Building a framework to genetically characterize “feather spots” and understand demographic impacts of solar energy sites on migratory bird populations

The lack of data on the impact of utility-scale solar facilities on avian species and populations adds to the cost of siting and operation. As much as 32 percent of the avian biological material (feathers and carcasses) recovered from solar facilities remain unidentified, because they often take the form of “feather spots”. Feather spots are remains of impacted animals that can be separated into two broad categories: 1) those remains that may be visually identified to a species, or 2) those that cannot be visually identified to a species due to degradation from the environment and/or scavenger activity (listed as “unknown”). Even when feather spots can be identified to species, they cannot be visually assigned to particular breeding populations. In some cases, it is unknown whether multiple feather spots represent single or multiple individuals. This project’s objectives were to: 1. Use a developed, genetic-based technique to identify and determine the species, population of origin, and number of individuals found in feather spots recovered from solar facilities. 2. Implement collected data and resulting analyses to develop a publicly accessible web-based decision-making tool that can be used by the solar industry, regulators and other stakeholders to inform siting, mitigation, and conservation management efforts. 3. Establish a not-for-profit fee-for-service center at UCLA to ensure collection and identification of feather spots continue after the project period of performance. During the Project Period, we proposed to establish a pipeline for collecting, transporting, and storing of avian biological material collected at solar facilities and the collection and identification of feather spots to species and individual. We proposed the development of a genetic-based framework that would recover viable DNA from feather spots, amplify this DNA (i.e., make millions of copies of the original DNA), and use it to match the resulting sequences to a national database of known species of birds. The result would be the identification of feathers spots that were previously unidentified, and the incorporation of these samples into a larger database that included all samples recovered from solar facilities. The resulting report (below) details the result of this work and its alignment with proposed activities. We proposed the use of the data collected to assess the comparative risk to specific species or populations of species from solar facilities. For some species, we have already identified genomic markers of specific breeding populations and developed “genoscapes,” maps of unique genetic variation across the full breeding range of a species. We used these (previously and newly developed) genoscapes to probabilistically link a feather spot to the specific breeding populations from which it originated (assignment probabilities range from 75%-100% depending on species and population groups). For those species without genoscapes, we developed a vulnerability and susceptibility estimate that determines the relative local and regional risk to populations that are in geographic proximity to solar facilities, using citizen science data (Breeding Bird Survey (BBS) and eBird). These two feather spot processing pipelines (see Figure 1 below) provide quantitative estimates as to the numbers of individuals from a given population of origin that are affected by solar facilities, and ultimately can reduce costs to the consumer by reducing the industry costs associated with mitigation and siting strategies for future solar energy development.

14 SOLAR ENERGY

Decay of the proton-unbound superradiant state in 13 N

Here, the 12 C ( 3 He, 𝑑)⁢ 13 N* reaction is studied in an experiment with a high-resolution magnetic spectrograph, in coincidence with protons detected in silicon detectors near the target. This allows for the observation of angular correlation patterns between the proton transfer and proton decays from populated unbound resonances. A formalism describing the spin polarization of direct reactions is developed to analyze these correlations, and is verified on the known directionally asymmetric decay distributions arising from parity mixing in the 13 N ⁡($\frac{3}{2}$ − , $\frac{5}{2}$ + ) doublet. The same formalism is used to study the decays from the continuum-aligned, broad 3/2 + resonance at 7.9 MeV excitation energy, which arises from superradiant coupling. The observed asymmetric angular correlation patterns are approximately reproduced by adding an “artificial” 3/2 − resonance with strength equal to that of the reaction formalism. This parity-mixing approach serves as a first approximation to a more advanced reaction model of rapid reaction and decay sequences.

13N

Single nucleotide variants drive evolutionary phage-host arms race in anaerobic carbon dioxide-converting microbiome

Microbial bioconversions are shaped by environmental perturbations and the adaptation of resident microbiomes. Prokaryotes coexist with bacteriophages, yet their coevolutionary trajectories remain underexplored. Here, we investigate the effects of a cultivation vessel leak on an anaerobic consortium performing carbon dioxide reduction. Using time-series shotgun metagenomic sequencing, we reconstruct microbial and viral genomes to track community shifts. We further apply single-nucleotide variant profiling and CRISPR array analysis to monitor viral microdiversity and host defense mechanisms. After bioaugmentation restores bioconversion efficiency, the consortium undergoes pronounced restructuring, with new dominant taxa emerging from the rare biosphere. We identify patterns consistent with phage predation selectively removing certain species, while others exhibit resilience to infection. This shift aligns with a widespread viral outbreak and a transient increased frequency of single nucleotide variants in bacterial CRISPR–Cas defense genes. Expansion of CRISPR spacers further supports that CRISPR-mediated processes influence microbial resilience. Concurrently, phages infecting resilient hosts exhibited adaptive evolution, marked by high genetic heterogeneity. Selective pressure varies across their genomes, targeting infectivity genes and protospacer-adjacent motifs. These findings highlight a dynamic evolutionary arms race driven by the selection of beneficial genetic variants, providing a mechanistic framework for multi-omics investigations, and informing biotechnological applications, including phage-based microbiome manipulation.

Ghiotto, G

Efficient 15 N hyperpolarization of [ 15 N 3 ]metronidazole antibiotic via spin-relayed pulsed SABRE-SHEATH

Signal Amplification by Reversible Exchange in SHield Enables Alignment Transfer to Heteronuclei (SABRE-SHEATH) is an NMR hyperpolarization technique that relies of the simultaneous exchange of parahydrogen and a to-be-hyperpolarized molecule on the metal center of a polarization-transfer catalyst in a microtesla magnetic field. Until recently, this method has been understood to perform hyperpolarization by establishing level anti-crossings between the nuclear spins of the parahydrogen derived hydrides (acting as a source of hyperpolarization) and those of the substrate. Recently, the application of highly non-intuitive pulse sequences (comprising pulses of microtesla DC fields) was predicted to hyperpolarize nuclear spins more efficiently than the canonical (static-field) SABRE-SHEATH approach. Here we show that by employing a basic “on-off” pulse sequence of rectangular microtesla pulses, it is possible to improve the hyperpolarization efficiency for SABRE-SHEATH of [ 15 N 3 ]metronidazole, an FDA-approved antibiotic (in non-enriched and non-hyperpolarized form) and potential hypoxia sensing molecule. Specifically, we demonstrate that 15N polarization of 18.5 % can be obtained in 80 s of parahydrogen bubbling parahydrogen through a solution containing 20 mM [ 15 N 3 ]metronidazole. In practice, (1.32 ± 0.14)-fold improvements in P 15N was obtained with the pulsed method described here compared to static field technique variant. These results show that pulsed SABRE-SHEATH was successfully applied to 15 N-labeled biologically relevant molecule. Moreover, we also demonstrate that although the pulsed SABRE-SHEATH sequence was designed for polarization transfer from parahydrogen derived hydrides to the metronidazole’s 15 N catalyst-binding site, all three 15 N sites of [ 15 N 3 ]metronidazole attained the hyperpolarized state. This spin-relayed polarization transfer becomes possible due to the 15 N relay network established by their spin-spin J-couplings. The feasibility of the spin-relayed polarization transfer is demonstrated here for the first time for pulsed SABRE-SHEATH (as opposed to the static-field SABRE-SHEATH reported previously) and it paves the way to broad applicability of the technique.

Hyperpolarization

Missing microbial eukaryotes and misleading meta-omic conclusions

Meta-omics is commonly used for large-scale analyses of microbial eukaryotes, including species or taxonomic group distribution mapping, gene catalog construction, and inference on the functional roles and activities of microbial eukaryotes in situ. Here, we explore the potential pitfalls of common approaches to taxonomic annotation of protistan meta-omic datasets. We re-analyze three environmental datasets at three levels of taxonomic hierarchy in order to illustrate the crucial importance of database completeness and curation in enabling accurate environmental interpretation. We show that taxonomic membership of sequence clusters estimates community composition more accurately than returning exact sequence labels, and overlap between clusters can address database shortcomings. Clustering approaches can be applied to diverse environments while continuing to exploit the wealth of annotation data collated in databases, and selecting and evaluating these databases is a critical part of correctly annotating protistan taxonomy in environmental datasets. We argue that ongoing curation of genetic resources is crucial in accurately annotating protists in in situ meta-omic datasets. Moreover, we propose that precise taxonomic annotation of meta-omic data is a clustering problem rather than a feasible alignment problem.

59 BASIC BIOLOGICAL SCIENCES

A Preferences Corpus and Annotation Scheme for Human-Guided Alignment of Time-Series GPTs

The process of time-series forecasting such as predicting trajectories of silicon content in blast furnaces is a difficult task. Most time-series approaches today focus on scalar-type MSE loss optimization. This optimization approach, while widely common, could benefit from the use of human expert or process-level preferences. In this paper, we introduce a novel alignment and fine-tuning approach that involves learning from a corpus of preferred and dis-preferred time-series prediction trajectories. Our contributions include (1) a preference annotation pipeline for time-series forecasts, (2) the application of Score-based Preference Optimization (SPO) to train decoder-only transformers from preferences, and (3) results showing improvements in forecast quality. The approach is validated on both proprietary blast furnace data and the UCI Appliances Energy dataset. The proposed preference corpus and training strategy offer a new option for fine-tuning sequence models in industrial settings.

DPO

AGFormer: Adaptive Spatiotemporal graph informed transformer for multi-reservoir inflow forecasting

Accurate reservoir inflow forecasting is crucial for effective water resource management, yet most machine learning models focus on single-reservoir prediction and overlook spatial dependencies among hydrologically connected reservoirs. Here, we propose AGFormer (Adaptive Graph-Informed Transformer), an end-to-end framework that integrates adaptive graph learning with temporal sequence modeling for multi-reservoir inflow forecasting. A shared encoder and graph attention mechanism generate reservoir-specific embeddings, which are then processed by the Transformer-based encoder–decoder for multi-step inflow forecasting. We also introduce a pretraining paradigm to learn robust temporal embeddings from misaligned historical records. Evaluated on 30 reservoirs in the Upper Colorado River Basin, AGFormer achieves superior seven-day-ahead forecasts, with NSE > 0.75 for 20 reservoirs—outperforming Encoder–Decoder LSTM, GCN+LSTM, and Transformer baselines. Adaptive graph learning captures dynamic inter-reservoir dependencies, and feature attribution aligns with snowmelt-driven hydrology. Incorporating forecasted meteorological inputs further enhances accuracy, demonstrating AGFormer’s potential to support reservoir management under dynamic hydrological conditions.

Adaptive graph learning

Leverage modern artificial intelligence (AI) enabled systems for waste reduction

Manufacturing industries continue to face challenges in reducing waste, as upstream strategies such as source reduction and product redesign require a deeper understanding of processes compared to conventional recycling methods. Recent advancements in artificial intelligence (AI) and machine learning (ML) have opened new opportunities to integrate modern computational techniques with traditional waste minimization strategies. This paper explores AI-enabled approaches for product redesign, source reduction, and recycling that can significantly reduce waste generation while improving efficiency and sustainability. AI-driven material substitution and lightweighting in product design enable discovery of novel materials with optimized properties, reducing waste without compromising performance. Reinforcement learning models optimize process parameters, raw material specifications, and machine sequencing to minimize production losses, while Industrial Internet of Things (IIoT) systems paired with AI analytics enhance real-time waste tracking, predictive maintenance, and quality inspection. Furthermore, AI-based demand forecasting and production planning reduce overproduction and excess inventory, as demonstrated in industrial applications. In recycling, ML-powered pattern recognition and robotic sorting technologies achieve higher accuracy in waste segregation, directly improving recycling efficiency. Complementary solutions such as smart bins and AI-enabled waste pickup scheduling optimize collection logistics, reducing both costs and emissions. Although implementation requires upfront investment in infrastructure and training, the long-term benefits include higher material efficiency, reduced waste, improved product quality, and stronger sustainability outcomes across the supply chain. By leveraging AI-enabled systems, manufacturers can align waste minimization efforts with circular economy principles, creating scalable solutions for both industry and society.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Tuning Shinkarev’s Bicycle: Separating the Parallel Cycles of Photosystem II Using Empirical Wavelet Transform

The oxygen-evolving complex (OEC) of Photosystem II (PSII) catalyzes light-driven water oxidation, a process necessary to sustain Earth’s atmospheric oxygen. Oxygen yields measured during single-turnover flash sequences exhibit period-four oscillations, which form the basis of the Joliot–Kok (S-state) model. However, when the oscillations of other processes contribute to the measured oxygen yield, fitting methods can conflate these signals and distort estimates of inefficiencies and initial S-state populations. To address this, we applied the empirical wavelet transform (EWT) as a model-independent method to separate overlapping oscillators and capture damping dynamics that are not well represented in Fourier analysis. We tested this framework on polarographic flash-oxygen traces from both our Synechocystis sp. PCC 6803 thylakoid membrane preparations and archival datasets on Chlorella and isolated chloroplasts. EWT consistently resolves the expected period-four component alongside a distinct binary oscillation. Simulations suggest that fitting this isolated period-four signal recovers VZAD parameters more accurately than analysis of raw traces, yielding different estimates for S-state distributions and transition probabilities. Notably, this binary oscillation aligns closely with semiquinone dynamics predicted solely from period-four fit parameters. These findings indicate that EWT can effectively distinguish complex signals in oxygen evolution, offering a framework potentially applicable to other spectroscopic probes of the S-state cycle.

Ferrari, Nicholas [Louisiana State Univ., Baton Ro

Depth-dependent links between microbial taxa and nitrous oxide emissions in a long-term cotton cropping system employing soil health practices

Long-term management practices can shape soil microbial communities in ways that influence nitrogen (N) dynamics and nitrous oxide (N 2 O) emissions. We leverage a 41-year continuous cotton cropping experiment with contrasting tillage, cover cropping, and N fertilization regimes to investigate how these long-term strategies influence soil microbial communities and their associations with N 2 O fluxes during the cotton growing season. Using 16S rRNA gene metabarcoding, we assessed microbial composition in surface and subsurface soils and evaluated its relationship with temporal N 2 O emissions. Among the management practices, N fertilization – a known driver of N 2 O emissions – had the strongest effect on microbial community composition and was linked to a greater number of taxa correlated to N 2 O emissions, particularly in surface soils. Soil pH emerged as a key variable influencing microbial structure across depth and was negatively associated with both N 2 O emissions and microbial composition in the surface layers of fertilized soils. In total, 57 archaeal/bacterial taxa were correlated with N 2 O fluxes, but only seven were shared across depths, suggesting distinct microbial contributors in surface and subsurface soils. Several of these taxa have been previously reported to be associated with N and C cycling processes such as nitrate respiration or carbon turnover, indicating functional context to their correlation with N 2 O fluxes. Temporal shifts in the abundance of key taxa aligned with seasonal peaks in N 2 O emissions, notably in early and late August, and were most pronounced under conventional tillage, hairy vetch cover cropping, and N fertilization. While 16S-based associations cannot confirm functional gene presence or activity, these findings demonstrate that long-term fertilization and associated soil acidification are dominant drivers of microbial shifts linked to N 2 O emissions and highlight the importance of accounting for depth-specific and seasonal microbial dynamics when evaluating management impacts on greenhouse gas emissions.

16S rRNA gene sequencing

Supervisory Control and Data Acquisition for Electrochemical Separation Experimentation

The Python-based program is a laboratory automation tool designed to control and monitor electrochemical systems. The tool was developed for capacitive deionization (CDI) experiments, but it can be used for any system that requires controlled voltage or current segments and multi-parameter monitoring. The program integrates hardware components to run user-defined experimental parameters, providing operational control of a programmable power supply, peristaltic pump, and data acquisition devices. Currently, the program is structured with a workflow that includes an initialization (or pre-run) phase, a main loop, and a post-experiment stabilization (or post-run) phase. The initialization phase prepares and stabilizes the cell, ensuring that the electrodes and solution reach a baseline state before the experiment begins. The main loop consists of multiple voltage segments that repeat, controlling the experiment while recording key parameters such as time, voltage, current, pH, and conductivity. Finally, the post-experiment stabilization phase allows the system to stabilize after the experiment, returning the cell and solution to equilibrium conditions before ending the sequence. The program is designed with four variations, each tailored to different experimental needs. All variations include both the initialization and post-experiment stabilization stages, which run for a set amount of time, voltage, current, and flow rate before and after the main experiment block. The main loop runs for a set number of cycles, as defined by the user input, and each cycle is composed of 2 or 4 segments. The 4 program variations are described as follows: Program 1: The main program includes 2 segments. Each segment is defined to have a set duration, flow rate, voltage, and current. This program measures conductivity, flow rate, voltage, and current. Program 2: The main program expands Program 1 to include 4 segments. Each segment has a specified duration, flow rate, voltage, and current. Like Program 1, it measures conductivity, flow rate, voltage, and current. Program 3: The main program consists of 2 segments, each defined by time, flow rate, voltage, and current. In addition to conductivity, flow rate, voltage, and current, Program 3 collects pH and temperature data through a 4-channel data acquisition device. Program 4: This program independently controls two channels of a multi-channel power supply simultaneously. While conductivity can only be measured for one cell at a time, the dual-channel control makes it possible to operate two cells simultaneously under different voltage/current conditions. The main program includes 2 segments.For each program, all measurements are automatically logged and integrated into a single Excel output file. Data are displayed in numerical format and plotted, both in real time, to track system performance. A key feature of the program is its ability to synchronize all outputs so that every measurement shares a single timestamp, ensuring accurate alignment of voltage, current, pH, conductivity, and pH data.By combining hardware control, real-time monitoring, and unified data collection, this program significantly reduces manual workload and minimizes errors, making it a reliable platform for researchers, engineers, and laboratory technicians conducting CDI experiments, among other electrochemical tests.

Valentino, Lauren [Argonne National Laboratory (AN

Beyond sequence similarity: toward function-based screening of nucleic acid synthesis

Synthetic nucleic acids are a key input to modern biotechnology, yet they represent dual-use materials that require robust screening to mitigate biosecurity risks. The prevailing screening paradigm, which identifies sequences of concern (SoCs) through sequence similarity to controlled pathogens and toxins, may not fully capture risks posed by AI tools that can decouple biomolecular function from reliance on known sequences. Rapidly advancing biodesign capabilities enable the generation of genes and proteins that might evade sequence-based detection. We highlight the critical need for function-based screening approaches that can detect sequences capable of hazardous biological functions, regardless of similarity to known SoCs. We examine the feasibility of function-based screening with an initial focus on proteins, arguing that, while protein sequence space is vast, biologically functional proteins are significantly constrained by biophysical and biochemical requirements that can be learned and modeled. We propose a concrete implementation framework organized along a continuum of complexity, starting with toxins as the most tractable targets before expanding to more complex pathogenic functions. We then discuss open challenges and describe a research and development strategy to address them.

59 BASIC BIOLOGICAL SCIENCES