Search NASA⌕ Search

SEARCH · Search NASA

Results for “Base Sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

A real-time energy and cost efficient vehicle route assignment neural recommender system

Here, this paper presents a neural network recommender system algorithm for assigning vehicles to routes based on energy and cost criteria. In this work, we applied this new approach to efficiently identify the most cost-effective medium and heavy duty truck (MDHDT) powertrain technology, from a total cost of ownership (TCO) perspective, for given trips. We employ a machine learning based approach to efficiently estimate the energy consumption of various candidate vehicles over given routes, defined as sequences of links (road segments), with little information known about internal dynamics, i.e. using high level macroscopic route information. A complete recommendation logic is then developed to allow for real-time optimum assignment for each route, subject to the operational constraints of the fleet. We show how this framework can be used to (1) efficiently provide a single trip recommendation with a top-k vehicles star ranking system, and (2) engage in more general assignment problems where n vehicles need to be deployed over m (m ≤ n) trips. This new assignment system has been deployed and integrated into the POLARIS. Transportation System Simulation Tool for use in research conducted by the Department of Energy's Systems and Modeling for Accelerated Research in Transportation (SMART) Mobility Consortium (SMART, 2024).

Energy consumption↗

torch-einshard v1.0

torch-einshard is a Python library for describing local and distributed PyTorch tensor computations with compact, einsum-like notation. Its expressions name logical axes, specify how they are sharded across a PyTorch DeviceMesh, and represent partial reductions. The library automatically performs contractions, permutations, reshaping, splitting, gathering, reduction, reduce-scatter, and repartitioning while preserving autograd. Additional features include sharding-aware FFTs, tensor rolls, halo exchange, sliding windows, 1D–3D convolutions, uneven-shard handling, parameter initialization and gradient management, and cost-based execution planning. It is designed for scientific machine learning and large-model workloads, including tensor-, sequence-, and spatial-parallel MLPs, attention, convolutions, and spectral operations. Compared with manually combining torch.einsum and distributed collectives, torch-einshard expresses both the mathematical operation and data placement in one readable formula. This reduces boilerplate and synchronization errors, keeps forward and backward communication consistent, and allows the library to select optimized collective strategies without changing model code.

Morozov, Dmitriy [Lawrence Berkeley National Labor↗

Assessing Options for Particulate Measurement Tools to Advance Wood Heaters (CRADA Final Report)

For approximately 40 years, particulate matter (PM) from residential wood heaters has been certified in North America using an EPA dilution tunnel with a filter-based gravimetric measurement. While reasonably repeatable, this approach reports only a single value for the entire burn sequence and does not provide information on PM at different points within the sequence, which would support improved wood heater designs. PM monitors that report with high time resolution offer a potential solution, but their performance for various types of heaters and across heater operating conditions has not been extensively evaluated. In this 18-month collaborative project, the Hearth, Patio & Barbecue Association (HPBA) and Lawrence Berkeley National Laboratory (LBNL) tested five continuous PM instruments spanning three measurement principles—light-scattering, inertial mass (TEOM), and triboelectric—against gravimetric and size-resolved references on three residential heater types (cordwood, catalytic cordwood, and pellet) across startup, high-burn, and low-burn operating phases.

42 ENGINEERING↗

A Generic and Multifunctional Electromagnetic Transient Model for Grid-Following Inverters

This article presents a generic and multi-functional electromagnetic transient (EMT) dynamic model of grid following (GFL) inverter-based resource (IBR) using the PSCAD software platform. The features of the developed model includes the flexibility in selecting various types and combinations of DC sources covering PV modules, battery modules, ideal DC source module, as well as flexibility in selecting either switched or averaged model of inverter. This model also covers exhaustive lists of controller logic covering open-loop/closed-loop PQ dispatch control, DC voltage and AC terminal voltage control along with the conventional current control designed in dq-domain, ate- domain and positive-negative sequence domain. Moreover, this model is equipped with the flexibility in selecting various types of current limiting schemes that includes saturation-based as well as latching-based current limiter, anti-windup protection. Moreover, the EMT model is agnostic to the MVA rating and is suitable for interfacing transmission systems by being complaint with the IEEE Std. 2800. The generality in the power circuits and the multi-functional options in operation and control of the developed EMT model makes it suitable for both academia and industry to study various power system aspects not limited to but such as fault behavior of GFL IBR and impacts on protection system, transient stability of a system interfaced with large number of GFL IBRs etc.

14 SOLAR ENERGY↗

A Generic and Multi-Functional Electromagnetic Transient Model for Grid-Following Inverter: Preprint

This article presents a generic and multi-functional electromagnetic transient (EMT) dynamic model of grid following (GFL) inverter-based resource (IBR) using the PSCAD TM software platform. The features of the developed model includes the flexibility in selecting various types and combinations of DC sources covering PV modules, battery modules, ideal DC source module, as well as flexibility in selecting either switched or averaged model of inverter. This model also covers exhaustive lists of controller logic covering open-loop/closed-loop PQ dispatch control, DC voltage and AC terminal voltage control along with the conventional current control designed in dq-domain, aB- domain and positive-negative sequence domain. Moreover, this model is equipped with the flexibility in selecting various types of current limiting schemes that includes saturation-based as well as latching-based current limiter, anti-windup protection. Moreover, the EMT model is agnostic to the MVA rating and is suitable for interfacing transmission systems by being complaint with the IEEE Std. 2800. The generality in the power circuits and the multi-functional options in operation and control of the developed EMT model makes it suitable for both academia and industry to study various power system aspects not limited to but such as fault behavior of GFL IBR and impacts on protection system, transient stability of a system interfaced with large number of GFL IBRs etc.

grid following inverter↗

Populus VariantDB v3.2 facilitates CRISPR and functional genomics research

The success of CRISPR genome editing studies depends critically on the precision of guide RNA (gRNA) design. Sequence polymorphisms in outcrossing tree species pose design hazards that can render CRISPR genome editing ineffective. Despite recent advances in tree genome sequencing with haplotype resolution, sequence polymorphism information remains largely inaccessible to various functional genomics research efforts. The Populus VariantDB v3.2 addresses these challenges by providing a user-friendly search engine to query sequence polymorphisms of heterozygous genomes. The database accepts short sequences, such as gRNAs and primers, as input for searching against multiple poplar genomes, including hybrids, with customizable parameters. We provide examples to showcase the utilities of VariantDB in improving the precision of gRNA or primer design. The platform-agnostic nature of the probe search design makes Populus VariantDB v3.2 a versatile tool for the rapidly evolving CRISPR field and other sequence-sensitive functional genomics applications. The database schema is expandable and can accommodate additional tree genomes to broaden its user base.

59 BASIC BIOLOGICAL SCIENCES↗

Optimizing genomic prediction for complex traits via investigating multiple factors in switchgrass

Genomic prediction has accelerated breeding processes and provided mechanistic insights into the genetic bases of complex traits. To further optimize genomic prediction, we assess the impact of genome assemblies, genotyping approaches, variant types, allelic complexities, polyploidy levels, and population structures on the prediction of 20 complex traits in switchgrass (Panicum virgatum L.), a perennial biofuel feedstock. Surprisingly, short read-based genome assembly performs comparably to or even better than long read-based assembly. Due to higher gene coverage, exome capture and multi-allelic variants outperform genotyping-by-sequencing and bi-allelic variants, respectively. Tetraploid models show higher prediction accuracy than octoploid models for most traits, likely due to the greater genetic distances among tetraploids. Depending on the trait in question, different types of variants need to be integrated for optimal predictions. Furthermore, our study provides insights into the factors influencing genomic prediction outcomes, guiding best practices for future studies and for improving agronomic traits in switchgrass and other species through selective breeding.

60 APPLIED LIFE SCIENCES↗

Massive compression for high data rate macromolecular crystallography (HDRMX): impact on diffraction data and subsequent structural analysis

New higher-count-rate, integrating, large-area X-ray detectors with framing rates as high as 17400 images per second are beginning to be available. These will soon be used for specialized macromolecular crystallography experiments but will require optimal lossy compression algorithms to enable systems to keep up with data throughput. Some information may be lost. Can we minimize this loss with acceptable impact on structural information? To explore this question, we have considered several approaches: summing short sequences of images, binning to create the effect of larger pixels, use of JPEG-2000 lossy wavelet-based compression, and use of Hcompress, which is a Haar-wavelet-based lossy compression borrowed from astronomy. We also explore the effect of the combination of summing, binning, and Hcompress or JPEG-2000. In each of these last two methods one can specify approximately how much one wants the result to be compressed from the starting file size. These provide particularly effective lossy compressions that retain essential information for structure solution from Bragg reflections.

47 OTHER INSTRUMENTATION↗

HYBRID COMPOSITES VIA CO-EXTRUSION ADDITIVE MANUFACTURING-COMPRESSION MOLDING FOR PERFORMANCE OPTIMIZATION

The growing demand for hybrid polymer composites with multifunctional properties has led to the development of various hybridization techniques, such as multi-material compounding and controlled laminate stacking sequence. In this study, a novel hybrid manufacturing approach was used by integrating a multiplexing extrusion system (MExS) based on additive manufacturing with subsequent compression molding process. This technique enabled the co-extrusion of different materials during the additive manufacturing process to fabricate composites with tailored performance. The developed hybrid composite featured a skin layer of glass fiberreinforced polycarbonate (PC/GF) encapsulating a carbon fiber-reinforced acrylonitrile butadiene styrene (CF/ABS) core. The structure was engineered to promote improved thermal and impact resistance at the surface, supported by a stiff core for enhanced overall mechanical integrity. Mechanical, thermal and morphological properties of the hybrid composites were investigated to understand trade-offs in performance compared to a single-material system. The results demonstrate that this approach enables the production of multifunctional composites suitable for applications such as automotive body panels and protective housings, where a balance of weight, mechanical strength, and thermal performance is essential.

Wasti, Sanjita [ORNL]↗

Cyote-attack Chain Estimator

Attack Chain Estimator (ACE) Application Overview The Attack Chain Estimator (ACE) Application is a sophisticated tool designed for the ingestion, classification, sequencing, and enrichment of cybersecurity threat reports. This application leverages advanced machine learning models and extensive historical data to provide comprehensive insights into cyber threats, specifically targeting Industrial Control Systems (ICS). Purpose The primary functions of the ACE Application include: Ingestion of Cybersecurity Threat Reporting: Capable of ingesting text-based threat reports in markdown or text file format. Supports ingestion of structured data from other sources in STIX/JSON format. Classification of Report’s Text-Based Events: Utilizes a DeBERTa classifier, specifically trained on cybersecurity data, to map the events to MITRE ATT&CK for ICS Tactics and Techniques. Classification is performed using multiple Jupyter notebooks and machine learning workflows hosted as FastAPI microservices: regex_data deberta_base_35_train_hft_classifier_mlflow.ipynb hft_regex_classifier_mlflow.ipynb param_train_hft_classifier_mlflow.ipynb regex_tactic_tech.ipynb Ordering of Tactics, Techniques, and Observable Events: Sequences the identified tactics, techniques, and events to form a coherent attack chain. Enrichment with Historical Attack Chain Details: Enhances the attack chain with details from historical attacks using a Markov model developed from CyOTE Precursor Analysis Report data. The Markov model is available as a FastAPI endpoint for seamless integration. Enrichment with Adversary Emulation Capabilities Data: Integrates adversary emulation capabilities data using MITRE Caldera for OT adversary abilities UUIDs. Export of Output Files: Provides options to export the enriched attack chain in JSON or CSV formats. Routing of Output to Other Applications: Facilitates routing of output to various platforms and applications, including: Threat Intelligence Platforms COREII Scout for Threat Intelligence Analysis COREII Modeling and Simulation for Adversary Emulation Technical Description The ACE Application is an advanced cybersecurity tool designed to provide detailed threat analysis and sequence generation. It is built on a robust architecture that integrates natural language processing, machine learning, and historical data modeling. Key Components: Data Ingestion Module: Handles the input of threat reports and data from various formats, ensuring flexibility in data sources. Classification Engine: Employs DeBERTa-based classifiers hosted as FastAPI microservices to analyze and classify threat report events in accordance with the MITRE ATT&CK framework for ICS. Sequence Generator: Orders the classified events into a logical attack chain, providing clear insight into the sequence of tactics and techniques used in the threat. Enrichment Engine: Integrates historical data and adversary emulation capabilities to enhance the attack chain with valuable context and additional details. The historical data enrichment is powered by a Markov model, which is available as a FastAPI endpoint. Export and Routing Module: Facilitates the export of the enriched attack chain in multiple formats and routes the output to designated applications for further analysis or emulation.

Paul, Tony [Idaho National Laboratory (INL), Idaho↗

Active learning path-dependent properties using a cloud-based materials acceleration platform

Solid state materials are central to many modern technologies in which a given material may be exposed to a variety of environments. The material properties often vary with the sequence of environments in an irreversible manner, resulting in a quintessential path-dependency in experimental observables. While sequential learning techniques have been effectively deployed for accelerating learning of state properties of materials, they often use a consistent environment path in all experiments. To elevate such techniques for making optimal decisions in experimental investigations of path-dependent properties, we introduce an iterated expected information gain acquisition function that optimizes over entire experimental trajectories. This approach is implemented within a cloud-based Materials Acceleration Platform architecture utilizing an event-driven stateful broker coupled with remote HELAO (Hierarchical Experimental Laboratory Automation and Orchestration) instances and an AI science manager. The platform's efficacy was demonstrated through a case study optimizing multi-step spectro-electrochemical experiments to identify optically stable potential windows in (Co–Ni–Sb)O z metal oxides. The system successfully integrated AI-driven experiment design, remote laboratory automation, and cloud-based data infrastructure, validating the platform's capability for managing complex, adaptive, path-dependent workflows in materials discovery.

Guevarra, Dan [California Institute of Technology ↗

Identification of shared viral sequences in peat moss metagenomes reveals elements of a possible Sphagnum core virome

Viruses are an understudied component of plant microbiomes. Identifying viruses that are shared between individual plants, or members of the “core virome”, could reveal stable viral populations with the potential to modulate the composition and function of the microbiome. Here, we examined the virome associated with Sphagnum mosses, a keystone species that has direct influence over the fate of peatland carbon stores. We analyzed bulk metagenomes and metatranscriptomes generated from Sphagnum field samples collected over a ten-month period to identify virus-like sequences shared among plants. Individual Sphagnum samples harbored distinct DNA and RNA viromes where only a small percentage (< 1%) of the total number of identified viral contigs were shared among all samples. Based on taxonomic classification, the shared viral contigs represent bacterial viruses, or phage (Caudoviricetes), as well as viruses of eukaryotes, namely nucleocytoplasmic large DNA viruses (Nucleocytoviricota) and RNA viruses (Riboviria). We linked the shared phage-like contigs to viral regions within sequenced genomes of bacterial taxa that are members of the Sphagnum core microbiome, suggesting that these contigs represent temperate phage or degraded prophage. The putative nucleocytoplasmic large DNA viruses and RNA viruses were phylogenetically diverse and showed sequence similarity to viruses associated with a broad range of hosts and environmental sources. The identification of shared viral contigs suggested that, despite the compositional heterogeneity between samples, Sphagnum mosses may harbor a core virome. Future work validating the presence of the core virome is warranted as it may aid in understanding how persistent viruses impact microbiome ecology and symbiont evolution within this climatically relevant keystone species.

Metagenomics↗

Luteolibacter sp. strain Populi

Luteolibacter sp. strain Populi is bacterium from the phylum Verrucomicrobiota, isolated from the rhizosphere of a black cottonwood tree, Populus trichocarpa, from the Cascade mountains in Washington. Its 6.6 Mb chromosome was completely sequenced using Oxford Nanopore long-reads and is predicted to encode 5301 proteins and 60 RNAs. The bacteria was isolated from the rhizosphere of a mature Populus trichocarpa from the Tieton riverwatershed of Washington state, USA (Lat: 46°42’9” N, Lon: 120°25 39’36” W). A rhizosphere sample (fine roots and adhering soil) was used to obtain a microbial fraction by centrifugation on Histodenz (12) and stained with 5µM Syto59 (Thermo Fisher Scientific Inc). A Cytopeia Influx cell sorter (BD, Franklin Lakes, NJ) was used to sort and array single cells (100 per plate) based on forward-side scatter and fluorescence intensity on asparagine-glucose nutrient agar (ATCC medium 184). The Luteolibacter sp. Populi genome sequence has been deposited in GenBank under the accession number CP161812. A draft genome annotated with Prokka and DRAM is available in this Narrative as Luteolibacter_sp_Prokka.240711.

59 BASIC BIOLOGICAL SCIENCES↗

Design and validation of a nanosecond short-wave infrared streak camera using gaseous detonation experiments

A custom short-wave infrared (SWIR) galvanometric streak camera was developed to provide nanosecond- scale, time-resolved optical diagnostics. The system was designed as a compact, mechanically driven alternative to traditional tube-based streak cameras, using a galvanometer mirror and multiple-reflection optical path to achieve high temporal resolution while maintaining electronic synchronization for trigger sequencing. Temporal resolution achieved 139.5 ns/px, allowing direct conversion of image-slope gradients to physical wave velocities. The instrument was validated using hydrogen-oxygen-argon detonations within a modular detonation tube, capturing transient emission structures and post-detonation decay behavior with microsecond precision. Although signal levels in filtered configurations were limited by lens transmission and detector sensitivity, the system successfully resolved intensity rise and decay characteristics across 10-60 % argon mixtures. The observed trends demonstrate the feasibility of compact galvanometer-based streak systems for high-speed SWIR imaging in reactive environments!

42 ENGINEERING↗

Amino Acid Sequence Controls Enhanced Electron Transport in Heme-Binding Peptide Monolayers

Metal-binding proteins have the exceptional ability to facilitate long-range electron transport in nature. Despite recent progress, the sequence-structure–function relationships governing electron transport in heme-binding peptides and protein assemblies are not yet fully understood. In this work, the electronic properties of a series of heme-binding peptides inspired by cytochrome bc1 are studied using a combination of molecular electronics experiments, molecular modeling, and simulation. Self-assembled monolayers (SAMs) are prepared using sequence-defined heme-binding peptides capable of forming helical secondary structures. Following monolayer formation, the structural properties and chemical composition of assembled peptides are determined using atomic force microscopy and X-ray photoelectron spectroscopy, and the electronic properties (current density–voltage response) are characterized using a soft contact liquid metal electrode method based on eutectic gallium–indium alloys (EGaIn). Our results show a substantial 1000-fold increase in current density across SAM junctions upon addition of heme compared to identical peptide sequences in the absence of heme, while maintaining a constant junction thickness. These findings show that amino acid composition and sequence directly control enhancements in electron transport in heme-binding peptides. Overall, this study demonstrates the potential of using sequence-defined synthetic peptides inspired by nature as functional bioelectronic materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

GenomeFace v1.0

GenomeFace is meta-genome binning software. Metagenomic binning, the process of grouping DNA sequences into taxonomic units, is critical for understanding the functions, interactions, and evolutionary dynamics of microbial communities. We propose a deep learning approach to binning using two neural networks, one based on composition and another on environmental abundance, dynamically weighting the contribution of each based on characteristics of the input data. Trained on over 43,000 prokaryotic genomes, our network for composition-based binning is inspired by metric learning techniques used for facial recognition. Using a task-specific, multi-GPU accelerated algorithm to cluster the embeddings produced by our network, our binner leverages marker genes observed to be universally present in nearly all taxa to grade and select optimal clusters of sequences from a hierarchy of candidates. We evaluate our approach on four simulated datasets with known ground truth. Our linear time integration of marker genes recovers more near complete genomes than state of the art but computationally infeasible solutions using them, while being over an order of magnitude faster. Finally, we demonstrate the scalability and acuity of our approach by testing it on three of the largest metagenome assemblies ever performed. Compared to other binners, we produced 47%-183% more near complete genomes. From these datasets, we find over the genomes of over 3000 new candidate species which have never been previously cataloged, representing a potential 4% expansion of the known bacterial tree of life.

Lettich, Richard [Lawrence Berkeley National Labor↗

EC-Bench: A Benchmark for Enzyme Commission Number Prediction

Enzymes are proteins that catalyze specific biochemical reactions in cells. Enzyme Commission (EC) numbers are used to annotate enzymes in a four-level hierarchy that classifies enzymes based on the specific chemical reactions they catalyze. Accurate EC number prediction is essential for understanding enzyme functions. Despite the availability of numerous methods for predicting EC numbers from protein sequences, there is no unified framework for evaluating and studying such methods systematically. This gap limits the ability of the community to identify the most effective approaches for enzyme annotation. We introduce EC-Bench, a benchmark for EC number prediction, consisting of 1) an initial representative set of existing methods (including homology-based, deep learning, contrastive learning, and language model methods), 2) existing and novel accuracy and efficiency performance metrics, and 3) selected datasets to allow for comprehensive comparative study. EC-Bench is open-source and provides a framework for researchers to not only compare among existing methods objectively under uniform conditions, but also to introduce and effectively evaluate performance of new methods in a comparative framework. To demonstrate the utility of EC-Bench, we perform extensive experimentation to compare the existing EC number prediction methods and establish their advantages and disadvantages in a variety of prediction tasks, namely “exact EC number prediction”, “EC number completion” and (partial or additional) “EC number recommendation”. We find wide variation in the performance of different methods, but also subtle but potentially useful differences in the performance of different methods across tasks and for different parts of the EC hierarchy.

59 BASIC BIOLOGICAL SCIENCES↗

Pixel-Resolved Long-Context Learning for Turbulence at Exascale: Resolving Small-scale Eddies Toward the Viscous Limit

Turbulence plays a crucial role in multiphysics applications, including aerodynamics, fusion, and combustion. Accurately capturing turbulence's multiscale characteristics is essential for reliable predictions of multiphysics interactions, but remains a grand challenge even for exascale supercomputers and advanced deep learning models. The extreme-resolution data required to represent turbulence, ranging from billions to trillions of grid points, pose prohibitive computational costs for models based on architectures like vision transformers. To address this challenge, we introduce a multiscale hierarchical Turbulence Transformer that reduces sequence length from billions to a few millions and a novel RingX sequence parallelism approach that enables scalable long-context learning. We perform scaling and science runs on the Frontier supercomputer. Our approach demonstrates excellent performance up to 1.1 EFLOPS on 32,768 AMD GPUs, with a scaling efficiency of 94\%. To our knowledge, this is the first AI model for turbulence that can capture small-scale eddies down to the dissipative range in three-dimensional turbulence at high Reynolds numbers.

Yin, Junqi [ORNL] (ORCID:0000000338435520)↗