Search NASASearch

SEARCH · Search NASA

Results for “sequence development”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A dynamic solvent chamber propagation estimation framework using RNN for warm solvent injection in heterogeneous reservoirs

Warm solvent injection (WSI), injecting low-temperature solvent into formations to reduce the viscosity of heavy oil, is a clean technology for heavy oil production through reducing greenhouse gas emissions and water usage. The success of WSI operation depends on the uniform development and propagation of solvent chambers in reservoirs. However, reservoir heterogeneity stemming from shale barriers plays a detrimental role in the conformance of solvent chamber development and oil production rate. In this work, we developed a novel recurrent neural network (RNN)-based framework with the capability of efficiently tracking and estimating the solvent chamber positions in heterogeneous reservoirs based on only production time-series data. The developed estimation model utilizes the “sequence-to-sequence" mapping methodology to correlate observed production time-series sequence and solvent chamber edge sequence via a long short-term memory (LSTM) algorithm. The trained RNN models exhibit high accuracy, evidenced by the predicted dynamic solvent chamber locations match the corresponding true locations from numerical simulation, with a high coefficient of determination (R 2 ) and a low mean squared error. Specifically, the achieved R 2 values exceed 0.98 on both the training and testing data. The developed RNN-based workflow was tested via several cases from both regularly- and irregularly-shaped shale barriers, and the results were promising. The predicted solvent chambers showed strong agreement with those obtained from numerical simulations. The major benefits of this workflow include reducing computational time and saving overall monitoring and tracking costs for conventional techniques. In conclusion, the present work would provide a good demonstration of the capability of practical integration of machine learning methods in solving engineering problems.

58 GEOSCIENCES

Barcoded overexpression screens in gut Bacteroidales identify genes with roles in carbon utilization and stress resistance

Abstract A mechanistic understanding of host-microbe interactions in the gut microbiome is hindered by poorly annotated bacterial genomes. While functional genomics can generate large gene-to-phenotype datasets to accelerate functional discovery, their applications to study gut anaerobes have been limited. For instance, most gain-of-function screens of gut-derived genes have been performed in Escherichia coli and assayed in a small number of conditions. To address these challenges, we develop Barcoded Overexpression BActerial shotgun library sequencing (Boba-seq). We demonstrate the power of this approach by assaying genes from diverse gut Bacteroidales overexpressed in Bacteroides thetaiotaomicron . From hundreds of experiments, we identify new functions and phenotypes for 29 genes important for carbohydrate metabolism or tolerance to antibiotics or bile salts. Highlights include the discovery of a d -glucosamine kinase, a raffinose transporter, and several routes that increase tolerance to ceftriaxone and bile salts through lipid biosynthesis. This approach can be readily applied to develop screens in other strains and additional phenotypic assays.

59 BASIC BIOLOGICAL SCIENCES

Bottom-Up Simulation, Reconstruction, and Quantification of Macromolecule Sequences from Experimental Polymerizations

Motivated by the canonical sequence–structure–function paradigm, tools to characterize chemical patterning in natural biomacromolecules, from proteins to nucleic acids, have grown exponentially in recent years. However, analogous strategies for synthetic macromolecules remain in nascent stages, complicated by sequence polydispersity and analytical limitations. To address this, we have developed a comprehensive and open-source Python package, PRISM (polymer rate insights and sequence modeling), an end-to-end workflow that provides a path from experimental kinetics measurements to quantitative and qualitative metrics for describing chemical patterning in stochastic polymers. First, a numerical integration strategy was constructed to simulate and fit experimental data from reversible addition–fragmentation chain transfer (RAFT) polymerization kinetics, enabling the facile estimation of relevant reactivity ratios. These ratios were then used in a mechanism-specific stochastic kinetic simulation strategy to simulate sequence ensembles corresponding to model systems spanning experimental copolymers, classes of statistical polymers (e.g., alternating, block, and gradient), and multiblock copolymers. Lastly, inspired by sequence homology metrics from bioinformatics, we introduce visualization strategies and quantitative metrics to facilitate comparisons of different sequence ensembles. As the sequence–structure–function paradigm becomes increasingly central in de novo design of synthetic macromolecules, this toolkit provides a first step toward accurate and representative sequence description and featurization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Multiomics and deep learning dissect regulatory syntax in human development

Transcription factors establish cell identity during development by binding regulatory DNA in a sequence-specific manner, often promoting local chromatin accessibility and regulating gene expression1. Mapping accessible chromatin offers critical insights into transcriptional control, but available datasets for human development are restricted to bulk tissue, single organs or single modalities2. Here we present the Human Development Multiomic Atlas, a single-cell atlas of chromatin accessibility and gene expression from 817,740 fetal cells across 12 organs, spanning 203 cell types and more than 1 million candidate cis-regulatory elements, many of which exhibit organ-specific in vivo enhancer activity. Deep learning models trained to predict accessibility from local DNA sequence unravel a comprehensive lexicon of motifs that influence accessibility, including composite motifs exhibiting distinct syntactic constraints that are predicted to mediate transcription factor cooperativity. We identify ‘hard’ syntactic rules requiring precise motif spacing and orientation, ‘soft’ rules allowing flexible motif arrangements, and ubiquitous motifs inhibiting accessibility. Model-based interpretation of genetic variants reveals that disruption of motifs with positive and negative effects is associated with concordant effects on gene expression. Our work delineates how motif syntax governs cell-type-specific chromatin accessibility and provides a foundational resource for decoding cis-regulatory logic and interpreting genetic variation during human development.

59 BASIC BIOLOGICAL SCIENCES

Large-Scale Inverter Integration in Bulk Power Grids Using Heterogeneous Grid-Forming Control Strategies

The conventional generation within the grid is gradually being replaced by inverter based resources. However, the integration of inverters into realistic transmission grid models has not been thoroughly explored. To address this gap, this work develops an automation framework aimed at facilitating the integration of a large number of inverter-based resources into a large-scale grid. Two types of grid-forming controlled inverters—namely droop-controlled inverters and virtual synchronous machine-controlled inverters—are developed for use in positive-sequence simulation packages. Various penetration levels ranging from 23% to 100% of grid-forming inverters are tested in a grid setup in the US Western Interconnection. The dynamic performance of inverters under different control schemes is compared in response to various disturbances.

Lyu, Xue [BATTELLE (PACIFIC NW LAB)]

Geology of the One Earth Energy Site

The One Earth Energy site is one of two sites in the Illinois Storage Corridor (ISC) project. The objectives of the ISC project is to accelerate commercial deployment of carbon capture utilization and storage at two individual sites and receive approvals for Underground Injection Control (UIC) Class VI permits for construction at each site. At the One Earth Energy site, an extensive data collection program was undertaken, which included the drilling of a test well (One Earth Energy #1 [OEE #1]), four 2D seismic lines, and a small 3D seismic survey. The OEE #1 well was drilled in 2022 and acquired extensive core, log, and testing data to characterize the subsurface geology of the site. Coring was focused on the storage interval, the Mt. Simon Sandstone, and the confining interval, the Eau Claire Formation. The core and log data were used to evaluate the sedimentology and sequence stratigraphy, as well as to develop the conceptual geologic model. This report includes the geological summaries of the Mt. Simon Sandstone and the Eau Claire Formation. The extensive analysis of the log data is included in the petrophysical section, showing ranges of porosity, estimated pore size, and the mineral content of selected zones in the well. The separate petrographic technical report entitled “Petrographic and Advanced Geologic Characterization Report on One Earth Energy #1 (API# 1211325373)”, report number DOE-UIUC-0031892-04, details thin section point-counting analysis that includes mineralogical and pore space analysis, including grain size analysis, annotated thin section photomicrographs, scanning electron microscopy (SEM) with energy dispersive X-ray spectroscopy (EDS), and statistics of grain size analysis on Mt. Simon thin sections from OEE #1. The final OEE #1 well data to be included in this geology report is the routine core analysis of both whole core plugs and rotary sidewall core plugs. In addition to the OEE #1 well, four 2D seismic lines and a small 3D survey were acquired as part of the overall subsurface geological characterization. This geology report references the seismic interpretation report, entitled “One Earth Energy Site Seismic Interpretation Task 5.0”, report number DOE-UIUC-0031892-07. This report details the stratigraphic and structural interpretation of the 2D and 3D seismic data acquired at the One Earth Energy site. The 2D seismic data was acquired in 2019 and 2021, and the 3D survey was acquired in 2022. The objectives of the seismic programs were to contribute to the subsurface characterization of the Mt. Simon-Eau Claire Storage Complex by evaluating the continuity of potential storage reservoirs and containment intervals across the project area, and to determine if any geologic features are present that would increase containment risk to the proposed carbon storage project.

09 BIOMASS FUELS

Advances in genetic tools for metabolic engineering of non-conventional yeasts

Non-conventional yeasts are emerging as powerful alternatives to Saccharomyces cerevisiae for metabolic engineering, owing to their innate stress tolerance, broad substrate utilization, and distinctive metabolic capabilities. These attributes position them as promising chassis for producing biofuels, pharmaceuticals, and specialty chemicals. This review synthesizes recent advances in genetic toolkits for four such species—Pichia kudriavzevii (Issatchenkia orientalis), Starmerella bombicola, Debaryomyces hansenii, and Pachysolen tannophilus—highlighting progress across plasmid architectures (episomal and integrative), identification of autonomously replicating sequences and centromeric elements, and the development of safe-harbor genomic loci. We summarize promoter and terminator libraries enabling tunable expression, the expansion of auxotrophic and antifungal selection markers with recycling strategies, and the rapid adaptation of CRISPR-based systems (Cas9 and Cas12a) with optimized guide RNA expression, multiplex editing, and approaches that enhance homologous recombination (e.g., KU70/80 disruption). We also review landing-pad platforms for modular, repeated integrations and transposon-based tools (e.g., piggyBac) that facilitate multigene pathway assembly. Collectively, these innovations are accelerating design-build-test-learn cycles and enabling precise, scalable engineering of non-conventional yeasts. Remaining challenges—including limited species-specific episomal systems, variable transformation efficiencies, genome-stability concerns, and alternative codon usage—define clear priorities for future toolkit development. Together, these advances and open needs chart a path toward robust, sustainable biomanufacturing using diverse non-conventional yeast chassis.

59 BASIC BIOLOGICAL SCIENCES

Digitizing Today’s Buildings in the Real World: Lessons from Field Demonstrations

Digital twins, created by generating a virtual replica of a building, enable safe evaluation of operational scenarios and applications like fault detection and diagnosis and advanced controls. However, a prerequisite is the creation of a machine-readable digital representation of a building, currently hindered by fragmented information scattered across mechanical drawings, point lists, and natural language sequences. As a result, digital twin development remains labor-intensive, error-prone, and difficult to validate. To address these challenges, two efforts from ASHRAE aim to support the digitalization of buildings. ASHRAE s223 establishes a semantic model of buildings, representing system components, configuration, and data sources. ASHRAE s231 defines a vendor-neutral programming language for expressing their control logic. As the industry evaluates implementing them in their products, understanding the challenges that vendors and implementers may face is crucial. In this paper, we present findings and lessons learned from field demonstrations in five buildings that implemented control applications using ASHRAE s223 and s231. The demonstrations highlight how semantic modeling and formalized control descriptions can significantly reduce software development time, manual point mapping, and hard-coding. Beyond time efficiency, they enable reliable automation by minimizing human interpretation and providing a means for consistency across projects. We describe the processes and best practices for model creation and model usage, from translating heterogeneous building documentation into semantic representations to implementing control logic in real-world systems. Finally, we discuss the challenges that persist, including integration with legacy software environments, gaps in interoperability, and the level of expertise still required to effectively leverage semantic models.

Prakash, Anand Krishnan

A PSCAD Library Component Featuring a Reduced-Order IBR Model for EMT-Based Fault Studies

This paper presents a fully implemented inverter reduce-order-model (ROM) in an EMT simulation (PSCAD) library component for direct user utilization in protection studies. The developed inverter ROM has the following features: Equivalent to a full IBR inverter model with positive- and negative-sequence current formulation and representation A python script is developed to fully automate this process, including training data generation, ROM parameter training, updating parameters, and model verification and validation. With this PSCAD ROM library component, protection engineers can utilize a trustworthy, accurate ROM for protection studies in an easy-to-use and streamlined manner.

24 POWER TRANSMISSION AND DISTRIBUTION

A Chemoselective and Stereodivergent Platform of Heme‐Nitrene Transferases to Access Chiral Aryl‐β‐Amino Esters and An Investigation of the Sequence‐Activity Landscape

Engineered biocatalysts can utilize nitrene precursors to access enantioenriched amination products, yet they have not been applied to produce valuable, enantiomerically enriched noncanonical β-amino esters. Current approaches to synthesizing β-amino acids rely on pre-oxidized precursors and multistep synthetic approaches involving various protecting groups. We engineered a platform of heme enzymes for stereoselective C–H bond amination of readily available carboxylic ester derivatives to install primary amines. A directed evolution campaign coupled with sequencing of over 1000 variants enabled us to develop engineered variants that use either O-pivaloylhydroxylamine triflic acid (PONT) or hydroxylamine hydrochloride (H 2 NOH∙HCl) as aminating reagents. An analysis of the resulting sequence–activity dataset revealed additional improvements that could be made to the final variant, highlighting the utility of sequencing data to guide future steps in directed evolution campaigns. Furthermore, the evolved nitrene transferases expand the scope of accessible chiral β-amino acid building blocks for peptidomimetic applications and provide new starting points for the design and synthesis of enantioenriched β-amino acid motifs.

amino ester building blocks

Evaluation of a Reduced-Order Model for IBR Fault Response Representation via OEM Blackbox Models

This paper presents a fully implemented inverter reduced-order-model (ROM) in an EMT simulation (PSCAD) library component for direct user utilization in protection studies. The developed inverter ROM has the following features: Equivalent to a full inverter-based resource (IBR) inverter model with positive- and negative-sequence current formulation and representation. A Python script is developed to fully automate this process, including training data generation, ROM parameter training, updating parameters, and model verification and validation. The ROM is validated using both IEEE 2800-compliant and non-compliant OEM modes in a real-world system, building confidence of its usability by protection engineers.

24 POWER TRANSMISSION AND DISTRIBUTION

Next-Generation Sequencing Data from a CUT&RUN Study of R. toruloides IFO0880 Cse4 and Orc1 Binding Sites

Rhodotorula toruloides has been increasingly explored as a host for bioproduction of lipids, fatty acid derivatives and terpenoids. Various genetic tools have been developed, but neither a centromere nor an autonomously replicating sequence (ARS), both necessary elements for stable episomal plasmid maintenance, has yet been reported. In this study, cleavage under targets and release using nuclease (CUT&RUN), a method used for genome-wide mapping of DNA–protein interactions, was used to identify R. toruloides IFO0880 genomic regions associated with the centromeric histone H3 protein Cse4, a marker of centromeric DNA. Fifteen putative centromeres ranging from 8 to 19 kb in length were identified and analyzed, and four were tested for, but did not show, ARS activity. These centromeric sequences contained below average GC content, corresponded to transcriptional cold spots, were primarily nonrepetitive and shared some vestigial transposon-related sequences but otherwise did not show significant sequence conservation. Future efforts to identify an ARS in this yeast can utilize these centromeric DNA sequences to improve the stability of episomal plasmids derived from putative ARS elements.

Genome Engineering

CRAGE-RB-PI-seq reveals transcriptional dynamics of plant-associated bacteria during root colonization

Plant roots release a wide array of metabolites into the rhizosphere, shaping microbial communities and their functions. While metagenomics has expanded our understanding of these communities, little is known about the physiology of their members in host environments. Transcriptome analysis via RNA sequencing is a common approach to learning more, but its use has been challenging because of low bacterial biomass and interference from plant RNA. To overcome this, we developed a randomly-barcoded promoter-library insertion sequencing (RB-PI-seq) combined with chassis-independent recombinase-assisted genome engineering (CRAGE). Using Pseudomonas simiae WCS417 as a model rhizobacterium, this method enabled targeted amplification of barcoded transcripts, bypassing plant RNA interference and allowing measurement of thousands of promoter activities during Arabidopsis root colonization. Our analysis revealed temporally resolved transcriptional regulation, including those associated with cell growth, chemotaxis, plant immune suppression, biofilm formation, and stress responses, reflecting the coordinated physiological adaptation to the root environment. Additionally, we discovered that transcriptional activation of xanthine dehydrogenase and a lysozyme inhibitor is crucial for evading plant immune systems. This framework is scalable to other bacterial species and provides new opportunities for understanding rhizobacterial gene regulation in native environments.

59 BASIC BIOLOGICAL SCIENCES

Unveiling a pervasive DNA adenine methylation regulatory network in the early-diverging fungus Rhizopus microsporus

Development of the DNA affinity purification and sequencing (DAP-seq) technique has allowed genome-scale studies of transcription factor (TF)-binding sites with high reproducibility. Here, we apply this technique to the human opportunistic pathogen Rhizopus microsporus, a mucoralean fungus belonging to the understudied group of early-diverging fungi. We characterize genome-wide binding sites of 58 TFs encoded by genes regulated through adenine methylation and representing major TF families. This analysis reveals their binding profiles and recognized sequences, expanding and diversifying the catalog of known fungal motifs. By integrating this data with DNA 6-methyladenine profiling, we uncover the extensive direct and indirect impact of this epigenetic modification on the regulation of gene expression. Furthermore, we use the generated data to identify TFs involved in biologically relevant processes such as zinc metabolism and light response. Our work enhances our understanding of regulatory mechanisms in R. microsporus and provides broader insights into gene regulation across the fungal kingdom.

Lax, Carlos [Universidad de Murcia (Spain)] (ORCID

Surfactant-like peptide gels are based on cross-β amyloid fibrils

Surfactant-like peptides, in which hydrophilic and hydrophobic residues are encoded within different domains in the peptide sequence, undergo facile self-assembly in aqueous solution to form supramolecular hydrogels. These peptides have been explored extensively as substrates for the creation of functional materials since a wide variety of amphipathic sequences can be prepared from commonly available amino acid precursors. The self-assembly behavior of surfactant-like peptides has been compared to that observed for small molecule amphiphiles in which nanoscale phase separation of the hydrophobic domains drives the self-assembly of supramolecular structures. Here, we investigate the relationship between sequence and supramolecular structure for a pair of bola-amphiphilic peptides, Ac-KLIIIK-NH 2 (L2) and Ac-KIIILK-NH 2 (L5). Despite similar length, composition, and polar sequence pattern, L2 and L5 form morphologically distinct assemblies, nanosheets and nanotubes, respectively. Cryo-EM helical reconstruction was employed to determine the structure of the L5 nanotube at near-atomic resolution. Rather than displaying self-assembly behavior analogous to conventional amphiphiles, the packing arrangement of peptides in the L5 nanotube displayed steric zipper interfaces that resembled those observed in the structures of β-amyloid fibrils. Like amyloids, the supramolecular structures of the L2 and L5 assemblies were sensitive to conservative amino acid substitutions within an otherwise identical amphipathic sequence pattern. This study highlights the need to better understand the relationship between sequence and supramolecular structure to facilitate the development of functional peptide-based materials for biomaterials applications.

Das, Abhinaba [Emory University, Atlanta, GA (Unit

Creb5 controls its own expression and directly induces the joint interzone regulatory program

Prior studies have indicated that the transcription factor Creb5 is expressed in the joint interzone, which contains the progenitors for all synovial joint tissues in both mouse and human embryos. In the absence of Creb5 function, most synovial joint interzones fail to form and the cartilage templates in the long bones remain fused. This earlier work did not clarify whether Creb5 initiates a cascade of signaling molecules, such as growth and differentiation factor 5 (Gdf5) and Wnt-family members, that in turn induce the formation of the joint interzone, or instead directly activates the expression of joint interzone markers. In the present study, an integrative analysis of the transcriptome, chromatin accessibility, and Creb5-occupancy in joint progenitors revealed that Creb5 directly binds to both its own two promoters and to the regulatory regions of Gdf5 and Sfrp2, each of whose expression in the joint interzone is Creb5-dependent. Functional enhancer analysis indicated that Creb5 binding sites in either the two Creb5 promoters, or in Gdf5 and Sfrp2 regulatory elements are necessary for these sequences to drive transgene expression in the developing synovial joints. While Creb5 directly drives Gdf5 and Sfrp2 expression in the inner joint interzone, Creb5 activates Barx1 expression specifically in the outer joint interzone. Our findings indicate that Creb5 initiates a regulatory network that both promotes the formation of synovial joints, and subsequently activates distinct transcriptional targets in the inner versus the outer regions of the joint interzone, thus regionalizing gene expression in the developing joint.

Zhang, Cheng-Hai

Identifying impacts of contact tracing on HIV epidemiological inference from phylogenetic data

Abstract Robust sampling methods are foundational to inferences using phylogenies. Yet the impact of using contact tracing, a type of non-uniform sampling used in public health applications such as infectious disease outbreak investigations, has not been investigated in the molecular epidemiology field. To understand how contact tracing influences a recovered phylogeny, we developed a new simulation tool called SEEPS (Sequence Evolution and Epidemiological Process Simulator) that allows for the simulation of contact tracing and the resulting transmission tree, pathogen phylogeny, and corresponding virus genetic sequences. Importantly, SEEPS takes within-host evolution into account when generating pathogen phylogenies and sequences from transmission histories. Using SEEPS, we demonstrate that contact tracing can significantly impact the structure of the resulting tree, as described by popular tree statistics. Contact tracing generates phylogenies that are less balanced than the underlying transmission process, less representative of the larger epidemiological process, and affects the internal/external branch length ratios that characterize specific epidemiological scenarios. We also examined real data from a 2007–2008 Swedish HIV-1 outbreak and the broader 1998–2010 European HIV-1 epidemic to highlight the differences in contact tracing and expected phylogenies. Aided by SEEPS, we show that the data collection of the Swedish outbreak was strongly influenced by contact tracing even after downsampling, while the broader European Union epidemic showed little evidence of universal contact tracing, agreeing with the known epidemiological information about sampling and spread. Overall, our results highlight the importance of including possible non-uniform sampling schemes when examining phylogenetic trees. For that, SEEPS serves as a useful tool to evaluate such impacts, thereby facilitating better phylogenetic inferences of the characteristics of a disease outbreak. SEEPS is available at https://github.com/MolEvolEpid/SEEPS.

Virology

Quantitative Assessment of Parent Well Effect on Hydraulic Fracture Propagation at HFTS2: Insights from Cross-Well Strain Measurements and Microseismic Data

Understanding fracture propagation behavior is essential for optimizing hydraulic fracturing in unconventional reservoirs. This study demonstrates the value of integrating Low-Frequency Distributed Acoustic Sensing (LF-DAS) and microseismic data, which together provide a more complete picture of fracture growth. Using data from Hydraulic Fracturing Test Site 2 (HFTS2), we identify stress changes in depletion zones induced by parent wells as a key factor influencing fracture propagation. This result is shown by new measurements of in-situ fracture propagation velocity and fracture-hit volume (fluid volume at fracture hit?) from LF-DAS and event density from microseismic. These findings highlight the importance of considering parent well effects, well spacing, and stimulation sequencing in completion design to improve reservoir development and production efficiency.

depletion zones