Search NASA⌕ Search

SEARCH · Search NASA

Results for “sequence development”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Utilizing Machine Learning to Improve Neutralization Potency of an HIV-1 Antibody Targeting the gp41 N-Heptad Repeat

The N-heptad repeat (NHR) of the HIV-1 gp41 prehairpin intermediate (PHI) is an attractive potential vaccine target with high sequence conservation across diverse strains. However, despite the potency of NHR-targeting peptides and clinical efficacy of the NHR-targeting entry inhibitor enfuvirtide, no potently neutralizing NHR-directed monoclonal antibodies (mAbs) nor antisera have been identified or elicited to date. The lack of potent NHR-binding mAbs both dampens enthusiasm for vaccine development efforts at this target and presents a barrier to performing passive immunization experiments with NHR-targeting antibodies. To address this challenge, we previously developed an improved variant of the NHR-directed mAb D5, called D5_AR, which is capable of neutralizing diverse tier-2 viruses. Building on that work, here we present the 2.7Å-crystal structure of D5_AR bound to NHR mimetic peptide IQN17. We then utilize protein language models and supervised machine learning to generate small (n < 100) libraries of D5_AR variants that are subsequently screened for improved neutralization potency. We identify a variant with 5-fold improved neutralization potency, D5_FI, which is the most potent NHR-directed monoclonal antibody characterized to date and exhibits broad neutralization of tier-2 and −3 pseudoviruses as well as replicating R5 and X4 challenge strains. Additionally, our work highlights the ability of protein language models to efficiently identify improved mAb variants from relatively small libraries.

Biopolymers↗

Depth-dependent Metagenome-Assembled Genomes of Agricultural Soils under Managed Aquifer Recharge

Abstract Managed Aquifer Recharge (MAR) systems, which intentionally replenish groundwater aquifers with excess water, are critical for addressing water scarcity exacerbated by demographic shifts and climate variability. To date, little is known about the functional diversity of the soil microbiome at different soil depth inhabiting agricultural soils used for MAR. Knowing the functional diversity is pivotal in regulating nutrient cycling and maintaining soil health. Metagenomics, particularly Metagenome-Assembled Genomes (MAGs), provide a powerful tool to explore the diversity of uncultivated soil microbes, facilitating in-depth investigations into microbial functions. In a field experiment conducted in a California vineyard, we sequenced soil DNA before and after water application of MAR. Through this process, we assembled 146 medium and 14 high-quality MAGs, uncovering a wide array of archaeal and bacterial taxa across different soil depths. These findings advance our understanding of the microbial ecology and functional diversity of soils used for MAR, contributing to the development of more informed and sustainable land management strategies.

Science & Technology - Other Topics↗

Streamlined spatial and environmental expression signatures characterize the minimalist duckweed Wolffia australiana

Single-cell genomics permits a new resolution in the examination of molecular and cellular dynamics, allowing global, parallel assessments of cell types and cellular behaviors through development and in response to environmental circumstances, such as interaction with water and the light–dark cycle of the Earth. Here, we leverage the smallest, and possibly most structurally reduced, plant, the semiaquaticWolffia australiana, to understand dynamics of cell expression in these contexts at the whole-plant level. We examined single-cell-resolution RNA-sequencing data and foundWolffiacells divide into four principal clusters representing the above- and below-water-situated parenchyma and epidermis. Although these tissues share transcriptomic similarity with model plants, they display distinct adaptations thatWolffiahas made for the aquatic environment. Within this broad classification, discrete subspecializations are evident, with select cells showing unique transcriptomic signatures associated with developmental maturation and specialized physiologies. Assessing this simplified biological system temporally at two key time-of-day (TOD) transitions, we identify additional TOD-responsive genes previously overlooked in whole-plant transcriptomic approaches and demonstrate that the core circadian clock machinery and its downstream responses can vary in cell-specific manners, even in this simplified system. Distinctions between cell types and their responses to submergence and/or TOD are driven by expression changes of unexpectedly few genes, characterizingWolffiaas a highly streamlined organism with the majority of genes dedicated to fundamental cellular processes.Wolffiaprovides a unique opportunity to apply reductionist biology to elucidate signaling functions at the organismal level, for which this work provides a powerful resource.

Biochemistry & Molecular Biology↗

Synchronic Web Digital Identity: Speculations on the Art of the Possible

As search, social media, and artificial intelligence continue to reshape collective knowledge, the preservation of trust on the public infosphere has become a defining challenge of our time. Given the breadth and versatility of adversarial threats, the best—and perhaps only—defense is an equally broad and versatile infrastructure for digital identity. This document discusses the opportunities and implications of building such an infrastructure from the perspective of a national laboratory. The technical foundation for this discussion is the emergence of the Synchronic Web, a Sandia-developed infrastructure for asserting cryptographic provenance at Internet scale. As of the writing of this document, there is ongoing work to develop the underlying technology and apply it to multiple mission-specific domains within Sandia. The primary objective of this document to extend the body of existing work toward the more public-facing domain of digital identity. Our approach depends on a non-standard, but philosophically defensible notion of identity: digital identity is an unbroken sequence of states in a well-defined digital space. From this foundation, we abstractly describe the infrastructural foundations and applied configurations that we expect to underpin future notions of digital identity.

97 MATHEMATICS AND COMPUTING↗

Automated Programmable Logic Controller Memory Forensics Using RGB Image Analysis and Deep Learning

The introduction of Industry 4.0 and Internet-based technologies has enhanced industrial control system operations but have inadvertently increased their vulnerabilities to cyber attacks. When an industrial control system is compromised, security analysts need to identify the root cause quickly to start the recovery process and develop mitigation strategies. Memory forensics is critical in the incident analysis process to ascertain what occurred. Approaches for analyzing the persistent memory in industrial control devices are limited and almost nonexistent for volatile memory. This chapter proposes an automated methodology for programmable logic controller memory dump analysis using computer vision and deep learning techniques. The methodology converts the sequences of bytes in a programmable logic controller memory dump to red-green-blue pixels and employs a deep learning model that learns the underlying patterns and features of pre-labeled forensic artifacts in images and segments them into distinct regions. The trained model is employed to automatically segment new memory images and identify forensic artifacts. Evaluation of the methodology on a Schneider Electric Modicon M221 programmable logic controller under code injection and code modification attacks demonstrates its ability to detect attack artifacts in memory dumps.

Asmar Awad, Rima [ORNL] (ORCID:0000000233407742)↗

Structural insights into RNase H catalytic mechanism from room-temperature X-ray and neutron crystallography of apo- and RNA/DNA hybrid-bound enzyme

RNase H enzymes are sequence-nonspecific endonucleases that cleave RNA strands in RNA/DNA hybrid duplexes, an enzymatic process essential in DNA replication and repair in both prokaryotes and eukaryotes. Also, RNase H activity of the reverse transcriptase in human immunodeficiency viruses (HIV-1 and HIV-2) is indispensable for the viral replication cycle. RNase H enzymes play an central role in the development of gene therapies and are targets for novel antivirals. It is therefore of great importance to gain a detailed understanding of the RNase H catalytic mechanism to improve drug design. We utilized Bacillus halodurans RNase H1 (BhRNase H1) to shed light on its function and catalytic mechanism. Room-temperature neutron crystallography of the wild-type and inactive D132N mutant enzymes revealed that E109, belonging to the catalytic DEDD motif, can change its protonation state, allowing us to propose its role in the protonation of the leaving O3′ hydroxyl group of RNA. X-ray crystallography has demonstrated the ability of the RNA/DNA duplex to slide along the protein surface upon metal ion binding at site M A , transforming a product mimic into a Michaelis-like complex, which confirms an essential role of the M A metal ion in catalysis.

Enzyme mechanisms↗

DFT-based insight into finite-temperature properties of ferroelectric perovskites with lone-pair: the case of CsGeX 3 (X = Cl, Br, I)

Ferroelectrics remain in the focus of scientific attention for decades owing to their fundamental and practical appeal. Recently, ferroelectricity has been demonstrated in semiconducting halide perovskites (Zhang et al 2022 Sci. Adv. 8 eabj5881), offering both a rare combination of ferroelectricity and semiconductivity in the same material and a possible alternative to the prevailing perovskite oxide ferroelectrics. We propose a route to simulating such materials at finite temperatures capable of reproducing key experimental and first-principle data, such as Curie temperature, phase transition sequence, spontaneous polarization, and soft mode frequencies. The key methodological finding is the superior performance of hybrid exchange correlation functionals in parametrization of effective Hamiltonians for ferroelectrics with lone pair. The parametrization for effective Hamiltonians for CsGeX 3 (X = Cl, Br, I) is reported. The application of methodology to study polarization reversal in CsGeX 3 allows for the development of a ‘minimalistic’ model for polarization reversal in ferroelectrics that provides an insight into the mechanisms of polarization reversal and its key features, such as the relationship between the coercive field, temperature, and AC field frequency. Importantly, the model reveals the origin of the well-known and ever-puzzling overestimation of coercive fields in computations. Furthermore, we report a variety of finite-temperature properties of CsGeX 3 ferroelectrics, such as dielectric susceptibility, pyroelectric coefficients, and energy storage density, which reveal that these halide perovskites possess properties comparable to their oxide counterparts. Here, we believe that our work provides significant methodological advancements, deepens fundamental understanding of ferroelectrics, and reveals the potential of halide perovskite ferroelectrics.

effective Hamiltonian↗

Beyond sequence similarity: toward function-based screening of nucleic acid synthesis

Synthetic nucleic acids are a key input to modern biotechnology, yet they represent dual-use materials that require robust screening to mitigate biosecurity risks. The prevailing screening paradigm, which identifies sequences of concern (SoCs) through sequence similarity to controlled pathogens and toxins, may not fully capture risks posed by AI tools that can decouple biomolecular function from reliance on known sequences. Rapidly advancing biodesign capabilities enable the generation of genes and proteins that might evade sequence-based detection. We highlight the critical need for function-based screening approaches that can detect sequences capable of hazardous biological functions, regardless of similarity to known SoCs. We examine the feasibility of function-based screening with an initial focus on proteins, arguing that, while protein sequence space is vast, biologically functional proteins are significantly constrained by biophysical and biochemical requirements that can be learned and modeled. We propose a concrete implementation framework organized along a continuum of complexity, starting with toxins as the most tractable targets before expanding to more complex pathogenic functions. We then discuss open challenges and describe a research and development strategy to address them.

59 BASIC BIOLOGICAL SCIENCES↗

RLGBS: Reinforcement Learning-Guided Beam Search for process optimization in a paper machine dryer section

Paper drying is responsible for over two-thirds of energy consumption in the U.S. pulp and paper industry, presenting significant potential for energy savings through optimization of process parameters. Current approaches often assume fixed operating conditions, neglecting dynamic ambient and process variations that limit achievable savings and real-world applicability. To this end, we develop a physics-based simulation environment for a paper machine dryer section and propose a reinforcement learning (RL) framework to minimize overall energy consumption by optimizing drying process parameters under diverse operating conditions. To mitigate overdrying and numerical instabilities caused by suboptimal local RL actions, we introduce Reinforcement Learning-Guided Beam Search (RLGBS), which explores multiple action sequences in parallel using beam search. Instead of making step-by-step decisions, RLGBS prioritizes solutions based on cumulative probability, reducing the impact of individual suboptimal actions. Experiments demonstrate that RLGBS achieves consistent energy savings under unseen operating conditions not encountered during training, outperforming conventional RL methods. While validated in drying optimization, this framework is broadly applicable to other RL-based industrial process control problems.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Evaluation of Cross-Protection of African Swine Fever Vaccine ASFV-G-ΔI177L Between ASFV Biotypes

Background/Objectives: Vaccine development for the prevention of ASF has been very challenging due to the extensive genetic and largely unknown antigenic diversity. Inactivated vaccines, using different inactivation methods and a variety of adjuvants, have been consistently inefficacious. Historically, animals recovering from an infection with an attenuated virus became protected from the development of a clinical disease caused by an antigenically related strain. Therefore, immunization of susceptible animals with attenuathe ted virus strains has become a common method of vaccination with the first two commercially available vaccines based on recombinant live-attenuated viruses (LAVs). An important limitation is that the efficacy of the LAV is restricted to those strains that are antigenically related and, in most cases, only provide protection against homologous strains. Due to the unknown antigenic heterogeneity among all ASFV field isolates, the development of broad-spectrum vaccines is a challenge. Besides the anecdotal data, there is not a large amount of information describing patterns of cross-protection between different ASFV strains. Methods: We evaluated the cross-protection induced by the ASFV live-attenuated vaccine ASFV-G-ΔI177L against different biotypes of ASFV and compared their genomic sequences to determine potential genetic mutations that could cause the lack of cross-protection. Results: Results presented here demonstrate different patterns of protection when ASFV-G-ΔI177L vaccinated pigs were challenged with six different ASFV field isolates belonging to different biotypes. Conclusions: The presence of cross-protection cannot be predicted solely by the classical methodology for genotyping-based B646L ORF only. Biotyping, considering the entire virus proteome, appears to be a more promising prediction tool, although additional gathering of experimental data will be necessary to fully validate it; until then, the presence of cross-protection needs to be confirmed in efficacy trials challenging vaccinated animals.

Immunology↗

Automatic speech recognition predicts contemporaneous earthquake fault displacement

Abstract Significant progress has been made in probing the state of an earthquake fault by applying machine learning to continuous seismic waveforms. The breakthroughs were originally obtained from laboratory shear experiments and numerical simulations of fault shear, then successfully extended to slow-slipping faults. Here we apply the Wav2Vec-2.0 self-supervised framework for automatic speech recognition to continuous seismic signals emanating from a sequence of moderate magnitude earthquakes during the 2018 caldera collapse at the Kīlauea volcano on the island of Hawai’i. We pre-train the Wav2Vec-2.0 model using caldera seismic waveforms and augment the model architecture to predict contemporaneous surface displacement during the caldera collapse sequence, a proxy for fault displacement. We find the model displacement predictions to be excellent. The model is adapted for near-future prediction information and found hints of prediction capability, but the results are not robust. The results demonstrate that earthquake faults emit seismic signatures in a similar manner to laboratory and numerical simulation faults, and artificial intelligence models developed for encoding audio of speech may have important applications in studying active fault zones.

58 GEOSCIENCES↗

Purification and expression of a novel bacteriocin, JUQZ-1, against Pseudomonas syringae pv. Actinidiae (PSA), secreted by Brevibacillus laterosporus Wq-1, isolated from the rhizosphere soil of healthy kiwifruit

Kiwifruit canker, caused by Pseudomonas syringae pv. actinidiae (PSA), has led to significant losses in the kiwifruit industry each year. Due to the drug resistance feature of PSA, biological control is currently the most promising method. Developing biocontrol bacteria against PSA could help solve the issue of drug resistance generated during the chemical control of PSA to a certain extent. In this research, a Wq-1 strain that demonstrated excellent inhibitory activity against PSA was isolated from the rhizosphere soil of healthy kiwifruit. Based on the morphological characteristics and phylogenetic analysis of the 16S rRNA gene sequence, the isolated strain was identified as Brevibacillus laterosporus Wq-1. Bacteriostatic proteins were isolated from the cell-free culture filtrate of strain Wq-1 and were found to have a molecular weight of approximately 12 kDa, as determined by sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE). Liquid chromatography–tandem mass spectrometry (LC–MS/MS) detection revealed that there were several peptides in the target band that were consistent with protein 01021 in the genome. The gene of the 01021 protein was cloned into the plasmid pPICZa, and the recombinant bacteriocin was successfully expressed using the Pichia pastoris X33 expression system. The recombinant protein 01021 effectively inhibited the growth of PSA. This is the first report of the protein’s antimicrobial activity, distinguishing it from previously identified bacteriocins. Therefore, we named this bacteriocin JUQZ-1. In addition, our results showed that the protein JUQZ-1 not only exhibited a broad bacteriostatic spectrum but also high thermal and pH stability suitable for harsh environmental conditions., JUQZ-1, a protein with antimicrobial properties and strong environmental tolerance, may serve as a promising alternative to antibiotics.

Shuai, Yang↗

Coherent Transfer of Lattice Entropy via Extreme Nonlinear Phononics in Metal Halide Perovskites

Entropy transfer in metal halide perovskites, characterized by significant lattice anharmonicity and low stiffness, underlies the remarkable properties observed in their optoelectronic applications, ranging from solar cells to lasers. The conventional view of this transfer involves stochastic processes occurring within a thermal bath of phonons, where the lattice arrangement and energy flow from higher- to lower-frequency modes. Here, we unveil a comprehensive chronological sequence detailing a conceptually distinct coherent transfer of entropy in a prototypical perovskite CH 3 NH 3 Pbl 3 . The terahertz periodic modulation imposes vibrational coherence into electronic states, leading to the emergence of mixed (vibronic) quantum beat between approximately 3 and 0.3 THz. We highlight a well-structured bidirectional time-frequency transfer of these diverse phonon modes, each developing at different times and transitioning from high to low frequencies from 3 to 0.3 THz, before reversing direction and ascending to around 0.8 THz. First-principles molecular dynamics simulations disentangle a complex web of coherent-phononic coupling pathways and identify the salient roles of the initial modes in shaping entropy evolution at later stages. Capitalizing on coherent entropy transfer and dynamic anharmonicity presents a compelling opportunity to exceed the fundamental thermodynamic (Shockley-Queisser) limit of photoconversion efficiency and to pioneer novel optoelectronic functionalities. Published by the American Physical Society 2024

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Pooled PPIseq: Screening the SARS-CoV-2 and human interface with a scalable multiplexed protein-protein interaction assay platform

Protein-Protein Interactions (PPIs) are a key interface between virus and host, and these interactions are important to both viral reprogramming of the host and to host restriction of viral infection. In particular, viral-host PPI networks can be used to further our understanding of the molecular mechanisms of tissue specificity, host range, and virulence. At higher scales, viral-host PPI screening could also be used to screen for small-molecule antivirals that interfere with essential viral-host interactions, or to explore how the PPI networks between interacting viral and host genomes co-evolve. Current high-throughput PPI assays have screened entire viral-host PPI networks. However, these studies are time consuming, often require specialized equipment, and are difficult to further scale. Here, we develop methods that make larger-scale viral-host PPI screening more accessible. This approach combines the mDHFR split-tag reporter with the iSeq2 interaction-barcoding system to permit massively-multiplexed PPI quantification by simple pooled engineering of barcoded constructs, integration of these constructs into budding yeast, and fitness measurements by pooled cell competitions and barcode-sequencing. We applied this method to screen for PPIs between SARS-CoV-2 proteins and human proteins, screening in triplicate >180,000 ORF-ORF combinations represented by >1,000,000 barcoded lineages. Our results complement previous screens by identifying 74 putative PPIs, including interactions between ORF7A with the taste receptors TAS2R41 and TAS2R7, and between NSP4 with the transmembrane KDELR2 and KDELR3. We show that this PPI screening method is highly scalable, enabling larger studies aimed at generating a broad understanding of how viral effector proteins converge on cellular targets to effect replication.

60 APPLIED LIFE SCIENCES↗

Microbial spies and bloggers: programming cells to convert environmental information into discernible signals

Microbes regulate their dynamic behaviors using the chemical and physical characteristics of their environment. The ability of microbes to continuously convert this physicochemical information into biochemical information and to use organic matter in the environment as a power source makes these organisms attractive as chassis for building sensors. However, most biosensors have severe limitations when considering applications in hard-to-image settings like soils, sediments, and wastewater. Emerging technologies at the interface of biomolecular design, microbiome engineering, and synthetic biology offer new tools to program cells and communities as biosensors for these settings. Here, in this review, we describe innovations in biosensor outputs that are enabling new applications in complex environments, including reporters that are read out using electrochemical, gas chromatography, hyperspectral imaging, and next-generation sequencing methods. We also discuss computational advances that are accelerating the diversification of sensing components by mining metagenomics data for new transcriptional regulators and by designing allosteric protein switches that directly regulate reporter outputs using analytes. We highlight emerging opportunities for programming undomesticated microbes in communities to function as distributed sensors in the environment. Finally, we discuss the need for responsible biosensor development and to modernize regulatory frameworks to support evidence-based assessment of environmental biosensors.

analyte↗

Recent Progress on Surface Water Quality Models Utilizing Machine Learning Techniques

Surface waterbodies are heavily exposed to pollutants caused by natural disasters and human activities. Empowering sensor technologies in water quality monitoring, sufficient measurements have become available to develop machine learning (ML) models. Numerous ML models have quickly been adopted to predict water quality indicators in various surface waterbodies. This paper reviews 78 recent articles from 2022 to October 2024, categorizing water quality models utilizing ML into three groups: Point-to-Point (P2P), which estimates the current target value based on other measurements at the same time point; Sequence-to-Point (S2P), which utilizes previous time series data to predict the target value at one time point ahead; and Sequence-to-Sequence (S2S), which uses previous time series data to forecast sequential target values in the future. The ML models used in each group are classified and compared according to water quality indicators, data availability, and model performance. Widely used strategies for improving performance, including feature engineering, hyperparameter tuning, and transfer learning, are recognized and described to enhance model effectiveness. The interpretability limitations of ML applications are discussed. This review provides a perspective on emerging ML for surface water quality models.

machine learning (ML)↗

Long-read sequencing transcriptome quantification with lr-kallisto

RNA abundance quantification has become routine and affordable thanks to high-throughput “short-read” technologies that provide accurate molecule counts at the gene level. Similarly accurate and affordable quantification of definitive full-length, transcript isoforms has remained a stubborn challenge, despite its obvious biological significance across a wide range of problems. “Long-read” sequencing platforms now produce data-types that can, in principle, drive routine definitive isoform quantification. However some particulars of contemporary long-read datatypes, together with isoform complexity and genetic variation, present bioinformatic challenges. We show here, using ONT data, that fast and accurate quantification of long-read data is possible and that it is improved by exome capture. To perform quantifications we developed lr-kallisto, which adapts the kallisto bulk and single-cell RNA-seq quantification methods for long-read technologies.

Loving, Rebekah K. (ORCID:0000000187250376)↗

Randomized Algorithms for Symmetric Nonnegative Matrix Factorization

Symmetric Nonnegative Matrix Factorization (SymNMF) is a technique in data analysis and machine learning that approximates a matrix with a product of a nonnegative, low-rank matrix and it transpose. To design faster and more scalable algorithms for SymNMF we develop two randomized algorithms for its computation. The first method uses randomized matrix sketching to compute an initial low-rank approximation to the input matrix and proceeds to uses this as a low-rank input to rapidly compute a SymNMF. The second methods uses randomized leverage score sampling to approximately solve constrained least squares problems. Many successful methods for SymNMF rely on (approximately) solving sequences of constrained least squares problems. Here, we prove theoretically that leverage score sampling can approximately solve constrained least squares problems to e-accuracy. Finally we demonstrate both methods work in practice by applying them to graph clustering tasks on large real world data sets. These experiments show that our methods approximately maintain solution quality and achieve significant speed ups for both large dense and large sparse problems.

97 MATHEMATICS AND COMPUTING↗