Search NASA⌕ Search

SEARCH · Search NASA

Results for “encoding”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Dense autoencoders, clustering techniques, and semi-supervised learning for HPGe $γ$-spectra

Classifying high-resolution gamma spectra by their isotopic content is an essential task in nuclear forensics and other applications. Traditional analysis methods are often time-intensive, but machine learning (ML) may help analysts quickly process many spectra. Such methods tend to rely on abundant, well-labeled data for training. Historical gamma data exists in various fields but is not uniformly useful for supervised ML due to inconsistent labeling. Here, to address some of these challenges, we present a method to classify and organize unlabeled data from high-purity germanium detectors using an autoencoding neural network (autoencoder). We trained dense autoencoders to compress gamma data into latent representations that enable efficient data characterization. By clustering the encoded spectra or lower-dimensional mappings of them, we identified and removed portions of over-abundant data categories, resulting in a more balanced dataset and improved autoencoder performance. This encoding and clustering pipeline also enabled the organization of spectra into self-consistent categories. Finally, we found that encoded representations showed potential as inputs for semi-supervised learning of nuclide identification (NID) labels, achieving an average F1 score of 0.85 ± 0.03 when mapping encodings to a set of 65 isotope labels.

Autoencoders↗

Genomics and physiology of Catenibacillus, human gut bacteria capable of polyphenol C-deglycosylation and flavonoid degradation

The genusCatenibacillus(familyLachnospiraceae, phylumBacillota) includes only one cultivated species so far,Catenibacillus scindens,isolated from human faeces and capable of deglycosylating dietary polyphenols and degrading flavonoid aglycones. Another human intestinalCatenibacillusstrain not taxonomically resolved at that time was recently genome-sequenced. We analysed the genome of this novel isolate, designatedCatenibacillus decagia, and showed its ability to deglycosylateC-coupled flavone and xanthone glucosides andO-coupled flavonoid glycosides. Most of the resulting aglycones were further degraded to the corresponding phenolic acids. Including the recently sequenced genome ofC. scindensand ten faecal metagenome-assembled genomes assigned to the genusCatenibacillus, we performed a comparative genome analysis and searched for genes encoding potentialC-glycosidases and other polyphenol-converting enzymes. According to genome data and physiological characterization, the core metabolism ofCatenibacillusstrains is based on a fermentative lifestyle with butyrate production and hydrogen evolution. BothC. scindensandC. decagiaencode a flavonoidO-glycosidase, a flavone reductase, a flavanone/flavanonol-cleaving reductase and a phloretin hydrolase. Several gene clusters encode enzymes similar to those of the flavonoidC-deglycosylation system ofDoreastrain PUE (DgpBC), while separately located genes encode putative polyphenol-glucoside oxidases (DgpA) required forC-deglycosylation. The diversity ofdgpAanddgpBCgene clusters might explain the broadC-glycoside substrate spectrum ofC. scindensandC. decagia. The otherCatenibacillusgenomes encode only a few potential flavonoid-converting enzymes. Our results indicate that severalCatenibacillusspecies are well-equipped to deglycosylate and degrade dietary plant polyphenols and might inhabit a corresponding, specific niche in the gut.

Genetics & Heredity↗

Variational Simulation of the Lipkin-Meshkov-Glick Model on a Neutral Atom Quantum Computer

We simulate the Lipkin-Meshkov-Glick model using the variational-quantum-eigensolver algorithm on a neutral atom quantum computer. We test the ground-state energy of spin systems with up to 15 spins. Two different encoding schemes are used: an individual spin encoding where each spin is represented by one qubit, and an efficient Gray code encoding scheme that only requires a number of qubits that scales with the logarithm of the number of spins. This more efficient encoding, together with zero-noise extrapolation techniques, is shown to improve the fidelity of the simulated energies with respect to exact solutions.

97 MATHEMATICS AND COMPUTING↗

3D Deep Learning Joint Inversion of Active Seismic Full Waveform and Passive Seismic Traveltime Data for Reservoir Imaging and Uncertainty Quantification

Here, we present deep learning (DL) networks for three-dimensional (3D) joint inversion of active seismic full waveform and passive seismic traveltime data to image reservoirs and their properties and quantify imaging uncertainties. Active seismic full-waveform data can provide high-resolution monitoring images but are collected only intermittently because of their high acquisition cost. In contrast, passive seismic data can be gathered at relatively low cost between regular active surveys, although their imaging quality can be compromised by factors such as low signal-to-noise ratios and limited ray coverage of the target. Although these datasets are routinely acquired together at CO 2 storage sites, their combined inversion within a 3D DL framework has not been previously demonstrated. To our knowledge, this is the first study to address this gap, combining the strength of both data types. For efficient data storage and DL training with large 3D seismic datasets, we use a 3D data matrix in which a random number of passive seismic traveltime data are stored as parabolic envelopes using one-hot encoding and a 3D full-waveform data matrix in which multiple shot gathers are summed. Two network architectures are evaluated: a single-encoder U-Net for single-data type inversion and a dual-encoder U-Net for joint inversion of active and passive seismic data. We also evaluate the single-encoder U-Net for joint inversion by concatenating full-waveform data and traveltime data. We propose a systematic approach for selecting an optimal dropout rate that balances regularization during training and Monte Carlo dropout-based uncertainty quantification during prediction by examining the correlation coefficient between standard deviation and prediction error, along with the training misfit, across a range of dropout rates. 3D DL inversion experiments include five different network configurations, with evaluations under ideal, noisy and dropout-enabled conditions. Both model and data uncertainties are assessed, as well as their combined effects. Across all conditions, the networks consistently predict accurate CO 2 saturation models with low prediction errors, such as a structural similarity index of 0.993 and CO 2 difference of 1.1%. Uncertainty estimates show strong spatial correlation with prediction errors, confirming the effectiveness of the proposed dropout selection approach. The results demonstrate that our DL approach, utilizing compact data representations and appropriate uncertainty quantification, yields accurate subsurface images under various inversion conditions and provides valuable insights into the reliability of predictions.

Um, Evan Schankee [Lawrence Berkeley National Labo↗

Tonoplast Sucrose Transporter SUT4-Dependent Sugar Partitioning Modulates Phenological Transitions and Reproductive Success in Poplar

Climate uncertainty is intensifying the need for greater plasticity in carbohydrate reserve utilization to support winter survival and spring growth in woody perennials. In poplar, the single-copy SUT4, which encodes a tonoplast-localized sucrose transporter, and the SUT5/SUT6 genome duplicates, which encode plasma membrane-localized transporters, are expressed year-round, with SUT4 showing the highest expression during cool seasons. Given its role in vacuolar sucrose efflux and winter-predominant expression, SUT4 may play a key role in modulating seasonal carbohydrate dynamics. While SUT4-knockdown and knockout effects have been studied under greenhouse conditions, their impact under field conditions remains unexplored. Here, we report a field-based study comparing CRISPR knockout mutants of winter-expressed SUT4 and SUT5/SUT6 in Populus tremula x alba. We show that sut4, but not sut5/6, mutants exhibited earlier autumn leaf senescence, delayed spring bud flush, reduced stem growth, and altered sugar partitioning in winter xylem and bark relative to controls. After 2 years in the field, all genotypes flowered before leaf flush in early spring; however, sut4 mutants produced sterile ovules despite developing normal-looking catkins. Metabolic profiling revealed disrupted sucrose and raffinose dynamics in elongating sut4 catkins. This was accompanied by transcriptomic signatures of elevated stress and downregulation of proanthocyanidin biosynthesis and circadian clock genes. These findings highlight the critical role of SUT4 in coordinating sugar allocation, stress responses, and seasonal development in poplar.

09 BIOMASS FUELS↗

The mevalonate pathway of isoprenoid biosynthesis supports metabolic flexibility in Mycobacterium marinum

ABSTRACT Isoprenoids are a diverse class of natural products that are essential in all domains of life. Most bacteria synthesize isoprenoids through either the methylerythritol phosphate (MEP) pathway or the mevalonate (MEV) pathway, while a small subset encodes both pathways, including the pathogen Mycobacterium marinum (Mm). It is unclear whether the MEV pathway is functional in Mm, or why Mm encodes seemingly redundant metabolic pathways. Here, we show that the MEP pathway is essential in Mm, while the MEV pathway is dispensable in culture, with the ΔMEV mutant having no growth defect in axenic culture but a competitive growth defect compared to WT Mm. We found that the MEV pathway does not play a role in ex vivo or in vivo acute infection but does play a role in survival of peroxide stress. Metabolite profiling revealed that modulation of the MEV pathway causes compensatory changes in the concentration of MEP intermediates DOXP and CDP-ME, suggesting that the MEV pathway is functional and that the pathways interact at the metabolic level. Finally, the MEV pathway is upregulated early in the shift down to hypoxia, suggesting that it may provide metabolic flexibility to this bacterium. Interestingly, we found that our complemented strains, which vary in copy number of the polyprenyl synthetase idsB2 , responded differently to peroxide and UV stresses, suggesting a role for this gene as a determinant of downstream prenyl phosphate metabolism. Together, these findings suggest that MEV may serve as an anaplerotic pathway to make isoprenoids under stress conditions. IMPORTANCE Organisms from all domains of life utilize isoprenoids to carry out thousands of critical and auxiliary cellular processes, including signaling, maintaining membrane integrity, stress response, and host-pathogen interactions. The common precursor of all isoprenoids is synthesized via one of two biosynthetic pathways. Importantly, some bacteria encode both pathways, including M. marinum . We found that only one pathway is essential in M. marinum , while the nonessential pathway may confer metabolic flexibility to help the bacterium better adapt to various environmental conditions. We also found that the polyprenyl synthetase IdsB2 plays an important role in driving such phenotypes. Further, we demonstrate metabolic interplay between both functional pathways. These insights represent the first characterization of isoprenoid biosynthesis in dual pathway-encoding mycobacteria.

Qabar, Christine M. [Department of Plant and Micro↗

Floodplain nitrifiers harbor the genetic potential for utilizing a wide range of organic nitrogen compounds

Organic compounds such as urea and cyanate can serve as nitrogen (N) sources for nitrifying microorganisms, including ammonia-oxidizing archaea (AOA) and bacteria (AOB), complete ammonia-oxidizing (comammox) bacteria, and nitrite-oxidizing bacteria (NOB). Here we investigated metagenome-assembled genomes (MAGs) for all four nitrifier guilds generated from hydrologically variable floodplain sediments of the Wind River Basin (WRB; Riverton, WY, USA) for their genetic potential to utilize organic N compounds. A vast majority of WRB nitrifier MAGs harbored urease (ure) and at least one urea transporter ( utp, urt, dur3 ). AOA were the most abundant and phylogenetically diverse nitrifiers in WRB floodplain sediments. Several AOA MAGs encoded cyanase ( cynS ), nitrilase ( nit1 ), omega-amidase ( nit2 ), nitrile hydratase ( nthA ), and genes related to purine degradation, including biuret hydrolase ( biuH ), oxamic transcarbamylase ( allFGH ), and catabolic carbamate kinase ( allK ). AOA often encoded an uncharacterized amidohydrolase collocated with biuH , rather than allophanate hydrolase ( atzF ). A small number of AOA encoded atzF , functioning in an unknown pathway. AOB and comammox were of relatively low abundance and taxonomic diversity and were present only at certain depths in WRB; however, they encoded triuret/biuret degradation genes ( trtA, biuH , and atzH ), and in comammox, these genes were also collocated with allFGHK . The genetic potential of ammonia oxidizers in the WRB floodplain suggests that organic N may support nitrification in this system. The proposed pathways for utilizing purine degradation products other than urea potentially expand the known metabolic capabilities of AOA, AOB, and comammox bacteria and reveal the possibility for cryptic N cycling between microbial community members.

floodplain↗

LLNL FESP Theory Highlights: October 2024

I. Novikau, I. Y. Dodin, E. A. Startsev, I. Joseph, Quantum algorithms for simulating dissipative linear and nonlinear dynamics of plasmas. Invited talk at the 66th Annual Meeting of the APS Division of Plasma Physics, Atlanta, Georgia. Novikau I., Dodin I.Y., Startsev E.A., Encoding of linear kinetic plasma problems in quantum circuits via data compression, Journal of Plasma Physics. 2024;90(4):805900401, doi:10.1017/S0022377824000795. We propose an algorithm for encoding linear kinetic plasma problems in quantum circuits. The focus is on modelling electrostatic linear waves in a one-dimensional Maxwellian electron plasma. The waves are described by the linearized Vlasov–Ampère system with a spatially localized external current that drives plasma oscillations. This system is formulated as a boundary-value problem and cast in the form of a linear vector equation to be solved by using the quantum signal processing algorithm. The latter requires encoding of a matrix in a quantum circuit as a sub-block of a unitary matrix. We propose how to encode in a circuit in a compressed form and discuss how the resulting circuit scales with the problem size and the desired precision.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

The Baby Universe is Fine and the CFT Knows It: On Holography for Closed Universes

Big bang/big crunch closed universes can be realized in AdS/CFT, even though they lack asymptotically AdS boundaries. With enough bulk entanglement, the bulk Hilbert space of a closed universe can be holographically encoded in the CFT. We clarify the relation of this encoding to observer-clone proposals and refute recent arguments about the breakdown of semiclassical physics in such spaces. In the limit of no bulk entanglement, the holographic encoding breaks down. The oft-cited one-dimensional nature of the closed universe Hilbert space represents the limitation of the external (CFT) Hilbert space to access the quantum information in the closed universe, similar to the limitations imposed on observers outside a perfectly isolated quantum lab. We advocate that the CFT nevertheless continues to determine the physical properties of the closed universe in this regime, showing how to interpret this relationship in terms of a final state projection in the closed universe. We provide a dictionary between the final state wavefunction and CFT data. We propose a model of the emergence of an arrow of time in the universe with a given initial or final state projection. Finally, we show that the conventional EFT in the closed universe, without any projection, can be recovered as a maximally ignorant description of the final state. This conventional EFT is encoded in CFT data, and it can be probed by computing coarse-grained observables. We provide an example of one such observable. Taken together, these results amount to a clean bill of health for baby universes born of AdS/CFT.

FOS: Physical sciences↗

NLR HPC Eagle GPU Node Metrics

Ganglia node metrics and iLO (Integrated Lights Out) power data captured from six representative Eagle GPU nodes The Eagle HPC operated at NLR from 2019 through 2024. Eagle was a 2,000-node, 8-petaflop system. This dataset is a representative sample of metrics for 6 of the GPU nodes. Each GPU node contained 2 CPUs and 2 GPUs. Data provided in compressed CSV format. Ganglia and iLO Power Time Series Fields ts: Timestamp dv: Device / Node - Rack and Unit - r103u17 == r(ack)103u(nit)17 mt: Metric (only present for Ganglia) vl: Value - Value in watts for iLO power (instantaneous value at sampling time) or specified Ganglia metric below Ganglia Metrics Metric name -- Metric description -- Unit cpu_aidle -- Percent of time since boot idle CPU -- Percent cpu_idle -- Percent CPU idle -- Percent cpu_nice -- Percent CPU nice -- Percent cpu_speed -- Speed in MHz of CPU -- MHz cpu_user -- Percent CPU user -- Percent cpu_wio -- The percentage of CPU Wait I/O -- Percent gpu0_bar1_memory -- Used GPU bar1 memory -- MB gpu0_decoder_util -- GPU decoder utilization -- Percent gpu0_ecc_db_error -- Total ECC error counts for the GPU -- Number gpu0_encoder_util -- GPU encoder utilization -- Percent gpu0_fan -- Fan speed -- RPM gpu0_fb_memory -- Used GPU framebuffer memory -- MB gpu0_graphics_clock_report -- Current clock speeds for the device -- MHz gpu0_mem_total -- Memory total -- MB gpu0_mem_util -- Memory utilization -- Percent gpu0_power_usage_report -- Power usage report -- Watts gpu0_temp -- GPU 1 temperature -- Celsius gpu1_bar1_memory -- Used GPU bar1 memory -- MB gpu1_decoder_util -- GPU decoder utilization -- Percent gpu1_ecc_db_error -- Total ECC error counts for the GPU -- Number gpu1_encoder_util -- GPU encoder utilization -- Percent gpu1_fan -- Fan speed -- RPM gpu1_fb_memory -- Used GPU framebuffer memory -- MB gpu1_graphics_clock_report -- Current clock speeds for the GPU -- MHz gpu1_mem_total -- Memory total -- MB gpu1_mem_util -- Memory utilization -- MB gpu1_power_usage_report -- Power usage report -- Watts gpu1_temp -- GPU 1 temperature -- Celsius ipmi_cpu1_temp -- CPU 1 temperature -- Celsius ipmi_cpu2_temp -- CPU 2 temperature -- Celsius ipmi_inlet_ambient_temp -- Temperature measured at intake -- Celsius ipmi_vr_p1_temp -- CPU 1 voltage regulator temperature -- Celsius ipmi_vr_p2_temp -- CPU 2 voltage regulator temperature -- Celsius mem_buffers -- Amount of buffered memory -- Bytes mem_cached -- Amount of cached memory -- Bytes mem_free -- Amount of available memory -- Bytes mem_shared -- Amount of shared memory -- Bytes mem_total -- Amount of available memory -- Bytes

97 MATHEMATICS AND COMPUTING↗

Deformable phrase level attention: A flexible approach for improving AI based medical coding

Objective: Improving the AI-driven automated medical encoding of clinical text plays a vital role in gathering information on the occurrence of diseases to improve population-level health. This work presents a novel attention mechanism designed to enhance text classification models and ensure appropriate classification of medical concepts in unstructured electronic health records. Materials and Methods: We developed a deformable, phrase-level attention mechanism to identify important lexical word-level and contextual phrase-level information from clinical text documents. We evaluated conventional and transformer-based deep learning models that we extended with our attention mechanism on the extraction of critical cancer information (e.g., site, subsite, laterality, histology, behavior) from 629,908 electronic pathology reports and on the automated medical encoding of 52,722 hospital discharge summaries. Results: Transformer-based models with the deformable, phrase-level attention mechanism achieved the best performance on the extraction of critical cancer information from pathology reports. Conventional- and transformer-based models show similar or better performance than their baseline counterparts on the automated medical encoding of clinical documents. Discussion: The addition of phrase-level information allowed models extended with our proposed method to outperform standard word-level attention. Our method showed favorable properties for the real-world application in terms of model robustness and phenotyping. These results indicate that our method is promising for automated data harmonization for common data models. Conclusion: This work proposes a novel deformable, phrase-level attention mechanism that enhances text classification models in the extraction of medical concepts from clinical text documents. We demonstrate strong performances on two clinical text datasets and showcase real-world deployability of our method.

Automated medical encoding↗

Chlamydomonas cells transition through distinct Fe nutrition stages within 48 h of transfer to Fe-free medium

Low iron (Fe) bioavailability can limit the biosynthesis of Fe-containing proteins, which are especially abundant in photosynthetic organisms, thus negatively affecting global primary productivity. Understanding cellular coping mechanisms under Fe limitation is therefore of great interest. For this paper, we surveyed the temporal responses of Chlamydomonas ( Chlamydomonas reinhardtii ) cells transitioning from an Fe-rich to an Fe-free medium to document their short and long-term adjustments. While slower growth, chlorosis and lower photosynthetic parameters are evident only after one or more days in Fe-free medium, the abundance of some transcripts, such as those for genes encoding transporters and enzymes involved in Fe assimilation, change within minutes, before changes in intracellular Fe content are noticeable, suggestive of a sensitive mechanism for sensing Fe. Promoter reporter constructs indicate a transcriptional component to this immediate primary response. With acetate provided as a source of reduced carbon, transcripts encoding respiratory components are maintained relative to transcripts encoding components of photosynthesis and tetrapyrrole biosynthesis, indicating metabolic prioritization of respiration over photosynthesis. In contrast to the loss of chlorophyll, carotenoid content is maintained under Fe limitation despite a decrease in the transcripts for carotenoid biosynthesis genes, indicating carotenoid stability. These changes occur more slowly, only after the intracellular Fe quota responds, indicating a phased response in Chlamydomonas, involving both primary and secondary responses during acclimation to poor Fe nutrition.

59 BASIC BIOLOGICAL SCIENCES↗

AGFormer: Adaptive Spatiotemporal graph informed transformer for multi-reservoir inflow forecasting

Accurate reservoir inflow forecasting is crucial for effective water resource management, yet most machine learning models focus on single-reservoir prediction and overlook spatial dependencies among hydrologically connected reservoirs. Here, we propose AGFormer (Adaptive Graph-Informed Transformer), an end-to-end framework that integrates adaptive graph learning with temporal sequence modeling for multi-reservoir inflow forecasting. A shared encoder and graph attention mechanism generate reservoir-specific embeddings, which are then processed by the Transformer-based encoder–decoder for multi-step inflow forecasting. We also introduce a pretraining paradigm to learn robust temporal embeddings from misaligned historical records. Evaluated on 30 reservoirs in the Upper Colorado River Basin, AGFormer achieves superior seven-day-ahead forecasts, with NSE > 0.75 for 20 reservoirs—outperforming Encoder–Decoder LSTM, GCN+LSTM, and Transformer baselines. Adaptive graph learning captures dynamic inter-reservoir dependencies, and feature attribution aligns with snowmelt-driven hydrology. Incorporating forecasted meteorological inputs further enhances accuracy, demonstrating AGFormer’s potential to support reservoir management under dynamic hydrological conditions.

Adaptive graph learning↗

An efficient explicit implementation of a near-optimal quantum algorithm for simulating linear dissipative differential equations

We propose an efficient block-encoding technique for the implementation of the Linear Combination of Hamiltonian Simulations (LCHS) for simulating dissipative initial-value problems. This algorithm approximates a target nonunitary operator as a weighted sum of Hamiltonian evolutions, thereby emulating a dissipative problem by mixing various time scales. We introduce an efficient encoding of the LCHS into a quantum circuit based on a simple coordinate transformation that turns the dependence on the summation index into a trigonometric function. Classically, this method is equivalent to the use of a highly accurate Fejér-Clenshaw-Curtis quadrature formula. Quantumly, this significantly simplifies block-encoding of a dissipative problem and allows one to perform an exponential number of Hamiltonian simulations by a single Quantum Signal Processing (QSP) circuit. The resulting LCHS circuit has high success probability and the selector scales logarithmically with the number of terms in the LCHS sum and linearly with time. Careful analysis of error convergence proves that this method is more efficient than other LCHS circuits that have recently appeared in the literature. We verify the quantum circuit and its scaling by simulating it on a digital emulator of fault-tolerant quantum computers and, as a test problem, solve the advection-diffusion equation. The proposed algorithm can be used for simulating a wide class of nonunitary initial-value problems including the Liouville equation with added dissipation and linear embeddings of nonlinear systems, such as the Koopman-von Neumann and Carleman embeddings.

Novikau, I [Lawrence Livermore National Laboratory↗

Quantum frequency resampling

In signal processing, resampling algorithms can modify the number of resources encoding a collection of data points. Downsampling reduces the cost of storage and communication, while upsampling interpolates new data from limited one, e.g., when resizing a digital image. We present a toolset of quantum algorithms to resample data encoded in the probabilities of a quantum register, using the quantum Fourier transform to adjust the number of high-frequency encoding qubits. We discuss advantage over classical resampling algorithms.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Schrödinger cat states of a nuclear spin qudit in silicon

High-dimensional quantum systems are a valuable resource for quantum information processing. They can be used to encode error-correctable logical qubits, which has been demonstrated using continuous-variable states in microwave cavities or the motional modes of trapped ions. For example, high-dimensional systems can be used to realize ‘Schrödinger cat’ states, which are superpositions of widely displaced coherent states that can be used to illustrate quantum effects at large scales. Recent proposals have suggested encoding qubits in high-spin atomic nuclei, which are finite-dimensional systems that can host hardware-efficient versions of continuous-variable codes. Here, in this study, we demonstrate the creation and manipulation of Schrödinger cat states using the spin-7/2 nucleus of an antimony atom embedded in a silicon nanoelectronic device. We use a multi-frequency control scheme to produce spin rotations that preserve the symmetry of the qudit, and we constitute logical Pauli operations for qubits encoded in the Schrödinger cat states. Our work demonstrates the ability to prepare and control non-classical resource states, which is a prerequisite for applications in quantum information processing and quantum error correction, using our scalable, manufacturable semiconductor platform.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

An expanded registry of candidate cis -regulatory elements

Mammalian genomes contain millions of regulatory elements that control the complex patterns of gene expression. Previously, the ENCODE consortium mapped biochemical signals across hundreds of cell types and tissues and integrated these data to develop a registry containing 0.9 million human and 300,000 mouse candidate cis-regulatory elements (cCREs) annotated with potential functions. Here we have expanded the registry to include 2.37 million human and 967,000 mouse cCREs, leveraging new ENCODE datasets and enhanced computational methods. This expanded registry covers hundreds of unique cell and tissue types, providing a comprehensive understanding of gene regulation. Functional characterization data from assays such as STARR-seq, massively parallel reporter assay, CRISPR perturbation and transgenic mouse assays have profiled more than 90% of human cCREs, revealing complex regulatory functions. We identified thousands of novel silencer cCREs and demonstrated their dual enhancer and silencer roles in different cellular contexts. Integrating the registry with other ENCODE annotations facilitates genetic variation interpretation and trait-associated gene identification, exemplified by the identification of KLF1 as a novel causal gene for red blood cell traits. This expanded registry is a valuable resource for studying the regulatory genome and its impact on health and disease.

Moore, Jill E. [Univ. of Massachusetts, Worchester↗

Structural insights into VRC01-class bnAb precursors with diverse light chains elicited in the IAVI G001 human vaccine trial

The development of germline-targeting vaccines represents a potentially transformative strategy to elicit broadly neutralizing antibodies (bnAbs) against HIV and other antigenically diverse pathogens. Here, we report on structural characterization of vaccine-elicited VRC01-class bnAb precursors in the IAVI G001 Phase 1 clinical trial with the eOD-GT8 60mer nanoparticle as immunogen. High-resolution X-ray structures of eOD-GT8 monomer complexed with Fabs of five VRC01-class bnAb precursors with >90% germline identity revealed a conserved mode of binding to the HIV CD4-binding site via IGHV1-2-encoded heavy chains, mirroring mature bnAb interactions. The light-chain V-gene diversity emulated VRC01 bnAbs and stabilized antigen engagement, while their conserved five-residue LCDR3 motifs prevented steric clashes. Notably, the VRC01-class bnAb precursors accommodated the N276 glycan, a key barrier in HIV Env recognition, through structural rearrangements in HCDR3 or LCDR1, despite its absence in the immunogen. Surface plasmon resonance analysis showed that 87% of elicited antibodies retained glycan binding capacity, albeit with reduced affinity. These findings validate the ability of eOD-GT8 60mer nanoparticles to prime VRC01-class bnAb precursors with native-like paratopes but with intrinsic glycan adaptability. Structural mimicry of mature bnAbs was observed even with limited somatic hypermutation, indicating that critical features are encoded in the germline repertoire. The structures highlight how germline-encoded features drive bnAb-like recognition at early stages. This work provides molecular evidence supporting germline targeting in humans and provides guidance for designing booster immunogens to shepherd affinity maturation toward broad neutralization.

Science & Technology - Other Topics↗