Search NASA⌕ Search

SEARCH · Search NASA

Results for “sequence development”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

scPlantAnnotate: an accurate and robust transformer-based model for plant cell type annotation

Accurate cell type annotation remains a major bottleneck in plant single-cell RNA sequencing (scRNA-seq), where existing tools are often adapted from animal studies and perform sub-optimally on plant data. The lack of plant-specific computational frameworks limits the construction of plant cell atlases and downstream biological discovery. We develop and evaluate scPlantAnnotate, a Transformer-based reference annotation framework tailored for plant scRNA-seq data, and benchmark it against state-of-the-art deep learning and conventional methods across multiple plant species. Species-specific scPlantAnnotate models were trained using curated datasets from Arabidopsis thaliana, Zea mays, Oryza sativa, and Glycine max. We compared scPlantAnnotate with leading baselines under both standard random-split evaluation and a more stringent leave-one-dataset-out setting, which tests robustness to completely unseen datasets and tissue types. scPlantAnnotate consistently outperforms existing approaches across all four species under random-split evaluation. In the leave-one-dataset-out setting for A. thaliana, where performance drops markedly for all methods due to strong batch effects and dataset heterogeneity, scPlantAnnotate nonetheless achieves the highest Accuracy, Macro-F1, Balanced Accuracy, and Macro-AUROC on average and ranks first on most held-out datasets. These results demonstrate improved robustness to dataset shifts, a critical yet underexplored challenge in plant scRNA-seq analysis. A freely accessible web server enables users to annotate their own datasets using pretrained models. scPlantAnnotate provides a plant-specific, Transformer-based framework for single-cell annotation that delivers state-of-the-art performance and enhanced robustness to unseen datasets. By addressing limitations of existing tools and enabling scalable reference-based annotation, scPlantAnnotate supports the development of comprehensive plant cell atlases and facilitates broader use of single-cell genomics in plant biology.

Bioinformatics↗

Detecting Reactive Products in Carbon Capture Polymers with Chemical Shift Anisotropy and Machine Learning

Aminopolymers are attractive sorbents for CO 2 direct air capture applications due to their high density of amine groups, which can readily react with atmospheric levels of CO 2 to form chemisorbed species. The identity of these chemisorbed species and the functional groups that form upon oxidative degradation depends on both material properties and processing conditions, forming a variety of carbonyl-type sites such as ammonium carbamates, bicarbonates, carbonates, carbamic acids, ureas, and amides. 13 C solid-state nuclear magnetic resonance (NMR) is often used to help elucidate the identity of these reacted species, but it is challenging due to the narrow chemical shift range of carbonyl sites. Herein, we demonstrate the application of a two-dimensional (2D) chemical shift anisotropy (CSA) recoupling pulse sequence (ROCSA) to obtain CSA tensor values at each isotropic chemical shift, overcoming limitations of isotropic peak resolution. CSA tensor values describe the local chemical environment and can readily differentiate between the chemisorbed and degradation products. To aid identification, we also developed a k-nearest neighbor (kNN) classification model to distinguish the functional groups via their CSA tensor parameters. This methodology was demonstrated on poly(ethylenimine) in γ-Al 2 O 3 exposed to CO 2 and showed that the chemisorbed products are ammonium carbamate and a mixed carbamate–carbamic acid species. The sample was analyzed again after desorption at 100 °C inducing mild degradation, and the remaining products were strongly bound carbamate and urea species. In conclusion, the combination of 2D CSA measurements coupled with a kNN classification model enhances the ability to accurately identify chemisorbed or degradation products in complex carbon capture materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Local Chain Dynamics in Sequence-Controlled Polymers as a Tunable Handle for Rare Earth Sequestration

Chain dynamics govern the intricate behaviors of proteins, underpinning functions such as catalysis, recognition, and stimulus response, and are an increasingly appreciated aspect of structure–function relationships. Analogously, manipulating chain dynamics and structure in abiotic polymers via sequence control is an exciting, yet underexplored, strategy for improving material functions. In this work, we report a systematic study relating the sequence of polymeric sequestrants to their structure and dynamics, as well as to their binding affinity and selectivity for model substrates, rare earth elements (REEs). A series of sequence-controlled polymers with metal chelating, solubilizing, and structure forming monomers was synthesized via multiblock polymerization, yielding compositionally identical polymers with spectroscopically resolved domains and distinct morphologies. Using a combination of small-angle X-ray scattering and 19 F NMR relaxometry measurements, we connected differences in polymer structure and dynamics to polymer sequence variables such as the patchiness (density) of the structure forming monomer and the location of the chelating monomer. Furthermore, we found that, relative to calcium, all polymers in the series collapse more and have slower dynamics when binding REEs (lanthanum and lutetium) , though the extent of these effects were sequence-dependent and localized to specific domains within the polymer. Notably, sequence-controlled polymers that exhibited the largest conformational and dynamic changes upon binding REEs also bound REEs with the greatest affinity and modest selectivity. Collectively, these results correlate monomer patterning with dynamics, morphology, and REE binding performance en route to the development of efficient and selective macromolecular chelators.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

TRACE Input Modernization

This work presents a Tom’s Obvious Minimal Language (TOML)-based representation of input for the US Nuclear Regulatory Commission’s TRAC/RELAP Advanced Computational Engine (TRACE) thermal hydraulics code. Implemented using the Workbench Analysis Sequence Processor (WASP), the approach maps traditional TRACE input structures to a hierarchical format composed of named parameters, typed values, and native data collections. The resulting representation preserves TRACE’s existing modeling capabilities while providing a modern, structured interface for model development and management. WASP further extends TOML through a file import directive that supports modular model composition and reusable input organization. In addition, WASP provides extended array data entry convenience with various data repeat and interpolation capabilities. Examples of the new TOML syntax are provided for major TRACE input categories, including hydraulic components, heat structures, control systems, and trip logic. The TOML representation establishes a foundation for improved validation, tooling, automation, and model maintainability while remaining compatible with existing TRACE workflows. To facilitate migration to the TOML-based input format, the TRACE executable now supports conversion of native TRACE input into an intermediate JSON representation. A Python utility subsequently transforms the JSON data into an equivalent TOML model. Lastly, the TRACE executable now supports execution using TOML-formatted input.

Lefebvre, Robert A. [Oak Ridge National Laboratory↗

Bacterial Bioleaching and Biorecovery for Biomining Unconventional Rare Earth Element Feedstocks

Bacterial metabolic interactions with rare earth elements (REEs) can be harnessed for biomining unconventional feedstocks like abandoned coal-mine drainage (AMD). Pennsylvania has ~500 AMD passive remediation systems that can precipitate REE rich solids. REEs include yttrium and the lanthanide series that are used in modern energy and technology. Bacteria that metabolically interact with REEs can be used for biomining in an affordable efficient process that does not require hazardous chemical additives. Currently, the microbial metal mechanisms that contribute to REE biorelease and biorecovery are poorly understood. Our work shows acidogenic bacterial isolates (Bacillus mycoides JR07 and Bacillus pseudomycoides KB7) successfully bioleach a mixed REE solution from AMD solids by their organic acid production and biofilm formation. Further, our work shows the potential for bacterial lanthanide-dependent enzymes to recover lanthanides from a mixed REE solution; here we have bacterial isolate Methylobacterium sp. B3 that can recover soluble lanthanum. Whole genome sequencing of Methylobacterium sp. B3 predict lanthanide-dependent methanol dehydrogenase XoxF. Understanding the microbial metabolism and genes involved in the REE release and recovery is crucial to optimize the biomining of AMD solids. Our work addresses the growing need to develop novel REE mining methods from unconventional feedstocks.

biogeochemistry↗

Multi-modal dynamic radiography using short-pulse laser-generated probe beams

Radiography is an important tool for the interrogation of dynamic experiments in the fields of dynamic properties of materials, and in condensed matter, high explosive, and high-energy-density physics. Multi-modal radiography advances the hypothesis that combining the information delivered by multiple radiographic modalities can lead to more constrained (improved) “reconstruction” of the scene than can be obtained from a single probe. We identify four modalities: multi-probe, time sequence, multi-view, and multi-messenger. Multi-probe radiography is a promising candidate for a next-generation dynamic radiographic facility. High-energy X-rays are the most frequently used probe for dynamic radiography, although recent developments show the utility of proton (pRad), electron (eRad), and neutron probe beams. Because each probing species interacts with material in the radiographic scene through quantitatively different mechanisms, each returns independent information about the scene, which can add extra constraints to the reconstruction process. How to conduct detailed, quantitative “co-analysis” of multiple data streams remains an area of active research. Multi-beam, short-pulse, laser-generated probes offer sufficient dose, an appropriate spectrum, and appropriate spatio-temporal resolution to produce high-quality dynamic radiographs. This paper reports on technology development to advance the state of the art of multi-modal/multi-probe radiography and the pursuit of both deterministic and inferential (AI/ML assisted) co-analysis methodologies to produce more constrained reconstructions from multi-modal data.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Using Calibrated Sodium Data for Preliminary Validation of the SRT Code for Advanced Reactors

Various types of non-light water reactors are currently engaged in the U.S. licensing process. Because of inherent differences compared with well-established large light water reactors, appropriate assessment tools are needed. Specifically, source term analysis, which determines environmental dose impacts from potential accident scenarios, is a crucial part of design and licensing. The U.S. Nuclear Regulatory Commission has emphasized the importance of mechanistic source term analysis for advanced reactor deployments. To align with these needs, Argonne National Laboratory has developed the Simplified Radionuclide Transport (SRT) source term analysis code for metal fuel Sodium-cooled Fast Reactors (SFRs) and microreactors. SRT conducts time-dependent radionuclide transport and retention in SFRs for core and ex-core radionuclide source accident sequences. The main objective of SRT is to provide rapid sensitivity and uncertainty analyses, incorporating parametric uncertainties and summarizing probabilistic results. As part of the code validation process, a study focused on the bubble scrubbing module was performed using an experiment recently carried out by the University of Wisconsin-Madison. Based on the analysis, the modeling approach in SRT provides accurate results for small and large aerosols, while slight underprediction of radionuclide aerosol removal are observed for medium sized aerosols. However, the deviation is minor, considering the highly uncertain phenomenon and range of results, and is in the conservative direction. In addition, uncertainty information derived from the experiments is further implemented, reflecting the actual span of parameters, which leads to enhanced agreement with code predictions. The results demonstrate that SRT provides reasonable predictions for the bubble scrubbing process in sodium pool.

Kam, Dong Hoon↗

Comparing Tandem Cell Designs for Electrochemical CO 2 Reduction to Ethylene

Electrochemical carbon dioxide reduction (CO 2 R) is a promising approach for the decentralized production of fuels such as ethylene (C 2 H 4 ). However, the use of Cu, the most efficient metal CO 2 R catalyst for the generation of C 2 H 4 known to date, generally yields a product stream with poor selectivity. In an effort to increase selectivity, the reaction from CO 2 to C 2 H 4 can be broken down into two steps using tandem CO 2 R electrolyzers: formation of CO from CO 2 and subsequent reduction of CO to C 2 H 4 . Here, in this study, we present two novel tandem electrolyzer architectures that closely integrate two cathodes, one for CO generation and one for conversion to C 2 H 4 , while still enabling independent electrical control of the cathodic surfaces. Cathode segmentation in each of these designs also permits the controlled sequencing of mass flow of chemical intermediates in the order of Au to Cu cathode catalysts, in contrast to earlier work relying on uncontrolled, passive diffusion to facilitate the flow of chemical intermediates between catalysts. When comparing the performance of the newly developed electrolyzer cell designs with a dual electrolyzer system, we found that the dual electrolyzer system yields the highest C 2 H 4 faradaic efficiencies (FEs) of 31% and C 2 H 4 concentrations (∼8 mol %). However, a single Cu-containing electrolyzer outperformed all three tandem systems in terms of C 2 H 4 FE (34%). Our findings, enabled by independent control of the two tandem cathode surfaces, indicate that tandem CO 2 R systems need to be evaluated carefully by testing them at various relevant current densities.

C2H4↗

Automated Signal Timing Plan Reconstruction Using High-Resolution Event-Based Controller Data for Digital Twins

Transportation digital twins are essential tools for evaluating emerging technologies such as connected and automated vehicles, adaptive traffic signal control, and mobility optimization strategies. Realistic digital twins require accurate emulation of real-world signal controllers and detailed signal timing plans. However, signal timing plans are often unavailable or difficult to access, forcing researchers and modelers to rely on assumed fixed timings or halt their analysis. To overcome this challenge, we present a method that directly estimates signal timing plan parameters using high-resolution, event-based data from traffic signal controllers. The proposed method extracts key parameters, including cycle length, offset, phase sequence, coordinated phases, phase-specific minimum and maximum green durations, vehicle extensions, and splits under coordination. A rule-based deterministic signal timing reconstruction algorithm based on traffic signal operation rules, such as those outlined in the Signal Timing Manual, is developed and validated. We evaluate this method, which uses high-resolution controller event logs and verified signal timing plans, on 94 signalized intersections in Nashville, Tennessee, demonstrating their ability to generate accurate, simulation-ready signal timing plans for tools such as SUMO and Vissim.

Saroj, Abhilasha [ORNL] (ORCID:0000000191178063)↗

Multi-scale Simulation, Calibration, and Optimization of Calcium Carbonate Precipitation in Microbial Communities

Ensuring the efficient engineering of microbially induced calcium carbonate precipitation (MICP) is crucial for a variety of environmental and civil engineering applications, such as soil stabilization and carbon sequestration. Addressing this need, we present a comprehensive multi-scale workflow that begins with the isolation of calcium carbonate-producing microbes from soil samples, followed by metagenomic sequencing and metabolic reconstruction. We then characterize microbial growth phenotypes under diverse nutrient conditions, compare observed growth with metabolic model predictions, and apply the Consistent Reproduction of Phenotype (CROP) algorithm to refine these models. Furthermore, we analyze metabolite consumption and production, and develop a consumer-resource model that is calibrated using time-series measurements of growth rates, pH levels, and calcium carbonate precipitation. The primary benefit of our approach lies in its ability to predict and control MICP outcomes, facilitated by a Bayesian methodology that incorporates priors on initial conditions and parameters. This allows us to compute posteriors by integrating experimental data, and to solve a risk optimization problem under uncertainty to identify nutrient conditions that maximize calcium carbonate production. In contrast to non-Bayesian methods, which fail to quantify uncertainty accurately, our approach provides a more reliable pathway to optimizing nutrient conditions, enhancing the likelihood of achieving desired MICP outcomes. This positions our method as a superior alternative in the quest to improve MICP through engineered microbial consortia.

54 ENVIRONMENTAL SCIENCES↗

Unified Universal Control and Coordination of Inverter-Based Resources, and Validation for a PV + Battery Hybrid Plant

As renewable energy deployment grows, hybrid power plants (HPPs) combining photovoltaic (PV) and battery systems must evolve to offer both energy and grid stability services. These systems typically include a mix of grid-following (GFL) and grid-forming (GFM) inverters, presenting unique coordination and control challenges. This Department of Energy–funded project developed and validated a Unified Universal Control and Coordination (UUCC) framework for such PV + battery hybrid plants, enabling seamless and stable operation, including ultrafast black start, autonomous synchronization, and robust frequency and voltage regulation, under different grid conditions. The project significantly advanced the understanding of inverter-based resource (IBR) control by developing and validating three complementary system-level approaches for hybrid GFL/GFM operation: 1. A combined Virtual Resistance (VR)-based GFL and Virtual Oscillator Control (VOC)-based GFM method, where each inverter type is governed by a specialized control strategy. Together, these achieve stable, fast-response coordination, eliminating inrush current and enabling smooth black start and grid synchronization across a wide range of grid strengths. 2. A Deadbeat-based UUCC strategy, which uses discrete-time, switching-cycle-level control for both GFL and GFM inverters. This approach replaces traditional PI/PLL control with a control parameter-free, high-bandwidth framework that supports stable LVRT and instantaneous synchronization under all conditions. 3. A benchmark comparison with Siemens’ commercial GFM microgrid controller, which provided a fast baseline platform. The commercial approach decoupled v & f control was implemented on a commercial microgrid controller.The baseline commercial benchmark helped highlight superior transient response and black start performance offered by the deadbeat and VOC approaches. These technical contributions offer substantial improvements over conventional inverter control schemes, which often rely on slow phase-locked loop (PLL)-based synchronization, require careful control parameters tuning, and prone to unstable in weak grids with GFL inverters and in stiff grid with GFM inverters therefore challenging for hybrid GFL+GFM under all grid conditions. The deadbeat-based UUCC framework enables simpler, faster, and more robust operation of hybrid IBR systems using wide-bandgap (WBG) devices such as SiC power semiconductors. The rapid expansion of hybrid distributed energy resources (DERs), including residential and commercial PV-BESS installations such as Tesla Powerwall, PV with vehicle-to-grid (V2G) capability, and other integrated configurations, presents complex operational challenges for medium-voltage radial distribution feeders. These networks are subject to frequent disturbances such as faults, switching operations, rapid reclosing sequences, and feeder reconfigurations, all of which introduce dynamic stress on IBRs. In addition, planned feeder segmentation and deliberate islanding for resilience will require DERs that can autonomously perform blackstart, establish voltage and frequency references, and resynchronize with the main grid. The advanced deadbeat-based UUCC control and blackstart functionalities developed in this project directly address these requirements, enabling decentralized and autonomous operation of inverter-dominated DERs in distribution systems under a wide range of fault and reconfiguration scenarios. From a public benefit perspective, these innovations enable more reliable and cost-effective integration of renewable energy into distribution networks. The ability to autonomously black start and stabilize grids under varying grid conditions support accelerates recovery from outages and support decentralized resilient energy systems. By reducing system complexity and improving performance, this project lays critical groundwork for future inverter-dominated power grids that are clean, reliable, and accessible to all.

14 SOLAR ENERGY↗

Domestication of Algae for Increasing Biomass Productivity

Microalgae cultivation processes have been developed for the production of a variety of bioproducts, however currently only a few species are used in commercial applications. Their domestication, that is strain improvements, is still in its infancy, with major advances required, specifically to maximize biomass productivity a limiting factor in microalgae production. This requires a deep understanding of algal biology, in particular to develop superior strains without the need of genetic technologies that would require lengthy regulatory permits, and often limit consumer acceptance. Adaptive Laboratory Evolution techniques, alone or in conjunction with sexual recombination, can allow for rapid develop of improved strains and their industrial production. Light harvesting antenna reduction has been a major approach to achieve increased photon utilization efficiency by cultures operating under full sunlight conditions due to higher light saturation levels, allowing for higher productivities under outdoor conditions. Decades of research yielded some promising results under controlled conditions with a few specific mutant strains. However, these failed to achieve the anticipated higher productivities in actual algal mass cultures, in part due to the inability of single mutations to overcome photoinhibition, reactive oxygen species, and other pleiotropic impacts on the complex metabolic processes of photosynthesis. Higher productivity strains will require multiple genetic improvements. We report on recent Adaptive Laboratory Evolution with the green alga Scenedesmus obliquus resulting in higher biomass productivity in open pond cultivation. Coupling our approach with sexual recombination and genome sequencing provides a path to algal domestication suitable for large-scale, low-cost biomass production.

09 BIOMASS FUELS↗

Machine learning prediction of enzyme optimum pH

The relationship between pH and enzyme catalytic activity, especially the optimal pH (pH opt ) at which enzymes function, is critical for biotechnological applications. Hence, computational methods to predict pH opt will enhance enzyme discovery and design by facilitating accurate identification of enzymes that function optimally at specific pH levels, and by elucidating sequence-function relationships. Here, in this study, we proposed and evaluated various machine learning methods for predicting pH opt , conducting extensive hyperparameter optimization and training over 11,000 model instances. Our results demonstrate that models utilizing language model embeddings markedly outperform other methods in predicting pHopt. We present EpHod, the best-performing model, to predict pHopt, making it publicly available to researchers. From sequence data, EpHod directly learns structural and biophysical features that relate to pH opt , including proximity of residues to the catalytic centre and the accessibility of solvent molecules. Overall, EpHod presents a promising advancement in pH opt prediction and will potentially speed up the development of enzyme technologies.

97 MATHEMATICS AND COMPUTING↗

Inventory of Composable Elements (ICE) v6.0.0

The Inventory of Composable Elements (ICE) is an open source registry software platform for managing information about biological parts. It is capable of recording information about plasmids, microbial host strains and seeds, as well as DNA parts. Includes features such as DNA sequence visualization, editing and annotation, auto-aligning sequencing trace files against reference templates, SBOL XML/RDF support, and web-of-registries functionality. The web of registries functionality provides strong support for distributed interconnected use and enables sharing and transfer of biological parts across various independent ICE instances. ICE adopts modern software development principles, leveraging component-base frameworks, offering a REST API for convenient third-party integration and emphasizing scalability, security, and service integrations for dynamic content availability. The source code is hosted at https://github.com/JBEI/ice. A public instance is available at public-registry.jbei.org, where users can try out features, upload parts or simply use it for their projects.

Plahar, Hector↗

Human RNome Project draft human RNome sequence of GM12878, B-cell line, obtained by mass-spectrometry sequencing, long-read sequencing and short-read sequencing.

Here we report the first draft of the human RNome sequence, a reference map of RNA chemical modifications in a human B-cell line. RNA carries a diverse repertoire of chemical modifications that regulate gene expression, cellular function, and responses to physiological and pathological cues. Yet, unlike the genome, no reference map of RNA modifications is available for any human cell. To generate this resource, the Human RNome Project Consortium analyzed a shared RNA preparation from the well-characterized GM12878 B-cell line using short-read sequencing, long-read direct RNA sequencing, and mass spectrometry, generating more than 7.1 billion sequencing reads spanning approximately 1.2 trillion nucleotides. The resulting maps of the human RNome reveal that RNA modifications are organized according to function, transcript architecture, and cellular identity. Modifications concentrate at functional centers of ribosomal and transfer RNAs, follow the canonical topology of N6-methyladenosine in coding transcripts, and form coordinated hotspots in immune regulatory genes. This first reference human RNome provides a foundation for understanding how RNA chemistry shapes cellular identity, human disease, and the development of RNA-based therapeutics.

59 BASIC BIOLOGICAL SCIENCES↗

HERO CarbonSAFE Phase 2 Project in the Columbia River Basalt Group

The Hermiston, Oregon Basalt CarbonSAFE Phase II project (HERO CarbonSAFE) seeks to accelerate the deployment of commercial carbon dioxide (CO2) storage projects in basaltic rocks. Basalt CO2 storage has several advantages to conventional saline storage reservoirs including 1. The potential for rapid mineralization of CO2, 2. Associated decreases in pressure and CO2 migration risks, 3. Reduced long-term monitoring requirements with respect to plume tracking, 4. Widespread geographic distribution and, 5. Large storage potential due to thickness, porosity, and CO2 interactions with basalt. And for locations such as the Pacific Northwest, Hawaii, Iceland, India and Japan, basalts may offer the only economically feasible option for local CO2 storage. However, there are limited field-scale assessments of CO2 storage in basalt, and current carbon capture utilization and storage (CCUS) permitting and regulatory frameworks were developed for conventional saline reservoirs. HERO CarbonSAFE is designed to address research gaps and uncertainties associated with basalt storage. Specifically, the project will assess the feasibility of CO2 injection in the deep layered basalts, long-term storage (mineralization), practical approaches for large-scale implementation (50+ million metric tons of CO2 over 30 years), lithology-specific risks, and the technoeconomic potential for CO2 storage in basalts. The HERO CarbonSAFE project will assess feasibility of developing a commercial-scale (50+ million metric tons of CO2) geological storage complex within the Columbia River Basalt Group (CRBG), a layered continental flood basalt complex that underlies Calpine’s natural gas-fired Hermiston Power Project (HPP) in Hermiston, OR (Figure 1). Under this 2-year CarbonSAFE Phase II project, the HERO team will conduct a data acquisition campaign that includes drilling a stratigraphic well to a total depth of ~1,500 m into the thick layered basalts proximal to HPP. A comprehensive well logging and hydrologic testing program will be augmented with new core collected from flow zones and sealing units, and comprehensive laboratory testing to help refine the kinetic rates of mineralization. The newly acquired information will be integrated with existing data from regional wells to correlate basalt injection zone properties to develop storage hub/commercial-scale models. Using these models, the project team will evaluate injection scenarios to define the technical and economic potential for storing a minimum of 50 million metric tons of CO2 over a 30-year period, along with a robust sensitivity analysis on key parameters governing reservoir viability for sustainable injection over a commercial project lifetime. Specific technical objectives of HERO are: (1) assessing the reservoir response of a series of stacked layered reservoir flowtop sequences occurring in this area of the CRBG to commercial-scale injection volumes; (2) extending prior efforts by the project team to characterize the deep layered basalts encountered in regional studies, to leverage prior investments by U.S. Department of Energy’s (DOE) Carbon Storage program; (3) leveraging DOE’s mineralization characterization efforts to advance model parametrization for commercial scale injection of CO2 in basalts; (4) conducting risk assessments associated with scaling up to commercial storage hub injection goals, while validating DOE’s National Risk Assessment Partnership (NRAP) tools, to identify potential constraints that would prevent the CRBG from serving as a commercial-scale storage complex; (5) developing mitigation plans to address identified risks; (6) developing a commercial-scale injection and monitoring, verification and accounting (MVA) strategy; (7) utilizing computational models to define and minimize, if possible, the Area of Review (AoR) under Class VI regulations; and (8) developing a robust CO2 management strategy for CRBG that also considers a regional source/sink approach that is responsive to stakeholder needs and industrial demand. Specific institutional objectives are: (1) identifying and developing plans to mitigate the nontechnical challenges associated with the build-out of a commercial-scale storage complex within the CRBG with integrated CO2 sources; (2) implementing the community outreach plan; (3) conducting regulatory research, including a survey of issues related to pore space ownership, MVA and long-term assurance of mineralization-based storage, to support an eventual application for a UIC Class VI permit; (4) advancing the project’s plan for CO2 liability management; and (5) continuing to refine and update the project’s economic model. The final objective is the preparation of a comprehensive Site Characterization Plan that draws upon the technical and institutional feasibility assessments to prepare the project for future commercialization efforts.

58 GEOSCIENCES↗

Spatial top-down proteomics for the functional characterization of human kidney

Background: The Human Proteome Project has credibly detected nearly 93% of the roughly 20,000 proteins which are predicted by the human genome. However, the proteome is enigmatic, where alterations in amino acid sequences from polymorphisms and alternative splicing, errors in translation, and post-translational modifications result in a proteome depth estimated at several million unique proteoforms. Recently mass spectrometry has been demonstrated in several landmark efforts mapping the human proteoform landscape in bulk analyses. Herein, we developed an integrated workflow for characterizing proteoforms from human tissue in a spatially resolved manner by coupling laser capture microdissection, nanoliter-scale sample preparation, and mass spectrometry imaging. Results: Using healthy human kidney sections as the case study, we focused our analyses on the major functional tissue units including glomeruli, tubules, and medullary rays. After laser capture microdissection, these isolated functional tissue units were processed with microPOTS (microdroplet processing in one-pot for trace samples) for sensitive top-down proteomics measurement. This provided a quantitative database of 616 proteoforms that was further leveraged as a library for mass spectrometry imaging with near-cellular spatial resolution over the entire section. Notably, several mitochondrial proteoforms were found to be differentially abundant between glomeruli and convoluted tubules, and further spatial contextualization was provided by mass spectrometry imaging confirming unique differences identified by microPOTS, and further expanding the field-of-view for unique distributions such as enhanced abundance of a truncated form (1-74) of ubiquitin within cortical regions. Conclusions: We developed an integrated workflow to directly identify proteoforms and reveal their spatial distributions. Where of the 20 differentially abundant proteoforms identified as discriminate between tubules and glomeruli by microPOTS, the vast majority of tubular proteoforms were of mitochondrial origin (8 of 10) where discriminate proteoforms in glomeruli were primarily hemoglobin subunits (9 of 10). These trends were also identified within ion images demonstrating spatially resolved characterization of proteoforms that has the potential to reshape discovery-based proteomics because the proteoforms are the ultimate effector of cellular functions. Applications of this technology have the potential to unravel etiology and pathophysiology of disease states, informing on biologically active proteoforms, which remodel the proteomic landscape in chronic and acute disorders.

59 BASIC BIOLOGICAL SCIENCES↗

Beyond microbial abundance: metadata integration enhances disease prediction in human microbiome studies

Multiple studies have highlighted the interaction of the human microbiome with physiological systems such as the gut, immune, liver, and skin, via key axes. Advances in sequencing technologies and high-performance computing have enabled the analysis of large-scale metagenomic data, facilitating the use of machine learning to predict disease likelihood from microbiome profiles. However, challenges such as compositionality, high dimensionality, sparsity, and limited sample sizes have hindered the development of actionable models. One strategy to improve these models is by incorporating key metadata from both the human host and sample collection/processing protocols. This remains challenging due to sparsity and inconsistency in metadata annotation and availability. In this paper, we introduce a machine learning-based pipeline for predicting human disease states by integrating host and protocol metadata with microbiome abundance profiles from 68 different studies, processed through a consistent pipeline. Our findings indicate that metadata can enhance machine learning predictions, particularly at higher taxonomic ranks like Kingdom and Phylum, though this effect diminishes at lower ranks. Our study leverages a large collection of microbiome datasets comprising 11,208 samples, therefore enhancing the robustness and statistical confidence of our findings. This work is a critical step toward utilizing microbiome and metadata for predicting diseases such as gastrointestinal infections, diabetes, cancer, and neurological disorders.

Mathematics and Computing↗