Search NASA⌕ Search

SEARCH · Search NASA

Results for “Object Store”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Integration of RNTuple in ATLAS Athena

After using ROOT’s TTree I/O subsystem for over two decades and storing more than an exabyte of compressed High Energy Physics (HEP) data, advances in technology have motivated a complete redesign, RNTuple, which breaks backward-compatibility to take better advantage of these storage options. The RNTuple I/O subsystem has been designed to address performance bottlenecks and other shortcomings of TTree. Specifically, RNTuple comes with an updated, more compact binary data format that can be stored both in ROOT files and natively in object stores. It is designed for modern storage hardware (e.g. high-throughput low-latency NVMe SSDs), and provides robust and easy to use interfaces. The binary format of RNTuple is scheduled to become production grade in 2024, and recently has become mature enough to start exploring the integration into software used by HEP experiments. In this contribution, we discuss the developments to support the features as required by the ATLAS analysis Event Data Model (EDM) in RNTuple, which will enable its integration into the Athena software framework. With these developments in place, we evaluate the performance of the current most recent versions of RNTuple-based ATLAS data sets and compare this to that of TTree.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

DoCeph: DPU-Offloaded Messaging in Ceph for Reduced Host CPU Utilization

Ceph is a widely used distributed object store, but its messenger layer imposes substantial CPU overhead on the host. To address this limitation, we propose DoCeph, a DPU-offloaded storage architecture for Ceph that disaggregates the system by offloading the communication-intensive messaging component to the DPU while retaining the storage backend on the host. The DPU efficiently manages communication, using lightweight RPC for metadata operations and DMA for data transfer. Moreover, DoCeph introduces a pipelining technique that overlaps data transmission with buffer preparation, mitigating hardware-imposed transfer size limitations. We implemented DoCeph on a Ceph cluster with NVIDIA BlueField-3 DPUs. Evaluation results indicate that DoCeph cuts host CPU usage by up to 92% while sustaining stable throughput and providing larger performance benefits for object writes over 1 MB.

Park, Kuri [Sogang University]↗

ROOT RNTuple and EOS: The Next Generation of Event Data I/O

For several years, the ROOT team is developing the new RNTuple I/O subsystem in preparation of the next generation of collider experiments. Both HL-LHC and DUNE are expected to start data taking by the end of this decade. They pose unprecedented challenges to event data I/O in terms of data rates, event sizes, and event complexity. At the same time, the I/O landscape is becoming more diverse. HPC cluster file systems and object stores, NVMe disk cache layers in analysis facilities, and S3 storage on cloud resources are mixing with traditional XRootD-managed spinning disk pools.The ROOT team will finalize a first production version of the RNTuple binary format by the end of 2024. After this point, ROOT will provide backward compatibility for RNTuple data. This contribution provides an overview of the RNTuple feature set, the related R&D activities and the long-term vision for RNTuple. We report on performance, interface design, tooling, robustness, integration with experiment frameworks, and validation results, as well as recent R&D on parallel reading and writing and exploitation of modern hardware and storage systems. We will give an outlook on possible future features after a first production release.Collaboratively, the IT and EP departments at CERN have launched a formal project within the Research and Computing sector to evaluate the novel data format for physics analysis data utilized in LHC experiments and other fields. This part of the project focuses on validating the scalability of the EOS storage backend during the transition from the over 25 years old TTree production format to the newly developed RNTuple format, using both replicated and erasure-coded storage profiles.

Blomer, Jakob [CERN]↗

BULKI-Store v0.3.2

BULKI-Store is a distributed object storage system optimized for high-performance computing environments. Built with a Rust core and Python bindings, it efficiently manages scientific and machine learning datasets across HPC clusters. The system employs a client-server architecture with MPI integration, enabling seamless scaling on supercomputers like Perlmutter. BULKI-Store's object-oriented approach provides intuitive data organization with rich metadata support, contrasting with traditional file-based solutions. Key optimizations include selective checkpoint loading, unified checkpoint files, and object chunking for large data transfers. For machine learning workloads, BULKI-Store offers advantages through fine-grained access patterns, dynamic data sharing between training instances, and reduced memory pressure. Memory management features include strategic Python GC calls, minimized data copies, and batch processing capabilities. The system leverages Rayon's thread pool for asynchronous data prefetching and supports multiple CPU architectures (ARM64, x86, AMD, RISC-V). By combining performance optimizations with developer-friendly APIs, BULKI-Store addresses the complex data management challenges of modern HPC applications while maintaining compatibility across heterogeneous computing environments.

Zhang, Wei [Lawrence Berkeley National Laboratory ↗

Adoption of ROOT RNTuple for the next main event data storage technology in the ATLAS production framework Athena

Since the start of LHC in 2008, the ATLAS experiment has relied on ROOT to provide storage technology for all its processed event data. Internally, ROOT files are organized around TTree structures that are capable of storing complex C++ objects. The capabilities of TTrees developed over the years and are now offering support for advanced concepts like polymorphism, schema evolution and user defined collections and ATLAS makes use of these features to handle its EDM. But some original TTrees concepts, like the POSIX file model and sequential writing, remain unchanged since the beginning and could be an obstacle to achieving the performance required for High Luminosity LHC. With the HL-LHC performance goals in mind, the ROOT project developed a new storage format - the RNTuple. RNTuple, with its accompanying user API, is now in the final development stage and is planned to be production-ready at the end of 2024. Soon after that, the TTree will become a legacy format. ATLAS intends to have its main Event processing framework Athena ready to use RNTuple in the production environment as early as possible. The work on adopting RNTuple as another ROOT storage technology in Athena started already in 2021 and is now nearly complete. Although the initial goal was to focus on derived-AOD products (PHYS and PHYSLITE), with a little added effort all ATLAS data products: RDO, HITS, ESD, AOD and DAOD can be now stored in RNTuple format and transparently read back. In this paper we will describe the current state of RNTuple adoption in the Athena framework and explain the ATLAS EDM requirements that had to be met on the ROOT side to successfully integrate both environments. We will demonstrate the ability to run standard ATLAS production workflows, based on RNTuple as the Event data storage technology, and point out key advantages of the new format.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Search for HH → bbτ⁺τ⁻ Using Run 3 Scouting Data Analyze b-tagging and tau-tagging Performance with Unified Particle Transformer

B-tagging and tau-tagging performances play an important role in the search for the rare event HH → bbτ⁺τ⁻. A transformer-based neural network, Unified Particle Transformer, is applied for both tagging tasks, and Run 3 proton–proton collision scouting data at center-of-mass energy of 13.6 TeV is used. The scouting data stream accepts events at a much higher rate compared to traditional triggers, but stores only the objects reconstructed in the trigger, no low-level detector information. Therefore, existing taggers trained for the offline event reconstruction cannot be used. Analysis of the SoftMax plots, ROC/AUC curves, confusion matrix, accuracy and losses are used to evaluate model performance. Specifically, the tagging efficiency of the signal and misidentification probability across multiple background processes are compared for varying working points. Different training samples with distinct distributions of jet flavors are utilized and related model performances are analyzed. Interpretability methods, such as Integrated Gradients, may further be applied to study the input features’ influence on the model’s decisions, providing insights into potential improvements.

Chen, Blair [Purdue U., West Lafayette; Fermilab]↗

What exactly does Bekenstein bound?

The Bekenstein bound posits a maximum entropy for matter with finite energy confined to a spatial region. It is often interpreted as a fundamental limit on the information that can be stored by physical objects. In this work, we test this interpretation by asking whether the Bekenstein bound imposes constraints on a channel's communication capacity, a context in which information can be given a mathematically rigorous and operationally meaningful definition. We study specifically the Unruh channel that describes a stationary Alice exciting different species of free scalar fields to send information to an accelerating Bob, who is confined to a Rindler wedge and exposed to the noise of Unruh radiation. We show that the classical and quantum capacities of the Unruh channel obey the Bekenstein bound that pertains to the decoder Bob. In contrast, even at high temperatures, the Unruh channel can transmit a significant number of zero-bits , which are quantum communication resources that can be used for quantum identification and many other primitive protocols. Therefore, unlike classical bits and qubits, zero-bits and their associated information processing capability are generally not constrained by the Bekenstein bound. However, we further show that when both the encoder and the decoder are restricted, the Bekenstein bound does constrain the channel capacities, including the zero-bit capacity.

Hayden, Patrick [Stanford University, CA (United S↗

ForceFinder

SAND2025-11750O ForceFinder extends the Structural Dynamics Python Libraries (SDynPy) with comprehensive tools for inverse source estimation (ISE) tasks via frequency response function (FRF) matrix inversion. The software is designed for transfer path analysis and multiple-input/multiple-output (MIMO) vibration control problems. It allows users to estimate sources through various algorithms, from the basic Moore-Penrose pseudo-inverse to statistical learning methods such as Tikhonov regularization via an L-curve and elastic net regularization via an information criterion. ForceFinder uses an object-oriented framework, where all components of the ISE problem—such as FRFs, responses, and transformations—are stored in a "SourcePathReceiver" object. This software can be applied to any noise and vibration problem. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Carter, Steven [Sandia National Lab. (SNL-CA), Liv↗

Accelerated CO2 Storage Optimization Using Multi-Resolution Fourier Neural Operator at the Illinois Basin Decatur Project (IBDP)

This paper presents a deep learning-based approach for optimizing CO2 injection in carbon capture and storage (CCS) operations. We developed a multi-resolution machine learning model to significantly reduce data generation costs. Utilizing this proxy model, we implemented a multi-objective genetic algorithm to optimize well control during the CO2 injection process. The proposed approach was applied to the Illinois Basin Decatur Project (IBDP), successfully optimizing the CO2 injection schedule based on three key objectives: maximizing the amount of CO2 stored, maximizing sweep efficiency, and minimizing pressure increase. The use of the proxy model accelerated the optimization workflow by two orders of magnitude, while the cost of data generation for the proxy model was reduced by 90% by utilizing a coarse-scale model.

accelerated CO2 storage optimization↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗

A randomized sketching trust-region secant method for low-memory dynamic optimization

The numerical solution of dynamic optimization problems is often limited by the memory required to store the state trajectory, which is used to evaluate the objective function and its derivatives. Recently, [R. Muthukumar et al., SIAM Journal on Optimization 31(2), pp. 1242–1275 (2021)] introduced a trust-region method for dynamic optimization that employs randomized sketching to compress the state trajectory, resulting in inexact derivative computations. By adaptively learning the sketch rank, the trust-region algorithm achieves rigorous convergence guarantees. Here, we extend this approach to use secant Hessian approximations. Due to the randomness introduced by the sketch, the traditional secant update formulae can produce poor Hessian approximations. In particular, the difference of two gradients, computed from two different sketches, may be inconsistent. To overcome this, we employ a sketched approximation of the Hessian application, in lieu of computing the gradient difference. We numerically demonstrate the improved stability of this approach on an example from PDE-constrained optimization.

dynamic optimization↗

Spatialyze: A Geospatial Video Analytics System with Spatial-Aware Optimizations

Videos that are shot using commodity hardware such as phones and surveillance cameras record various metadata such as time and location. We encounter suchgeospatial videoson a daily basis and such videos have been growing in volume significantly. Yet, we do not have data management systems that allow users to interact with such data effectively. In this paper, we describe Spatialyze, a new framework for end-to-end querying of geospatial videos. Spatialyze comes with a domain-specific language where users can construct geospatial video analytic workflows using a 3-step, declarative,build-filter-observeparadigm. Internally, Spatialyze leverages the declarative nature of such workflows, the temporal-spatial metadata stored with videos, and physical behavior of real-world objects to optimize the execution of workflows. Our results using real-world videos and workflows show that Spatialyze can reduce execution time by up to 5.3×, while maintaining up to 97.1% accuracy compared to unoptimized execution.

Computer Science↗

CHESS 2025: Waveform LiDAR data from NEON AOP surveys

This dataset provides Level 1 (L1) full-waveform light detection and ranging (LiDAR) data collected for the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS). These data were acquired to enable characterization of vegetation structure and other three-dimensional features of the land surface, and to evaluate structural changes that may have occurred between a prior LiDAR acquisition in 2018 and the 2025 overflight. Waveform LiDAR data can provide more detailed information about objects on the ground than discrete point clouds typically do, and they are often used for granular target segmentation and characterization of subcanopy vegetation. The data were acquired over three study domains in the Upper Gunnison river basin: the upper East River watershed (CRBU); Almont Triangle and Taylor Canyon (ALMO); and Upper Taylor River watershed (UPTA) between 2025-06-13 and 2025-07-15. LiDAR data were acquired using the Optech Galaxy Prime Airborne LiDAR Terrain Mapper onboard the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP). These are the primary waveform LiDAR data delivered by NEON and are provided per flightline in compressed Pulsewaves format, an open-source binary file standard. A Pulsewaves object comprises a two files: a pulse (.pls) file, which stores the geographic origin, outgoing vector, and metadata for every laser pulse emitted by the scanner, and a wave file (.wvs), which stores the sequential amplitude samples of the outgoing pulse and the returning signals. The files are published here in their compressed forms (.plz, .wvz). All waveform data were processed following the theoretical workflow described in the NEON L0-to-L1 Waveform LiDAR Algorithm Theoretical Basis Document (Krause and Goulden 2022a); however, the Pulsewaves output format differs from a legacy format described in that document. Waveform amplitude samples are recorded at 1 nanosecond intervals. All coordinates are provided in meters. Horizontal coordinates are referenced in Universal Transverse Mercator (UTM) zone 13N and the World Geodetic System (WGS) 1984 ensemble datum. Elevations are referenced to Geoid12A. Waveform data for the UPTA survey area were collected without incident and the published records are complete. However, both the ALMO and CRBU collections experienced issues that resulted in incomplete data for those areas. On collection day 2018-06-16 a hardware failure caused the waveform digitizer to lose data from the eastern edge of the ALMO site (Figure 22). The waveform data for flightlines 2–20 could not be extracted from the digitizer, and the data proved unrecoverable. As a result, a portion of the site does not have coverage with waveform data. Although no hardware failure was observed during collection over the CRBU area, final waveform files generated by vendor software contained only ~25% of the expected number of return pulses. After discovery, NEON initiated troubleshooting with the vendor. The root cause of the data ablation had not been identified at the time of publication. Additional data will be published in an update to this package if further recovery proves successful. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgement: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

DE-FE0023919 Phase 5 Scientific/Technical Report

Phase 5 of the Deepwater Methane Hydrate Characterization and Scientific Assessment research project (DOE Award No. DE-FE0023919) occurred from Oct. 1, 2020 to Nov. 15, 2023. Throughout Phase 5, UT performed all aspects of project management and planning according to the award, project management plan, and statement of project objectives (Task 1). UT maintained and augmented the capability to transport, store, manipulate and analyze pressure cores (Task 13). UT’s hydrate core effective stress chamber can now run tests at effective stresses up to 20 MPa. A benchmark study was conducted and confirmed that the K0 permeameter accurately estimates geomechanical and petrophysical properties of geomaterials under uniaxial strain conditions. UT continued to analyze remaining UT-GOM2-1 pressure cores from GC955 (Task 10).

03 NATURAL GAS↗

Elucidating the Link Between Alkali Metal Ions and Reaction-Transport Mechanisms in Cathode Electrodes for Alkali-ion Batteries

Our long-term goal is to improve the reliability of electrode materials and their ability to transport and store various metal ions for electrochemical energy storage applications. The main objective of this work was to investigate the intrinsic relationship between the role of alkali metal ions and electrochemically driven mechanical stability and kinetic properties of battery materials. The overall question was “What is the role of alkali metal ions on the electrochemical and mechanical behavior of cathode electrodes? Our guiding hypothesis was that intercalation of larger alkali metal ions (Na and K) inevitably alters the coupled transport-reaction processes during battery operation in organic electrolytes, leading to more intensive chemo-mechanical instabilities in cathode electrodes, resulting in rapid capacity fade. To validate the hypothesis, we experimentally characterized the reaction-transport processes and governing forces driving the instability of electrode materials in different alkali metal-ion environments. The project had three main tasks. The first one was to investigate intercalation-induced strains and associated stress generation, and their impact on structural deformations in composite cathode electrodes. The second task focused on identifying potential-dependent dynamic changes in the electrode-electrolyte interface in alkali metal ion batteries. The last task was focused on determining how larger alkali metal ions with slower diffusivity affect the transport-mechanics coupling at faster scan rates, compared to smaller ions with faster diffusivity in electrodes. We shortly provided the outcome of each task in the accomplishment section. This project produced 10 peer-reviewed publications (9 research papers and one review manuscript) and supported two Ph.D. students, who graduated from Oklahoma State University.

25 ENERGY STORAGE↗

MRCI Subtask 2.4/2.5: Regional/Subregional Analysis and Risk Assessment Final Technical Summary Report

The objective of the Midwest Regional Carbon Initiative (MRCI) project is to implement a collaborative Regional Initiative (RI) to accelerate the deployment of carbon capture, and storage (CCS) in the Midwest-Northeastern quadrant of the United States. This report is a Technical Summary report describing work performed on Tasks 2.4 (Conducting Regional/Subregional Analysis) and 2.5 (Assessing and Managing Risk) during the MRCI project. In Task 2.4, detailed numerical reservoir simulation models were developed for selected carbon storage (CS) systems identified under Task 2.1 (Battelle, 2021a) in the MRCI study area. The objective of Task 2.4 is to demonstrate a dynamic modeling methodology for evaluating the suitability of the (selected) CS systems in the MRCI region for hosting a commercial-scale storage project. In this study, CO2 injectivity was evaluated for different CS systems with an annual injection rate of 1 million metric tonnes (MMT) of CO2 considered as the minimum requirement for a commercial scale project. The objective of Task 2.5 is to assess key risks associated with storing CO2 in the different CS systems across the MRCI region and to demonstrate a method(s) for assessing these risks that may be used by developers of future CO2 storage projects in the region. The risk analysis was limited to evaluating two types of leakage risks (i.e., wellbore leakage, flow across unfractured caprock) at three modeled sites considered in Task 2.4.

MRCI,Report,Summary,Technical Challenges,dynamic m↗

ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis

Objective: Electronic health record (EHR) systems contain a wealth of clinical data stored as both codified data and free-text narrative notes (NLP). The complexity of EHR presents challenges in feature representation, information extraction, and uncertainty quantification. Here, to address these challenges, we proposed an efficient Aggregated naRrative Codified Health (ARCH) records analysis to generate a large-scale knowledge graph (KG) for a comprehensive set of EHR codified and narrative features. Methods: Using data from 12.5 million Veterans Affairs patients, ARCH first derives embedding vectors and generates similarities along with associated p-values to measure the strength of relatedness between clinical features with statistical certainty quantification. Next, ARCH performs a sparse embedding regression to remove indirect linkage between features to build a sparse KG. Finally, ARCH was validated on various clinical tasks, including detecting known relationships between entity pairs, predicting drug side effects, disease phenotyping, as well as sub-typing Alzheimer’s disease patients. Results: ARCH produces high-quality clinical embeddings and KG for over 60,000 codified and narrative EHR concepts. The KG and embeddings are visualized in the R-shiny powered web-API.3 ARCH achieved high accuracy in detecting EHR concept relationships, with AUCs of 0.926 (codified) and 0.861 (NLP) for similar EHR concepts, and 0.810 (codified) and 0.843 (NLP) for related pairs. It detected drug side effects with a 0.723 AUC, which improved to 0.826 after fine-tuning. Using both codified and NLP features, the detection power increased significantly. Compared to other methods, ARCH has superior accuracy and enhances weakly supervised phenotyping algorithms’ performance. Notably, it successfully categorized Alzheimer’s patients into two subgroups with varying mortality rates. Conclusion: The proposed ARCH algorithm generates large-scale high-quality semantic representations and knowledge graph for both codified and NLP EHR features, useful for a wide range of predictive modeling tasks.

Electronic health records↗

Fundamental Studies of the Vibrational, Electronic, and Photophysical Properties of Tetrapyrrolic Architectures

The ability to capture and utilize light in the near-ultraviolet (NUV), visible and near-infrared (NIR-I and NIR-II) spectral regions (i.e., 320–400, 400–700, 700–1000, 1000–1700 nm) is essential for any solar-energy conversion scheme. Nature employs chlorophylls and bacteriochlorophylls in light-harvesting architectures to absorb light in the blue and red/NIR regions. Accessory pigments (carotenoids, bilins) augment absorption of the (bacterio)chlorophylls in the green region. The harvested energy is funneled to a reaction center protein, where charge separation occurs. Subsequent migration of the electron and the hole stabilizes and stores the energy from light via redox chemistry. The long-term objective of the Bocian/Holten&Kirmaier/Lindsey research program under this DOE grant has been to develop tetrapyrrole-based molecular architectures that absorb sunlight, funnel energy and separate charge with high efficiency. Integral to the program has been iterative cycles of design, synthesis and characterization that provided deep insights into the relationships between chemical composition, electronic structure, and key static and dynamic properties (vibrational, redox, photophysical, energy/charge transfer) of tetrapyrrolic systems. Such architectures included monomers, dyads, larger arrays, and complexes with accessory components. The objective was to develop molecular designs and guiding principles to enhance current and future energy-conversion schemes. Molecular arrays targeted to address one or more fundamental questions concerning light harvesting and energy/charge transfer were constructed from analogues of the naturally occurring hemes, chlorophylls and bacteriochlorophylls. Diverse, tunable synthetic building blocks were prepared that spanned the three respective tetrapyrrole families, which are the porphyrins, chlorins and bacteriochlorins. Thus, the research focused on porphyrins as well as synthetic surrogates for chlorophylls (chlorins, 13 1 -oxophorbines and chlorin-imides) and bacteriochlorophylls (bacteriochlorins, bacterio-13 1 -oxophorbines and bacteriochlorin-imides), generically termed hydroporphyrins. Although the three tetrapyrrole classes (porphyrins, chlorins and bacteriochlorins) absorb light strongly in the violet-blue spectral region, the long-wavelength absorption band typically lies in the green-orange, red, and NIR regions, respectively, with increasing intensity. Understanding the spectra, electronic structure, and energy/charge-transfer properties of such tetrapyrrolic macrocycles is of central importance for the rational design of molecular architectures for solar-energy conversion. Our integrated program of molecular design and synthesis coupled with a variety of spectroscopic, electrochemical, and computational studies have probed from first principles how structural and electronic properties of tetrapyrrolic macrocycles dictate spectral properties as well as the rates of ground-state hole/electron transfer and excited-state energy flow in multicomponent architectures. Individual molecules and multicomponent architectures were designed to test ideas of fundamental importance, often requiring the development of new synthetic methodology. The members of the collaborative team had almost daily discussions by phone and/or e-mail concerning design of molecules, flow of compounds between the labs, planning of physical characterization studies, discussing results and analysis and integrating into design of next generation architectures, and the preparation of manuscripts. Furthermore, students and postdocs in the different labs routinely communicated with one another to facilitate the advancement of the research activities. In short, a highly integrated and collaborative research program was well established among the groups. The research effort involved molecular design and synthesis of synthetic molecular architectures by the Lindsey group integrated with physicochemical and photophysical characterization by the Bocian group and the Holten&Kirmaier group (Figure 2). The Bocian group carried out electrochemical, electron paramagnetic resonance (EPR), resonance Raman (RR), and Fourier-transform infrared (FT-IR) studies, as well as density functional theory (DFT) calculations and the time-dependent extension (TDDFT) to gain insight into excited-state properties. The Holten&Kirmaier group carried out static and time-resolved absorption and fluorescence spectroscopy studies and simulated absorption spectra using molecular orbital (MO) energies from DFT as input to the four-orbital model to complement the TDDFT calculations. The combined measurements provided understanding of the vibrational/electronic properties of the individual molecules and the changes that occur upon incorporation into multicomponent architectures. This information underpinned elucidating the mechanisms and timescales of ground-state hole/electron transfer and excited-state energy and charge transfer.

14 SOLAR ENERGY↗