Search NASA⌕ Search

SEARCH · Search NASA

Results for “pipeline data processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

The "PVLib" of Degradation: PVDeg

The Photovoltaic (PV) industry constantly aims for lower costs through higher-efficiency cells, improved module designs, and improvements in durability. This leads to the use of new materials, designs, and manufacturing processes, and not always with a sufficient amount of durability testing. To help drive down costs there is a desire to create modules that will last for up to 50 years of service life. To accomplish this, every degradation mode and mechanism must be identified and either eliminated or otherwise mitigated. This involves the extrapolation of laboratory results to the field conditions. There is a need to organize the existing degradation data into an accessible format and to provide industry relevant tools for extrapolation from laboratory to field conditions. While the basic equations used to model degradation are sometimes very simple, the full analysis involves calculations are cumbersome but ubiquitous for many degradation processes. A simplified, modeling framework to accomplish these repetitive processes will facilitate the analysis to help researchers keep up with the rapid pace of technological changes. In this talk, we will describe our progress creating the open-source tool PVDeg. This tool can be used to search for and analyze degradation information and extrapolate PV module performance and durability to field exposure. PVDeg simplifies many of the common foundational computational operations for obtaining meteorological data and using it to generate a model of the PV deployment. This prediction tool repository also contains various degradation models as well as a library of material parameters suitable for estimating the durability assessment of materials and components. We use an integration pipeline approach that allows us to leverage weather data from the National Solar Radiation Database, and other weather sources, to perform geospatial degradation analysis in the US and worldwide. We hope to become a repository that can be used for weathering and degradation analysis for various applications beyond the PV industry. During the talk, we will provide the PVPMC attendees the opportunity to interact with the tool via a Google Collab tutorial they can run on their phones or laptops.

durability↗

Dense autoencoders, clustering techniques, and semi-supervised learning for HPGe $γ$-spectra

Classifying high-resolution gamma spectra by their isotopic content is an essential task in nuclear forensics and other applications. Traditional analysis methods are often time-intensive, but machine learning (ML) may help analysts quickly process many spectra. Such methods tend to rely on abundant, well-labeled data for training. Historical gamma data exists in various fields but is not uniformly useful for supervised ML due to inconsistent labeling. Here, to address some of these challenges, we present a method to classify and organize unlabeled data from high-purity germanium detectors using an autoencoding neural network (autoencoder). We trained dense autoencoders to compress gamma data into latent representations that enable efficient data characterization. By clustering the encoded spectra or lower-dimensional mappings of them, we identified and removed portions of over-abundant data categories, resulting in a more balanced dataset and improved autoencoder performance. This encoding and clustering pipeline also enabled the organization of spectra into self-consistent categories. Finally, we found that encoded representations showed potential as inputs for semi-supervised learning of nuclide identification (NID) labels, achieving an average F1 score of 0.85 ± 0.03 when mapping encodings to a set of 65 isotope labels.

Autoencoders↗

Roughrider Carbon Storage Hub (Final Report)

The Roughrider Carbon Storage Hub was a 2-year project (October 2023 – September 2025) conducted by the Energy & Environmental Research Center (EERC) focused on advancing the feasibility of a commercial-scale carbon dioxide (CO 2 ) geologic storage hub in McKenzie County, North Dakota. The project’s objective was to investigate the potential that stacked storage complexes (multiple deep saline formations) can safely and economically store at least 50 million tonnes of CO 2 within 30 years. The captured CO 2 would be sourced from industrial emitters including project partner ONEOK, Inc.’s gas-processing plants and a planned gas-to-liquids facility. Drilling of the Roughrider 1 stratigraphic test well (14,979-ft total depth) was completed in November 2024. The wellbore intersected four candidate storage formations: Inyan Kara, Broom Creek, Mission Canyon, and Black Island–Deadwood. Operational challenges, including a stuck drill string, were resolved without long-term impact. A comprehensive logging and coring program was conducted, followed by successful well abandonment and site reclamation. Over 660 ft of 4-in. whole core was retrieved. Core plug samples were processed and analyzed for petrophysical and geochemical properties. Results confirmed promising porosity and permeability in the Inyan Kara and Broom Creek Formations and removal of the Mission Canyon and Black Island–Deadwood horizons from further investigation. Data derived from the logging and coring program were used to improve initial geologic models built from legacy data. CO 2 injection simulations showed that the Inyan Kara alone can feasibly store the target mass of CO 2 . Because of subtle differences in geologic structure and porosity trends between the formations, a stacked storage scenario using the Broom Creek and Inyan Kara Formations resulted in a larger overall plume area than using the Inyan Kara alone. Preliminary CO 2 pipeline routes from the industrial sources were mapped utilizing existing rights of way and evaluated for capacity and cost using U.S. Department of Energy Office of Fossil Energy and Carbon Management/National Energy Technology Laboratory models and U.S. Environmental Protection Agency emissions data. Integrating capture, transport, and storage cost estimates with policy incentives (e.g., 45Q credits) provided a total cost-per-ton analysis. Results indicate that the small scale of the volumes to be transported over the cumulative large distances does not support the project’s financial viability. However, the groundwork laid during this project from geological, regulatory, and social perspectives positions the Roughrider hub site as a promising candidate for commercial carbon storage in North Dakota, especially if the economy of scale is introduced for CO 2 transportation to the hub site.

01 COAL, LIGNITE, AND PEAT↗

White-Rabbit-Disciplined FPGA Readout for Fermilab Timing Events

Fermilab's accelerator timing links broadcast short event codes to thousands of devices at once, but the links themselves carry no absolute notion of time; that comes separately from a White Rabbit reference. This work builds the piece that ties the two together on a single board. On a Xilinx Kria KR260 (Zynq UltraScale+), the programmable logic decodes a real Fermilab TCLK link, stamps every event with an absolute White-Rabbit \{sec, ns\} UTC time, and reads the timestamped stream out over AXI4-Lite; a thin Linux process on the same die publishes each event into a Redis stream on the control network. To exercise the full chain on one board, the decoded events are re-encoded as gigabit ACLK, transmitted out an SFP+ optical port, looped back over a short fiber jumper, and decoded again on the same timeline, and are additionally mirrored as an ACLK-Lite Manchester waveform for benchtop probing. Across sustained, multi-day testing against real Fermilab TCLK, the pipeline has decoded, timestamped, and published hundreds of millions of events with practically zero loss, and folding the timestamped stream on the 60-second accelerator supercycle recovers the machine's periodic structure directly from the published data.

Rossel, Jacob [Fermilab; UC, Berkeley (main)] (ORC↗

A Single-Board Fermilab Timing Pipeline on the Xilinx KR260: Decode, Timestamp, Publish, Mirror

Full-width abstract \renewcommand{\maketitlehookd}{% \begin{abstract} \noindent Fermilab's accelerator timing links broadcast short event codes to thousands of devices at once, but the links themselves carry no absolute notion of time; that comes separately from a White Rabbit reference. This work builds the piece that ties the two together on a single board. On a Xilinx Kria KR260 (Zynq UltraScale+), the programmable logic decodes a real Fermilab TCLK link, stamps every event with an absolute White-Rabbit \{sec, ns\} UTC time, and reads the timestamped stream out over AXI4-Lite; a thin Linux process on the same die publishes each event into a Redis stream on the control network. To exercise the full chain on one board, the decoded events are re-encoded as gigabit ACLK, transmitted out an SFP+ optical port, looped back over a short fiber jumper, and decoded again on the same timeline, and are additionally mirrored as an ACLK-Lite Manchester waveform for benchtop probing. Across sustained, multi-day testing against real Fermilab TCLK, the pipeline has decoded, timestamped, and published hundreds of millions of events with practically zero loss, and folding the timestamped stream on the 60-second accelerator supercycle recovers the machine's periodic structure directly from the published data.

Rossel, Jacob [UC, Berkeley; Fermilab] (ORCID:0009↗

Mapping Hsp104 interactions using cross‐linking mass spectrometry

Molecular machines from the AAA+ (ATPases Associated with diverse cellular Activity) superfamily of protein disaggregases play important roles in protein folding, disaggregation and DNA processing. Recent cryo-EM structures of AAA+ molecular machines have uncovered nuanced changes in their conformation that underlie their specialized functions. Structural knowledge of these molecular machines in complex with substrates begins to explain their mechanism of activity. Here, we explore how cross-linking mass spectrometry (XL-MS) can be used to interpret changes in conformation induced by ATP in Hsp104 and how a substrate may interact with Hsp104. We applied a panel of cross-linking reagents to produce cross-linking maps of Hsp104 and interpret our data on previously determined X-ray and cryo-EM structures of Hsp104 from a thermophilic yeast, Calcarisporiella thermophila. We developed an analysis pipeline to differentiate between intra-subunit and inter-subunit contacts within the hexameric homo-oligomer. We identify cross-links that break the asymmetry that is present in Hsp104 in an ATP-hydrolysis competent conformation but is absent in an ATP-hydrolysis-defective mutant. Finally, we identify contacts between Hsp104 and a selected protein (proprotein convertase subtilisin/kexin type 9 PCSK9) to reveal contacts on the central channel of Hsp104 across the length of this protein indicating that we might have trapped interactions consistent with its translocation. Our simple and robust XL-MS-based experiments and methods help interpret how these molecular machines change conformation and bind to other proteins even in the context of homo-oligomeric assemblies enabling coupling state-of-the-art modeling approaches with XL-MS.

60 APPLIED LIFE SCIENCES↗

CORE-BFS: Communication-Optimized REctangular-partitioned BFS Achieving 160.845 TeraTEPS on Frontier Supercomputer

Distributed Breadth-First Search (BFS) is fundamental to many large-scale graph applications, but its performance on parallel systems is often limited by high communication overhead. This paper presents CORE-BFS, an extremely scalable GPU-based BFS implementation that introduces a unique rectangular 2D partitioning-based design for Frontier supercomputer. To further improve performance, we propose four key optimizations: (1) Rectangular 2D-partition specific data formats that use two compressed row and one compressed column status array bitmaps combined with a Double Compressed Sparse Row (DCSR) format per partition, reducing memory footprint and inter-rank traffic; (2) Adaptive frontier & communication strategy that unifies top-down and bottom-up traversal on the rectangular layout, uses lazy synchronization in top-down levels, and switches variants based on frontier size to minimize communication overhead; (3) Frontier-split degree-aware update that maps frontier vertices to thread-centric, wavefront-centric, and block-centric kernels based on their degree to improve GPU utilization and memory coalescing; (4) Row-reduction pipeline that overlaps bottom-up adjacency list processing with row-wise bitmap reduction to hide inter-rank latency. Together, these techniques increase parallelism while reducing memory and communication overhead. On the Graph500 benchmark, CORE - BFS scales up to 9,248 Frontier nodes with scale-42 graphs and reaches 160.845 TTEPS, delivering a 5.42 × speedup over our previous Frontier implementation.

Yang, Haoshen [Rutgers University]↗

Towards a Robust Adaptive Digital Twin for Fusion Applications

The development of a digital twin system for fusion applications is essential for enhancing the prediction, analysis, and optimization of complex plasma processes. Machine learning (ML), particularly deep learning has demonstrated strong capabilities in modeling such highly nonlinear and intricate systems. However, two critical challenges limit the deployment of deep learning-based digital twins: Uncertainty Quantification (UQ) and data drift. UQ is vital for ensuring trustworthy predictions, especially in decision-support scenarios. Additionally, data-driven models are often sensitive to changes in the underlying data distribution, such as shot-to-shot variations in fusion experiments, which can lead to performance degradation over time. To address these challenges, we are developing an uncertainty-aware, adaptive digital twin framework. Our approach incorporates deep learning models enhanced with Gaussian Process approximations for predictive uncertainty estimation, coupled with an online learning mechanism that enables continuous model adaptation to new experimental data. This adaptive capability allows the data driven models to respond effectively to evolving plasma behaviors and equipment conditions. Specifically, to mitigate the effects of shot-to-shot drift, our system updates itself incrementally as new data becomes available, improving both robustness and fidelity. Our vision is to evolve this data driven model into a self-sustaining digital twin system that leverages UQ based feedback to continuously refine itself and potentially support real-time decision making. This presentation will cover a brief background on uncertainty quantification for ML, our ongoing effort on development of UQ capabilities for ML, our data science pipeline from data collection to model development and analysis and online learning framework for modeling coil deflection at DIII-D. I will also briefly touch upon opportunities and challenges in development of digital twin framework.

Sammuli, Brian [General Atomics]↗

Vulcan-Forge: Architecture and Design of a Multi-Modal Forensic Analysis Plugin for CALDERA

Forge and VULCAN together describe an open-architecture cybersecurity analysis ecosystem that unifies forensic artifact processing, detection engineering, and vulnerability intelligence within integrated platforms. Forge operates as a plugin for MITRE CALDERA, ingesting diverse evidence formats—including EVTX, PCAP/PCAPNG, CSV, JSON, YAML, XML, binaries, and archives—to construct a unified artifact graph enriched with severity scoring, TLP classification, and audit trails. It provides subsystems for artifact parsing, streaming structured-data visualization, NetworkMiner-based packet inspection, PE/.NET binary analysis, and LLM-assisted triage and rule generation, with outputs validated against CCCS-YARA and pySigma schemas. VULCAN complements this by serving as a cybersecurity analyst platform that integrates a Neo4j knowledge graph, Qdrant vector retrieval, SSVC-based triage, and a local LLM to deliver CVE intelligence and forensic analysis through a multi-source ingest pipeline drawing from NVD, CISA KEV, EPSS, MITRE ATT&CK, and CAPEC. Together, they bridge structured threat intelligence with automated forensic analysis and detection workflows.

97 MATHEMATICS AND COMPUTING↗

Benchmark Tracking System for Performance Monitoring

Benchmarking is essential for high-performance software development, particularly for monitoring performance across code iterations. This project focused on enhancing the benchmarking process for Lamellar, an asynchronous runtime for High-Performance Computing (HPC) systems developed at Pacific Northwest National Laboratory. Prior to this work, benchmark results were difficult to track and compare across code versions, presenting significant challenges in identifying performance regressions and long-term trends. The primary objective was to establish a systematic, reproducible approach for measuring performance and detecting regressions following code commits. Our methodology involved three key components: standardizing benchmark outputs, implementing data versioning, and developing analysis tools. We standardized the benchmark output format to JSON Line records containing specific fields (execution time, hardware specifications, and environmental variables). To address data management challenges, we evaluated several options and eventually chose a git repository dedicated to benchmark data. We developed a suite of Python tools that processed benchmark results, enriched them with metadata, and facilitated search in the repository. The resulting system enables more efficient filtering and comparison of performance metrics across commit histories, hardware configurations, and benchmark variants through a unified query interface. Our implementation reduces computational overhead by first checking for existing results through configuration matching before initiating new benchmark runs, thereby conserving resources. The system has been validated by Lamellar developers. It organizes results by benchmark type and build configurations for efficient retrieval. Future developments include a planned Large Language Model interface for predicting benchmark performance, incorporating the criterion package for statistical analysis, which will enable automated detection of statistically significant performance changes, and integration with continuous integration pipelines. Despite these enhancements being reserved for future work, this project has successfully provided the Lamellar development team with a framework for maintaining consistent performance standards and identifying optimization opportunities across workloads and hardware environments.

97 MATHEMATICS AND COMPUTING↗

Identifying Adversarial Cyber-Activity in Operational Technology Environments Using Bayesian Networks

Critical infrastructure and other operational technology (OT) environments face increasing cybersecurity risks from adversarial behavior. This paper describes the development of a risk model using a Bayesian network to enhance the comprehension of observable cyber events caused by malicious activity in OT environments. The core of the Bayesian network is a process model that describes the stages of adversary behavior. The remainder of the model is based on the MITRE ATT&CK® for Industrial Control Systems (ICS) taxonomy, which includes tactics and techniques that may be used by the adversary. The observables provide evidence for adversary behavior through the intermediary technique and tactic nodes. One challenge in constructing this model is a lack of open-source data from cyber-attacks on OT systems. This paper discusses learning from limited data, the elicitation of expert opinion to construct the conditional probability tables when data is scarce, and the refinement of the most difficult conditional probabilities tables using several forms of sensitivity analyses. Finally, the Bayesian network is demonstrated using two historical case studies: the DarkSide ransomware attack on the Colonial Pipeline and the destructive cyberattack targeting the ThyssenKrupp blast furnace. Index Terms—Cybersecurity, industrial control systems, operational technology

97 - MATHEMATICS AND COMPUTING↗

MPACT Safeguards Modeling: FY25 Update

Sandia National Laboratories develops and maintains several open-source software packages to support material accountancy analyses. This includes the Material Accountancy Performance Indicator Toolkit (MAPIT), the Fissile Facility Flow Modeler (F3M) and the Separation and Safeguards Performance Model Library (SSPM-L). MAPIT is responsible for performing statistical safeguards analyses on bulk and itemized data from nuclear fuel cycle facilities and can operate on real or synthetic data. MAPIT is the only open-source software for such analyses. F3M is a library of modules, built in MATLAB Simulink, that contain pre made blocks to represent different generic fuel cycle processes. These blocks can be used together in a modular fashion to represent and simulate nuclear fuel cycle processes with the goal of improving facility-level accountancy during the design phase. F3M is also an open-source library. Finally, the SSPM-L library is a series of completed models built from F3M. The library includes facility models such as a generic PUREX facility and a fuel fabrication facility. The SSPM-L library is not open source, but is available to collaborators with a relevant use case. These tools include modeling and simulation pipelines to simulate nuclear fuel cycle facilities and the underlying software needed to simulate measurement uncertainty and perform statistical analyses. Together, these tools can perform end-to-end nuclear material accountancy analyses. This report documents the various improvements made to these tools in FY25. Specifically, we added new statistical test, new statistical modeling capabilities, new fuel cycle facility models, and launched a new open-source model component library.

97 MATHEMATICS AND COMPUTING↗

Global Archaeal Diversity Revealed Through Massive Data Integration: Uncovering Just Tip of Iceberg

The domain of Archaea has gathered significant interest for its ecological and biotechnological potential and its role in helping us to understand the evolutionary history of Eukaryotes. In comparison to the bacterial domain, the number of adequately described members in Archaea is relatively low, with less than 1000 species described. It is not clear whether this is solely due to the cultivation difficulty of its members or, indeed, the domain is characterized by evolutionary constraints that keep the number of species relatively low. Based on molecular evidence that bypasses the difficulties of formal cultivation and characterization, several novel clades have been proposed, enabling insights into their metabolism and physiology. Given the extent of global sampling and sequencing efforts, it is now possible and meaningful to question the magnitude of global archaeal diversity based on molecular evidence. To do so, we extracted all sequences classified as Archaea from 500 thousand amplicon samples available in public repositories. After processing through our highly conservative pipeline, we named this comprehensive resource the ‘Global Archaea Diversity’ (GAD), which encompassed nearly 3 million molecular species clusters at 97% similarity, and organized it into over 500 thousand genera and nearly 100 thousand families. Saline environments have contributed the most to the novel taxa of this previously unseen diversity. The majority of those 16S rRNA gene sequence fragments were verified by matches in metagenomic datasets from IMG/M. These findings reveal a vast and previously overlooked diversity within the Archaea, offering insights into their ecological roles and evolutionary importance while establishing a foundation for the future study and characterization of this intriguing domain of life.

59 BASIC BIOLOGICAL SCIENCES↗

Importance of Higher Fidelity Model Geometries during Optimization of Critical Experiments

PARADIGM, PARallel Approach of Differential and InteGral Measurements, is a cross-collaborative effort at Los Alamos National Laboratory between nuclear data theorists, differential and integral experimenters, as well as machine learning statisticians to tackle uncertainties in the intermediate region of 239 Pu. In essence, the idea behind PARADIGM is to remove the linear conceptualization of the nuclear data pipeline, shown in Figure 1, and replace it with a far more parallelized approach. The novel approach leverages machine learning to guide which differential measurements and integral experiments will result in the largest decrease in uncertain ties for a nuclide reaction pair in a given energy range. The concept builds off earlier work, EUCLID, which focused on the fast region of 239 Pu. The practical benefit of having evaluation, differential measurement, and integral experiment personnel in collaboration with machine learning is to represent the entire nuclear data in one snapshot. This enable large reduction in the time to deliver improved nuclear data, which using the PARADIGM approach could be done in 3 years. A general outline of PARADIGM and specific topics are available in other papers. The discussion here will pertain directly to the integral experiment design. More specifically, the process of taking a rough design and transforming it into a finalized neutronic model will be discussed.

97 MATHEMATICS AND COMPUTING↗

Artificial intelligence to unlock real-world evidence in clinical oncology: A primer on recent advances

Purpose: Real world evidence is crucial to understanding the diffusion of new oncologic therapies, monitoring cancer outcomes, and detecting unexpected toxicities. In practice, real world evidence is challenging to collect rapidly and comprehensively, often requiring expensive and time-consuming manual case-finding and annotation of clinical text. In this Review, we summarise recent developments in the use of artificial intelligence to collect and analyze real world evidence in oncology. Methods: We performed a narrative review of the major current trends and recent literature in artificial intelligence applications in oncology. Results: Artificial intelligence (AI) approaches are increasingly used to efficiently phenotype patients and tumors at large scale. These tools also may provide novel biological insights and improve risk prediction through multimodal integration of radiographic, pathological, and genomic datasets. Custom language processing pipelines and large language models hold great promise for clinical prediction and phenotyping. Conclusions: Despite rapid advances, continued progress in computation, generalizability, interpretability, and reliability as well as prospective validation are needed to integrate AI approaches into routine clinical care and real-time monitoring of novel therapies.

60 APPLIED LIFE SCIENCES↗

Leveraging unlabeled SEM datasets with self-supervised learning for enhanced particle segmentation

Scanning Electron Microscopes (SEMs) are widely used in experimental science laboratories, often requiring cumbersome and repetitive user analysis. Automating SEM image analysis processes is highly desirable to address this challenge. In particle sample analysis, Machine Learning (ML) has emerged as the most effective approach for particle segmentation. However, the time-intensive process of manually annotating thousands of SEM images limits the applicability of supervised learning approaches. Self-Supervised Learning (SSL) offers a promising alternative by enabling knowledge extraction from raw, unlabeled data. This study presents a framework for evaluating SSL techniques in SEM image analysis, focusing on novel methods leveraging the ConvNeXtV2 architecture for particle detection. A dataset comprising 25,000 SEM images is curated to benchmark these proposed SSL methods. The results demonstrate that ConvNeXtV2 models, with varying parameter counts, consistently outperform other techniques in particle detection across different length scales, achieving up to a 34% reduction in relative error compared to established SSL methods. Furthermore, an ablation study explores the relationship between dataset size and SSL performance, providing actionable insights for practitioners regarding model selection and resource efficiency. This research advances the integration of SSL into autonomous analysis pipelines and supports its application in accelerating materials science discovery.

Rettenberger, Luca↗

Trisodium Phosphate Phases and Solubility in Alkaline Solutions Relevant to Radioactive Waste Processing

The U.S. Department of Energy’s Hanford Site faces significant challenges in managing millions of gallons of legacy radioactive waste, where phosphate precipitation can obstruct pipelines during retrieval and processing. To resolve long-standing inconsistencies in reported solubility and clarify factors governing trisodium phosphate hydrates, we examined the Na3PO4:NaOH:H2O system. Powder and single-crystal X-ray diffraction revealed that commercial precursors undergo transformations that produce multiple hydrates, including three previously unreported phases comprising an ordered polymorph of Na3(PO4)·12H2O·1/6NaOH, Na3PO4·5H2O, and Na3PO4·9H2O. Computational modeling indicated that the ordered dodecahydrate is more stable than its disordered counterpart, suggesting kinetic persistence of structural disorder. Solubility measurements were conducted with solutions prepared either from anhydrous Na3PO4 or from Na3(PO4)·12H2O·1/6NaOH and revealed that release of interstitial NaOH from the hydrate precursor elevated solution alkalinity, thereby significantly reducing phosphate solubility relative to solutions prepared with anhydrous Na3PO4 across 20–44 °C. These results begin to reconcile inconsistencies in prior solubility data and clarify how phase composition dictates phosphate precipitation under alkaline conditions.

Graham, Trenton R. (ORCID:0000000189078004)↗

Carbon Management Projects (CONNECT) Database and Explorer

Overview The Carbon Management Projects (CONNECT) Toolkit is an online exploratory visualization tool developed by the U.S. Department of Energy's (DOE) Office of Fossil Energy and Carbon Management (FECM) with support from other federal agencies such as the U.S. Environmental Protection Agency (EPA) and the U.S. Department of Transportation (DOT). It provides a single point of access to authoritative information on federal agency investment in a portfolio of research, development, and demonstration (RD&D) projects that have been publicly announced to advance technologies for point source carbon capture, carbon dioxide removal, transport, storage, and conversion, collectively referred to as carbon management. The RD&D programs covered in this tool are authorized by annual congressional appropriations ("Base Program") and the 2021 Infrastructure Investment and Jobs Act (IIJA). The tool also incorporates public information on other federal initiatives, such as the Regional Clean Hydrogen Hubs, and public information released by other government agencies, such as the Environmental Protection Agency's (EPA) and Primacy States’ Underground Injection Control Class VI permits and EPA’s facility level greenhouse gas (GHG) emissions. Developed in a geographic information system, the tool organizes carbon management projects into five groups based on the primary technology that a project aims to advance, each visually represented as a digital layer ("carbon management project layer"). Only federally funded projects are included, which can be awarded projects that are completed or ongoing, or projects that have been selected but are currently under negotiation. Project information can be viewed in the map or in the attribute table below it when turned on. In the map view, each project is displayed at either its host site (for field work), where available, or its performer site (project lead's location, further explained in the table below). Host sites and performer sites are represented in distinct icons. Several reference layers offer additional public information on infrastructural and natural resource environment for carbon management. These reference layers, combined with multiple geographical basemaps, enable users to visualize the carbon management project layers in context. Carbon management project information will be updated monthly based on feedback and information availability. Carbon management project layers Point Source Carbon Capture (PSC) This layer contains DOE-funded projects focused on capturing carbon dioxide (CO2) from power plants or industrial facilities. Carbon Dioxide Removal (CDR) This layer contains DOE-funded projects focused on capturing CO2 from the atmosphere, including direct air capture (DAC) and DAC hubs, direct ocean capture, enhanced mineralization, and biomass carbon removal and storage. For projects with multiple host sites, each of the sites are displayed individually with the project cost and cost sharing information representing the total for the entire project. Carbon Transport This layer contains DOE- and DOT-funded projects focused on CO2 transport. The Transport Research and Development sublayer contains projects that do not involve physical infrastructure; the Proposed Transport Corridor sublayer contains projects for which either a route for the transport infrastructure has been proposed or a general area for the transport infrastructure has been identified. Carbon Storage This layer contains DOE-funded key projects focused on CO2 storage. For projects with multiple field-work sites, each of the sites are displayed individually on the map with the project cost and cost sharing information representing the overall total for the entire project. Carbon Conversion This layer contains DOE-funded projects focused on converting CO2 into economically valuable products. Reference layers The following layers provide additional information in the geographic proximity of carbon management projects. Users should reference the original sources for more details (weblinks provided below and in pop-up windows on the map). Regional Clean Hydrogen Hub and Facility These layers illustrate the approximate areas of the Regional Clean Hydrogen Hubs announced by DOE's Office of Clean Energy Demonstrations (OCED) and the approximate locations of individual facilities that constitute the hubs (see "Where are the H2Hubs located?" on the webpage linked above). EPA Facility Level GHG Emissions (direct emitter) This layer shows direct CO2 emissions from stationary sources in 2022, using data extracted from EPA's Facility Level Information on GreenHouse gases Tool (FLIGHT). Captured and injected CO2 are not deducted from direct emitters’ total emissions. Contact EPA for additional details. Underground Injection Control Class VI permit/permit application This layer shows the locations of CO2 injection wells that are granted or in the process of applying for an Underground Injection Control Class VI permit by EPA or a Primacy State (currently Louisiana, North Dakota, and Wyoming). The URLs for the permits or permit applications are provided in the pop-up windows associated with the well locations. Contact EPA for additional details. Carbon Storage Resource This layer contains information on prospective CO2 storage resources in saline formations and oil and gas reservoirs provided by the National Carbon Sequestration Database and Geographic Information System (NATCARB) spatial database. Contact NETL for additional details. Existing CO2 pipeline This layer shows active CO2 pipelines based on information digitized from the map issued by the Pipeline and Hazardous Materials Safety Administration (PHMSA). Contact PHMSA for additional details.

Carbon Conversion↗