Search NASA⌕ Search

SEARCH · Search NASA

Results for “pipeline data processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Data Selection Improvement For MicroBooNE

Data selection is an extremely important part of data analysis for any experiment. Finding a physics result is often the result of sifting through a massive amount of data, keeping data that we believe to be signal, and throwing out data we do not. This process is called data selection. Creating a selection algorithm is an intensive process that must balance keeping enough data to have statistics and maximizing the signal purity of that data. We also need to choose the right reconstruction method, a tool to take raw data from the detector and convert it into physics results. In this study, we used three different reconstruction tools, Pandora, WireCell, and LANTERN, for the MicroBooNE experiment in conjunction to improve the selection algorithm for analysis. For the case of this study, we look into the charged current N proton 0 pions (CCNp0$\pi$) interaction channel. This is the dominant channel for the Short Baseline Neutrino (SBN) program and is expected to be a large contributor to the Deep Underground Neutrino Experiment (DUNE). We first investigated each of the three tools to find out more about their strengths and weaknesses as reconstructions, and compared them to the truth information directly from the MicroBooNE simulation pipeline. We then put together a direct comparison of the three methods to find which method or combination of methods would return the best result for us. While the study is ongoing, we have learned a lot about data selection for the experiment and the differences between the reconstruction tools.

Dillon, Brayden [Fermilab]↗

Window convolution of the galaxy clustering bispectrum

In galaxy survey analysis, the observed clustering statistics do not directly match theoretical predictions but rather have been processed by a window function that arises from the survey geometry including the sky footprint, redshift-dependent background number density and systematic weights. While window convolution of the power spectrum is well studied, for the bispectrum with a larger number of degrees of freedom, it poses a significant numerical and computational challenge. In this work, we consider the effect of the survey window in the tripolar spherical harmonic decomposition of the bispectrum and lay down a formal procedure for their convolution via a series expansion of configuration-space three-point correlation functions, which was first proposed by Sugiyama et al. (2019). We then provide a linear algebra formulation of the full window convolution, where an unwindowed bispectrum model vector can be directly premultiplied by a window matrix specific to each survey geometry. To validate the pipeline, we focus on the Dark Energy Spectroscopic Instrument (DESI) Data Release 1 (DR1) luminous red galaxy (LRG) sample in the South Galactic Cap (SGC) in the redshift bin 0.4 ≤ z ≤ 0.6. We first perform convergence checks on the measurement of the window function from discrete random catalogues, and then investigate the convergence of the window convolution series expansion truncated at a finite of number of terms as well as the performance of the window matrix. This work highlights the differences in window convolution between the power spectrum and bispectrum, and provides a streamlined pipeline for the latter for current surveys such as DESI and the Euclid mission.

79 ASTRONOMY AND ASTROPHYSICS↗

Systems Analysis of Biomass and Coal Co-firing Power Plants with Deep Carbon Capture Toward Net-zero Emissions

Achieving a net-zero emission economy in the United States requires integrating diverse low-carbon and negative-emission technologies into the existing fossil fuel-dominant power fleet. Potential technologies from the low-carbon portfolio include renewable power, fossil power with carbon capture and storage (CCS), bioenergy with CCS (BECCS), and direct air capture (DAC). Renewable power is a clean energy source but has to pair with costly battery storage to provide dispatchable electricity. Fossil power with CCS offers dispatchable electricity yet still relies on DAC to offset residual emissions, even when deploying deep CCS with more than 90% CO2 capture. Coal-biomass co-firing with CCS, a subset of BECCS, is a reliable energy production technology that can be retrofitted from existing electricity generation units (EGUs). Power plant retrofit maximizes the use of the current U.S. coal power fleet without the need for large-scale deployment of new renewable power, battery storage, or DAC. Retrofitting coal-biomass co-firing with deep CCS in EGUs is a promising option, but not a universal solution. Biomass co-firing at a power plant introduces economic challenges and indirectly poses pressure on land and water resources. Meanwhile, retrofitting deep CCS affects plant efficiency and raises electricity generation costs. Overall, the technical feasibility and economic viability of plant retrofits vary across EGUs, as they are contingent upon the regional availability of biomass, unit-specific characteristics, site-specific fuel supply costs, and adjacent CO2 storage potential. Government incentives like 45Q can improve the retrofit viability, though the impact requires further quantification. A comprehensive analysis at the unit level is essential to address the question regarding the fate of the U.S. coal-fired electricity generation fleet toward the net-zero emission goal. This study conducts a systematic techno-economic-environmental assessment of EGUs to identify the viability of biomass co-firing and deep CCS retrofits in the U.S. coal-fired power fleet. Specifically, it characterizes the techno-economic performance of deep carbon capture, estimates life cycle greenhouse gas (GHG) emissions, and conducts a fleet-level assessment on retrofit viability. The key objectives are (1) to estimate the unit-specific performance and retrofitted cost under various biomass co-firing levels and CO2 capture rates; (2) to determine the possibility of reaching net-zero emission at the fleet level; (3) to quantify the cumulative capacities that are suitable for plant retrofits under current and future biomass supply scenarios; and (4) to improve the understanding of policy impacts on such retrofits to help the power sector’s transition to a net-zero economy. Techno-economic Model of Deep Carbon Capture. This study develops the performance and economic models for Monoethanolamine-based post-combustion CO2 capture at 95–99% capture rates. The process is simulated in Aspen Plus, analyzing the performance of carbon capture technology by varying the plant sizes, solvent lean loading, CO2 concentrations, and flue gas inlet temperature. Based on the key inputs and output parameters of CO2 capture, a reduced-order performance model of deep carbon capture is formulated. In addition, an engineering-economic model integrating the performance metrics is developed to estimate the capital as well as operation and maintenance (O&M) costs. Capital cost estimations follow the framework of the Integrated Environmental Control Model (IECM) and incorporate data regressions from three technical reports by IECM, the National Energy Technology Laboratory (NETL), and the National Renewable Energy Laboratory. The O&M cost estimation utilizes the actual inventory consumption rate and labor requirements. Both performance and cost models are embedded into IECM v13.0-beta, a fossil-fuel power plant modeling tool. Life Cycle Assessment of Power Plants. This study estimates the GHG emissions of power plants through life cycle assessment (LCA). The LCA scope includes fuel supply, combustion-based power generation, and CO2 transport and storage. The fuel-based life cycle module is designed following the framework of the NETL Unit Process Library and CO2U LCA Guidance Toolkit. The module is then incorporated into IECM v13.0-beta. The process-based LCA is applied to estimate the GHG emissions of coal and biomass supply, coal- and coal-biomass co-firing power plant operation, as well as CO2 pipeline transport and geographical sequestration. An uncertainty analysis is conducted to quantify the variability and uncertainty associated with the LCA using the Latin Hypercube Sampling (LHS) method. Fleet-level Assessment. This study evaluates the technical and economic feasibility of selected coal-fired EGUs, examines the role of tax credits in retrofit viability, and assesses the competitiveness of retrofitted units against other low-carbon options. Unit screening identifies EGUs for the study, focusing on new, efficient baseload units with air pollution controls. The power plant databases are then established to organize unit-specific information on performance and operating conditions from the relevant public databases. Biomass for co-firing retrofits is selected based on home and neighboring county availability, ensuring sustained operation with at least a 5% co-firing level. The CO2 storage site is determined by state-level storage potential, with ArcGIS Pro and NETL CO2 Saline Storage Cost Model used to identify the optimal balance between the nearest transport distances and affordable storage costs. The latest IECM v13.0-beta is then employed to configure and evaluate the eligible EGUs with or without the deployment of deep CCS and biomass co-firing. A supply curve is established to illustrate the cumulative installed capacity suitable for retrofits at different cost levels. A sensitivity analysis on tax credits for carbon sequestration is performed. Finally, a unit-level cost comparison is conducted among retrofitted plants, renewable power with battery storage, and abated fossil fuels with DAC. Expected Results. This study evaluates the technical, economic, and environmental metrics of each EGU across an array of CO2 capture rates and biomass co-firing level scenarios. Unit-level comparisons will identify critical factors influencing technical performance. The supply curves with and without tax incentives will provide insights into the impact of tax credits on biomass co-firing and CCS deployment. The cost comparisons with renewables and DAC-retrofit will assess the competitiveness of the retrofitted units. Life cycle emissions from each unit will be assessed to identify the scenarios under which net-zero emissions can be achieved. These analyses are expected to determine the total coal-fired capacity suitable for serving as a low-carbon energy source with or without tax incentives. The study results are novel in identifying optimal unit-specific strategies for producing carbon-neutral power, whether through retrofitting EGUs with deep CCS, biomass co-firing, DAC, or installing renewable power with battery. The findings will provide insight into nationwide efforts to ensure reliable, affordable, and low-carbon electricity. It also will inform investment decisions and policies in the deployment of deep carbon capture and negative emission technologies for a net-zero energy future.

Biomass Co-firing↗

The Atacama Cosmology Telescope: DR6 power spectrum foreground model and validation

We discuss the model of astrophysical emission at millimeter wavelengths used to characterize foregrounds in the multi-frequency power spectra of the Atacama Cosmology Telescope (ACT) Data Release 6 (DR6), expanding on Louis et al. (2025) (2503.14452). We detail several tests to validate the capability of the DR6 parametric foreground model to describe current observations and complex simulations, and show that cosmological parameter constraints are robust against model extensions and variations. We demonstrate consistency of the model with pre-DR6 ACT data and observations from Planck and the South Pole Telescope. We evaluate the implications of using different foreground templates and extending the model with new components and/or free parameters. In all scenarios, the DR6 ΛCDM and ΛCDM+N eff cosmological parameters shift by less than 0.5σ relative to the baseline constraints. Some foreground parameters shift more; we estimate their systematic uncertainties associated with modeling choices. From our constraint on the kinematic Sunyaev-Zel'dovich power, we obtain a conservative limit on the duration of reionization of Δz rei < 4.4, assuming a reionization midpoint consistent with optical depth measurements and a minimal low-redshift contribution, with varying assumptions for this component leading to tighter limits. Finally, we analyze realistic non-Gaussian, correlated microwave sky simulations containing Galactic and extragalactic foreground fields, built independently of the DR6 parametric foreground model. Processing these simulations through the DR6 power spectrum and likelihood pipeline, we recover the input cosmological parameters of the underlying cosmic microwave background field, a new demonstration for small-scale CMB analysis. These tests validate the robustness of the ACT DR6 foreground model and cosmological parameter constraints.

CMBR experiments↗

Enabling the Broader Use of MOOSE for Nuclear Energy and Other Simulation

This Final Scientific and Technical Report summarizes work performed under the Phase IIA SBIR project “Enabling the Broader Use of MOOSE for Nuclear Energy and Other Simulation” (DE-SC0020906) from August 2023 through August 2025. The objective of the Phase IIA effort was to mature and harden capabilities developed during Phase II, with the goal of enabling practical interoperability between Coreform’s isogeometric analysis (IGA) technologies and the Multiphysics Object-Oriented Simulation Environment (MOOSE), while improving robustness, performance, and scalability for complex, nuclear-relevant geometries. Over the course of Phase IIA, the project established and validated an extraction-based interoperability pathway between Coreform tools and MOOSE. A combined mesh and matrix format was defined collaboratively with MOOSE developers and integrated into the solver, enabling standard MOOSE workflows to operate on data exported from Coreform’s IGA and Flex Representation Method (FRM) pipelines. Early demonstrations validated architectural compatibility using linear solid mechanics problems, while later efforts focused on benchmark testing and external use. By the end of the project period, engineers at BWXT were able to independently set up and execute a simulation using the Coreform–MOOSE workflow and provide direct feedback that informed further refinement. In parallel, substantial effort was devoted to improving the robustness of trimmed U-spline construction for complex CAD geometries. A growing test suite of nuclear-relevant models was compiled through collaboration with multiple stakeholders and used to drive extensive bug fixing and reliability improvements. These efforts resulted in improved robustness and performance, including the addition of fallback capabilities that enhance reliability when the underlying commercial CAD kernel fails. Performance-oriented work progressed later in the project, with the development and demonstration of methods to decompose complex geometries into structured subregions and updated data representations to support more efficient solver processing. Additionally, extensive enhancements to threadsafe parallel data structures and trimming operations established a foundation for scalable processing of large assemblies. Collaboration with Sandia National Laboratories on the SGM geometric modeling kernel advanced to a functioning interface test case, positioning the workflow for future kernel integration. Overall, the Phase IIA effort successfully transitioned the project from architectural proof-of-concept to externally exercised, solver-integrated capability, while clarifying remaining technical challenges related to standardization, performance optimization, and kernel integration.

42 ENGINEERING↗

Archetype-based Redshift Estimation for the Dark Energy Spectroscopic Instrument Survey

We present a computationally efficient galaxy archetype-based redshift estimation and spectral classification method for the Dark Energy Survey Instrument (DESI) survey. The DESI survey currently relies on a redshift fitter and spectral classifier using a linear combination of principal component analysis–derived templates, which is very efficient in processing large volumes of DESI spectra within a short time frame. However, this method occasionally yields unphysical model fits for galaxies and fails to adequately absorb calibration errors that may still be occasionally visible in the reduced spectra. Our proposed approach improves upon this existing method by refitting the spectra with carefully generated physical galaxy archetypes combined with additional terms designed to absorb data reduction defects and provide more physical models to the DESI spectra. We test our method on an extensive data set derived from the survey validation (SV) and Year 1 (Y1) data of DESI. Our findings indicate that the new method delivers marginally better redshift success for SV tiles while reducing catastrophic redshift failure by 10%–30%. At the same time, results from millions of targets from the main survey show that our model has relatively higher redshift success and purity rates (0.5%–0.8% higher) for galaxy targets while having similar success for QSOs. These improvements also demonstrate that the main DESI redshift pipeline is generally robust. Additionally, it reduces the false-positive redshift estimation by 5%–40% for sky fibers. We also discuss the generic nature of our method and how it can be extended to other large spectroscopic surveys, along with possible future improvements.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Graph-Based Modeling for the Detection and Tracking of Sarin-Surrogate-Induced Neurotoxicity Using a Human-Relevant, In-Vitro Brain Model

Organophosphorus (OP) nerve agents are a chemical threat to the United States, to the civilian population (e.g., pesticides) and historically weaponized (e.g., sarin) as chemical warfare agents. The unprecedented, accelerated process from “bench-to-bedside” during the SARSCov2 pandemic has made it clear that technology and tools need to be readily available for immediate response. Advances in human organ tissue mimetic systems are a promising technology to evaluate the human-relevant response in vitro for basic and applied research and drug screening. In particular, current brain microphysiological systems (MPS) have the capability to monitor and detect changes in engineered human neural circuit activity. However, current data analytics approaches for these systems lack the granularity to functionally detect and distinguish the different mechanisms that occur in the brain following neurotoxicity, injury, and disease. The goal of this project was to advance the computational analytical capabilities of the brain MPS to detect functional changes in neural circuit structure at different stages of Sarin surrogate-induced neurotoxicity. We developed graph-based models to (1) identify the composition of the neural circuit structure; (2) detect and monitor how this structure changes following sarin-induced neurotoxicity; and (3) evaluate the analytical pipeline using known/promising oxime reactivators. Through experiments on the bMPS where in vitro neuronal cultures were exposed to a sarin surrogate, we demonstrated the capabilities of our computational pipeline to identify different responses in the functional networks of brain cells exposed to low and high concentrations of the nerve agent. We identified a biphasic response of human neural network activity following exposure to a sarin-surrogate that had not been reported in the literature before. The graph-based models and software developed in this project can be used for future studies that leverage the brain MPS technology, such as treatment efficacy assessment.

59 BASIC BIOLOGICAL SCIENCES↗

BM3DORNL

BM3DORNL is a high-performance, open-source library for removing streak and ring artifacts from computed-tomography (CT) data, developed for neutron imaging at Oak Ridge National Laboratory's Spallation Neutron Source (VENUS beamline) and applicable to X-ray CT as well. Ring artifacts — concentric rings in reconstructed slices caused by detector pixel-to-pixel response non-uniformities — appear as vertical streaks in the sinogram and degrade both image quality and quantitative analysis. BM3DORNL operates in the sinogram domain using an adaptation of the BM3D (block-matching and 3D collaborative filtering) algorithm (Dabov et al., 2007). It provides a dedicated streak-removal mode, a true multi-scale BM3D variant (after Mäkinen et al., 2021) that suppresses wide streaks single-scale methods miss, and an alternative Fourier–SVD method (~2.6× faster) combining FFT-based energy detection with rank-1 SVD. The computationally intensive core is implemented in Rust with parallel (Rayon) block matching, integral-image pre-screening, and optimized transforms, and is exposed through a simple Python API (with an optional GUI) so it integrates directly into existing tomography reconstruction pipelines. It processes both 2D sinograms and 3D sinogram stacks, is pip-installable for Linux and macOS, and is documented at https://bm3dornl.readthedocs.io.

Zhang, Chen [Oak Ridge National Laboratory (ORNL),↗

Custom Accessors: Enabling Scalable Data Ingestion, (Re-)Organization, and Analysis on Distributed Systems

The emerging class of high velocity and high volume data analytic workflows comprise interwoven data ingestion, organization, and processing stages, with ingestion and organization steps often contributing comparable or even higher computational costs than actual processing steps. Since complex workflows consist of a variety of phases that view and use data differently, being able to construct efficient, scalable, distributed data structures (arrays, vectors, sets, maps, and multi-maps) is essential and requires custom methods to extend and shrink containers, analyze and position data, and, maintain globallyconsistent meta-data. In this paper, we propose a novel datastructure access paradigm based on the concept of Accessors. At a high level, accessors are customizable callable objects that can modify the behavior of insert, read, update, and delete operations for distributed containers while preserving atomicity guarantees. Accessors provide a very clean and natural way to implement a variety of programming patterns, e.g., conditional insertion/deletion and cascading computations, which would be otherwise hard (or even impossible) to express in parallel and distributed settings without using locks. We demonstrate the practicality and usefulness of our approach with two representative use cases and study the performance of these applications on a distributed High-Performance Computing system. Our analysis highlights that our proposed abstraction allows for an effective overlapping and concurrent execution of different workflow steps (e.g., data ingestion and analysis), which in a conventional analytics pipeline would execute sequentially, contributing cumulatively to the overall latency.

Castellana, Vito G. [BATTELLE (PACIFIC NW LAB)] (O↗

The Novel Charfuel® Coal Refining Process 18 TPD Pilot Plant Project for Co- Producing an Upgraded Coal Product, and Commercially Valuable Co- Products: Area of Interest #3 – Coal Beneficiation Pilot Plant Testing (Final Report)

Operation of Carbon Fuels, LLC’s (“CF”) existing, permitted 18 TPD pilot plant located in Golden, Colorado using two individually ranked (ASTM D 388) coal types (two campaigns), employing the novel Charfuel® coal refining process to produce an upgraded coal product and a number of high-valued organic and inorganic coproducts (for which there presently exists large commercial markets) in order to produce engineering and product data which will then be utilized toward the design of a commercial scale integrated facility (pre-feed document). Carbon Fuels, LLC has developed the Charfuel® Coal Refining Process which refines domestically abundant, raw coal (in the same manner as crude oil is refined) to produce the identical, high value co-products that are refined from crude oil. Thus, gasoline, jet fuel, “green diesel”, fuel oil, and marine fuels, as well as petrochemicals such as benzene, toluene, xylene, and methanol are refined from raw coal using this process. The Charfuel® Coal Refining Process is not a coal conversion process, like pyrolysis, or indirect liquefaction. Nor is it an alternative energy system. Rather it is a coal refining process that has the ability to economically produce products traditionally associated with the refining of crude oil but using only abundant, raw coal as the refinery feed stock. The Charfuel® Coal Refining Process is more economical than crude oil refining and is environmentally benign. Therefore, this value added process yields a return on investment well above 50% for a commercial facility. Furthermore, the Charfuel® process, unlike alternatives such as ethanol and hydrogen, can utilize the existing transportation, delivery, and other petroleum based systems. Hence, there is no need for new engines, pipelines, tankers, or product acceptance. As a result, the profitability of the process is increased. Objectives: (1) Operation of the integrated 18 tpd pilot plant, using two coal types (ranks); (2) Demonstration of process flexibility in being able to produce different products (gas, liquid, and char), as well as determination of operating parameters for identifying scale up criteria for two coal types (ranks); (3) Generation of engineering and design information (process specifications) for use in designing a commercial scale plant (scale-up); (4) Determination of important environmental issues surrounding the process and the products such as fate of trace elements (mercury and other heavy metals) and distributions of SO2, NOx, and CO 2 by analysis of effluent streams; (5) Production of sufficient product to allow reliable commercial economic evaluation of both the refined coal product and the coproducts; and, (6) Assessment of longer-term reliability of unit operations. Period 1: reconfiguration of the 18 TPD plant to meet specific FOA requirements and to qualify the facility for operation; and, Period 2: operation of the 18 TPD plant for two campaigns using two coals types (ranks) which are widely commercially used and abundant - the first being a subbituminous (Powder River Basin (“PRB”)) coal, and the second a bituminous (Illinois #6) coal.

01 COAL, LIGNITE, AND PEAT↗

CERF: IM3 Projected Western US Power Plant Locations

Overview The Capacity Expansion Regional Feasibility (CERF) model is an open-source geospatial python package that provides new power plant locations at a 1km resolution. The model ingests U.S. state or regional-scale electricity system capacity expansion plans, such as those produced by the Global Change Analysis Model (GCAM-USA), and identifies feasible, site-specific locations for individual new power plants (renewable and non-renewable). CERF combines high-resolution geospatial suitability analyses with an economic algorithm that selects individual plant siting locations based on grid interconnection costs and the locational marginal value of new generation. The model incorporates a wide range of dynamic constraints and opportunities, such as protected lands, population density, existing infrastructure, and water availability. This dataset provides CERF power plant siting results for IM3 Phase 2 simulations across eight different scenarios for the Western US through 2055. The scenarios include combinations of two Shared Socioeconomic Pathways (SSP3 and SSP5) with four high-resolution climate projections specific to the United States (see, https://tgw-data.msdlive.org/). These climate projections include "hotter" and "cooler" variants for two Representative Concentration Pathways (RCP4.5 and RCP8.5). The resulting eight simulations are: rcp45cooler_ssp3 rcp45cooler_ssp5 rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85cooler_ssp3 rcp85cooler_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 CERF siting results in this dataset correspond to capacity expansion plans in the GCAM-USA IM3 Phase 2 simulation data and are available for each of the above scenarios. Data Details Temporal Range: 2015-2055 in 5-year timesteps. Note that 2015 is the experiment base year and 2020 and beyond represent model simulation years. Spatial Range: Plant locations are provided for the eleven states in the Western US including Arizona, California, Colorado, Idaho, Montana, New Mexico, Nevada, Oregon, Utah, Washington, and Wyoming. Spatial Resolution: 1 km-squared, provided in x and y coordinates Geospatial Projection: Albers Equal Area Conic (ESRI:102003) File Type: csv The dataset contains subdirectories for each of the eight scenarios described in the overview. Each scenario folder contains two subfolders with the following information: 1. Power Plant Data This directory contains a single .csv file of power plant locations for both pre-existing (non-CERF sited plants in operation in 2015) and new (CERF-sited) power plants across the temporal range along with additional CERF model output parameters for CERF-sited plants. Plant with a siting year earlier than 2020 correspond to facilities that are operational leading into the first timestep CERF simulation. For a more detailed description of CERF model output parameters, see the CERF model documentation. Note that the cerf_plant_id parameter is unique within each scenario file but not across scenario files. Parameter Descriptions scenario - Name of scenario cerf_plant_id - Unique siting identifier cerf_sited - If True, indicates that plant was sited by CERF model. If False, indicates pre-existing facility region_name - Name of region (state) tech_id - Technology ID tech_name - Full generation technology name inclusive of cooling type (if applicable) and additional characteristics tech_simple - Simplified generation technology type unit_size_mw - Power plant unit size (MW) xcoord - X coordinate in the default CRS (meters) ycoord - Y coordinate in the default CRS (meters) index - Index position in the flattend 2D array buffer_in_km - Exclusion buffer around site (km) sited_year - Year of siting retirement_year - Year of retirement lmp_zone - Locational marginal price (LMP) zone ID locational_marginal_price_usd_per_mwh - Locational marginal price ($/MWh) generation_mwh_per_year - Generation output (MWh/yr) operating_cost_usd_per_year - Cost of plant operations ($/yr) net_operational_value - Net operational value based on LMP and and operating costs ($/yr) interconnection_cost - Cost of interconnection for transmission & gas pipeline (if applicable) net_locational_cost -- Difference of interconnection cost and operating value ($/yr) capacity_factor_fraction - Capacity factor (fraction) carbon_capture_rate_fraction - Carbon capture rate (fraction) fuel_co2_content_tons_per_btu - Fuel CO2 content (tons/Btu) fuel_price_usd_per_mmbtu - Fuel price ($/MMBtu) fuel_price_esc_rate_fraction - Fuel price escalation rate (fraction) heat_rate_btu_per_kWh - Heat rate (Btu/kWh) lifetime_yrs - Technology lifetime for annuity (years) operational_life_yrs - Operational lifetime for retirement (years) variable_om_usd_per_mwh - Variable operation and maintenance costs of yearly capacity use ($/MWh) variable_om_esc_rate_fraction - Variable operation and maintenance costs escalation rate (fraction) carbon_tax_usd_per_ton - Carbon tax ($/ton) carbon_tax_esc_rate_fraction - Carbon tax escalation rate (fraction) 2. Storage Data This directory contains information on new and pre-existing energy storage facilities operational in each timestep along with various storage operational parameters. The 2015 timestep provides pre-existing energy storage data and corresponds with facilities that are operational leading into the first model simulation timestep. Note that coordinates in the storage files correspond to the interconnection point on the grid (substation location), not individual energy storage locations. Energy storage is added in a cumulative process at each given interconnection point. That is, each individual file provides the total operational storage capacity interconnected to the specified substation for the given timestep, inclusive of previously installed storage at that location and new storage installed in that timestep at that location. Parameters scenario - Name of scenario timestep - Simulation timestep name - Unique storage identifier s_typ - Type of energy storage technology (battery or pumped storage hydro) s_node - Node ID of interconnecting substation xcoord - X coordinate in the default CRS (meters) ycoord - Y coordinate in the default CRS (meters) charge_rate - Maximum charge rate (power capacity) of storage system (MW) discharge_rate - Maximum discharge rate (power capacity) of storage system (MW) duration - Duration of storage system (hours) max_SoC - Allowed maximum state of charge (energy capacity) of storage system (MWh) min_SoC -Allowed minimum state of charge (energy capacity) of storage system (MWh) charge_eff - Efficiency of charge (fraction between 0 and 1) discharge_eff - Efficiency of discharge (fraction between 0 and 1) Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program.

CERF↗

Automated Gold Nanorod Spectral Morphology Analysis Pipeline

The development of a colloidal synthesis procedure to produce nanomaterials with high shape and size purity is often a time-consuming, iterative process. This is often due to quantitative uncertainties in the required reaction conditions and the time, resources, and expertise intensive characterization methods required for quantitative determination of nanomaterial size and shape. Absorption spectroscopy is often the easiest method for colloidal nanomaterial characterization. However, due to the lack of a reliable method to extract nanoparticle shapes from absorption spectroscopy, it is generally treated as a more qualitative measure for metal nanoparticles. This work demonstrates a gold nanorod (AuNR) spectral morphology analysis tool, called AuNR-SMA, which is a fast and accurate method to extract quantitative structural information from colloidal AuNR absorption spectra. To demonstrate the practical utility of this model, we apply it to three distinct applications. First, we demonstrate this model's utility as an automated analysis tool in a high-throughput AuNR synthesis procedure by generating quantitative size information from optical spectra. Second, we use the predictions generated by this model to train a machine learning model to predict the resulting AuNR size distributions under specified reaction conditions. Third, we apply this model to spectra extracted from the literature where no size distributions are reported and impute unreported quantitative information on AuNR synthesis. This approach can potentially be extended to any other nanocrystal system where absorption spectra are size dependent, and accurate numerical simulation of absorption spectra is possible. In addition, this pipeline could be integrated into automated synthesis apparatuses to provide interpretable data from simple measurements, help explore the synthesis science of nanoparticles in a rational manner, or facilitate closed-loop workflows.

36 MATERIALS SCIENCE↗

Geometric GNNs for charged particle tracking at GlueX

Nuclear physics experiments are aimed at uncovering the fundamental building blocks of matter. The experiments involve high-energy collisions that produce complex events with many particle trajectories. Tracking charged particles resulting from collisions in the presence of a strong magnetic field is critical to enable the reconstruction of particle trajectories and precise determination of interactions. It is traditionally achieved through combinatorial approaches that scale worse than linearly as the number of hits grows. Since particle hit data naturally form a point cloud and can be structured as graphs, graph neural networks (GNNs) emerge as an intuitive and effective choice for this task. In this study, we evaluate the GNN model for track finding on the data from the GlueX experiment at Jefferson Lab. We use simulation data to train the model and test on both simulation and real GlueX measurements. We demonstrate that GNN-based track finding outperforms the currently used traditional method at GlueX in terms of segment-based efficiency at a fixed purity while providing faster inferences. We show that the GNN model can achieve significant speedup by processing multiple events in batches, which exploits the parallel computation capability of graphical processing units (GPUs). Finally, we compare the GNN implementation on GPU and field-programmable gate array and describe the trade-off.

batched GNN pipeline↗

Data for "Evaluating the industrial potential of emerging biomass pretreatment technologies in bioethanol production and lipid recovery from transgenic sugarcane"

The selection of pretreatment methods is critical to achieving high product yields during bioconversion of lignocellulosic biomass. Hydrothermal, soaking-in-aqueous ammonia, and ionic liquid pretreatment methods are viable candidates for minimizing sugar decomposition, permitting the effective hydrolysis of structural carbohydrates, and producing a fermentable substrate suitable for achieving industrial ethanol titers and yields. In this study, the effect of these three pretreatment methods on non-modified sugarcane cultivar CP88-1762 and two transgenic lipid-accumulating sugarcane lines, oilcane 1565 and oilcane 1566, were investigated and compared in terms of lipid recovery, sugar yield, and ethanol yields within the lignocellulosic biomass conversion pipeline. Fed-batch enzymatic hydrolysis at high solid loading yielded hydrolysates capable of supporting industrial bioethanol titers across all conditions. The highest sugar yields were obtained on ammonia-pretreated biomass hydrolysate (253.73 g L−1), followed by hydrothermally pretreated hydrolysate (213.10 g L−1) and ionic liquid-pretreated hydrolysate (154.20 g L−1). Commercially viable ethanol titers of 100.62 g L−1, 64.47 g L−1, and 52.95 g L−1 were achieved from ammonia, hydrothermal, and ionic liquid pretreated hydrolysate with the corresponding ethanol productivities of 2.08 g L−1 h−1, 0.53 g L−1 h−1, and 0.36 g L−1 h−1. The lower acetic acid concentration in ammonia-pretreated hydrolysate may have enhanced its fermentability relative to the hydrothermal pretreatment condition, as indicated by the differences in ethanol titer and productivity. Lower sugar yields and ethanol productivities under the ionic liquid conditions likely resulted from the inhibitory effect of cholinium lysinate. Oilcane 1565 and oilcane 1566 bagasse accumulated over 16- and 3 times higher lipids than the non-modified sugarcane CP88-1762. The total fatty acid content in the oilcane samples was reduced in ammonia and ionic liquid-pretreated bagasse relative to the hydrothermal pretreatment condition. While all pretreatment techniques tested are industrially viable, the observed differences in titer, productivity, and lipid content indicate that careful selection and validation of upstream processing methods can contribute to improved economic and environmental outcomes.

biomass analytics↗

Montane Conifer, Aspen, Meadow, and Sagebrush Metagenome Resolved Genomes and Traits in East River Watershed, Colorado, USA

Climate change is driving vegetation shifts in mountain watersheds, with unknown impacts on biogeochemical cycles. We hypothesize that these shifts will reshape soil microbiomes and associated biogeochemical processes. As a part of Lawrence Berkeley National Laboratory (LBNL) Watershed Science Focus Area (SFA), we assessed microbiome and microbial functional trait differences between soils under conifer, aspen, forby meadows, and sagebrush across the East River Watershed, CO, controlling for elevation and aspect.Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from soils 0-20cm in depth across three locations in the watershed—Headwaters, Upper Reaches, and Lower Reaches from August 3-11th 2016. Each location was further subdivided into two blocks, with one block on a west facing aspect, and two on the east aspect of the valley. Within blocks, two samples per vegetation type were taken (one at each depth). This resulted in 66 samples, which were sequenced at JGI and can be found under the Joint Genome Institute (JGI) Genomes Online Database (GOLD) sequencing project Gs0118068. Metagenomes were assembled through an inhouse pipeline (see methods), binned using four autobinners (concoct, maxbin2, metabat2, and vamb) and consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>75%) and contamination (<25%), and dereplicated at 95% ANI using drep. The dataset includes a zip file of 687 genomes (Vegtype_MAGS.zip), the accession numbers for the underlying metagenomes, a csv file with MAG quality metrics and taxonomy from Genome Taxonomy Database (GTDB) and National Center for Biotechnology Information (NCBI) taxonomic representative genome proteins (EastRiver_Vegtype_drep_genome_info.csv), and a file containing MAG quality metrics and taxonomy (gtdb_drep_bin_taxonomy.csv). The dataset additionally includes a sample metadata file (EastRiver_Vegtype_sample_metadata.csv), a metadata file used to register associated samples with IGSNs (International Generic Sample Numbers) (samples.csv), a Google KML file for the sampled locations (sample_collection_sites.kml), a location metadata file (locations.csv), a file-level metadata file (flmd.csv), and a data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Describing Point Defect Topology in 2D Energy Materials Through Computer Vision

Point defects such as vacancies and impurity atoms strongly impact the performance of 2D materials. Traditional efforts often rely on manual detection, a process that is time-intensive, prone to human error, and challenging to scale. Here we leverage machine learning (ML) methods to identify and quantify vacancies within 2D transition metal carbides (Ti3C2, MXenes), aiming to expedite detection while improving accuracy. MXenes exhibit valuable defect-defined electrochemical properties, but we currently lack statistical understanding of defect topology needed to fully harness these materials. Here we employ a convolutional neural network for semantic segmentation of experimental MXene images, opening an opportunity to conduct a rigorous statistical study on defect hierarchy while investigating local relaxation in the lattice. We show how the integration of ML can yield fundamental insight into point defects, providing a powerful tool that will play an increasingly crucial role in the future of materials science. ML is often not just a matter of straightforward application, and pretrained models proved ineffective in this case. Instead, we trained our own neural network (NN) and applied data augmentation techniques and fine-tuning to the training dataset. Since labeled microscopy data is often scarce, we developed training data from a previously published wide-frame MXene image, using customized Gaussian fitting to locate atomic positions. Our trained model was then applied to a large dataset of experimental images, enabling a statistical study of defect configurations across three samples prepared with different HF etchant concentrations (5%, 9.1%, and 12.5%), as shown in Fig. 1. This also allowed us to investigate local strain around vacancies, though we find that we are limited by the precision of measurements using high-angle annular dark field (HAADF) images, as shown in Fig. 2. This study demonstrates how ML enables large-scale, quantitative analysis of atomic defects - an otherwise infeasible task with traditional methods. While our NN was specialized for Ti3C2 MXenes, the pipeline we developed provides a foundation for future ML models tailored to other materials. Ultimately, we envision embedding the NN onto the microscope to give real-time feedback to the user. To make this a reality, continued work is necessary to fully understand the NN's capabilities and limitations. This study gets one step closer to our goals of automated experimentation moving away from traditional methods of manual labeling. As ML capabilities advance, we hope to continue adapting and applying these techniques in microscopy.

2D materials↗

Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection

Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection Description This dataset contains input and output data for the manuscript Mongird, K. et al. (under review) titled "Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection". Input data corresponds to gridded spatial siting attributes that are necessary to conduct a random forest machine learning analysis of siting feature importance. Output data includes SHAP feature analysis outputs, and classification report values. For data on power plant siting results referred to in the manuscript, please refer to the CERF: IM3 Projected Western US Power Plant Locations data download page. The downloadable data includes values for eight different future scenarios for the Western US. The scenarios include combinations of two Shared Socioeconomic Pathways (SSP3 and SSP5) with four high-resolution climate projections specific to the United States (see, https://tgw-data.msdlive.org/). These climate projections include "hotter" and "cooler" variants for two Representative Concentration Pathways (RCP4.5 and RCP8.5). The resulting eight simulations are: rcp45cooler_ssp3 rcp45cooler_ssp5 rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85cooler_ssp3 rcp85cooler_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 Technical Information The dataset includes two sets of data files: (1) CERF gridded siting parameters and (2) Feature analysis outputs and classification reports. All downloadable data is in csv file format. Files with x/y coordinate information use the Albers Equal Area Conic projection (ESRI:102003). 1. CERF Gridded Siting Parameters This directory provides a balanced sample of gridded CERF siting parameters data for eight different scenarios for the Western US through 2055, seven different technologies, and eight timesteps. This data serves as input to the feature analysis. It contains the following parameters. region_name - name of region (i.e., state) sited - binary value representing whether the grid cell received a siting of that technology type (1=True) rcp - binary value representing scenario resource concentration pathway (0 = RCP4.5, 1 = RCP8.5) ssp - binary value representing scenario shared socioeconomic pathway (0 = SSP3, 1 = SSP5) climate - binary value representing cooler (0) or hotter (1) GCM forcing tech_name - generation technology name sited_year - year that values correspond to transmission_cost - cost of transmission interconnection pipeline_cost - cost of natural gas pipeline interconnection interconnection_cost - total interconnection cost (sum of transmission cost and gas pipeline cost) lmp - associated locational marginal value ($/MWh) associated with the grid cell, timestep, scenario, and technology xcoord - x-coordinate of location ycoord - y-coordinate of location 2a. Feature Analysis Output The dataset includes the feature analysis shap output for locational marginal price and interconnection cost. It contains the following parameters. technology - generator technology name scenario - name of scenario feature - name of feature, either locational_marginal_price or interconnection_cost value - the mean of absolute value of SHAP values for given feature 2b. Feature Analysis Classification Report This download includes the classification report associated with each random forest model. The dataset contains the following parameters. technology - generation technology name scenario - name of scenario test - one of precision (the proportion of predicted positives that are actually correct), recall (the proportion of actual positives that were correctly identified), f1-score (the harmonic mean of precision and recall) 0.0 - value of test for classification of 0 (grid cell not chosen for siting) 1.0 - value of test for classification of 1 (grid cell chosen for siting) accuracy - accuracy of model (i.e., fraction of all predictions that were right) macro avg - Simple average of test values for all classes weighted avg - Weighted average of test values for all classes, weighted based on Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall [Pacific Northwest National Labor↗

Data-Driven Modeling and Control of Systems with Plasma-Surface Interactions (Final Technical Report)

This final technical report summarizes the activities and accomplishments in the period from February 2023 thru January 2026. The objective of the proposed research is to investigate the physical mechanisms and processes underlying the formation of structures and patterns in systems with plasma-surface interactions. In the past decades, there have been extensive studies on the interaction of glow discharges, dielectric barrier discharges, and arc discharges with confining or intervening surfaces. The advancement of the understanding of these phenomena is not only of fundamental scientific interest and relevance to the knowledge of the plasma state, but also with profound implications in various technological applications. The research will integrate theoretical, computational, and experimental work within an innovative framework of data assimilation, i.e., optimally combining model predictions with measurements. The scientific merit of this research has three aspects. Firstly, it extends the studies of plasma-surface interactions to systems with insulator surfaces and multi-layer systems, while existing studies are predominantly on electrode surfaces. Secondly, it expects to develop a novel data-driven modeling approach based on data assimilation to enhance the predictive and control capabilities, which could make transformative contributions to basic plasma research. Thirdly, it will shed new light on outstanding problems related to formation of patterns interfacing plasmas. This project also aims to launch an education and outreach initiative at Texas A&M University-Kingsville, a non-R1, minority-serving institution in South Texas. The initiative is structured as a four-tier pyramid. Tier one will be a webinar series for culture and capacity building to inform broader audience in the region about the research fields of plasma science and engineering. Tier two will be the creation and offering of an upper-level undergraduate course on introductory plasma physics, which will help with the recruitment for the upper tiers. On tier three, we will engage and mentor senior design students to conduct work toward the research goal of this project. There will also be a certificate program on general plasma science for undergrad and graduate students, part of which will be lab training at Princeton University. Tier four will be the supervision and mentoring of Ph.D. students. Therefore, this project will systematically expand the talent pipeline, broaden participation from communities historically and geographically underrepresented in DOE SC research portfolio, significantly improve the research and education capacity at the PI’s institution, and contribute to developing a diverse workforce in plasma science and engineering.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗