Search NASA⌕ Search

SEARCH · Search NASA

Results for “analysis and statistical methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Field To Farm Aggregation For Agricultural Systems

The Fields to Farms methodology illustrates the generation of farm parcels from the Crop Data Layer (CDL), a raster dataset containing 133 categories representing various crop types and land uses. This methodology involves two primary steps: Field Delineation and Farm Aggregation. The code specifically addresses the aggregation of pre-delineated fields within a county to form farms, adhering to predefined criteria for farm size categories. It is assumed that the field delineation process, which involves creating vector polygons from CDL raster, has been completed beforehand, possibly through external tools or methods. Upon initialization, the script processes county-level fields, preparing them for farm aggregation. In the Farm Aggregation phase, the code iteratively combines delineated fields into farms based on specified criteria, continuing until the aggregated farm size meets predefined thresholds derived from data from the 2017 National Agricultural Statistics Service (NASS) census. Throughout this iterative process, the script dynamically adjusts the aggregation to ensure alignment with the desired distribution reported by NASS. The resulting output of the script is a GeoDataFrame containing classified farms, which are subsequently saved as GeoPackage files. These files enable further analysis and visualization, facilitating comprehensive exploration of the farm landscape generated through the methodology.

Paudel, Rajiv [Idaho National Laboratory (INL), Id↗

Extracting Topological Orders of Generalized Pauli Stabilizer Codes in Two Dimensions

In this paper, we introduce an algorithm for extracting topological data from translation invariant generalized Pauli stabilizer codes in two-dimensional systems, focusing on the analysis of anyon excitations and string operators. The algorithm applies to Z d qudits, including instances where d is a nonprime number. This capability allows the identification of topological orders that differ from the Z d toric codes. It extends our understanding beyond the established theorem that Pauli stabilizer codes for Z p qudits (with p being a prime) are equivalent to finite copies of Z p toric codes and trivial stabilizers. The algorithm is designed to determine all anyons and their string operators, enabling the computation of their fusion rules, topological spins, and braiding statistics. The method converts the identification of topological orders into computational tasks, including Gaussian elimination, the Hermite normal form, and the Smith normal form of truncated Laurent polynomials. Furthermore, the algorithm provides a systematic approach for studying quantum error-correcting codes. We apply it to various codes, such as self-dual CSS quantum codes modified from the two-dimensional honeycomb color code and non-CSS quantum codes that contain the double semion topological order or the six-semion topological order. Published by the American Physical Society 2024

Physics↗

Improved Subseasonal Forecasting of Extreme Polar Vortices Using Machine Learning

Our research was focused on forecasting the position and shape of the winter stratospheric polar vortex at a subseasonal timescale of 15 days in advance. To achieve this, we employed both statistical and neural network machine learning techniques. The analysis was performed on 42 winter seasons of reanalysis data provided by NASA giving us a total of 6,342 days of data. The state of the polar vortex for determined by using geometric moments to calculate the centroid latitude and the aspect ratio of an ellipse fit onto the vortex. Timeseries for thirty additional precursors were calculated to help improve the predictive capabilities of the algorithm. Feature importance of these precursors was performed using random forest to measure the predictive importance and the ideal number of precursors. Then, using the precursors identified as important, various statistical methods were tested for predictive accuracy with random forest and nearest neighbor performing the best. An echo state network, a type of recurrent neural network that features sparsely connected hidden layer and a reduced number of trainable parameters that allows for rapid training and testing, was also implemented for the forecasting problem. Hyperparameter tuning was performed for each methods using a subset of the training data. The algorithms were trained and tuned on the first 41 years of data, then tested for accuracy on the final year. In general, the centroid latitude of the polar vortex proved easier to predict than the aspect ratio across all algorithms. Random forest outperformed other statistical forecasting algorithms overall but struggled to predict extreme values. Forecasting from echo state network suggested a strong predictive capability past 15 days, but further work is required to fully realize the potential of recurrent neural network approaches.

54 ENVIRONMENTAL SCIENCES↗

Inference of the linear matter power spectrum at z = 0 using DESI DR1 Full-Shape data

Measurements of galaxy distributions at large cosmic distances capture clustering from the past. In this study, we use a cosmological model to translate these observations into the present-day galaxy distribution. Specifically, we reconstruct the 3D linear matter power spectrum at redshift z = 0 using Dark Energy Spectroscopic Instrument (DESI) Year 1 (DR1) galaxy clustering data and Cosmic Microwave Background (CMB) observations, assuming the ΛCDM model, and compare it to the result assuming the w 0 w a CDM model. Building on previous state-of-the-art methods, we apply Effective Field Theory (EFT) modelling of the galaxy power spectrum to account for small-scale effects in the 2-point statistics of galaxy data. Implementation of the EFT approach improves the modelling of the galaxy power spectrum, providing a more robust consistency test of the assumed cosmological model. By casting both CMB and galaxy clustering observations, spanning distinct redshift regimes, into k-space, we can identify discrepancies between the datasets of different redshifts, which would indicate potential inaccuracies in the assumed expansion history. While previous studies have shown consistency with ΛCDM, this work extends the analysis with higher-quality data to further test the expansion histories of both ΛCDM and w 0 w a CDM. Our findings show that both ΛCDM and w 0 w a CDM provide consistent fits to the linear matter power spectrum recovered from DESI DR1 data.

cosmological parameters from LSS↗

Development and implementation of high-throughput proteomic and metabolomics assays by using advanced chromatographic and mass spectrometric systems (CRADA Final Report)

The mission of this CRADA with Agilent was to couple powerful MS platforms (QQQ, IM-QTOFMS) with Agilent’s novel Ultra-High-Performance Liquid Chromatography (UHPLC) fast metabolomic workflows and perform ABF Machine Learning (ML) to generated datasets. Agilent transferred UHPLC methods to PNNL and LBNL and methods were implemented and demonstrated in both labs, achieving total acquisition times of < 10 min. Metabolites analyzed using Agilent’s shared methods included metabolites from central carbon metabolism, common across hosts, and metabolites unique to engineered strains. Standards were acquired in an UHPLC-Drift Tube Ion Mobility Mass Spectrometer (DTIMS) system for the first time within the context of ABF and methods were optimized based on Agilent’s protocols. Samples from ABF hosts Pseudomonas putida, Aspergillus pseudoterreus, Aspergillus niger and Rhodosporidium toruloides were analyzed using the UHPLC-DTIMS platform for a total of 276 runs. A data analysis workflow compatible with the Experimental Data Depot (EDD) and completely shareable was developed for the acquired UHPLC-DTIMS data. Samples were analyzed using a Data Independent Acquisition Approach (DIA), which for most of the standards provided more transitions therefore increasing detection confidence. Using the data acquired by PNNL, LBNL, and Agilent’s specifications from previous ML projects, SNL applied an ensemble ML strategy to pick the best performing model for automated LC-method selection. Finally, with the contribution of the participant labs and Agilent, SNL developed an Automated Method Selection (AMS) software tool to predict the best liquid chromatography method for analysis of any new molecules of interest. Samples with novel pathways and new metabolite targets of interest are generated at a high pace in the ABF. Overall, the project advanced rapid metabolomics by combining liquid chromatography, ion mobility spectrometry, and data-independent mass spectrometry with machine learning. This multidimensional approach uses retention time, collision cross-section, precursor mass, and fragment-ion information to distinguish chemically similar metabolites that can be difficult to resolve using conventional liquid- or gas-chromatography methods. The resulting workflow also provided automated metabolite-identification error estimates, addressing a recognized need for statistical confidence measures in metabolomics.

Petzold, Christopher [Lawrence Berkeley National L↗

Machine Learning-Based Anomaly Detection for PMT Data Quality Monitoring in the SBN and DUNE

Maintaining high-quality detector data is essential for achieving the scientific objectives of the Short-Baseline Neutrino (SBN) Program at Fermilab. Current data quality monitoring (DQM) procedures rely primarily on threshold-based metrics and manual inspection of detector monitoring plots, making the detection of subtle or gradually developing anomalies both time-consuming and dependent on expert interpretation. This project developed and evaluated a machine-learning workflow for automatically identifying anomalous photomultiplier tube (PMT) channels in the Short-Baseline Near Detector (SBND) using optical-hit amplitude data. A Python-based analysis program was developed to process ROOT files, extract statistical features describing individual PMT amplitude distributions, and generate feature vectors for anomaly detection. These features were used to train an Isolation Forest model using data representing normal detector operation. The trained model was subsequently applied to independent detector runs to identify channels exhibiting statistically unusual behavior relative to the learned reference response. To support expert interpretation, the workflow generated complementary diagnostic products, including anomaly score distributions, normalized amplitude comparisons, decision-tree visualizations, and principal component analysis (PCA) projections. This project demonstrated the feasibility of integrating unsupervised machine learning into detector data-quality monitoring and developed a complete workflow for automated PMT performance assessment to aid expert-driven review. Beyond its technical contributions, the VFP appointment fostered a research collaboration between Aurora University and Fermilab and provided direct workforce development benefits by training the visiting faculty member in detector-scale machine-learning methods that are now being incorporated into undergraduate coursework and research. The methodology developed here provides a foundation for future applications to ProtoDUNE and other liquid argon time projection chamber (LArTPC) detectors, contributing to ongoing efforts to improve detector reliability, reduce manual monitoring requirements, and enable scalable data quality monitoring for future large-scale neutrino experiments, including the Deep Underground Neutrino Experiment (DUNE).

Colón Santana, Juan A. [Unlisted, US, IL]↗

Search for muon neutrino disappearance with multi-topologically splitted Inclusive events at SBN

The Short-Baseline Neutrino (SBN) program at Fermilab searches for signatures of sterile neutrinos with mass-squared splittings at the O(1) eV2 scale, motivated by anomalies previously reported by the LSND and MiniBooNE experiments. We present a search for such signatures using the muon neutrino disappearance channel. This analysis exploits the large statistics of fully inclusive charged-current interactions collected by the SBND and ICARUS detectors to maximize sensitivity. The inclusive samples in both detectors are further subdivided into multiple exclusive topological channels, a strategy designed to enhance sensitivity by isolating differences in interaction-mode composition and reconstructed energy response. This multi-topology approach represents a novel analysis technique within SBN, enabling improved constraints on systematic uncertainties arising from neutrino flux, interaction modeling, and detector response. We will present comparisons of data with detailed predictions incorporating a comprehensive treatment of systematic uncertainties and constraints on the systematics obtained by the multi-topology method.

Yadav, Shweta [Texas U., Arlington]↗

SCALE inventory and reactivity analysis as part of the Hermes 2021 PSAR review

The readiness of SCALE for comprehensive studies of pebble-bed reactors has been demonstrated through detailed analysis of a fluoride salt–cooled, high-temperature pebble-bed reactor (PB-FHR). The methods developed for pebble-bed reactor modeling in SCALE, particularly for inventory generation, have proven effective in gaining insights into the reactor physics of this advanced reactor. Excellent agreement with another code package has been observed, further highlighting SCALE’s strong performance. The SCALE results supported the US Nuclear Regulatory Commission’s construction permit application review of the Hermes low-power PB-FHR demonstration reactor. A SCALE model of the Hermes reactor was developed at Oak Ridge National Laboratory using information from the Preliminary Safety Analysis Report (PSAR) and supplemented with publicly available data. SCALE reactivity coefficient simulations reproduced PSAR results within 1σ statistical uncertainties. Sensitivity studies emphasized the importance of graphite specifications for accurate keff predictions.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Machine Learning Approaches to Predicting Induced Seismicity and Imaging Geothermal Reservoir Properties

This project developed machine learning (ML) methods, lab data sets, and field data to advance geothermal exploration and geothermal energy production. The work had three focus areas. One involved the development of ML methods to use microearthquakes (MEQs) for imaging geothermal reservoir properties and improving subsurface characterization – most importantly the evolution of permeability within the evolving reservoir. This part of the work included development of ML approaches for automated MEQ location, focal mechanism determination and identification of earthquake precursors. The second area focused on using MEQ signals generated by geothermal exploration and production to predict the relationship between fluid injection and seismicity. Here, we extended to reservoir scale our success in using ML to predict laboratory earthquakes and fault zone stress state. The third focus area was on lab experiments. Here, we developed new ML models for lab earthquake prediction and identification of precursors to failure to improve earthquake forecasting and early warning in geothermal settings. Major outcomes of our work include ML models that learn from MEQ signals during geothermal exploration and production to predict induced seismicity. MEQs occur naturally in connection with drilling and energy production. We developed ML methods to use the seismic waves from these events to characterize the elastic, hydraulic and poromechanical properties of reservoirs. Our work illuminated fracture geometry and the evolution of fracture permeability by incorporating seismic coda wave analysis and ML methods to relate fluid injection and seismicity. We significantly expanded laboratory earthquake prediction to include methods that use both passive measurements of microearthquakes within the lab fault zones and also active source acoustic measurements of fault zone elastic properties. These methods can now predict fault zone stress state, time to failure and the magnitude of lab earthquakes. Our work showed that repetitive stick- slip failure events during frictional sliding (the lab equivalent of earthquakes) are preceded by a cascade of micro-failure events that radiate energy in a manner that foretells unstable failure – manifest as laboratory MEQs. We documented a mapping between fracture properties and statistical attributes of elastic radiation. We extended existing works to geothermal reservoir scale and developed ML methods to determine reservoir permeability, fracture properties, and their evolution during geothermal energy production. An attractive feature of ML algorithms is their ability to handle big datasets and reveal patterns and correlations that may remain invisible to conventional analyses. Our work connected data from field, laboratory and intermediate scales to study permeability, stress, strength, fracture stiffness and geometry. At the field scale we used data from the Newberry Volcano field site, UtahFORGE, EGS Collab, and also the Bedretto underground research lab in Switzerland. These data sets are bridging the gap between the lab scale, theory, and reservoir scale. Our work produced plain language summaries to improve public understanding of DOE research. We also developed openly distributed ML and seismicity datasets for use by all researchers and we published connections between induced seismicity in geothermal areas and reservoir properties including permeability, fracture properties, and stress state. Our models are designed for the large data sets of induced seismicity typically associated with geothermal sites. We produced labeled event catalogs and used them on geothermal data to assess how ML can facilitate geothermal production and exploration. All datasets are available on the GDR Productivity: The project produced 32 publications in peer reviewed journals (two are in review). It supported the work of 6 PhD students, 40 conference presentations, 6 keynote talks at national meetings, and mentoring and professional development for 4 postdoctoral fellows.

15 GEOTHERMAL ENERGY↗

Lanczos algorithm for lattice QCD matrix elements

Recent work [M. L. Wagman, Lanczos, the transfer matrix, and the signal-to-noise problem, .] found that an analysis formalism based on the Lanczos algorithm allows energy levels to be extracted from Euclidean correlation functions with faster ground-state convergence than effective masses, convergent estimators for multiple states from a single correlator, and two-sided error bounds. After filtering out spurious eigenvalues and using outlier-robust estimators within a nested bootstrap framework, Lanczos estimators behave more like multistate fit results than effective masses—but without involving statistical fitting. We extend this formalism to the determination of matrix elements from three-point correlation functions and provide a physical picture of “spurious-state filtering” involving restriction to a Hermitian subspace. We demonstrate similar advantages for matrix elements as for spectroscopy through example applications to noiseless mock-data and (bare) forward matrix elements of the strange scalar current between both ground and excited states with the quantum numbers of the nucleon.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Density estimation via measure transport: Outlook for applications in the biological sciences

Abstract One among several advantages of measure transport methods is that they allow or a unified framework for processing and analysis of data distributed according to a wide class of probability measures. Within this context, we present results from computational studies aimed at assessing the potential of measure transport techniques, specifically, the use of triangular transport maps, as part of a workflow intended to support research in the biological sciences. Scenarios characterized by the availability of limited amount of sample data, which are common in domains such as radiation biology, are of particular interest. We find that when estimating a distribution density function given limited amount of sample data, adaptive transport maps are advantageous. In particular, statistics gathered from computing series of adaptive transport maps, trained on a series of randomly chosen subsets of the set of available data samples, leads to uncovering information hidden in the data. As a result, in the radiation biology application considered here, this approach provides a tool for generating hypotheses about gene relationships and their dynamics under radiation exposure.

gene expression data↗

The Profiled Feldman-Cousins Method for Confidence Interval Construction for the Nova 3-Flavor Oscillation Analysis

The small interaction cross-section of neutrinos makes experimental neutrino physics particularly responsive to technological advancements. A significant development leveraged by the NOvA experiment is large-scale parallel processing, enabling novel computational approaches to longstanding experimental challenges. Central to managing the resulting high-throughput data is NOvA’s implementation of the Freight Train model, designed for efficient data production and handling.This dissertation details the methodology and execution of the NOvA 2024 3-Flavor Oscillation Analysis, supported by a comprehensive dataset spanning ten years. It emphasizes frequentist results refined through the Feldman-Cousins (FC) technique, specifically addressing confidence interval corrections in parameter estimation. The computational intensity associated with Feldman-Cousins arises from extensive Monte Carlo simulations, which were substantially mitigated through parallel computing on the Perlmutter supercomputer at the National Energy Research Scientific Computing Center (NERSC), employing the MPI framework.To further enhance computational efficiency, an Importance Sampling method is introduced and evaluated, demonstrating significant potential to reduce complexity, particularly in exploring extreme parameter space regions. This thesis presents both the successful application of advanced computational resources and the development of sophisticated statistical techniques, aiming to enhance the precision and scope of neutrino oscillation analyses.

Dye ajdye11190@gmail.com, Andrew Joseph [Mississip↗

Velocity reconstruction in the era of DESI and Rubin/LSST. I. Exploring spectroscopic, photometric, and hybrid samples

Peculiar velocities of galaxies and halos can be reconstructed from their spatial distribution alone. This technique is analogous to the baryon acoustic oscillations reconstruction, using the continuity equation to connect density and velocity fields. The resulting reconstructed velocities can be used to measure imprints of galaxy velocities on the cosmic microwave background like the kinematic Sunyaev-Zel’dovich effect or the moving lens effect. As the precision of these measurements increases, characterizing the performance of the velocity reconstruction becomes crucial to allow unbiased and statistically optimal inference. In this paper, we quantify the relevant performance metrics: the variance of the reconstructed velocities and their correlation coefficient with the true velocities. We show that the relevant velocities to reconstruct for kSZ and moving lens are actually the halo—rather than galaxy—velocities. We quantify the impact of redshift-space distortions, photometric redshift errors, satellite galaxy fraction, incorrect cosmological parameter assumptions and smoothing scale on the reconstruction performance. Here, we also investigate hybrid reconstruction methods, where velocities inferred from spectroscopic samples are evaluated at the positions of denser photometric samples. We find that using exclusively the photometric sample is better than performing a hybrid analysis. The 2 Gpc/ℎ length simulations from abacussummit with realistic galaxy samples for DESI and Rubin LSST allow us to perform this analysis in a controlled setting. In the companion paper [B. Hadzhiyska, S. Ferraro, B. Ried Guachalla, and E. Schaan, companion paper, Phys. Rev. D 109, 103534 (2024).], we further include the effects of evolution along the light cone and give realistic performance estimates for DESI luminous red galaxies, emission line galaxies, and Rubin LSST-like samples.

79 ASTRONOMY AND ASTROPHYSICS↗

Increased inflammation as well as decreased endoplasmic reticulum stress and translation differentiate pancreatic islets from donors with pre-symptomatic stage 1 type 1 diabetes and non-diabetic donors

Aims/hypothesis Progression to type 1 diabetes is associated with genetic factors, the presence of autoantibodies and a decline in beta cell insulin secretion in response to glucose. Very little is known regarding the molecular changes that occur in human insulin-secreting beta cells prior to the onset of type 1 diabetes. Herein, we applied an unbiased proteomics approach to identify changes in proteins and potential mechanisms of islet dysfunction in islet-autoantibody-positive organ donors with pre-symptomatic stage 1 type 1 diabetes (HbA1c ≤42 mmol/mol [6.0%]). We aimed to identify pathways in islets that are indicative of beta cell dysfunction. Methods Multiple islet sections were collected through laser microdissection of frozen pancreatic tissues from organ donors positive for single or multiple islet autoantibodies (AAb + , n=5), and age (±2 years)- and sex-matched non-diabetic (ND) control donors (n=5) obtained from the Network for Pancreatic Organ donors with Diabetes (nPOD). Islet sections were subjected to MS-based proteomics and analysed with label-free quantification followed by pathway and functional annotations. Results Analyses resulted in ~4500 proteins identified with low false discovery rate (<1%), with 2165 proteins reliably quantified in every islet sample. We observed large inter-donor variations that presented a challenge for statistical analysis of proteome changes between donor groups. We therefore focused on only the donors with stage 1 type 1 diabetes who were positive for multiple autoantibodies (mAAb + , n=3) and genetic risk compared with their matched ND controls (n=3) for the final statistical analysis. Approximately 10% of the proteins (n=202) were significantly different (unadjusted p<0.025, q<0.15) for mAAb + vs ND donor islets. The significant alterations clustered around major functions for upregulation in the immune response and glycolysis, and downregulation in endoplasmic reticulum (ER) stress response as well as protein translation and synthesis. The observed proteome changes were further supported by several independent published datasets, including a proteomics dataset from in vitro proinflammatory cytokine-treated human islets and single-cell RNA-seq datasets from AAb + individuals. Conclusions/interpretation In situ human islet proteome alterations in stage 1 type 1 diabetes centred around several major functional categories, including an expected increase in immune response genes (elevated antigen presentation/HLA), with decreases in protein synthesis and ER stress response, as well as compensatory metabolic response. The dataset serves as a proteomics resource for future studies on beta cell changes during type 1 diabetes progression and pathogenesis. Data availability The LC-MS raw datasets that support the findings of this study have been deposited in the online repository: MassIVE (https://massive.ucsd.edu/ProteoSAFe/static/massive.jsp) with accession no. MSV000090212.

Autoantibody-positive↗

Mapping the structural–mechanical landscape of amorphous carbon with ReaxFF molecular dynamics

We use ReaxFF molecular dynamics (MD) to investigate the relationship between structural and mechanical properties in bulk and nanostructured amorphous carbon (a-C). The liquid-quench MD method is used to generate isotropic bulk samples with mass densities ranging from 0.96 to 3.29 g/cm3. Structural analysis identifies two types of structures with distinct short- and medium-range order: lower-density sp2-dominated a-C, which is characterized by a bimodal ring-size distribution, and higher-density sp3-dominated tetrahedral amorphous carbon (ta-C), exhibiting a unimodal ring-size distribution. Stress–strain MD simulations and analysis reveal how an atomistic structure impacts elastic properties and post-yield atomic rearrangements. All stretched structures demonstrate elastic isotropy and plasticity driven by a ring-size expansion mechanism reflected in changes in ring statistics. The plastic region is substantially larger in ta-C than in a-C due to the post-yield shift from sp3 to sp2 C dominant bonding. In both a-C and ta-C, ultimate failure occurs when a reactive crack, traversed by long sp chains, forms and propagates predominantly perpendicular to the direction of the applied strain. Oxygen infiltration into the fractured region significantly reduces stress resistance, primarily through the early rupture of long sp chains. MD simulations and analysis are extended to a-C slabs, a-C nanotubes, and partially a-C nanotubes. The latter nanostructure highlights the differences between the elastically isotropic a-C walls, which develop circumferential cracking, and the crystalline walls, which tear along crystallographic directions. These results provide a strong foundation for further computational characterization of a-C materials.

Dernov, A. (ORCID:0009000004220973)↗

Selection Algorithm Improvement for MicroBooNE

Data selection is an extremely important part of data analysis for any experiment. Finding a physics result is often the result of sifting through a massive amount of data, keeping data that we believe to be signal and throwing out data we do not. This process is called data selection. Creating a selection algorithm is an intensive process that must balance keeping enough data to have statistics and maximizing the signal purity of that data. In this study, we used three different reconstruction tools, Pandora, WireCell, and LANTERN, for the MicroBooNE experiment in conjunction to improve the selection algorithm for analysis. For the case of this study, we look into the charged current N proton 0 pions (CCNp0$\pi$) interaction channel. This is the dominant channel for the Short Baseline Neutrino (SBN) program and is expected to be a large contributor to the Deep Underground Neutrino Experiment (DUNE). We first investigated each of the three tools to find out more about their strengths and weaknesses as reconstructions. We then put together a direct comparison of the three methods to find which method or combination of methods would return the best result for us. While the study is ongoing, we have learned a lot about data selection for the experiment and the differences between the reconstruction tools.

Dillon, Brayden [Michigan State U.]↗

Comparative study of machine learning techniques for post-combustion carbon capture systems

Computational analysis of countercurrent flows in packed absorption columns, often used in solvent-based post-combustion carbon capture systems (CCSs), is challenging. Typically, computational fluid dynamics (CFD) approaches are used to simulate the interactions between a solvent, gas, and column's packing geometry while accounting for the thermodynamics, kinetics, heat, and mass transfer effects of the absorption process. These simulations can then be used explain a column's hydrodynamic characteristics and evaluate its CO 2 -capture efficiency. However, these approaches are computationally expensive, making it difficult to evaluate numerous designs and operating conditions to improve efficiency at industrial scales. In this work, we comprehensively explore the application of statistical ML methods, convolutional neural networks (CNNs), and graph neural networks (GNNs) to aid and accelerate the scale-up and design optimization of solvent-based post-combustion CCSs. We apply these methods to CFD datasets of countercurrent flows in absorption columns with structured packings characterized by several geometric parameters. We train models to use these parameters, inlet velocity conditions, and other model-specific representations of the column to estimate key determinants of CO 2 -capture efficiency without having to simulate additional CFD datasets. We also evaluate the impact of different input types on the accuracy and generalizability of each model. We discuss the strengths and limitations of each approach to further elucidate the role of CNNs, GNNs, and other machine learning approaches for CO 2 -capture property prediction and design optimization.

97 MATHEMATICS AND COMPUTING↗

A Parameter-masked Mock Data Challenge for Beyond-two-point Galaxy Clustering Statistics

The past few years have seen the emergence of a wide array of novel techniques for analyzing high-precision data from upcoming galaxy surveys, which aim to extend the statistical analysis of galaxy clustering data beyond the linear regime and the canonical two-point (2pt) statistics. We test and benchmark some of these new techniques in a community data challenge named “Beyond-2pt,” initiated during the Aspen 2022 Summer Program “Large-Scale Structure Cosmology beyond 2-Point Statistics,” whose first round of results we present here. The challenge data set consists of high-precision mock galaxy catalogs for clustering in real space, in redshift space, and on a light cone. Participants in the challenge have developed end-to-end pipelines to analyze mock catalogs and extract unknown (“masked”) cosmological parameters of the underlying ΛCDM models with their methods. The methods represented are density-split clustering, nearest neighbor statistics, BACCO power spectrum emulator, void statistics, LEFTfield field-level inference using effective field theory (EFT), and joint power spectrum and bispectrum analyses using both EFT and simulation-based inference. In this work, we review the results of the challenge, focusing on problems solved, lessons learned, and future research needed to perfect the emerging beyond-2pt approaches. The unbiased parameter recovery demonstrated in this challenge by multiple statistics and the associated modeling and inference frameworks supports the credibility of cosmology constraints from these methods. The challenge data set is publicly available, and we welcome future submissions from methods that are not yet represented.

Krause, Elisabeth [Univ. of Arizona, Tucson, AZ (U↗