Search NASA⌕ Search

SEARCH · Search NASA

Results for “massive data set analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Compactly‐Supported Nonstationary Kernels for Computing Exact Gaussian Processes on Big Data

The Gaussian process (GP) is a widely used method for analyzing large-scale data sets, including spatio-temporal measurements of nonlinear processes that are now commonplace in the environmental sciences. Traditional implementations of GPs involve stationary kernels (also termed covariance functions) that limit their flexibility, and exact methods for inference that prevent application to data sets with more than about 10,000 points. Modern approaches to address stationarity assumptions generally fail to accommodate large data sets, while all attempts to address scalability focus on approximating the Gaussian likelihood, which can involve subjectivity and lead to inaccuracies. In this work, we explicitly derive an alternative kernel that can discover and encode both sparsity and nonstationarity. We embed the kernel within a fully Bayesian GP model and leverage high-performance computing resources to enable the analysis of massive data sets. We demonstrate the favorable performance of our novel kernel relative to existing exact and approximate GP methods across a variety of synthetic data examples. Furthermore, we conduct space–time prediction based on more than 1 million measurements of daily maximum temperature and verify that our results outperform state-of-the-art methods in the Earth sciences. More broadly, having access to exact GPs that use ultra-scalable, sparsity-discovering, nonstationary kernels allows GP methods to truly compete with a wide variety of machine learning methods.

Gaussian processes↗

UMap: An application-oriented user level memory mapping library

Exploiting the prominent role of complex memories in exascale node architecture, the UMap page fault handler offers new capabilities to access large memory-mapped data sets directly. UMap provides flexible configuration options to customize page handling to each application, including analysis of massive observational and simulation data sets. The high-performance design features I/O decoupling, dynamic load balancing, and application-level controls. Page faults triggered by application threads and processes accessing data mapped to a UMapp’ed region are handled via the Linux userfaultfd protocol, an asynchronous message-oriented kernel-user communication mechanism that avoids the context switch penalty of traditional signal fault handlers. UMap is fully open source. In this paper, we give an overview of the UMap library architecture, its extensible plugin architecture, and the use/performance of UMap in emerging heterogeneous memory hierarchies such as near-node Non-volatile Memory (NVM) and network attached memories. We highlight new capabilities in two pagefault management plugins, the NetworkStore and SparseStore. We demonstrate the integration between UMap and multiple ECP products including Caliper, Metall, ZFP, Mochi, and Ripples.

97 MATHEMATICS AND COMPUTING↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗

Sensitivity to low-mass WIMPs with an improved liquid argon ionization response model within the DarkSide program

Dark matter detection experiments using liquid argon rely on a precise characterization of the ionization response to nuclear recoils, especially in the keV energy range relevant for light dark matter interactions. In this work, we present a comprehensive analysis that combines new measurements from the ReD setup, part of the DarkSide experimental program, with calibration data from DarkSide-50, as well as results from the ARIS and SCENE experiments. These combined datasets enable improved constraints on atomic screening effects in the modeling of the ionization response of liquid argon to nuclear recoils. The analysis is performed within the Thomas-Imel recombination framework adopted in previous DarkSide studies, and is here further constrained by the inclusion of ReD data, which allow the screening function to be determined from calibration measurements. By including the updated ionization model into the DarkSide-50 analysis framework, we obtain stronger exclusion limits on low-mass weakly interacting massive particle (WIMP) interactions, setting new world-leading constraints in the 1 – 3 GeV / c 2 WIMP mass range. Finally, we recast the sensitivity projections for the upcoming DarkSide-20k detector, demonstrating a significantly enhanced discovery potential for low-mass dark matter candidates.

Acerbi, F. [Fond. Bruno Kessler, Trento]↗

Comparing Compressed and Full-Modeling analyses with FOLPS: implications for DESI 2024 and beyond

The Dark Energy Spectroscopic Instrument (DESI) will provide unprecedented information about the large-scale structure of our Universe. In this work, we study the robustness of the theoretical modelling of the power spectrum of F OLPS , a novel effective field theory-based package for evaluating the redshift space power spectrum in the presence of massive neutrinos. We perform this validation by fitting the AbacusSummit high-accuracy N -body simulations for Luminous Red Galaxies, Emission Line Galaxies and Quasar tracers, calibrated to describe DESI observations. We quantify the potential systematic error budget of F OLPS finding that the modelling errors are fully sub-dominant for the DESI statistical precision within the studied range of scales. Additionally, we study two complementary approaches to fit and analyse the power spectrum data, one based on direct Full-Modelling fits and the other on the ShapeFit compression variables, both resulting in very good agreement in precision and accuracy. In each of these approaches, we study a set of potential systematic errors induced by several assumptions, such as the choice of template cosmology, the effect of prior choice in the nuisance parameters of the model, or the range of scales used in the analysis. Furthermore, we show how opening up the parameter space beyond the vanilla ΛCDM model affects the DESI observables. These studies include the addition of massive neutrinos, spatial curvature, and dark energy equation of state. We also examine how relaxing the usual Cosmic Microwave Background and Big Bang Nucleosynthesis priors on the primordial spectral index and the baryonic matter abundance, respectively, impacts the inference on the rest of the parameters of interest. This paper pathways towards performing a robust and reliable analysis of the shape of the power spectrum of DESI galaxy and quasar clustering using F OLPS .

79 ASTRONOMY AND ASTROPHYSICS↗

A Parametric, Data-Driven, Non-Intrusive Reduced-Order Model Framework for Crystal Plasticity Simulations of Voids

The influence of the internal structure at micrometer length scales on the deformation of polycrystalline materials can be effectively captured using crystal plasticity finite element methods (CPFEM). However, the complexity and nonlinearity of the deformation equations CPFEM solves demand significant computational power and resources to achieve accurate predictions, limiting its broader application. To address this challenge, we have identified a reduced-order representation of the complex data in order to establish a computationally efficient reduced-order models (ROM) and drastically reduce the computational expense of CPFEM. Specifically, in this work, we developed a parametric, data-driven, and non-intrusive ROM framework for CPFEM using proper orthogonal decomposition (POD) and sparse variational Gaussian process (SVGP) regression for single-crystal microstructures under tensile loading conditions. The developed protocol enables one to compress field into a latent/low-dimensional space described by principal component analysis (PCA) via the singular value decomposition (SVD) algorithm. As a result, the high-dimensional data are reduced to a significantly smaller amount of dimensions with POD bases and POD coefficients. Furthermore, we deployed an ensemble of SVGPs—extended from the classical Gaussian process (GP) regression for scalability and handling big data—in a massively parallel manner to train and predict latent POD coefficients using known POD bases from a set of previously obtained simulations results. Lastly, using the predicted POD coefficients, we reconstructed the full-field results and showed reasonable agreement compared with the true values obtained from running CPFEM. The developed framework is validated with a set of CPFEM simulations of a single embedded void in single-crystal aluminum alloy. While the framework is broadly applicable, this work specifically focuses on single-crystal microstructures, a single load case (e.g., tensile), and a specific void geometry (spherical).

Anisotropy↗

Search for a scalar or pseudoscalar dilepton resonance produced in association with a massive vector boson or top quark-antiquark pair in multilepton events at s = 13 TeV

A search for beyond the standard model spin-0 bosons, ϕ , that decay into pairs of electrons, muons, or tau leptons is presented. The search targets the associated production of such bosons with a W or Z gauge boson, or a top quark-antiquark pair, and uses events with three or four charged leptons, including hadronically decaying tau leptons. The proton-proton collision data set used in the analysis was collected at the LHC from 2016 to 2018 at a center-of-mass energy of 13 TeV, and corresponds to an integrated luminosity of 138 fb - 1 . The observations are consistent with the predictions from standard model processes. Upper limits are placed on the product of cross sections and branching fractions of such new particles over the mass range of 15 to 350 GeV with scalar, pseudoscalar, or Higgs-boson-like couplings, as well as on the product of coupling parameters and branching fractions. Several model-dependent exclusion limits are also presented. For a Higgs-boson-like ϕ model, limits are set on the mixing angle of the Higgs boson with the ϕ boson. For the associated production of a ϕ boson with a top quark-antiquark pair, limits are set on the coupling to top quarks. Finally, limits are set for the first time on a fermiophilic dilaton-like model with scalar couplings and a fermiophilic axion-like model with pseudoscalar couplings.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

L-PBF High-Throughput Data Pipeline Approach for Multi-modal Integration

Abstract Metal-based additive manufacturing requires active monitoring solutions for assessing part quality. Multiple sensors and data streams, however, generate large heterogeneous data sets that are impractical for manual assessment and characterization. In this work, an automated pipeline is developed that enables feature extraction from high-speed camera video and multi-modal data analysis. The framework removes the need for manual assessment through the utilization of deep learning techniques and training models in a weakly supervised paradigm. We demonstrate this pipeline’s capability over 700,000 high-speed camera frames. The pipeline successfully extracts melt pool and spatter geometries and links them to corresponding pyrometry, radiography, and processparameter information. 715 individual prints are examined to reveal melt pool areas that exceeds 0.07 mm 2 and pyrometry signal over a threshold (375 pyrometry units) were more likely to have defects. These automated processes enable massive throughput of characterization techniques.

36 MATERIALS SCIENCE↗

Antiproton bounds on dark matter annihilation from a combined analysis using the DRAGON2 code

Abstract Early studies of the AMS-02 antiproton ratio identified a possible excess over the expected astrophysical background that could be fit by the annihilation of a weakly interacting massive particle (WIMP). However, recent efforts have shown that uncertainties in cosmic-ray propagation, the antiproton production cross-section, and correlated systematic uncertainties in the AMS-02 data, may combine to decrease or eliminate the significance of this feature. We produce an advanced analysis using the DRAGON2 code which, for the first time, simultaneously fits the antiproton ratio along with multiple secondary cosmic-ray flux measurements to constrain astrophysical and nuclear uncertainties. Compared to previous work, our analysis benefits from a combination of: (1) recently released AMS-02 antiproton data, (2) updated nuclear fragmentation cross-section fits, (3) a rigorous Bayesian parameter space scan that constrains cosmic-ray propagation parameters.We find no statistically significant preference for a dark matter signal and set strong constraints on WIMP annihilation tobb̅, ruling out annihilation at the thermal cross-section for dark matter masses below ∼ 200 GeV. We do find a positive residual that is consistent with previous work, and can be explained by a ∼ 70 GeV WIMP annihilating below the thermal cross-section. However, our default analysis finds this excess to have a local significance of only 2.8σ, which is decreased to 1.8σwhen the look-elsewhere effect is taken into account.

Astronomy & Astrophysics↗

Observation of 𝑊⁢𝑍⁢𝛾 production and constraints on new physics scenarios in proton-proton collisions at $\sqrt{s}$ =13 TeV

A measurement of the 𝑊⁢𝑍⁢𝛾 triboson production cross section is presented. The analysis is based on a data sample of proton-proton collisions at a center-of-mass energy of $\sqrt{s}$ =13 TeV recorded with the CMS detector at the LHC, corresponding to an integrated luminosity of 138 fb −1 . The analysis focuses on the final state with three charged leptons, ℓ ± ⁢𝜈⁢ℓ + ⁢ℓ − , where ℓ = 𝑒 or 𝜇, accompanied by an additional photon. The observed (expected) significance of the 𝑊⁢𝑍⁢𝛾 signal is 5.4 (3.8) standard deviations. The cross section is measured in a fiducial region, where events with an ℓ originating from a tau lepton decay are excluded, to be 5.48 ± 1.11 fb, which is compatible with the prediction of 3.69 ± 0.24 fb at next-to-leading order in quantum chromodynamics. Exclusion limits are set on anomalous quartic gauge couplings and on the production cross sections of massive axionlike particles.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Search for heavy resonances decaying into two Higgs bosons in the $\text {b}{\bar{\text {b}}} \tau ^{+} \tau ^{-}$ final state in proton–proton collisions at $\sqrt{s} = 13\,\text {Te}\hspace{-.08em}\text {V}$

A search is presented for massive narrow-width resonances in the mass range of 1–4.5 TeV , decaying into pairs of Higgs bosons (HH). The search uses proton–proton collision data at a center-of-mass energy of 13 TeV collected with the CMS detector at the CERN LHC during 2016–2018, corresponding to an integrated luminosity of 138 fb -1 . The analysis targets final states where one Higgs boson decays into a pair of bottom quarks and the other into a pair of tau leptons, ${\text {X}} \rightarrow {\text {HH}} \rightarrow \text {b}{\bar{\text {b}}}\,\tau ^{+}\tau ^{-}$. It uses a single large radius jet to reconstruct the ${\text {H}} \rightarrow \text {b}{\bar{\text {b}}}$ decay, while the ${\text {H}} \rightarrow \tau ^{+}\tau ^{-}$ decay products can either be contained within a single large radius jet or appear as two isolated tau leptons. The observed data are consistent with standard model background expectations. Upper limits at 95% confidence level are set on the production cross section for resonant HH production for masses between 1 and 4.5 TeV. This analysis sets the most sensitive limits to date on ${\text {X}} \rightarrow {\text {HH}} \rightarrow \text {b}\bar{\hbox {b}}\, \tau ^{+}\tau ^{-}$ decays in the mass range of 1.4–4.5 TeV.

Hayrapetyan, A. [Yerevan Physics Institute]↗

Joint Analysis of Small-scale Galaxy Clustering and Galaxy–Galaxy Lensing from BOSS Galaxies

We present a joint analysis of galaxy clustering and galaxy–galaxy lensing measurements from BOSS galaxies using a simulation-based emulation method combined with a halo occupation distribution model. Our emulators are constructed with the Aemulus ν simulations, a suite of wνCDM N-body simulations with massive neutrinos as independent particle species. We combine small-scale analysis of clustering from 0.1 to 60.2 h −1 Mpc and lensing from 1.7 to 60.2 h −1 Mpc to perform cosmological constraints. We split the BOSS galaxies into three redshift bins to measure their clustering and employ galaxies from Dark Energy Camera Legacy Survey and Hyper Suprime-Cam as source galaxies to measure lensing separately. We find that the addition of lensing significantly improves the constraining power on $S_8 = σ_8(Ω_m/0.3)^{0.5}$, with a weak improvement for fσ 8 . Our results of fσ 8 indicate tensions of around 1σ−4σ below the results of the cosmic microwave background observations of Planck. For S 8 , our results are also lower than Planck, and the tension can be mitigated when considering possible systematics in lensing measurement. As a by-product, our analysis prefers a nonzero neutrino mass but without strong significance, with the constraining power dominated by the clustering. Given the accuracy and precision of our model and the observational data, it is anticipated that larger and higher-quality spectroscopic data sets will improve the constraints on this fundamental property in the near future.

Gao, Wenhao [Shanghai Jiao Tong University (China)↗

Reliable and Efficient Machine Learning (Final Technical Report)

Modern scientific experiments generate massive amounts of data at a pace much faster than humans can manually analyze. While machine learning has revolutionized commercial data analysis (such as recommending movies or recognizing faces), applying these tools to complex scientific discovery is challenging because scientific answers must be precise, interpretable, and adhere to physical laws. The research under this project aims to develop new mathematical tools and computer algorithms specifically designed for scientific applications. Major progress has been made in automatically cleaning and deconstructing messy experimental data, analyzing the visual information of physical phenomena, determining the underlying physical variables, and providing rig orous mathematical analysis of interesting algorithms and concepts widely used in machine learning. This project addressed the critical gap between our ability to generate massive scientific data and our ability to extract interpretable information from it. We established mathematical foundations for Scientific Machine Learning (SciML) aimed at effective data analytics and automated discovery. Our work focused on three core objectives: (1) developing reliable feature extraction methods for dynamic high-dimensional data, (2) establishing mathematical foundations for discovering dynamics via neural networks, and (3) creating rigorous optimization techniques for these models. Key outcomes come from two fronts. On the practical side, they include the development of algorithms that significantly enhance the extraction of signals from field data, as well as the capability to handle situations that exhibit smooth variations or physical stretching due to temperature changes. They also include the creation of an automated framework for discovering fundamental state variables from raw experimental data, demonstrating the ability to identify intrinsic physical dimensions without prior knowledge of the governing laws. On the theoretical front, the research results in theoretical advances in Optimal Transport, a widely used notion in SciML, specifically regarding functions with fixed-size nodal sets, provide sharp bounds relevant to uncertainty quantification. Meanwhile, the outcomes also include the establishment of convergence theories for nonlocal gradient descent methods, enabling robust optimization with noisy data in high-dimensional settings commonly encountered in scientific modeling. The project also helps creating opportunities to train the next generation of researchers, equipping them with the necessary technical skills for today’s workplace and preparing them for future advances.

97 MATHEMATICS AND COMPUTING↗

GLBRC Soil Yearlong Incubation 13C-SIP-Lipidomics

Data package for Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon This data is published under a CC0 license. The authors encourage data reuse and request attribution by referencing the below citations for the data packages and associated manuscript. Please cite as: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. GLBRC Soil Yearlong Incubation 13C-SIP-Lipidomics. [Data Set] PNNL DataHub. doi: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. MSV000097435: GLBRC soil yearlong incubation 13C-SIP-Lipidomics [Data Set] MassIVE. doi:10.25345/C57659T3K Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon. In Prep This data package consists of compound-specific 13C SIP-lipidomics data from a yearlong tracer incubation experiment designed to investigate microbial lipid persistence in switchgrass bioenergy crop soils. In order to explore how lipid structure may modulate the persistence of C in soil lipids, we leveraged soils from two sites (Michigan - sandy texture, Wisconsin - silty texture) operated by the U.S. Department of Energy-funded Great Lakes Bioenergy Research Center (GLBRC). These sites had comparable climates, identical management practices, but contrasting soil textures, allowing us to assess the variability of lipid accrual or degradation in soils as well as provide insight regarding the degree to which edaphic properties may regulate the retention of soil lipids. Untargeted lipidomics analyses were performed to identify 13C-labeled lipids in the soil microbiome after long-term incubation. Soils were supplemented with 100 micrograms glucose per gram dry soil (99 atom % 13C or natural abundance for paired control) and incubated; samples were collected two months and one year after glucose addition. Lipid extracts (MPLEx) were analyzed by LC-MS/MS and identified using LIQUID. Calculation of isotopic enrichment of lipids was performed by targeted approach using TarMet to quantify lipid isotopologues and IsoCorrectoR to correct for natural abundance isotopes. Contents: Data package contents reported here are the first version and contain downstream analysis files for the raw LC-MS mass spectrometry files (.mzXML) deposited at the MassIVE database repository under accession MSV000097435 (80 experimental runs; 5.85 GB) | MassIVE DOI: 10.25345/C57659T3K. Support files include the additional data download 'Read Me' file containing data descriptor information. Reported data download contents are structured for compliance with project data sharing guidelines, community standards initiatives, and sponsor stakeholder policies supporting FAIR data principles. Data processing software, analysis tools, and data workflows are listed below corresponding to the host repository long-term location. Available Data Downloads (0.3 GB): "GLBRC soil yearlong incubation 13C-SIP-Lipidomics_readme.txt" - 'Read Me' data package content file (txt) "GLBRC_DataPackage_analysis files" - Data processing files (Rmd) and saved intermediate data processing outputs (rds, csv, xlsx) "GLBRC_13C_lipidomics_dataset.xlsx" - processed data in tabular format (xlsx) Linked Software: LIQUID LC-MS Analysis Software | 10.5281/zenodo.6459462 Lipid Mini-On Software Tools | 10.5281/zenodo.1492803 pmartR Omics Statistical Software | 10.5281/zenodo.6108667 xcms (v4.3.3) TarMet (v1.1.1) IsoCorrectoR (1.24.0) Funding Acknowledgments: This research was supported by an Early Career Research Program award funded by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research (OBER) Genomic Science program under FWP 68292, FWP 07880 and EMSL Exploratory Research Project 51095. A portion of this work was performed in the William R. Wiley Environmental Molecular Sciences Laboratory, a national scientific user facility sponsored by OBER and located at Pacific Northwest National Laboratory (PNNL). PNNL is a multi-program national laboratory operated by Battelle for the DOE under Contract DE-AC05-76RLO1830.

Rempfert, Kaitlin R [Pacific Northwest National La↗

OzDES Reverberation Mapping of Active Galactic Nuclei: Final Data Release, Black-Hole Mass Results, & Scaling Relations

Over the last decade, the Australian Dark Energy (OzDES) collaboration has used Reverberation Mapping to measure the masses of high redshift supermassive black holes. Here we present the final review and analysis of this OzDES reverberation mapping campaign. These observations use 6-7 years of photometric and spectroscopic observations of 735 Active Galactic Nuclei (AGN) in the redshift range 0.13-3.85 and bolometric luminosity range 44.3 - 47.5 erg/s. Both photometry and spectra are observed in visible wavelengths, allowing for the physical scale of the AGN broad line region to be estimated from reverberations of the H{̱e̱ṯa̱}̱, MgII and CIV emission lines. We successfully use reverberation mapping to constrain the masses of 62 super-massive black holes, and combine with existing data to fit a power law to the lag-luminosity relation for the H{̱e̱ṯa̱}̱ and MgII lines with a scatter of ~0.25 dex, the tightest yet identified, fit specifically for consistency with high redshift AGN. We fit a similarly constrained relation for CIV, resolving a tension with the low luminosity literature AGN by accounting for selection effects arising from finite survey length. We also examine the impact of emission line width and luminosity (related to accretion rate) in reducing the scatter of these scaling relationships and find no significant improvement over the lag-only approach for any of the three lines. Using these relations, we further estimate the masses and accretion rates of 246 AGN with single epoch methods. We also use these relations to estimate the relative sizes of the H{̱e̱ṯa̱}̱, MgII and CIV emitting regions. In short, we provide a comprehensive benchmark of high redshift AGN reverberation mapping at the close of this most recent generation of surveys, including light curves, time-delays, and a set of significantly improved radius-luminosity relations for use with high-redshift populations.

McDougall, Hugh [Queensland U.]↗

Hardware acceleration for HPS algorithms in two and three dimensions

We provide a flexible, open-source framework for hardware acceleration, namely massively-parallel execution on general-purpose graphics processing units (GPUs), applied to the hierarchical Poincaré–Steklov (HPS) family of algorithms for building fast direct solvers for linear elliptic partial differential equations. To take full advantage of the power of hardware acceleration, we propose two variants of HPS algorithms to improve performance on two- and three-dimensional problems. In the two-dimensional setting, we introduce a novel recomputation strategy that minimizes costly data transfers to and from the GPU; in three dimensions, we modify and extend the adaptive discretization technique of Geldermans and Gillman [1] to greatly reduce peak memory usage. We provide an open-source implementation of these methods written in JAX, a high-level accelerated linear algebra package, which allows for the first integration of a high-order fast direct solver with automatic differentiation tools. We conclude with extensive numerical examples showing our methods are fast and accurate on two- and three-dimensional problems.

Fast direct solvers↗

MOOSE ProbML: Parallelizable Probabilistic Machine Learning and Uncertainty Quantification Capabilities

The Multiphysics Object Oriented Simulation Environment (MOOSE) is a widely used open- source finite element software for performing multiphysics multiscale simulations in a massively parallel fashion. Recently, the computational team at Idaho National Laboratory (INL) has implemented Probabilistic Machine Learning (ProbML) capabilities in MOOSE—in a parallelized fashion—and enable active learning with large-scale computational models for tasks such as surrogate model development, scale bridging, forward/inverse uncertainty quantification (UQ), Bayesian optimization, etc. This presentation summarizes these developments in MOOSE along with demonstrations on several real applications relevant to nuclear energy. At the fundamental level, samplers like Monte Carlo/Latin Hypercube, variance reduction, parallelized Markov Chain Monte Carlo (MCMC) support uncertainty propagation in both forward and inverse settings. These samplers can be integrated with the Gaussian processes (GP) suite in MOOSE, which offer several variants like scalar GPs, multi-output GPs, and deep GPs, to enable active learning. These GPs can be tuned using gradient-based optimization methods like Adam and its variants or gradient-free methods like the elliptical slice sampler (a variant of MCMC adept under Gaussian settings) for more complex covariance kernels or likelihoods whose gradient computations can be cumbersome. A variety of batch acquisition functions permit parallelized evaluation of the computational model and support different learning objectives with high efficiency like Bayesian inference, global surrogate development, optimization, etc. Furthermore, libtorch integration supports training, evaluation, and re-training of neural networks and other complex machine learning models in active learning settings. The impacts of these developments are shown on several real applications: (1) nuclear fuel inverse UQ and model inadequacy assessment using the Kennedy O’Hagan framework; (2) uncertainty aware surrogate modeling for additive manufacturing to predict field quantities; (3) nuclear reactor rare events analysis; and (4) complex fluid flow prediction using a global surrogate with quantified prediction uncertainty. Finally, the outlook of MOOSE ProbML is discussed for both outer-loop and inner-loop computations in the broad view to accelerate fuels and materials qualification, address gaps in knowledge and data, and assess new reactor/fuel systems.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Utility-Scale Solar, 2024 Edition: Empirical Trends in Deployment, Technology, Cost, Performance, PPA Pricing, and Value in the United States [Slides]

Berkeley Lab’s “Utility-Scale Solar, 2024 Edition” presents analysis of empirical plant-level data from the U.S. fleet of ground-mounted photovoltaic (PV), PV+battery, and concentrating solar-thermal power (CSP) plants with capacities exceeding 5 MWAC (PV plants of 5 MWAC or less, including residential rooftop systems, are covered separately in Berkeley Lab’s companion annual report, Tracking the Sun). Key findings from this year’s report include: -18.5 GWAC of new utility-scale PV capacity came online in 2023, bringing cumulative installed capacity to more than 80.2 GWAC across 47 states. Installed costs continued to fall in 2023. Relative to 2022, capacity-weighted averages decreased by 8% to -$\$1.43$/WAC (or $\$1.08$/WDC). Costs, based on a 7.1 GWAC sample of 76 plants completed in 2023, have fallen by 75% (averaging 10% annually) since 2010. Plant-level capacity factors vary widely, from 6% to 36% (on an AC basis), with a sample median of 24%. -Levelized cost of energy (LCOE) of new 2023 projects increased slightly to $\$46$/MWh prior to the application of tax credits but continued to fall to $\$31$/MWh when accounting for federal incentives. PPA prices have largely followed the decline in solar’s LCOE over time, but newly signed longer-term PPA prices have increased since 2021, to an average of $\$35$/MWh (levelized, in 2023 dollars). -Solar’s average energy and capacity value (i.e., ability to offset costs of other power generation sources) across the U.S. was $\$45$/MWh in 2023. Solar’s average market value was lowest in CAISO ($\$27$/MWh), the market with the greatest solar generation share, and highest in ERCOT ($\$67$/MWh). -Newer solar projects had greater market value in 2023 than their generation costs, yielding $\$1.1$ billion in benefits. Projects built in 2022 delivered on average $\$15$/MWh more market value than their costs in 2023. -Solar’s combined value from wholesale electricity markets, public health and climate damage reduction were greater than generation costs and incentives, yielding $\$13.7$ billion in net benefits in 2023. We estimate U.S. health benefits of $\$24$/MWh and reduced global climate damages of $\$101$/MWh. -Adding battery storage is one way to increase the value of solar. Deployment of 52 new PV+battery hybrid plants set a record with 5.3 GW installed in 2023. Our public data file tracks metadata and PPA prices from more than 100 PV+battery hybrid projects that are already online or that have secured offtake arrangements. -Looking ahead, a massive pipeline of at least 1,085 GW of solar capacity dominates the nation’s interconnection queues at the end of 2023. Nearly 571 GW, or 53%, of that total was paired with a battery – in CAISO it was a staggering 98%. Historically only 10% of the requested solar capacity is built. -For more information, and to explore related interactive data visualizations, go to utilityscalesolar.lbl.gov.

14 SOLAR ENERGY↗