Search NASA⌕ Search

SEARCH · Search NASA

Results for “Partitioned scheme”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

35 records · Page 2

Exposing and Reducing Biases of Simulating Mixed-Phase Clouds in the Convection-Permitting E3SM Atmosphere Model: Lessons From an Arctic Cold-Air Outbreak

Mixed-phase clouds modulate the water and energy cycles of high-latitude regions, yet their liquid-ice phase partitioning has long been poorly simulated in climate models. Here, simulations of Arctic mixed-phase clouds by the Simple Cloud-Resolving E3SM Atmosphere Model (SCREAM) are assessed against large-eddy simulations, satellite data, and ground-based observations during the Cold-Air Outbreaks in the Marine Boundary Layer Experiment field campaign. SCREAM simulates nearly completely frozen clouds, which is attributed largely to the unreasonably strong Wegener–Bergeron–Findeisen (WBF) process that converts liquid to ice excessively and partly to the early over-abundant ice production at cold temperatures from a temperature-deterministic deposition ice nucleation scheme. Assuming no subgrid variation for the WBF process in the original formulation particularly conflicts with the instantaneous saturation adjustment assumption in the condensation scheme that assumes subgrid variability, leading to exaggerated WBF process rates. A proposed simple physically-based improvement on the treatment of subgrid cloud overlap substantially increases supercooled liquid water content and notably improves cloud-top phase partitioning, aligning better with observations. Improvement of supercooled liquid water content also converges with increasing horizontal resolution. The deposition ice nucleation scheme is found responsible for a falsely-produced ice cloud aloft that is not observed, biasing the simulated cloud radiative effects and top-of-atmosphere radiative fluxes. This study identifies key deficiencies in cloud parameterizations that continue to challenge convection-permitting models.

Geosciences↗

Aeroelastic Modelling of Large Wind Turbines: Towards a Unified OpenFAST-SEAHOWL Approach

In recent years, the scale of wind turbines has significantly increased to maximize energy capture for a given site (particularly offshore), presenting new challenges in terms of structural design and dynamics. As towers grow taller and blades grow longer, flexion and torsion of the latter have a non-negligible impact on the behavior and performance of the turbine in terms of overall loads, power production, and control. When representing large-scale wind turbines numerically to capture these important effects, particular attention must therefore be given to the level of fidelity for representing each structural component, as well as the coupling scheme used between them to keep simulations accurate, stable, and efficient. To address this issue, we combine here the two following tools: (1) OpenFAST, the reference whole-turbine simulation tool from NREL with standalone modules covering each physics and the choice between loose coupling and a new tight coupling scheme for structural dynamics, and (2) SEAHOWL, the whole-turbine simulation tool from TotalEnergies with monolithic coupling of structural dynamics through Project Chrono and partitioned coupling for multiphysics interactions.

17 WIND ENERGY↗

Study of the connected four-point correlation function of galaxies from the DESI Data Release 1 luminous red galaxy sample

We present a measurement of the non-Gaussian four-point correlation function (4PCF) from the DESI DR1 luminous red galaxy (LRG) sample. For the gravitationally induced parity-even 4PCF, we detect a signal with a significance of 14.7⁢𝜎 using our fiducial setup. We assess the robustness of this detection through a series of validation tests, including auto and cross-correlation analyses, sky partitioning across multiple patch combinations, and variations in radial scale cuts. Due to the low completeness of the sample, we find that differences in fiber assignment implementation schemes can significantly impact estimation of the covariance and introduce biases in the data vector. After correcting for these effects, all tests yield consistent results. This is one of the first measurements of the connected 4PCF on the DESI LRG sample; the good agreement between the simulation and the data implies that the amplitude of the density fluctuation inferred from the connected 4PCF is consistent with the Planck Λ⁢ CDM cosmology. The methodology and diagnostic framework established in this work provide a foundation for interpreting parity-odd 4PCF.

Cosmology↗

Time-dependent-bases with local CUR decomposition method for accelerating turbulent combustion simulations

Here, this study presents a novel reduced-order modeling framework, Time-Dependent Bases with Local CUR decomposition (TDB-L-CUR), designed to efficiently and accurately approximate the species transport equations in reacting flow simulations. The method extends the existing TDB-CUR approach for chemically reacting flows (Jung et al. Comput. Methods Appl. Mech. Engrg. 437 (2025) 117758), which leverages matrix decomposition techniques to form a global-in-space, time-dependent low-dimensional manifold. While TDB-CUR performs well in homogeneous systems, it may be less well-suited to spatially heterogeneous systems such as turbulent flames, where higher-rank approximations are typically required. The proposed TDB-L-CUR framework introduces two methodological extensions to the baseline approach. First, it applies unsupervised clustering to partition the physical domain into distinct regions, enabling spatially localized manifold construction, thereby reducing the rank required for the reduced-order representation. Second, it incorporates a computational singular perturbation (CSP)-based scheme for identifying and penalizing fast species, allowing for spatio-temporally adaptive mitigation of chemical stiffness. The proposed framework is validated on a hierarchy of test cases, including a one-dimensional premixed flame, a two-dimensional nonpremixed ignition case with vortex interaction, and a three-dimensional turbulent premixed flame. TDB-L-CUR significantly improves accuracy over TDB-CUR while further reducing computational cost. The fully on-the-fly formulation of TDB-L-CUR (i.e., requiring no offline training or prior knowledge) makes it a robust and scalable tool for reduced-order modeling of reactive flows.

Local manifold↗

Graph-Based Representations and Applications to Process Simulation

Rapid and robust convergence of a process flowsheet is critical to enable large-scale simulations that address core scientific questions related to process design, optimization, and sustainability. However, due to the highly coupled and nonlinear nature of chemical processes, efficiently solving a flowsheet remains a challenge. In this work, we show that graph representations of the underlying physical phenomena in unit operations may help identify potential avenues to systematically reformulate the network of equations and enable more robust topology-based convergence of flowsheets. To this end, we developed graph abstractions of the governing equations of vapor-liquid and liquid-liquid equilibrium separation equipment. These graph abstractions consist of a mesh of interconnected variable nodes and equation nodes that are systematically generated through PhenomeNode, a new open-source library in Python developed in this study. We show that partitioning the graph into separate mass, energy, and equilibrium subgraphs can help decouple nonlinearities and guide decomposition algorithms. By employing the graph abstraction on an industrial separation process for separating glacial acetic acid from water, we implemented a new block decomposition scheme in BioSTEAM and demonstrated that this can accelerate convergence over a traditional sequential modular approach.

Distillation↗

Anomaly inflow, dualities, and quantum simulation of Abelian lattice gauge theories induced by measurements

Previous work [] has demonstrated that quantum simulation of Abelian lattice gauge theories (Wegner models including the toric code in a limit) in general dimensions can be achieved by local adaptive measurements on symmetry-protected topological (SPT) states with higher-form generalized global symmetries. The entanglement structure of the resource SPT state reflects the geometric structure of the gauge theory. In this work we explicitly demonstrate the anomaly inflow mechanism between the deconfining phase of the simulated gauge theory on the boundary and the SPT state in the bulk by showing that the anomalous gauge variation of the boundary state obtained by bulk measurement matches that of the bulk theory. Moreover, we construct the resource state and the measurement pattern for the measurement-based quantum simulation of a lattice gauge theory with a matter field (Fradkin-Shenker model), where a simple scheme to protect gauge invariance of the simulated state against errors is proposed. We further consider taking an overlap between the wave function of the resource state for lattice gauge theories and that of a parameterized product state, and we derive precise dualities between partition functions with insertion of defects corresponding to gauging higher-form global symmetries, as well as measurement-induced phases where states induced by a partial overlap possess different (symmetry-protected) topological orders. Measurement-assisted operators to dualize quantum Hamiltonians of lattice gauge theories and their noninvertibility are also presented. Published by the American Physical Society 2024

Okuda, Takuya↗

MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs

Graph Neural Networks (GNN) are indispensable in learning from graph-structured data, yet their rising computational costs, especially on massively connected graphs, pose significant challenges in terms of execution performance. To tackle this, distributed-memory solutions such as partitioning the graph to concurrently train multiple replicas of GNNs are in practice. However, approaches requiring a partitioned graph usually suffer from communication overhead and load imbalance, even under optimal partitioning and communication strategies due to irregularities in the neighborhood minibatch sampling. This paper proposes practical trade-offs for improving the sampling and communication overheads for representation learn- ing on distributed graphs (using popular GraphSAGE architecture) by developing a parameterized prefetch and eviction scheme on top of the state-of-the-art Amazon DistDGL distributed GNN framework, demonstrating about 15–40% improvement in end-to-end training performance on the NERSC Perlmutter supercomputer for various OGB datasets.

Machine Leanring, high performance comptuing, grap↗

A Vertically Resolved Canopy Improves Chemical Transport Model Predictions of Ozone Deposition to North Temperate Forests

Abstract Dry deposition is the second largest tropospheric ozone (O 3 ) sink and occurs through stomatal and nonstomatal pathways. Current O 3 uptake predictions are limited by the simplistic big‐leaf schemes commonly used in chemical transport models (CTMs) to parameterize deposition. Such schemes fail to reproduce observed O 3 fluxes over terrestrial ecosystems, highlighting the need for more realistic treatment of surface‐atmosphere exchange in CTMs. We address this need by linking a resolved canopy model (1D Multi‐Layer Canopy CHemistry and Exchange Model, MLC‐CHEM) to the GEOS‐Chem CTM and use this new framework to simulate O 3 fluxes over three north temperate forests. We compare results with in situ measurements from four field studies and with standalone, observationally constrained MLC‐CHEM runs to test current knowledge of O 3 deposition and its drivers. We show that GEOS‐Chem overpredicts observed O 3 fluxes across all four studies by up to 2×, whereas the resolved‐canopy models capture observed diel profiles of O 3 deposition and in‐canopy concentrations to within 10%. Relative humidity and solar irradiance are strong O 3 flux drivers over these forests, and uncertainties in those fields provide the largest remaining source of model deposition biases. Flux partitioning analysis shows that: (a) nonstomatal loss accounts for 60% of O 3 deposition on average; (b) in‐canopy chemistry makes only a small contribution to total O 3 fluxes; and (c) the CTM big‐leaf treatment overestimates O 3 ‐driven stomatal loss and plant phytotoxicity in these temperate forests by up to 7×. Results motivate the application of fully online vertically explicit canopy schemes in CTMs for improved O 3 predictions.

Vermeuel, Michael P. [Department of Soil, Water, a↗

Interplay Between Time and Energy in Bosonic Noisy Quantum Metrology

Quantum entanglement and coherence often allow for protocols that outperform classical ones in estimating a system’s parameter. When using infinite-dimensional probes (such as a bosonic mode), one could, in principle, obtain infinite precision in a finite time for both classical and quantum protocols, which makes it hard to quantify potential quantum advantage. However, such a situation is unphysical, as it would require infinite resources, so one needs to impose some additional constraint: typically the average energy employed by the probe is finite. Here we treat both energy and time as a resource, showing that, in the presence of noise, there is a nontrivial interplay between the average energy and the time devoted to the estimation. Our results are valid for the most general metrological schemes (e.g., adaptive schemes, which may involve entanglement with external ancillae or any kind of continuous measurement). We apply recently derived precision bounds for all parameters characterizing the paradigmatic case of a bosonic mode, subject to Lindbladian noise. We show how the time employed in the estimation should be partitioned in order to achieve the best possible precision. In most cases, the optimal performance may be obtained without the necessity of adaptivity or entanglement with ancilla. We compare results with classical strategies. Interestingly, for temperature estimation, applying a fast-prepare-and-measure protocol with Fock states provides better scaling with the number of photons than any classical strategy.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

TrustDER: Trusted, Private and Scalable Coordination of Distributed Energy Resources

In this project, the Stanford and SLAC Teams have developed a Trusted, Private and Scalable platform for coordinating Coordination of Distributed Energy Resources (TrustDER). This is a layered system that ensures private, trusted and scalable coordination and monitoring of DERs. It accommodates a variety of resources, such as solar generation, gensets and loads, with a particular focus on battery systems-based resources, as they are a transformational technology experiencing fast growth in adoption by large critical facilities. The platform can be used as standalone or added to existing aggregation systems to enable trust, privacy and resilience. TrustDER consists of layers that address each of the shortcomings of the existing state of the art. Each layer in the platform can operate independently but provides information to the layers above it to enable a novel form of overall coordination architecture. The project consists of several tasks, with each task dedicated to the design of each layer. Task 2 Resource Virtualization defined a software abstraction layer for distributed energy resources (DERs). The goal of this abstraction was to simplify the implementation of algorithms utilizing cooperation of DERs resources in a variety of use cases. Task 3 is on Secure ID for Asset Authentication. Identity Management Systems (IDMS) are a foundational infrastructure for interactions between entities (organizations, users, devices, and services). Secure ID is blockchain-based a distributed identity management system allowing (1) identity provisioning, (2) authentication, (3) authorization, and (4) identity data sharing for IoT-enabled assets on the electricity grid. In this project, the SLAC team focused on designing and testing Keymaker, a protocol for authenticating device identity managed by Secure ID. Task 5 Private and Safe Integration is focused on the design and evaluation of a DER cooperation scheme which allows for the aggregation of DERs without impacting network reliability. The approach is designed based on realistic assumptions regarding data availability, communication infrastructure limitations, and privacy. Task 6 Scalable Distributed Privacy for Information explored how virtualized batteries could be managed privately. Specifically, it examined the case in which a principal provides a partitioned battery to multiple clients. Task 7 Use Cases was to ensure that this technology was applied in relevant situations and scenarios. Primarily, this means that virtualization needed to be employed in a manner that either improved flexibility, bolstered security or privacy, or decreased costs.

25 ENERGY STORAGE↗

Batch Extraction Studies to Evaluate Trace Element Behavior in PUREX Conditions

The multilab Intentional Forensics Venture is working to identify which stable elements (i.e., taggants) at trace concentrations relative to U would persist throughout the nuclear fuel cycle in a voluntary fuel tagging scheme. A taggant would provide the nuclear forensics community with a “barcode” to help identify nuclear materials found outside of regulatory control. A portion of this project was focused on reprocessing effects and determining which, if any, elements would coextract with U(VI) in standard Pu–U reduction extraction (PUREX) conditions. Elements with a propensity to coextract could, in theory, be used as taggants from a PUREX perspective. Although retention is not a performance requirement, the taggant signature would need to partition predictably from the U stream after the PUREX process to maintain forensic utility. This report documents results from several batch extraction studies with numerous trace elements from HNO 3 (1.5–5 M), with and without U(VI), into 30% tri-n-butyl phosphate (TBP) in kerosene. Extraction and back-extraction tests were used to evaluate nearly 60 elements in surrogate conditions for PUREX, and distribution coefficients (i.e., D-values) for most species were <0.1, indicating few species are likely to co-extract with U through PUREX. Additional studies are needed to optimize sample volumes and dilutions to dial in these low D-values. The D-values (D) were determined for several of the more promising elements, including Re and Se. Ultimately, we conclude that only a limited number of the ~ 60 elements investigated are extractable in the U stream of PUREX, based on measured D values, meaning most candidate elemental taggants would likely be lost at this stage of the nuclear fuel cycle, even when considering a range of acid concentrations.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Assessing Clouds in GFDL's AM4.0 With Different Microphysical Parameterizations Using the Satellite Simulator Package COSP

We evaluate cloud simulations using satellite simulators against multiple observational data sets. These simulators have been run within the Geophysical Fluid Dynamics Laboratory's Atmosphere Model version 4.0 (AM4.0), as well as an alternative configuration where a fully two‐moment Morrison‐Gettelman cloud microphysical parameterization with prognostic precipitation (MG2) is applied, denoted as AM4‐MG2. The modeled cloud spatial distributions, vertical profiles, phase partitioning, cloud‐to‐precipitation transitions, and radiative effects compare reasonably well with satellite observations. Model biases include the under‐prediction of total and low‐level clouds, especially optically thin/intermediate clouds with cloud optical depth of less than 23, but the over‐prediction of thick clouds, indicating “too few, too bright” biases. These biases counteract each other, and give rise to reasonable estimates of cloud radiative effects. The underestimate of low‐level clouds is associated with too early and too frequent drizzle/precipitation formation. The precipitation bias is improved in AM4‐MG2, where the autoconversion scheme initiates the precipitation more realistically. There also exist discrepancies between models and observations for midlevel and high‐level clouds. Additional biases include the underestimate of liquid cloud fraction and the overestimate of ice cloud fraction.

54 ENVIRONMENTAL SCIENCES↗

The impact of aerosol mixing state on immersion freezing: insights from classical nucleation theory and particle-resolved simulations

Immersion freezing, initiated by ice-nucleating particles (INPs) in supercooled aqueous droplets, plays an important role in the formation of ice crystals within clouds. The efficiency of immersion freezing depends strongly on INP composition and, crucially, on the mixing state – how chemical species are distributed across the particle population. Here, we quantify the impact of aerosol mixing state on immersion freezing using a combined theoretical and particle-resolved modeling approach. We derive analytical expressions for the frozen fraction of internally and externally mixed INP populations based on classical nucleation theory, showing that the frozen fraction is sensitive to whether ice-active species are present in all particles or only in a subset of the population. We introduce a multi-species immersion freezing scheme into the particle-resolved model PartMC, using the water activity-based immersion freezing model (ABIFM) to compute freezing probabilities for mixed-composition particles. To improve computational efficiency, we implement a Binned Tau-Leaping algorithm and demonstrate an order-of-magnitude speedup with minimal accuracy loss. Simulations reproduce the analytical trends in limiting cases and extend the analysis to more general aerosol populations, where mixing state continues to exert a substantial control on frozen fraction. Sensitivity analyses across particle size, species type, and cooling condition reveal that the mixing state effect is most pronounced when small amounts of highly efficient INPs are mixed with less efficient materials. These findings underscore the need to represent aerosol mixing state explicitly in models of heterogeneous ice nucleation to reduce uncertainty in cloud-phase partitioning.

54 ENVIRONMENTAL SCIENCES↗

Restricted Partition Function Semiclassical Transition State Theory (RPF-SCTST): Applications to Reactions Involving Hydrogen Cyanide, Hydrogen Peroxide, and Formaldehyde Oxide

We demonstrate the efficiency of the numerical calculation of thermal semiclassical transition state theory (SCTST) rates across several representative chemical reactions using the restricted partition function (RPF)─viz a function that depends only on the imaginary action associated with a reaction. Here, we treat the potential energy surfaces (PESs) through fourth order expansions around the saddle point and integrate the Hamiltonian up to second order in employing vibrational perturbation theory. We apply this formalism to uni- and bimolecular reactions with different─though relatively high─barrier heights to uncover the influence of the barrier properties, minimal energies, and separability of the rotational motion on rate constants and tunneling corrections. Although all modes are coupled within the RPF, we found in our numerical examples that the rotational component in the absence of significant rotational distortions can be separated from the vibrational contributions. Moreover, a classical treatment of the rotational contribution is adequate over the temperature range from approximately 100 to 1000 K. We also found that the choice of DFT basis set can lead to variations in rate constants of up to an order of magnitude. Lastly, different schemes for counting eligible energy levels result in rate constants that differ by no more than a factor of 2, with the remaining discrepancies attributed to the treatment of near-convergent levels in perturbation theory.

74 ATOMIC AND MOLECULAR PHYSICS↗

Continuous Counter‐Current Microfluidic Liquid–Liquid Extraction Achieved Using a Pair of Wettable Screen Meshes

Continuous counter‐current microfluidic liquid–liquid extraction performs separations by flowing immiscible liquids in opposing directions within a single flow channel. In principle, this flow arrangement enables a large number of theoretical separation units in a small footprint, without using interstage valving, pumping, and phase separation. Despite its potential for excellent separation performance, this microfluidic scheme rarely appears in literature due to the requirement for capillary forces to be greater than hydrodynamic forces for stable flow. We present a novel microfluidic device and flow approaches that overcome this force‐balance challenge, enabling stable, long‐duration continuous counter‐current flow. Additionally, we cover a suite of methodologies for quantifying the performance of the microfluidic device, revealing the number of theoretical equilibrium stages achieved. The enabling technologies include a woven mesh screen‐based microfluidic device architecture that is easily fabricated outside of a clean room, surface functionalization strategies to promote conjugate (organic/aqueous) wettability, flow approaches to eliminate bubbles and carryover, and computer‐aided flow automation with optical measurement of extraction performance. The reported experiments lasted for over 36 h, terminated only at experiment conclusion, where the device still exhibited good performance. Automated Raman spectroscopy was used for solute quantitation of the ternary system tert‐butanol in a toluene/water matrix, a ternary system that was specifically chosen to analyze the device's performance with a small solute partition ratio and to enable in‐line Raman measurements of solute concentrations in both phases. The microfluidic device possessed a 55 mm contact length and a 38.5 µL internal volume. During counter‐current flow, we observed approximately 37 equilibrium stages (37 ± 13) based on a best‐fit of the solute fraction remaining in the aqueous phase using a Kremser Group Method analysis.

36 MATERIALS SCIENCE↗

EPCAPE-PT-LANL Measurements: Wideband Integrated Bioaerosol Sensor

Coastal cities offer a unique environment for studying aerosol-cloud interactions and the effects of urban emissions on cloud properties. As part of the Eastern Pacific Cloud Aerosol Precipitation Experiment (EPCAPE), the Partitioning Thrust by Los Alamos National Laboratory (EPCAPE-PT-LANL) was conducted. Our campaign focused on measuring the optical and chemical properties of aerosols and their interactions within marine stratocumulus clouds in La Jolla, California. EPCAPE-PT-LANL enhances the primary goals of EPCAPE through innovative observations of vapor-phase transitions between aerosols and cloud droplets, the impact of black carbon on aerosol-cloud dynamics, and the effects of cloud processing on aerosol optical properties. Instrument: Wideband Integrated Bioaerosol Sensor (Droplet Measurements Technology) Data Notes: The WIBS is an online single-particle measurement that detects FBAPs (within a size range of 0.5 - 30 microns in diameter) based on the excitation and emission wavelengths of the individual particles. Using two xenon lamps, the WIBS excites FBAPs at 280 nm and 370 nm. Their emission is detected across two wavebands of 310-400 nm and 420-650 nm. We classified the FBAPs into seven different categories (A, B, C, AB, BC, AC, and ABC) using the classification scheme in Perring et. al. (2015) [1]. Averaged number concentration of FBAPs (total and by category) and particles that non-fluorescent bioaerosols particles (NFBAPs). In separate files, we also present one-minute-averaged size distributions and the asymmetry factor (AF, a surrogate for shape) of all FBAPs and NFBAPs. The logarithmic bin width of the size bins are the same as the average bin width of the AOS's optical particle counter (OPC, Grimm) for the range of sizes in which they overlap (26 bins from 0.5 - 30 microns). AF of the particles ranges from 0-100 and is divided into five bins with a linear spacing at increments of 20. The smallest AF bin represents more spherical particles while the largest bin represents more rod-shaped particles. [1] Perring, A. E., et al. (2015), Airborne observations of regional variation in fluorescent aerosol across the United States, J. Geophys. Res. Atmos., 120, 1153–1170, doi:10.1002/2014JD022495. Abstract and description of the campaign can be found here : https://www.arm.gov/research/campaigns/amf2023epcape-pt-lanl. Files data_10min_WIBS_AFDist.csv Header: - FBAP_AFDist[/cm3]_Bin_1 to Bin_5: Concentration of fluorescent bioaerosol particles in the each of 5 AF bins, measured in particles per cubic centimeter. Each bin represents a specific range of particle AF, capturing the shapes of FBAPs detected during the measurement. - NFBAP_AFDist[/cm3]_Bin_1 to Bin_5: Concentration of non-fluorescent bioaerosol particles in the each of 5 AF bins, measured in particles per cubic centimeter. Each bin represents a specific range of particle AF, capturing the shapes of NFBAPs detected during the measurement. - CVI_Flag[bool]: A boolean flag indicating whether the Counterflow Virtual Impactor (CVI) was active (true) or inactive (false) during the measurement. AF Bins: • Bin 1: 0 – 20 [unitless] • Bin 2: 21 – 40 [unitless] • Bin 3: 41 – 60 [unitless] • Bin 4: 61 – 80 [unitless] • Bin 5: 81 – 100 [unitless] Files data_10min_WIBS_Conc.csv Header: - NumberConcentrationA[/cm3]: Number concentration of bioaerosol particles detected by fluorescence channel A, measured in particles per cubic centimeter. - NumberConcentrationB[/cm3]: Number concentration of bioaerosol particles detected by fluorescence channel B, measured in particles per cubic centimeter. - NumberConcentrationC[/cm3]: Number concentration of bioaerosol particles detected by fluorescence channel C, measured in particles per cubic centimeter. - NumberConcentrationAB[/cm3]: Combined number concentration of bioaerosol particles detected by both fluorescence channels A and B, measured in particles per cubic centimeter. - NumberConcentrationBC[/cm3]: Combined number concentration of bioaerosol particles detected by both fluorescence channels B and C, measured in particles per cubic centimeter. - NumberConcentrationAC[/cm3]: Combined number concentration of bioaerosol particles detected by both fluorescence channels A and C, measured in particles per cubic centimeter. - NumberConcentrationABC[/cm3]: Combined number concentration of bioaerosol particles detected by all three fluorescence channels A, B, and C, measured in particles per cubic centimeter. - NumberConcentrationNFBAP[/cm3]: Number concentration of non-fluorescent bioaerosol particles, measured in particles per cubic centimeter. - CVI_Flag[bool]: A boolean flag indicating whether the Counterflow Virtual Impactor (CVI) was active (true) or inactive (false) during the measurement. Files data_10min_WIBS_SizeDist.csv Header: - FBAP_SizeDist[/cm3]_Bin_1 to FBAP_SizeDist[/cm3]_Bin_26: Number concentrations of FBAP in each of 26 size bins, measured in particles per cubic centimeter. Each bin represents a specific range of particle sizes, capturing the size distribution of FBAPs detected during the measurement. - NFBAP_SizeDist[/cm3]_Bin_1 to NFBAP_SizeDist[/cm3]_Bin_26: Number concentrations of NFBAP in each of 26 size bins, measured in particles per cubic centimeter. Similar to FBAP, each bin covers a specific range of particle sizes, detailing the size distribution of NFBAPs detected. - CVI_Flag[bool]: A boolean flag indicating whether the Counterflow Virtual Impactor (CVI) was active (true) or inactive (false) during the measurement. Size Bins: • Bin 1: 0.48 to 0.57 μm • Bin 2: 0.57 to 0.67 μm • Bin 3: 0.67 to 0.79 μm • Bin 4: 0.79 to 0.93 μm • Bin 5: 0.93 to 1.1 μm • Bin 6: 1.1 to 1.29 μm • Bin 7: 1.29 to 1.52 μm • Bin 8: 1.52 to 1.8 μm • Bin 9: 1.8 to 2.11 μm • Bin 10: 2.11 to 2.5 μm • Bin 11: 2.5 to 2.94 μm • Bin 12: 2.94 to 3.46 μm • Bin 13: 3.46 to 4.08 μm • Bin 14: 4.08 to 4.81 μm • Bin 15: 4.81 to 5.67 μm • Bin 16: 5.67 to 6.68 μm • Bin 17: 6.68 to 7.88 μm • Bin 18: 7.88 to 9.29 μm • Bin 19: 9.29 to 10.96 μm • Bin 20: 10.96 to 12.92 μm • Bin 21: 12.92 to 15.23 μm • Bin 22: 15.23 to 17.96 μm • Bin 23: 17.96 to 21.17 μm • Bin 24: 21.17 to 24.96 μm • Bin 25: 24.96 to 29.43 μm • Bin 26: 29.43 to 34.70 μm

54 ENVIRONMENTAL SCIENCES↗

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning↗