Search NASA⌕ Search

SEARCH · Search NASA

Results for “Deep”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Accurate and uncertainty-aware multi-task prediction of HEA properties using prior-guided deep Gaussian processes

Surrogate modeling techniques have become indispensable in accelerating the discovery and optimization of high-entropy alloys (HEAs), especially when integrating computational predictions with sparse experimental observations. This study systematically evaluates the training and testing performance of four prominent surrogate models—conventional Gaussian processes (cGP), Deep Gaussian processes (DGP), encoder-decoder neural networks for multi-output regression and eXtreme Gradient Boosting (XGBoost)—applied to a hybrid dataset of experimental and computational properties of the 8-component HEA system Al-Co-Cr-Cu-Fe-Mn-Ni-V. We specifically assess their capabilities in predicting correlated material properties, including yield strength, hardness, modulus, ultimate tensile strength, elongation, and average hardness under dynamic/quasi-static conditions, alongside auxiliary computational properties. The comparison highlights the strengths of hierarchical deep modeling approaches in handling heteroscedastic, heterotopic, and incomplete data commonly encountered in materials science. Our findings illustrate that combined surrogate models such as DGPs infused with machine-learned priors outperform other surrogates by effectively capturing inter-property correlations and by assimilating prior knowledge. This enhanced predictive accuracy positions the combined surrogate models as powerful tools for robust and data-efficient materials design.

36 MATERIALS SCIENCE↗

Deep Gaussian process-based cost-aware batch Bayesian optimization for complex materials design campaigns

The accelerating pace and expanding scope of materials discovery demand optimization frameworks that efficiently navigate vast design spaces with complex response surfaces while judiciously allocating limited evaluation resources. We present a cost-aware, batch Bayesian optimization scheme powered by deep Gaussian process (DGP) surrogates and a heterotopic querying strategy. Our DGP surrogate, formed by stacking GP layers, models complex hierarchical relationships among high-dimensional compositional features and captures correlations across multiple target properties, propagating uncertainty through successive layers. We integrate evaluation cost into an upper-confidence-bound acquisition extension, which, together with heterotopic querying, proposes small batches of candidates in parallel, balancing exploration of under-characterized regions with exploitation of high-mean, low-variance predictions across correlated properties. Applied to refractory high-entropy alloys for high-temperature applications, our framework converges to optimal formulations in fewer iterations with cost-aware queries than conventional GP-based BO, highlighting the value of deep, uncertainty-aware, cost-sensitive strategies in materials campaigns.

36 MATERIALS SCIENCE↗

Dual interfacial H-bonding-enhanced deep-blue hybrid copper–iodide LEDs

Solution-processed light-emitting diodes based on non-toxic copper–iodide hybrids are a compelling solution for efficient and stable deep-blue lighting, owing to their tunability, high photoluminescence efficiency and environmental sustainability. Here we present a hybrid copper–iodide that shows near-unity photoluminescence quantum yield (99.6%) with an emission wavelength of 449 nm and colour coordinates (0.147, 0.087), alongside its emission mechanism and charge transport characteristics. Here, we use the thin film of this hybrid as the sole active emissive layer to fabricate deep-blue light-emitting diodes and subsequently enhance the device performance through a dual interfacial hydrogen-bond passivation strategy. This synergetic surface modification approach, integrating a hydrogen-bond-acceptor self-assembled monolayer with an ultrathin polymethyl methacrylate capping layer, effectively passivates both heterojunctions of the copper–iodide hybrid emissive layer and optimizes charge injections. We achieve a maximum external quantum efficiency of 12.57%, a maximum luminance of 3,970.30 cd m −2 with colour coordinates (0.147, 0.091) and an excellent operational stability (half-lifetime) of 204 hours under ambient conditions. We further showcase a large-area device of 4 cm 2 that maintains high efficiency. Our findings reveal the potential of copper–iodide-based hybrid materials for applications in solid-state lighting and display technologies, offering a versatile strategy for enhancing device performances.

14 SOLAR ENERGY↗

A non-intrusive framework using acoustic signals and deep learning for boiling diagnostics in visual-limited environments

Accurate monitoring of boiling heat transfer is critical for safeguarding high-power systems operating in environments where conventional optical diagnostics are hindered by radiation fields or restricted visual accessibility. This study presents a non-intrusive framework that integrates hydroacoustic sensing with deep learning to infer near-wall boiling characteristics and enable predictive thermal assessment without visual access. In a prototypical subcooled flow-boiling facility representative of the Isotope Production Facility (IPF) at Los Alamos, hydrophones capture boiling-induced acoustic emissions that are transformed into background-removed Short-Time Fourier Transform (STFT) spectrograms. A convolutional neural network (CNN) then regresses heat flux, wall superheat, and key bubble parameters directly from these spectrograms. The CNN achieved predictive accuracy under nominal conditions and demonstrated robustness and generalization under acoustic noise for Signal-to-Noise Ratios (SNRs) down to approximately 0 dB. When integrated into an ANSYS CFX wall-boiling model, the acoustically inferred parameters reproduced boiling curve and critical heat flux (CHF) values consistent with image-based benchmarks. Furthermore, the model retained reliable performance under moderate variations in bulk temperature, flow rate, and hydrophone placement, confirming its generalizability across practical boundary conditions. These results demonstrate the feasibility of hydroacoustic-based deep learning as a viable path toward real-time, radiation-tolerant boiling diagnostics and predictive thermal safety assessment in inaccessible systems such as the IPF.

42 ENGINEERING↗

Advected carbon younger than the sediment fuels microbial metabolism in a pumped deep aquifer

Ever deeper wells are drilled worldwide to pump potable groundwater. Recent studies argue that overpumping compresses clays and releases reactive dissolved organic carbon (DOC), which in turn drives a series of microbial reactions that affect groundwater potability including arsenic concentrations. Here, we use a novel method to measure the radiocarbon ages of microbial RNA to determine the source of reactive or metabolizable DOC and argue against the compression of clays as the sole source of carbon. We show that microbial RNA (5,230; 5,550; 6,250 yr; n = 3 wells), DOC (280-10,800 yr; n = 13), and methane (modern to 6,240 yr; n = 3), from an overpumped deep aquifer in Bangladesh are much younger than the overlying clay layers deposited during the Pleistocene over 12,000 years ago. Mass-balance indicates that at least half of the carbon incorporated into RNA has to come from reactive DOC or methane that is advected downward via vertical recharge. This metabolism of advected organic carbon could have implications for the quality of water pumped from deep aquifers.

Biological and medical sciences↗

Formation and optical properties of indium nanoparticle arrays for deep-UV plasmonics

We utilize a combined computational-experimental approach to examine the influence of indium nanoparticle (NP) array distributions on deep-ultraviolet (UV) plasmon resonances. For photon energies < 5.7 eV, analysis of ellipsometric spectra reveals an increase in silicon reflectance induced by indium NP arrays on silicon. For various energies in the range 5.7–7.0 eV, a decrease in reflectance is induced by the NP arrays. Similar trends in reflectance are predicted from finite-difference time-domain (FDTD) simulations using NP size distributions extracted from atomic-force micrographs as input. In addition, in the energy range of 7.4–9.2 eV, the FDTD simulations reveal reflectance minima, characteristic of localized surface plasmon resonances. Here, electron energy-loss spectroscopy collected from individual indium NPs reveals the presence of LSPR at ≈ 8 eV, further supporting the promise of indium NP arrays on silicon for deep-UV plasmonics.

36 MATERIALS SCIENCE↗

Comparison of time-resolved photoluminescence and deep-level transient spectroscopy defect evaluations in an InAs nBn detector subjected to in situ and ex situ 63 MeV proton irradiation

Deep-level transient spectroscopy and temperature-dependent time-resolved photoluminescence experiments are performed on identical InAs nBn photodetector structures as a function of in situ and ex situ 63 MeV proton irradiation to assess their generation and recombination dynamics. Pre-irradiation, the n-type InAs absorbing region, exhibits a steadily increasing minority carrier lifetime with increasing temperature, providing evidence that excited minority carriers may be recombining via shallow defect levels. From deep-level transient spectroscopy, two features are found between 10 and 275 K: a low temperature broad “shoulder,” which suggests emission from multiple shallow electron defect levels with energies <29 meV and a high temperature minimum occurring at ∼230 K with an activation energy of 539 meV, which suggests a defect in the barrier layer in the device. Two similar nBn detectors are then subjected to 63 MeV proton irradiation in step doses and measured between steps. One experiment is performed in situ with an nBn held at ∼10 K during dosing, and the other experiment is performed ex situ with a similar nBn held at room temperature for dosing. The ex situ dosing results in an evaluation of the defect introduction rate that is three to four times lower than in situ due to partial annealing of the proton-induced displacement damage at room temperature. The results of these two experiments are then compared with the dose-dependent recombination rate analysis, resulting in an estimated recombination defect cross section of 1.6 × 10 −13 cm 2 for the shallow shoulder defect.

Carrasco, Rigo A. [Air Force Research Laboratory (↗

Deep learning models map rapid plant species changes from citizen science and remote sensing data

Anthropogenic habitat destruction and climate change are reshaping the geographic distribution of plants worldwide. However, we are still unable to map species shifts at high spatial, temporal, and taxonomic resolution. Here, we develop a deep learning model trained using remote sensing images from California paired with half a million citizen science observations that can map the distribution of over 2,000 plant species. Our model— Deepbiosphere— not only outperforms many common species distribution modeling approaches (AUC 0.95 vs. 0.88) but can map species at up to a few meters resolution and finely delineate plant communities with high accuracy, including the pristine and clear-cut forests of Redwood National Park. These fine-scale predictions can further be used to map the intensity of habitat fragmentation and sharp ecosystem transitions across human-altered landscapes. In addition, from frequent collections of remote sensing data, Deepbiosphere can detect the rapid effects of severe wildfire on plant community composition across a 2-y time period. These findings demonstrate that integrating public earth observations and citizen science with deep learning can pave the way toward automated systems for monitoring biodiversity change in real-time worldwide.

Gillespie, Lauren E.↗

Methane emission hotspots in a boreal forest-fen mosaic potentially linked to deep taliks

Permafrost thaw is transforming boreal forests into mosaics of wetlands and drier uplands. Topographic controls on hydrological and ecological conditions impact methane (CH 4 ) fluxes, contributing to uncertainty in local and regional CH 4 budgets and underlying drivers. The objective of this study was to explore CH 4 fluxes and their drivers in a transitioning boreal forest-fen ecosystem (Goldstream Valley, Alaska, USA). This landscape is characterized by thawing discontinuous permafrost and heterogeneous mosaics of fens, collapse-scar channels, and small mounds of permafrost soils. From a survey in July 2021, observed chamber CH4 fluxes included fen areas with intermediate to very high emissions (29.8–635.3 mg CH 4 m −2 d −1 ), clustered locations with CH 4 uptake (−2.11 to −0.7 mg CH 4 m −2 d −1 ), and three anomalous emission hotspots (342.4–772.4 mg CH 4 m −2 d −1 ) that were located near samples with lower emissions. Some surface and near-surface variables partially explained the spatial variation in CH 4 flux. Log-transformed CH 4 flux had a positive linear relationship with soil moisture at 20 cm depth ( R 2 = 0.31, p -value < 1e-5) and negative linear relationships with microtopography ( R 2 = 0.13, p -value < 0.006) and slope ( R 2 = 0.28, p -value < 2e-5). Methane emissions generally occurred in flat, wet, graminoid-dominated fens, whereas CH 4 uptake occurred on permafrost mounds dominated by feather mosses and woody vegetation. However, the CH 4 hotspots occurred on drier, slightly sloped locations with low or undetectable near-surface methanogen abundance, suggesting that CH 4 was produced in deeper soils. When the hotspot samples were omitted, log-transformed CH 4 flux had a positive linear relationship with near-surface methanogen abundance ( R 2 = 0.29, p -value = 0.0023), and stronger linear relationships with soil moisture, slope, and soil macronutrient concentrations. Our findings suggest that some CH 4 emission hotspots could arise from CH 4 in deep taliks. The inference that methanogenesis occurs in deep taliks was strengthened by the identification of intrapermafrost taliks across the study area using low-frequency geophysical induction. This study assesses surface spatial heterogeneity in the context of subsurface permafrost conditions and highlights the complexity of CH 4 flux patterns in transitioning forest-wetland ecosystems. To better inform regional CH 4 budgets, further research is needed to understand the spatial distribution of terrestrial CH 4 hotspots and to resolve their surface, near-surface, and subsurface drivers.

boreal↗

Evaluating probabilistic deep learning methods for uncertainty quantification of temperature downscaling

Deep learning (DL) has emerged as a promising tool for downscaling coarse-resolution climate data to high-resolution outputs, enabling improved regional climate predictions. A critical aspect of DL-based downscaling is the incorporation of uncertainty quantification (UQ), which enhances the interpretability and reliability of predictions—key factors for climate risk assessment and decision-making. This study develops a DL model to downscale 2 m temperature across the contiguous United States using reanalysis datasets. We systematically evaluate three epistemic UQ methods—deep ensembles (DEns), Monte Carlo dropout (MCD), and Flipout—based on their probabilistic accuracy, downscaling performance, sensitivity to geographical features, and computational efficiency. Results indicate that MCD generally outperforms Flipout and DEns in terms of calibration and downscaling accuracy. However, DEns demonstrate lower calibration errors in coastal regions, indicating its higher confidence within these areas. Flipout, in contrast, is more sensitive to elevation gradients and exhibits higher calibration errors in mountainous regions. Hence, the choice of UQ method for this task depends on the specific requirements of the application. For applications that prioritize overall calibration, downscaling accuracy, and computational efficiency, MCD is a strong candidate. These findings highlight the importance of selecting UQ methods based on application-specific requirements, such as geographical context and computational constraints. By addressing the trade-offs between UQ methods, this study provides actionable insights for improving the reliability, scalability, and utility of DL-based downscaling in climate science.

Environmental sciences↗

Data imbalance in drug response prediction: multi-objective optimization approach in deep learning setting

Abstract Drug response prediction (DRP) methods tackle the complex task of associating the effectiveness of small molecules with the specific genetic makeup of the patient. Anti-cancer DRP is a particularly challenging task requiring costly experiments as underlying pathogenic mechanisms are broad and associated with multiple genomic pathways. The scientific community has exerted significant efforts to generate public drug screening datasets, giving a path to various machine learning models that attempt to reason over complex data space of small compounds and biological characteristics of tumors. However, the data depth is still lacking compared to application domains like computer vision or natural language processing domains, limiting current learning capabilities. To combat this issue and improves the generalizability of the DRP models, we are exploring strategies that explicitly address the imbalance in the DRP datasets. We reframe the problem as a multi-objective optimization across multiple drugs to maximize deep learning model performance. We implement this approach by constructing Multi-Objective Optimization Regularized by Loss Entropy loss function and plugging it into a Deep Learning model. We demonstrate the utility of proposed drug discovery methods and make suggestions for further potential application of the work to achieve desirable outcomes in the healthcare field.

Biochemistry & Molecular Biology↗

Medium-Range Order, Density Fluctuations, and Activated Relaxation in the Equilibrated Deep Glass Regime

A successful microscopic theory of activated relaxation in metastable supercooled liquids is extended to the equilibrated deep glass regime. Surprisingly, the predicted power-law scaling connections of the dynamic barrier with diverse scalar order parameters (medium-range order correlation length, dimensionless compressibility, shear modulus) remain unchanged up to astronomically long timescales, despite a fundamental crossover of equilibrium thermodynamics and structure near the laboratory kinetic vitrification point. Quantitative tests against experiments on aged to equilibrium glass-forming liquids up to nearly 20 decades in time scale reveal good agreement. This conflicts with the idea of a crossover from super-Arrhenius to literal Arrhenius relaxation around the laboratory glass transition temperature, and supports the robustness of the theoretical idea that ultraslow dynamics is causally related to medium-range structural order. Here, new avenues of experimental and theoretical research in the deep glass regime are suggested.

Amorphous materials↗

Production of neutron-rich heavy nuclei in deep-inelastic 208 Pb + 208 Pb collisions within the stochastic mean-field theory

In deep-inelastic collisions of heavy nuclei, reaction products with a wide range of mass and charge are produced. Such collisions been considered as a possible way to produce superheavy nuclei, as an alternative to fusion reactions. To provide reliable theoretical predictions, it is desired to develop microscopic approaches that correctly and accurately describe nucleon transfer processes in dissipative collisions of heavy nuclei. The purpose of the present work is (1) to investigate the mechanism of nucleon transfers in dissipative collisions of two heavy nuclei, and (2) to explore possible pathways to produce neutron-rich heavy nuclei, through detailed theoretical analyses of fluctuations and correlations in nucleon transfers in 208 Pb + 208 Pb reactions. Three-dimensional time-dependent Hartree-Fock (TDHF) calculations are performed for the collisions of 208 Pb + 208 Pb at 𝐸 c.m. = 832, 936, and 1040 MeV, using the Skyrme SLy4d energy density functional. To calculate fluctuations and correlations in nucleon transfers, we employ the stochastic mean-field (SMF) theory, and the results are compared with another theoretical framework currently available, the time-dependent random phase approximation (TDRPA). Primary and secondary production cross sections are calculated with the SMF theory combined with a statistical model, GEMINI ++ . Using information of nucleon flow across a neck of colliding nuclei in TDHF calculations, we solve quantal diffusion equations for fluctuations and correlations in nucleon transfers based on the SMF theory. From the SMF calculations, we obtain the time evolution of diffusion coefficients as well as fluctuations and correlations in nucleon transfers for a range of initial orbital angular momenta. We compare the results of the SMF calculations with those of TDRPA, showing that TDRPA tends to predict substantially larger fluctuations and correlations in strongly damped collisions of heavy nuclei, which exhibit complex initial angular momentum dependence, while the SMF results provide almost constant (stable) values. Using the obtained fluctuations and correlations, we calculate primary and secondary production cross sections for the 208 Pb + 208 Pb collisions. From the results, we find that both lighter and heavier reaction products as compared to 208 Pb are produced for a wide region in the 𝑁−𝑍 plane as primary products, thanks to the quantal diffusion mechanism in the dissipative collisions. However, we show that cross sections for production of heavy nuclei with 𝑍 ≳ 90 or 𝑁 ≳ 135 are washed out due to secondary particle evaporation and/or fission processes. On the other hand, we find that there remain sizable cross sections for production of neutron-rich nuclei along 𝑁 = 126 with 𝑍< 82, even after secondary disintegration processes. We demonstrate that the secondary production cross sections depend weakly on incident energies, but lower (higher) energy is slightly preferred for production of nuclei with smaller (larger) atomic numbers as compared to 𝑍 = 82. Here, based on the microscopic SMF calculations, it has been shown that deep-inelastic collisions of heavy nuclei, such as 208 Pb + 208 Pb examined in this study, can be a promising means to produce neutron-rich heavy nuclei along 𝑁=126. Discrepancies between the SMF and TDRPA approaches are left unsolved for future investigations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Semi-inclusive deep-inelastic scattering on a polarized spin-1 target. II. Deuteron and spectator nucleon tagging

We develop the theoretical framework for semi-inclusive deep-inelastic scattering on a polarized spin-1 target and apply it to scattering on the polarized deuteron with spectator nucleon tagging. In Part I (previous article), we present the general form of the semi-inclusive cross section and polarization observables for the spin-1 target. In Part II (this article), we consider deep-inelastic scattering on the polarized deuteron with spectator nucleon tagging as a special case of target fragmentation. Methods of light-front quantization are employed to separate nuclear and hadronic structure in the high-energy process and achieve a composite description. The light-front wave function of the polarized deuteron is obtained from a rotationally covariant three-dimensional wave function in the center-of-mass frame of the proton-neutron system. The tagged structure functions are computed in the impulse approximation. The momentum and spin distribution of the active nucleon are controlled by the deuteron polarization and the detected spectator momentum (𝐷/𝑆 wave ratio). The cross section and spin asymmetries are evaluated for general deuteron polarization (vector and tensor, longitudinal and transverse) as functions of the spectator momentum. Tensor-polarized spin asymmetries of order unity are achieved for spectator momenta of approximately 300 MeV, which select configurations with a large 𝐷 wave. Sum rules for the tagged spin structure functions are derived. The results can be used for simulations of spectator tagging in future polarized fixed-target experiments (Jefferson Lab) or at the Electron-Ion Collider.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Maximum Entropy Principle in Deep Thermalization and in Hilbert-Space Ergodicity

We report universal statistical properties displayed by ensembles of pure states that naturally emerge in quantum many-body systems. Specifically, two classes of state ensembles are considered: those formed by (i) the temporal trajectory of a quantum state under unitary evolution or (ii) the quantum states of small subsystems obtained by partial, local projective measurements performed on their complements. These cases, respectively, exemplify the phenomena of “Hilbert-space ergodicity” and “deep thermalization.” In both cases, the resultant ensembles are defined by a simple principle: The distributions of pure states have maximum entropy, subject to constraints such as energy conservation, and effective constraints imposed by thermalization. We present and numerically verify quantifiable signatures of this principle by deriving explicit formulas for all statistical moments of the ensembles, proving the necessary and sufficient conditions for such universality under widely accepted assumptions, and describing their measurable consequences in experiments. We further discuss information-theoretic implications of the universality: Our ensembles have maximal information content while being maximally difficult to interrogate, establishing that generic quantum state ensembles that occur in nature hide (scramble) information as strongly as possible. Our results generalize the notions of Hilbert-space ergodicity to time-independent Hamiltonian dynamics and deep thermalization from infinite to finite effective temperature. Our work presents new perspectives to characterize and understand universal behaviors of quantum dynamics using statistical and information-theoretic tools.

Eigenstate thermalization↗

Enhancing synchrotron radiation micro-CT images using deep learning: an application of Noise2Inverse on bone imaging

In bone-imaging research, in situ synchrotron radiation micro-computed tomography (SRµCT) mechanical tests are used to investigate the mechanical properties of bone in relation to its microstructure. Low-dose computed tomography (CT) is used to preserve bone's mechanical properties from radiation damage, though it increases noise. To reduce this noise, the self-supervised deep learning method Noise2Inverse was used on low-dose SRµCT images where segmentation using traditional thresholding techniques was not possible. Simulated-dose datasets were created by sampling projection data at full, one-half, one-third, one-fourth and one-sixth frequencies of an in situ SRµCT mechanical test. After convolutional neural networks were trained, Noise2Inverse performance on all dose simulations was assessed visually and by analyzing bone microstructural features. Visually, high image quality was recovered for each simulated dose. Lacunae volume, lacunae aspect ratio and mineralization distributions shifted slightly in full, one-half and one-third dose network results, but were distorted in one-fourth and one-sixth dose network results. Following this, new models were trained using a larger dataset to determine differences between full dose and one-third dose simulations. Significant changes were found for all parameters of bone microstructure, indicating that a separate validation scan may be necessary to apply this technique for microstructure quantification. Noise present during data acquisition from the testing setup was determined to be the primary source of concern for Noise2Inverse viability. While these limitations exist, incorporating dose calculations and optimal imaging parameters enables self-supervised deep learning methods such as Noise2Inverse to be integrated into existing experiments to decrease radiation dose.

Obata, Yoshihiro (ORCID:0000000303659129)↗

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning↗