Search NASA⌕ Search

SEARCH · Search NASA

Results for “DNN accelerator”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

MIND-MAC: Multi-Level In-memory Quasi Non-Destructive MAC Operation in Compact 2T-nC FeRAM for Efficient DNN Accelerator

We present MIND-MAC, a compact 2T-nC FeRAM architecture that performs multi-level, quasi-non-destructive in-memory multiply–accumulate (MAC) for deep neural networks. By exploiting voltage-controlled partial domain switching in MFM capacitors and read-transistor amplification, the cell stores multi-bit weights and gates bit-serial inputs to produce an accumulated current on shared lines. We combine TCAD-extracted parasitics with experimentally calibrated ferroelectric models in SPICE to validate device-/circuit-level behavior, and validate multi-level sensing and QNRO with measurements on a fabricated 2T-3C test vehicle. An analytical system model maps MIND-MAC to a 6-GB main-memory in-memory compute (IMC) architecture and benchmarks VGG13 inference in 61.08 ms at 964.99 mJ. Results indicate high density, reduced rewrite overhead, and energy efficiency, positioning 2T-nC FeRAM as a promising IMC candidate for next-generation AI hardware.

36 MATERIALS SCIENCE↗

Throughput-Oriented and Accuracy-Aware DNN Training with BFloat16 on GPU

Deep Neural Networks (DNNs) have transformed the field of artificial intelligence and achieved extraordinary success in many areas. The training of DNNs is commonly compute and memory-intensive, which has resulted in several optimizations in the training phase. Among them, reduced precision is a typical and widely used technique to accelerate DNN training and reduce memory requirements. However, applying a widely adopted reduced precision format such as Float16 to all involved operations in DNN training is not optimal as the use of Float16 in some operations can hurt model accuracy. Meanwhile, additional optimizations including loss scaling and autocast techniques can mitigate the accuracy loss but lead to inherent overhead and inadequate use of reduced precision. In this work, we leverage another reduced precision format, BFloat16, and introduce a throughput-oriented and accuracy-aware approach to maximize the performance potential of DNN training. Since the high throughput provided by BFloat16 format is accompanied by low precision of the floating-point representation, this approach achieves high throughput by using BFloat16 on all DNN operations and avoids the accuracy loss through a customized accuracy-aware normalization. Results show that our approach outperforms the state-of-the-art mixed-precision training by 1.21x on an NVIDIA A100 GPU.

Xie, Zhen↗

Improving Deep Neural Networks’ Training for Image Classification With Nonlinear Conjugate Gradient-Style Adaptive Momentum

Momentum is crucial in stochastic gradient-based optimization algorithms for accelerating or improving training deep neural networks (DNNs). In deep learning practice, the momentum is usually weighted by a well-calibrated constant. However, tuning the hyperparameter for momentum can be a significant computational burden. In this article, we propose a novel adaptive momentum for improving DNNs training; this adaptive momentum, with no momentum-related hyperparame- ter required, is motivated by the nonlinear conjugate gradient (NCG) method. Stochastic gradient descent (SGD) with this new adaptive momentum eliminates the need for the momentum hyperparameter calibration, allows using a significantly larger learning rate, accelerates DNN training, and improves the final accuracy and robustness of the trained DNNs. For example, SGD with this adaptive momentum reduces classification errors for training ResNet110 for CIFAR10 and CIFAR100 from 5.25% to 4.64% and 23.75% to 20.03%, respectively. Furthermore, SGD, with the new adaptive momentum, also benefits adversarial training and, hence, improves the adversarial robustness of the trained DNNs.

97 MATHEMATICS AND COMPUTING↗

Evaluating Spatial Accelerator Architectures with Tiled Matrix-Matrix Multiplication

There is a growing interest in custom spatial accelerators for machine learning applications. These accelerators employ a spatial array of processing elements (PEs) interacting via custom buffer hierarchies and networks-on-chip. The efficiency of these accelerators comes from employing optimized dataflow (i.e., spatial/temporal partitioning of data across the PEs and fine-grained scheduling) strategies to optimize data reuse. The focus of this work is to evaluate these accelerator architectures using a tiled general matrix-matrix multiplication (GEMM) kernel. To do so, we develop a framework that finds optimized mappings (dataflow and tile sizes) for a tiled GEMM for a given spatial accelerator and workload combination, leveraging an analytical cost model for runtime and energy. Our evaluations over five spatial accelerators demonstrate that the tiled GEMM mappings systematically generated by our framework achieve high performance on various GEMM workloads and accelerators.

43 PARTICLE ACCELERATORS↗

Evaluating Spatial Accelerator Architectures with Tiled Matrix-Matrix Multiplication.

There is a growing interest in custom spatial accelerators for machine learning applications. These accelerators employ a spatial array of processing elements (PEs) interacting via custom buffer hierarchies and networks-on-chip. The efficiency of these accelerators comes from employing optimized dataflow (i.e., spatial/temporal partitioning of data across the PEs and fine-grained scheduling) strategies to optimize data reuse. The focus of this work is to evaluate these accelerator architectures using a tiled general matrix-matrix multiplication (GEMM) kernel. To do so, we develop a framework that finds optimized mappings (dataflow and tile sizes) for a tiled GEMM for a given spatial accelerator and workload combination, leveraging an analytical cost model for runtime and energy. Finally, our evaluations over five spatial accelerators demonstrate that the tiled GEMM mappings systematically generated by our framework achieve high performance on various GEMM workloads and accelerators.

42 ENGINEERING↗

A Study on the Impact of Temperature-Dependent Ferroelectric Switching Behavior in 3D Memory Architecture

The flourishing development of neural networks that require exponentially growing amounts of data has presented an elevated demand for memory footprint. To address this, researchers have been exploring hardware accelerators with innovative memory architectures like 3D memory. These 3D memory architectures offer enhanced storage capacity and processing capabilities, at a cost of rising on-chip temperature during operation. Hafnium Zirconium Oxide (HZO) based Ferroelectric Random Access Memory (FeRAM) is a promising nonvolatile memory candidate in neural network hardware accelerators for its outstanding write performance and reliability. However, its implementation in the architecture regarding the temperature-dependent ferroelectric switching behavior has not been well studied. In this work, we study the thermal impacts on polarization switching through experimental devices and simulation results. We conduct the circuit and architecture-level simulations to showcase that one can exploit this temperature rise to reduce FeRAM's write voltage and write energy due to its unique temperature-activated polarization switching mechanisms. As the on-chip temperature increases to 351K (ambient temperature at 300K) due to neural network workloads, the access energy per bit can be reduced by 27.6% when a dynamic write voltage is applied.

36 MATERIALS SCIENCE↗

Thermal Experiments for Fractured Rock Characterization: Theoretical Analysis and Inverse Modeling

Abstract Field‐scale properties of fractured rocks play a crucial role in many subsurface applications, yet methodologies for identification of the statistical parameters of a discrete fracture network (DFN) are scarce. We present an inversion technique to infer two such parameters, fracture density and fractal dimension, from cross‐borehole thermal experiments data. It is based on a particle‐based heat‐transfer model, whose evaluation is accelerated with a deep neural network (DNN) surrogate that is integrated into a grid search. The DNN is trained on a small number of the heat‐transfer model runs and predicts the cumulative density function of the thermal field. The latter is used to compute fine posterior distributions of the (to be estimated) parameters. Our synthetic experiments reveal that fracture density is well constrained by data, while fractal dimension is harder to determine. Adding nonuniform prior information related to the DFN connectivity improves the inference of this parameter.

Zhou, Zitong↗

Scalable Risk Assessment of Rare Events in Power Systems With Uncertain Wind Generation and Loads

Risk assessment of rare events has become increasingly important in power system planning and operation with the increasing integration of renewable energy and the presence of system uncertainties. However, quantifying the risk posed by rare events via the traditional method, i.e., Monte Carlo sampling (MCS), incurs substantial computational expense stemming from the vast ensemble of power flow simulations. To accelerate the assessment, this paper proposes a Deep Neural Network (DNN)-kernelized vector-valued Gaussian Process (VVGP) approach with excellent computational efficiency while maintaining high accuracy. Consequently, serving as a surrogate model for the power flow solver, the DNN-kernelized VVGP enables significantly faster but accurate risk assessment compared to the power flow solver. The developed surrogate model evaluates low-order N - k events that contain more than 90% instances by adeptly capturing the topological features while the high-order N - k events are assessed via a power flow solver, thereby striking a balance between computational efficiency and uncertainty quantification accuracy. Moreover, the model incorporates a Support Vector Machine (SVM) classifier to resample concerning low-probability tail events to counteract the biases potentially introduced during the DNN-kernelized VVGP evaluations. Simulations conducted on the modified IEEE 24-bus, 118-bus, and European 1354-bus systems demonstrate that the proposed method maintains the accuracy benchmark set by MCS while significantly reducing computational demands in large-scale power systems as compared to other state-of-the-art methods.

17 WIND ENERGY↗

Understanding Mixed Precision GEMM with MPGemmFI: Insights into Fault Resilience

Emerging deep learning workloads urgently need fast general matrix multiplication (GEMM). Thus, one of the critical features of machine-learning-specific accelerators such as NVIDIA Tensor Cores, AMD Matrix Cores, and Google TPUs is the support of mixed-precision enabled GEMM. For DNN models, lower-precision FP data formats and computation offer acceptable correctness but significant performance, area, and memory footprint improvement. While promising, the mixed-precision computation on error resilience remains unexplored. To this end, we develop a fault injection framework that systematically injects fault into the mixed-precision computation results. We investigate how the faults affect the accuracy of machine learning applications. Based on the characteristics of error resilience, we offer lightweight error detection and correction solutions that significantly improve the overall model accuracy by 75% if the models experience hardware faults. The solutions can be efficiently integrated into the accelerator's pipelines.

Fang, Bo↗

Harnessing ML Privacy by Design Through Crossbar Array Non-idealities

Deep Neural Networks (DNNs), handling computeand data-intensive tasks, often utilize accelerators like Resistiveswitching Random-access Memory (RRAM) crossbar for energyefficient in-memory computation. Despite RRAM’s inherent nonidealities causing deviations in DNN output, this study transforms the weakness into strength. By leveraging RRAM non-idealities, the research enhances privacy protection against Membership Inference Attacks (MIAs), which reveal private information from training data. RRAM non-idealities disrupt MIA features, increasing model robustness and revealing a privacy-accuracy tradeoff. Empirical results with four MIAs and DNNs trained on different datasets demonstrate significant privacy leakage reduction with a minor accuracy drop (e.g., up to 2.8% for ResNet-18 with CIFAR-100).

artificial intelligence↗

Enhanced Ionization Charge Identification in the Short-Baseline Neutrino Program Neutrino Detectors with Deep Neural Networks

We present a deep neural net-based region of interest detection method (DNN ROI) for signal processing in the liquid argon time projection chambers of the Short-Baseline Neutrino (SBN) Program, SBND and ICARUS. DNN ROI addresses limitations of the traditional wire-by-wire thresholding algorithm by leveraging the full two-dimensional detector readout and cross-plane matching information. To account for detector performance variations, we explore training with augmented samples. We find that DNN ROI outperforms the traditional method in both low-level ROI identification performance and high-level reconstruction metrics for high-energy cosmic and accelerator neutrino interaction products, while also being more robust against detector variations, with or without sample augmentation.

Abratenko, P. [Tufts U.]↗

Generating An Advanced Cross-section Library For HTGR Pebble Bed Depletion Calculations Using Reduced-Order Model Generation Techniques

For code development, Advanced Reactor Technologies - Gas Cooled Reactors Program (ART-GCR) rely on a collaboration with the Nuclear Energy Advanced Modeling and Simulation (NEAMS) program, but the cross sections generation and the methodology definition is part of this program area goals. Based on previous studies in FY23, the size of microscopic cross section libraries increases rapidly with the number of tabulations, requiring significant amount of memory and drastically slowing down the Griffin calculations when evaluating cross sections via the multivariate linear interpolation approach. Rising to these challenges, this work investigates constructing Reduced-order Models (ROMs) for the multi-group microscopic cross sections to accelerate the cross section evaluation in Griffin. A database of multigroup cross sections is first collected considering all possible parameters that a designer could change for optimization. Down-selection of the ROM techniques afterward shows Deep Neural Network (DNN) as the best candidate when jointly consider memory efficiency, predictive accuracy, computational cost, scalability, flexibility and ease of implementation of the algorithms in comparison to the multidimensional interpolation. This work develops a specific interface that enables the cross section predictions using pre-trained DNN models into Griffin leveraging the existing ROM capabilities. DNNs have been trained for all isotopes for use in Griffin. Preliminary Griffin testing shows that DNNs exhibit exceptional predictive accuracy and the use of DNNs provides orders of magnitude improvement in memory efficiency compared to conventional interpolation techniques. With such ROM techniques, it holds great promise to further increase the fidelity of the Pebble Bed Reactor (PBR) simulation by increasing the number of tabulations/state variables during cross section evaluation, while maintaining the computational cost affordable in Griffin.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

LUNA: LUT-Based Neural Architecture for Fast and Low-Cost Qubit Readout

Qubit readout is a critical operation in quantum computing systems, which maps the analog response of qubits into discrete classical states. Deep neural networks (DNNs) have recently emerged as a promising solution to improve readout accuracy . Prior hardware implementations of DNN-based readout are resource-intensive and suffer from high inference latency, limiting their practical use in low-latency decoding and quantum error correction (QEC) loops. This paper proposes LUNA, a fast and efficient superconducting qubit readout accelerator that combines low-cost integrator-based preprocessing with Look-Up Table (LUT) based neural networks for classification. The architecture uses simple integrators for dimensionality reduction with minimal hardware overhead, and employs LogicNets (DNNs synthesized into LUT logic) to drastically reduce resource usage while enabling ultra-low-latency inference. We integrate this with a differential evolution based exploration and optimization framework to identify high-quality design points. Our results show up to a 10.95x reduction in area and 30% lower latency with little to no loss in fidelity compared to the state-of-the-art. LUNA enables scalable, low-footprint, and high-speed qubit readout, supporting the development of larger and more reliable quantum computing systems.

Farooq, M. A. [Arizona State U., Tempe]↗

Deep Learning Accelerated History Matching and Forecasting in Geologic CO 2 Sequestration [Slides]

We have successfully developed a predictive workflow for the Geologic CO 2 Sequestration using a deep learning model based on the Fourier Neural Operator. The workflow has high accuracy for predicting pressure & saturation during the injection and post-injection periods and has decent accuracy for predicting water and CO 2 production rates during injection period. It is necessary to train exclusive DNN models for pressure prediction in long term GCS, but for saturation prediction, we can use a single model to predict it.

58 GEOSCIENCES↗

XploreNAS : Explore Adversarially Robust and Hardware-efficient Neural Architectures for Non-ideal Xbars

Compute In-Memory platforms such as memristive crossbars are gaining focus as they facilitate acceleration of Deep Neural Networks (DNNs) with high area and compute efficiencies. However, the intrinsic non-idealities associated with the analog nature of computing in crossbars limits the performance of the deployed DNNs. Furthermore, DNNs are shown to be vulnerable to adversarial attacks leading to severe security threats in their large-scale deployment. Thus, finding adversarially robust DNN architectures for non-ideal crossbars is critical to the safe and secure deployment of DNNs on the edge. This work proposes a two-phase algorithm-hardware co-optimization approach called XploreNAS that searches for hardware efficient and adversarially robust neural architectures for non-ideal crossbar platforms. We use the one-shot Neural Architecture Search approach to train a large Supernet with crossbar-awareness and sample adversarially robust Subnets therefrom, maintaining competitive hardware efficiency. Our experiments on crossbars with benchmark datasets (SVHN, CIFAR10, CIFAR100) show up to ~8–16% improvement in the adversarial robustness of the searched Subnets against a baseline ResNet-18 model subjected to crossbar-aware adversarial training. We benchmark our robust Subnets for Energy-Delay-Area-Products (EDAPs) using the Neurosim tool and find that with additional hardware efficiency–driven optimizations, the Subnets attain ~1.5–1.6× lower EDAPs than ResNet-18 baseline.

97 MATHEMATICS AND COMPUTING↗

Machine learning for design principles for single atom catalysts towards electrochemical reactions

Machine learning (ML) integrated density functional theory (DFT) calculations have recently been used to accelerate the design and discovery of heterogeneous catalysts such as single atom catalysts (SACs) through the establishment of deep structure–activity relationships. Here, this review provides recent progress in the ML-aided rational design of heterogeneous catalysts with the focus on SACs in terms of structure–activity relationships, feature importance analysis, high-throughput screening, stability, and metal–support interactions for electrochemistry. Support vector machine (SVM), random forest regression (RFR), and deep neural networks (DNN) along with atomic properties are mainly used for the design of SACs. The ML results have shown that the number of electrons in the d orbital, oxide formation enthalpy, ionization energy, Bader charge, d-band center, and enthalpy of vaporization are mainly the most important parameters for the defining of the structure–activity relationships for electrochemistry. However, the black-box nature of ML techniques occasionally makes a physical interpretation of descriptors, such as the Bader charge, d-band center, and enthalpy of vaporization, non-trivial. At the current stage, ML application is limited by the lack of a large and high-quality database. Future prospects for the development of a large database and a generalized ML algorithm for SAC design are discussed to give insights for further studies in this field.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine Learning for Joint Quality Control

The use of lightweight material combinations has been highly demanded in manufacturing automotive structures. However, making robust dissimilar material joints of such lightweight materials is still challenging. A significant barrier to achieving high-quality and repeatable joint performance is a deficient understanding of the relationship between the welding process, joint attributes, and joint performance. In this context, welding factors refer to material, equipment, environment, and process parameters, while joint features comprise specific microstructural attributes of the weld such as nugget size, heat affected zone (HAZ) topology, intermetallic layer thickness, and sheet thickness reduction. Joint performance is quantified in terms of strength (e.g., tensile shear, coach peel, cross-tension), weld size, and hardness, among other factors. While there have been many attempts to establish this process-structure-property relationship by developing a model derived from the associated physics and first principles, the complexity of the joining processes compounded by the complex interactions with different materials in an automotive assembly line environment, has hindered the usefulness of such attempts. The complexity is further exacerbated using different stacking materials, especially comprising dissimilar material combinations. In practice, the common approach has been the laborious process of creating welds, characterizing them, and then physically testing them through experimentation. With the emergence of artificial intelligence (AI) methods, an alternative pathway to eliciting the desired process-structure-property relationship at an accelerated pace is to use a data-driven approach by employing machine-learning (ML) techniques. This approach is benefitted by the availability of large streams of data, generated through years of research and testing by original equipment manufacturers, in the form of material, process, environmental, equipment, microstructural, and bulk-scale performance information from multimodal, multiscale sensors making measurements from laboratory-scale to production-scale processes. During Phase I efforts, which ended in fiscal year (FY) 2021, the Oak Ridge National Laboratory and Pacific Northwest National Laboratory (ORNL/PNNL) team demonstrated the effectiveness of different ML/AI frameworks in modeling complex relationships between resistance spot welding (RSW) process parameters, weld attributes, and joint properties using a subset of data from General Motors (GM). In FY 2022, the project team further refined and expanded their respective ML models to analyze additional welds with new weld stack-ups and materials to enhance the ML model predictive capability. ORNL extended its unified deep neural networks (DNN) ML training and prediction framework with new data streams of process parameters, and PNNL extended its model describing RSW process parameters’ associations with weld attributes. In FY 2023, the project team completed the development of the AI/ML architecture for analyzing aluminum/steel joints manufactured by GM via RSW and transitioned into the inline welding quality monitoring task for steel/steel RSW joints provided by GM.

36 MATERIALS SCIENCE↗

Physics constrained learning for data-driven inverse modeling from sparse observations

Deep neural networks (DNN) have been used to model nonlinear relations between physical quantities. Those DNNs are embedded in physical systems described by partial differential equations (PDE) and trained by minimizing a loss function that measures the discrepancy between predictions and observations in some chosen norm. This loss function often includes the PDE constraints as a penalty term when only sparse observations are available. As a result, the PDE is only satisfied approximately by the solution. However, the penalty term typically slows down the convergence of the optimizer for stiff problems. We present a new approach that trains the embedded DNNs while numerically satisfying the PDE constraints. We develop an algorithm that enables differentiating both explicit and implicit numerical solvers in reverse-mode automatic differentiation. This allows the gradients of the DNNs and the PDE solvers to be computed in a unified framework. We demonstrate that our approach enjoys faster convergence and better stability in relatively stiff problems compared to the penalty method. Furthermore, our approach allows for the potential to solve and accelerate a wide range of data-driven inverse modeling, where the physical constraints are described by PDEs and need to be satisfied accurately.

97 MATHEMATICS AND COMPUTING↗