Search NASA⌕ Search

SEARCH · Search NASA

Results for “skip connection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Tailor : Altering Skip Connections for Resource-Efficient Inference

Deep neural networks use skip connections to improve training convergence. However, these skip connections are costly in hardware, requiring extra buffers and increasing on- and off-chip memory utilization and bandwidth requirements. In this article, we show that skip connections can be optimized for hardware when tackled with a hardware-software codesign approach. We argue that while a network’s skip connections are needed for the network to learn, they can later be removed or shortened to provide a more hardware-efficient implementation with minimal to no accuracy loss. We introduceTailor, a codesign tool whose hardware-aware training algorithm gradually removes or shortens a fully trained network’s skip connections to lower the hardware cost.Tailorimproves resource utilization by up to 34% for block random access memories (BRAMs), 13% for flip-flops (FFs), and 16% for look-up tables (LUTs) for on-chip, dataflow-style architectures.Tailorincreases performance by 30% and reduces memory bandwidth by 45% for a two-dimensional processing element array architecture.

Computer Science↗

Rethinking skip connections in Spiking Neural Networks with Time-To-First-Spike coding

Time-To-First-Spike (TTFS) coding in Spiking Neural Networks (SNNs) offers significant advantages in terms of energy efficiency, closely mimicking the behavior of biological neurons. In this work, we delve into the role of skip connections, a widely used concept in Artificial Neural Networks (ANNs), within the domain of SNNs with TTFS coding. Our focus is on two distinct types of skip connection architectures: (1) addition-based skip connections, and (2) concatenation-based skip connections. We find that addition-based skip connections introduce an additional delay in terms of spike timing. On the other hand, concatenation-based skip connections circumvent this delay but produce time gaps between after-convolution and skip connection paths, thereby restricting the effective mixing of information from these two paths. To mitigate these issues, we propose a novel approach involving a learnable delay for skip connections in the concatenation-based skip connection architecture. This approach successfully bridges the time gap between the convolutional and skip branches, facilitating improved information mixing. We conduct experiments on public datasets including MNIST and Fashion-MNIST, illustrating the advantage of the skip connection in TTFS coding architectures. Additionally, we demonstrate the applicability of TTFS coding on beyond image recognition tasks and extend it to scientific machine-learning tasks, broadening the potential uses of SNNs.

97 MATHEMATICS AND COMPUTING↗

Densely Connected G-invariant Deep Neural Networks with Signed Permutation Representations

We introduce and investigate, for finite groups G, G-invariant deep neural network (GDNN) architectures with ReLU activation that are densely connected- i.e., include all possible skip connections. In contrast to other G-invariant architectures in the literature, the preactivations of theG-DNNs presented here are able to transform by signed permutation representations (signed perm-reps) of G. Moreover, the individual layers of the G-DNNs are not required to be G-equivariant; instead, the preactivations are constrained to be G-equivariant functions of the network input in a way that couples weights across all layers. The result is a richer family of G-invariant architectures never seen previously. We derive an efficient implementation of G-DNNs after a reparameterization of weights, as well as necessary and sufficient conditions for an architecture to be "admissible"- i.e., nondegenerate and inequivalent to smaller architectures. We include code that allows a user to build a G-DNN interactively layer-by-layer, with the final architecture guaranteed to be admissible. We show that there are far more admissible G-DNN architectures than those accessible with the "concatenated ReLU" activation function from the literature. Finally, we apply G-DNNs to two example problems--(1) multiplication in --1, 1} (with theoretical guarantees) and (2) 3D object classification--finding that the inclusion of signed perm-reps significantly boosts predictive performance compared to baselines with only ordinary (i.e., unsigned) perm-reps.

97 MATHEMATICS AND COMPUTING↗

Comparative Assessment of U-Net-Based Deep Learning Models for Segmenting Microfractures and Pore Spaces in Digital Rocks

Segmentation of high-resolution X-ray microcomputed tomography (µCT) images is crucial in digital rock physics (DRP), affecting the characterization and analysis of microscale phenomena in the porous media. The complexity of geological structures and nonideal scanning conditions pose significant challenges to conventional image segmentation approaches. Motivated by the recent increasing popularity of deep learning (DL) techniques in image processing, this work undertakes a comparative study of DL models, specifically U-Net and its variants, for segmenting multiple targets with distinguished features in digital rocks, including discrete fracture networks (DFNs), pore spaces, and solid rock. Particularly, DFNs have a smaller volumetric fraction over others, bringing in a substantial challenge of imbalanced segmentation. The primary focus is to evaluate the architecture and feature enhancement strategies of various DL models, including U-Net, attention U-Net, residual U-Net, U-Net++, and residual U-Net++. The models were designed as 2.5D, utilizing a central 2D image and its two adjacent upper and lower 2D images as input to provide a pseudo-3D context. In addition, because the ground truth of segmentation was unknown for real-world digital rocks, we created a benchmark data set following the inverse operations of segmentation. The data synthesis started from the label images (i.e., solid rock, pore spaces, and DFNs), followed by simulating partial volume blurring, adding random background noise, and introducing ring artifacts to mimic real raw X-ray µCT images. The data set, which included various rock types (i.e., sandstone and artificial data), scanning resolution, and magnitudes of noise and artifacts, was divided into training and testing data sets with a 90% and 10% ratio, respectively. Moreover, in addition to the conventional pixel-wise evaluation metrics, the physics-based metric of the lattice-Boltzmann method (LBM) simulated permeability provided more comprehensive assessments. The results demonstrated that the residual connections, nested architectures, and redesigned skip connections contribute to the model performance and give the residual U-Net++ the highest accuracy. The improvements were mainly on the boundaries and small targets, especially the DFNs, which dominate the interconnectivity and therefore affect the permeability greatly. This study also rigorously evaluated the efficiency and generalization of each model, demonstrating that the sophisticated architectures achieved excellent practicability and maintained robust performance on completely unseen data, ensuring their suitability for diverse and challenging DRP applications.

58 GEOSCIENCES↗

A Comparative Study of Deep Learning Models for Fracture and Pore Space Segmentation in Synthetic Fractured Digital Rocks

This study focuses on the comparative study of deep learning (DL) models for pore space and discrete fracture networks (DFNs) segmentation in synthetic fractured digital rocks, specifically targeting low-permeability rock formations, such as shale and tight sandstones. Accurate characterization of pore space and DFNs is critical for subsequent property analysis and fluid flow modeling. Four DL models, SegNet, U-Net, U-Net-wide, and nested U-Net (i.e., U-Net++), were trained, validated, and tested using synthetic datasets, including input and label image pairs with varying properties. The model performance was assessed regarding pixel-wise metrics, including the F1 score and pixel-wise difference maps. In addition, the physics-based metrics were considered for further analysis, including sample porosity and absolute permeability. Particularly, We first simulated the permeability of porous media containing only pore space and then simulated the permeability of porous media with DFNs added. The difference between these two values is used to quantify the connectivity of segmented DFNs, which is an important parameter for low-permeability rocks. The pixel-wise metrics showed that the nested U-Net model outperformed the rest of the DL models in pore space and DFNs segmentation, with the SegNet model exhibiting the second-best performance. Particularly, nested U-Net enhanced segmentation accuracy for challenging boundary pixels affected by partial volume effects. The U-Net-wide model achieved improved accuracy compared to the U-Net model, which indicated the influence of parameter numbers. Similarly, nested U-Net has the closest match to the ground truth of physics-based metrics, including the porosity of pore space and DFNs, and the permeability difference quantifying the connectivity of DFNs. The findings highlight the effectiveness of DL models, especially the U-Net++ model with nested architecture and redesigned skip connections, in accurately segmenting pore spaces and DFNs, which are crucial for pore-scale fluid flow and transport simulation in low-permeability rocks.

Wang, Hongsheng↗

Simultaneously improving accuracy and computational cost under parametric constraints in materials property prediction tasks

Abstract Modern data mining techniques using machine learning (ML) and deep learning (DL) algorithms have been shown to excel in the regression-based task of materials property prediction using various materials representations. In an attempt to improve the predictive performance of the deep neural network model, researchers have tried to add more layers as well as develop new architectural components to create sophisticated and deep neural network models that can aid in the training process and improve the predictive ability of the final model. However, usually, these modifications require a lot of computational resources, thereby further increasing the already large model training time, which is often not feasible, thereby limiting usage for most researchers. In this paper, we study and propose a deep neural network framework for regression-based problems comprising of fully connected layers that can work with any numerical vector-based materials representations as model input. We present a novel deep regression neural network, iBRNet, with branched skip connections and multiple schedulers, which can reduce the number of parameters used to construct the model, improve the accuracy, and decrease the training time of the predictive model. We perform the model training using composition-based numerical vectors representing the elemental fractions of the respective materials and compare their performance against other traditional ML and several known DL architectures. Using multiple datasets with varying data sizes for training and testing, We show that the proposed iBRNet models outperform the state-of-the-art ML and DL models for all data sizes. We also show that the branched structure and usage of multiple schedulers lead to fewer parameters and faster model training time with better convergence than other neural networks. Scientific contribution: The combination of multiple callback functions in deep neural networks minimizes training time and maximizes accuracy in a controlled computational environment with parametric constraints for the task of materials property prediction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Toward ultra-efficient high-fidelity predictions of wind turbine wakes: Augmenting the accuracy of engineering models with machine learning

This study proposes a novel machine learning (ML) methodology for the efficient and cost-effective prediction of high-fidelity three-dimensional velocity fields in the wake of utility-scale turbines. The model consists of an autoencoder convolutional neural network with U-Net skipped connections, fine-tuned using high-fidelity data from large-eddy simulations (LES). The trained model takes the low-fidelity velocity field cost-effectively generated from the analytical engineering wake model as input and produces the high-fidelity velocity fields. The accuracy of the proposed ML model is demonstrated in a utility-scale wind farm for which datasets of wake flow fields were previously generated using LES under various wind speeds, wind directions, and yaw angles. Comparing the ML model results with those of LES, the ML model was shown to reduce the error in the prediction from 20% obtained from the Gauss Curl hybrid (GCH) model to less than 5%. In addition, the ML model captured the non-symmetric wake deflection observed for opposing yaw angles for wake steering cases, demonstrating a greater accuracy than the GCH model. The computational cost of the ML model is on par with that of the analytical wake model while generating numerical outcomes nearly as accurate as those of the high-fidelity LES.

Mechanics↗

Overcoming sparse datasets with multi-task learning as applied to high entropy alloys

Abstract The design of novel High Entropy Alloys for use in high-temperature applications is an area of active interest due to their potential to provide exceptional properties compared to conventional alloys. Since the increased popularity of machine learning, an important cog in the design process has been training surrogate models on alloy properties. However, these Single-Task models are trained on individual mechanical properties and do not take advantage of the relatedness between properties. Multi-Task models can capture the interdependencies between tasks, leading to potentially more accurate predictions for all tasks. In this paper, we investigate if Multi-Task models can show improvement over Single-Task models when used for predicting the mechanical properties of these alloys. To ensure fair evaluation between the models, we apply L 0 regularization and skip connections to the models, which allows them to adjust the number of model parameters and depth for optimal performance. We find that the Multi-Task models can leverage task relationships to perform better than Single-Task models, especially for high amounts of missing data in the tasks. Furthermore, adding simple auxiliary targets can boost Multi-Task performance even further despite not being effective as input descriptors to single-task models themselves. We anticipate that the proposed strategies can achieve more accurate predictions and consequently enable better design capabilities for such data-constrained domains without incurring much additional computational cost.

Debnath, Arindam (ORCID:0000000194274499)↗

Glass Refraction Distortion Object Detection via Abstract Features

Glass reflection and refraction lead to missing and distorted object feature data, affecting the accuracy of object detection. In order to solve the above problems, this paper proposed a glass refraction distortion object detection via abstract features. The number of parameters of the algorithm is reduced by introducing skip connections and expansion modules with different expansion rates. The abstract feature information of the object is extracted by binary cross-entropy loss. Meanwhile, the abstract feature distance between the object domain and source domain is reduced by a loss function, which improves the accuracy of object detection under glass interference. To verify the effectiveness of the algorithm in this paper, the GRI dataset is produced and made public on GitHub. The algorithm of this paper is compared with the current state-of-the-art Deep Face, VGG Face, TBE-CNN, DA-GAN, PEN-3D, LMZMPM, and the average detection accuracy of our algorithm is 92.57% at the highest, and the number of parameters is only 5.13 M.

Cai, Lei↗

Revising Apetrei’s bounding volume hierarchy construction algorithm to allow stackless traversal

Stackless traversal is a technique to speed up range queries by avoiding usage of a stack during the tree traversal. One way to achieve that is to transform a given binary tree to store a left child and a skip-connection (also called an escape index). In general, this operation requires an additional tree traversal during the tree construction. For some tree structures, however, it is possible to achieve the same result at a reduced cost. We propose one such algorithm for a GPU hierarchy construction algorithm proposed by Karras in Karras 2012. Furthermore, we show that our algorithm also works with the improved algorithm proposed by Apetrei in Apetrei 2014, despite a different ordering of the internal nodes. We achieve that by modifying Apetrei’s algorithm to restore the original Karras’ ordering of the internal nodes. Using the modified algorithm, we show how to construct a hierarchy suitable for a stackless traversal in a single bottom-up pass.

97 MATHEMATICS AND COMPUTING↗

Phonon-informed Neural Thermal Scattering (NeTS) Optimization for Crystalline Graphite and Beryllium Metal

Fast neutrons born from fission lose energy through scattering interactions in the process of slowing-down. As neutrons thermalize to the order of $k$ $b$ $T$ (where $k$ $b$ is the Boltzmann constant, and $T$ is the temperature of the medium), their de Broglie wavelength and energy approaches the order of inter-atomic spacing and quantized lattice vibrations, i.e., phonons. At thermal energies, the thermal scattering law (TSL), i.e., $S$($α, β$), captures crystal binding contributions to the total reaction rate, or cross section. This dimensionless material property describes the energy ($β$) and momentum ($α$) exchanges available in a medium. Currently, $S$($α, β$) is evaluated in the Full Law Analysis Scattering System Hub (FLASSH) code for discrete inputs and stored as ENDF/B File 7 for 0-phonon elastic (MT 2) and n-phonon inelastic (MT 4) processes. Further processing recasts $S$($α, β$) into cumulative distribution functions for sampling post-collision scattering kinematics. In practice, interpolation schemes are employed to access data between tabulated values. An improvement to this juncture of the nuclear data pipeline is supplying cross sections on-the-fly (OTF), as has been developed for the un-resolved resonance region to minimize non-physical interpolation errors. This capability may improve simulation accuracy for accident and transient analyses, where rapidly varying changes in temperature and pressure are difficult to predict beforehand. To do so, deep artificial neural networks (ANNs) can be employed which collapse non-linear, complex data into a lightweight dictionary of neural weights and biases. This has been successfully demonstrated for the hydrogen in light water $S$($α, β$) dataset in the form of a Neural Thermal Scattering (NeTS) module. In this work, the NeTS framework is extended to consider the impact of material-dependent dynamical features on optimal neural pre-processing and architecture design decisions, such as number of neurons per hidden layer, residual skip connections and neural depth. New NeTS modules for crystalline graphite and beryllium metal illuminate a novel correlation between dynamical nonlinearity and optimal neural parametrization when deploying $S$($α, β$) on-the-fly.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Holographic thermal correlators: a tale of Fuchsian ODEs and integration contours

We analyze real-time thermal correlation functions of conserved currents in holographic field theories using the grSK geometry, which provides a contour prescription for their evaluation. We demonstrate its efficacy, arguing that there are situations involving components of conserved currents, or derivative interactions, where such a prescription is, in fact, essential. To this end, we first undertake a careful analysis of the linearized wave equations in AdS black hole backgrounds and identify the branch points of the solutions as a function of (complexified) frequency and momentum. All the equations we study are Fuchsian with only regular singular points that for the most part are associated with the geometric features of the background. Special features, e.g., the appearance of apparent singular points at the horizon, whence outgoing solutions end up being analytic, arise at higher codimension loci in parameter space. Using the grSK geometry, we demonstrate that these apparent singularities do not correspond to any interesting physical features in higher-point functions. We also argue that the Schwinger-Keldysh collapse and KMS conditions, implemented by the grSK geometry, continue to hold even in the presence of such singularities. For charged black holes above a critical charge, we furthermore demonstrate that the energy density operator does not possess an exponentially growing mode, associated with ‘pole-skipping’, from one such apparent singularity. Our analysis suggests that the connection between the scrambling physics of black holes and energy transport has, at best, a limited domain of validity.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

New horizon symmetries, hydrodynamics, and quantum chaos

Abstract We generalize the formulation of horizon symmetries presented in previous literature to include diffeomorphisms that can shift the location of the horizon. In the context of the AdS/CFT duality, we show that horizon symmetries can be interpreted on the boundary as emergent low-energy gauge symmetries. In particular, we identify a new class of horizon symmetries that extend the so-called shift symmetry, which was previously postulated for effective field theories of maximally chaotic systems. Additionally, we comment on the connections of horizon symmetries with bulk calculations of out-of-time-ordered correlation functions and the phenomenon of pole-skipping.

Physics↗

NEXT Generation Energy Technologies for Connected and Automated On-Road Vehicles (NEXTCAR Phase I & II)

The Ohio State University’s ARPA-E NEXTCAR project was a multi-phase, multi-year research, development, and demonstration program focused on improving the energy efficiency of connected and automated vehicles (CAVs). The team developed and validated advanced vehicle motion and powertrain control algorithms that coordinate propulsion and automation systems to optimize energy use. Key technologies included Dynamic Skip Fire engine control, predictive eco-driving functions such as Eco-Approach and Departure (Eco-AND) and Eco-Adaptive Cruise Control (Eco-ACC), and powertrain-agnostic optimization frameworks for hybrid, plug-in hybrid, and battery electric vehicles. The project successfully demonstrated up to 30% energy-efficiency improvement during real-world testing at the Transportation Research Center and the American Center for Mobility. The outcomes provide a foundation for scalable, cost-effective deployment of energy-optimized CAV technologies across the automotive industry.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Stereoisomer-dependent unimolecular kinetics of 2,4-dimethyloxetanyl peroxy radicals

2,4,dimethyloxetane is an important cyclic ether intermediate that is produced from hydroperoxyalkyl (QOOH) radicals in the low-temperature combustion of n-pentane. However, the reaction mechanisms and rates of consumption pathways remain unclear. In the present work, the pressure- and temperature-dependent kinetics of seven cyclic ether peroxy radicals, which stem from 2,4,dimethyloxetane via H-abstraction and O 2 addition, were determined. The automated kinetic workflow code, KinBot, was used to model the complexity of the chemistry in a stereochemically resolved manner and solve the resulting master equations from 300–1000 K and from 0.01–100 atm. The main conclusions from the calculations include (i) diastereomeric cyclic ether peroxy radicals show significantly different reactivities, (ii) the stereochemistry of the peroxy radical determines which QOOH isomerization steps are possible, (iii) conventional QOOH decomposition pathways, such as cyclic ether formation and HO 2 elimination, compete with ring-opening reactions, which primarily produce OH radicals, the outcome of which is sensitive to stereochemistry. Ring-opening reactions lead to unique products, such as unsaturated, acyclic peroxy radicals, that form direct connections with species present in other chemical kinetics mechanisms through "cross-over" reactions that may complicate the interpretation of experimental results from combustion of n-pentane and, by extension, other alkanes. For example, one cross-over reaction involving 1-hydroperoxy-4-pentanone-2-yl produces 2-(hydroperoxymethyl)-3-butanone-1-yl, which is an iso-pentane-derived ketohydroperoxide (KHP). At atmospheric pressure, the rate of chemical reactions of all seven peroxy radicals compete with that of collisional stabilization, resulting in well-skipping reactions. However, at 100 atm, only one out of seven peroxy radicals undergoes significant well-skipping reactions. Here, the rates produced from the master equation calculations provide the first foundation for the development of detailed sub-mechanisms for cyclic ether intermediates. In addition, analysis of the complex reaction mechanisms of 2,4-dimethyloxetane-derived peroxy radicals provides insights into the effects of stereoisomers on reaction pathways and product yields.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗