Schizophrenia-Mimicking Layers Outperform Conventional Neural Network Layers
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
The combination of neural networks and quantum Monte Carlo methods has arisen as a promising path forward for highly accurate electronic structure calculations. Previous proposals have combined equivariant neural network layers with a final antisymmetric layer in order to satisfy the antisymmetry requirements of the electronic wavefunction. However, to date it is unclear if one can represent antisymmetric functions of physical interest, and it is difficult to precisely measure the expressiveness of the antisymmetric layer. Here, this work attempts to address this problem by introducing explicitly antisymmetrized universal neural network layers. This approach has a computational cost which increases factorially with respect to the system size, but we are nonetheless able to apply it to small systems to better understand how the structure of the antisymmetric layer affects its performance. We first introduce a generic antisymmetric (GA) neural network layer, which we use to replace the entire antisymmetric layer of the highly accurate ansatz known as the FermiNet. We demonstrate that the resulting FermiNet-GA architecture can yield effectively the exact ground state energy for small atoms and molecules. We then consider a factorized antisymmetric (FA) layer which more directly generalizes the FermiNet by replacing the products of determinants with products of antisymmetrized neural networks. We find, interestingly, that the resulting FermiNet-FA architecture does not significantly outperform the FermiNet. This strongly suggests that the sum of products of antisymmetries is a key limiting aspect of the FermiNet architecture. To explore this further, we investigate a slight modification of the FermiNet, called the full determinant mode, which replaces each product of determinants with a single combined determinant. We find that the full single-determinant FermiNet closes a large part of the gap between the standard single-determinant FermiNet and FermiNet-GA on small atomic and molecular problems. Surprisingly, on the nitrogen molecule at a dissociating bond length of 4.0 Bohr, the full single-determinant FermiNet can outperform the largest standard FermiNet calculation with 64 determinants.
Network communication has been proven to be a very important tool and a key factor in the recent development and progress of the power grid operation. It is also considered as the foundation for the smart grid because information and communication are integrated into electricity distribution to achieve reliable and accurate knowledge of the power grid. In previous years, absorbing energy from substations and delivering it to customers was the only type of interaction we knew between utility companies and customers. Presently, the growing connections of small distributed generation units caused by the cost reduction of most of the technologies used in generation and storage of electrical energy, along with the potential benefits of renewable energy have pushed many researchers to look into the improvement of information and communication technologies (ICT) in order to ensure a bidirectional flow of power and data. Moreover, the evolution of information and communication technologies and its applications to smart grid have converted the smart grid into a cyber-physical system where vulnerabilities and additional security challenges such as cyber-threats and cyber-attacks have emerged. Previously, we have demonstrated that using machine learning-based processing on data gathered from communication networks and the power grid was a promising solution for detecting cyber threats by implementing a co-simulation of cyber-security for cross-layer strategy. Since the majority of the challenges observed can only be solved in the network communication layer, we present in this work a physics-based state estimation model of the communication network system towards enhanced cyber-physical security of the smart grid. Information integration with the previously developed machine learning model is developed, providing a enhanced cyber-physical security application for the smart grid. Easy-to-implement model, without hard-to-derive parameters, highlight potential aspects of the model for real-life applications.
Not Available
Abstract not provided.
Explore the source record for details and available documents.
Abstract In complex networked systems theory, an important question is how to evaluate the system robustness to external perturbations. With this task in mind, I investigate the propagation of noise in multi-layer networked systems. I find that, for a two layer network, noise originally injected in one layer can be strongly amplified in the other layer, depending on how well-connected are the complex networks in each layer and on how much the eigenmodes of their Laplacian matrices overlap. These results allow to predict potentially harmful conditions for the system and its sub-networks, where the level of fluctuations is important, and how to avoid them. The analytical results are illustrated numerically on various synthetic networks.
Widespread integration of social media into daily life has fundamentally changed the way society communicates, and, as a result, how individuals develop attitudes, personal philosophies, and worldviews. The excess spread of disinformation and misinformation due to this increased connectedness and streamlined communication has been extensively studied, simulated, and modeled. Less studied is the interaction of many pieces of misinformation, and the resulting formation of attitudes. We develop a framework for the simulation of attitude formation based on exposure to multiple cognitions. We allow a set of cognitions with some implicit relational topology to spread on a social network, which is defined with separate layers to specify online and offline relationships. An individual’s opinion on each cognition is determined by a process inspired by the Ising model for ferromagnetism. We conduct experimentation using this framework to test the effect of topology, connectedness, and social media adoption on the ultimate prevalence of and exposure to certain attitudes.
Nonlinear complex network-coupled systems typically have multiple stable equilibrium states. Following perturbations or due to ambient noise, the system is pushed away from its initial equilibrium, and, depending on the direction and the amplitude of the excursion, it might undergo a transition to another equilibrium. It was recently demonstrated [M. Tyloo, J. Phys. Complex. 3 03LT01 (2022)] that layered complex networks may exhibit amplified fluctuations. Here, I investigate how noise with system-specific correlations impacts the first escape time of nonlinearly coupled oscillators. Interestingly, I show that, not only the strong amplification of the fluctuations is a threat to the good functioning of the network but also the spatial and temporal correlations of the noise along the lowest-lying eigenmodes of the Laplacian matrix. Finally, I analyze first escape times on synthetic networks and compare noise originating from layered dynamics to uncorrelated noise.
Full-waveform inversion (FWI) is an accurate imaging approach for modeling the velocity structure by minimizing the misfit between recorded and predicted seismic waveforms. However, the strong nonlinearity of FWI resulting from fitting oscillatory waveforms can trap the optimization in local minima. We have adopted a neural-network-based full-waveform inversion (NNFWI) method that integrates deep neural networks with FWI by representing the velocity model with a generative neural network. Neural networks can naturally introduce spatial correlations as regularization to the generated velocity model, which suppresses noise in the gradients and mitigates local minima. Furthermore, the velocity model generated by neural networks is input to the same partial differential equation (PDE) solvers used in conventional FWI. The gradients of the neural networks and PDEs are calculated using automatic differentiation, which back propagates gradients through the acoustic PDEs and neural network layers to update the weights of the generative neural network. Experiments on 1D velocity models, the Marmousi model, and the 2004 BP model determine that NNFWI can mitigate local minima, especially for imaging high-contrast features such as salt bodies, and it significantly improves the inversion in the presence of noise. Adding dropout layers to the neural network model also allows analyzing the uncertainty of the inversion results through Monte Carlo dropout. NNFWI opens a new pathway to combine deep learning and FWI for exploiting the characteristics of deep neural networks and the high accuracy of PDE solvers. Because NNFWI does not require extra training data and optimization loops, it provides an attractive and straightforward alternative to conventional FWI.
The emerging multipolar international security environment represents a fundamental restructuring of global nuclear balance of power to include two nuclear peer competitors, growing non-peer nuclear threats, and concerns of nuclear latency from both allies and adversaries. Conflicts in the grey zone, cyber operations, mis- and disinformation campaigns, and emerging disruptive technologies like drones, and hypersonic missiles are becoming more prevalent. These present a risk of cross-domain and multi-domain conflicts that may not follow known escalatory patterns. In order to prepare for the new deterrence environment, it is critical to have quantitative and qualitative understandings of these cross-domain conflicts, their potential for escalation, and which systems they may impact. To that end, our team created a Multi-Layer Network (MLN) model of ‘integrated deterrence’ where instruments of national power are modeled as individual network graph layers that include efforts from all domains. We then evaluate the potential for escalation against escalation scenarios. Analysis of the escalation scenarios is then used to identify insights of potential risk and escalation within integrated deterrence.
Recent developments integrating micromechanics and neural networks offer promising paths for rapid predictions of the response of heterogeneous materials with similar accuracy as direct numerical simulations. The deep material network is one such approaches, featuring a multi-layer network and micromechanics building blocks trained on anisotropic linear elastic properties. Once trained, the network acts as a reduced-order model, which can extrapolate the material’s behavior to more general constitutive laws, including nonlinear behaviors, without the need to be retrained. However, current training methods initialize network parameters randomly, incurring inevitable training and calibration errors. Here, we introduce a way to visualize the network parameters as an analogous unit cell and use this visualization to “quilt” patches of shallower networks to initialize deeper networks for a recursive training strategy. The result is an improvement in the accuracy and calibration performance of the network and an intuitive visual representation of the network for better explainability.
One or more aspects of the present disclosure are directed to network optimization solutions provided as software agents (applications) executed on network nodes in a heterogenous multi-vendor environment to provide cross-layer network optimization and ensure availability of network resources to meet associated Quality of Experience (QoE) and Quality of Service (QoS). In one aspect, a network slicing engine is configured to receive at least one request from at least one network endpoint for access to the heterogeneous multi-vendor network for data transmission; receive information on state of operation of a plurality of communication links between the plurality of nodes; determine a set of data transmission routes for the request; assign a network slice for serving the request; determine, from the set of data transmission routes, an end-to-end route for the network slice; and send network traffic associated with the request using the network slice and over the end-to-end route.
Not Available
Quantization, effective Neural Network architecture, and efficient accelerator hardware are three important design paradigms to maximize accuracy and efficiency. Mixed Precision Quantization is a process of assigning different precision to different Neural Network layers for optimized inference. Neural Architecture Search (NAS) is a process of automatically designing the neural network for a task and can also be extended to search for the precision of each weight and activation matrix. In this paper, we develop the following three methods: (i) Fast Differentiable Hardware-aware Mixed Precision Quantization Search method to find optimal precision, (ii) Joint Differentiable hardware-aware Architecture and Mixed Precision Quantization Co-search, (iii) Joint Accelerator, Architecture, and Precision triple co-search to find best possibilities in all the three worlds. We demonstrate the effectiveness of our proposed methods targeting Bitfusion accelerator by searching mixed precision models on MobilenetV2. We achieve better accuracy-latency trade-off models than the manually designed and previously proposed search methods.