Engineering topics
Thiagarajan, Jayaraman
Publications and source records attributed to Thiagarajan, Jayaraman.
Systems and methods for customizing kernel machines with deep neural networks
A method including receiving an input data set. The input data set can include one of a feature domain set or a kernel matrix. The method also can include constructing dense embeddings using: (i) Nyström approximations on the input data set when the input data set comprises the kernel matrix, and (ii) clustered Nyström approximations on the input data set when the input data set comprises the feature domain set. The method additionally can include performing representation learning on each of the dense embeddings using a multi-layer fully-connected network for each of the dense embeddings to generate latent representations corresponding to each of the dense embeddings. The method further can include applying a fusion layer to the latent representations corresponding to the dense embeddings to generate a combined representation. The method additionally can include performing classification on the combined representation. Other embodiments of related systems and methods are also disclosed.
Out of Distribution Detection with Neural Network Anchoring
This is code to reproduce and build on OOD detection from the paper "Out of Distribution Detection with Neural Network Anchoring". Our goal here is to exploit heteroscedastic temperature scaling as a calibration strategy for out of distribution (OOD) detection. Heteroscedasticity here refers to the fact that the optimal temperature parameter for each sample can be different, as opposed to conventional approaches that use the same value for the entire distribution. To enable this, we propose a new training strategy called anchoring that can estimate appropriate temperature values for each sample, leading to state-of-the-art OOD detection performance across several benchmarks. Using NTK theory, we show that this temperature function estimate is closely linked to the epistemic uncertainty of the classifier, which explains its behavior. In contrast to some of the best-performing OOD detection approaches, our method does not require exposure to additional outlier datasets, custom calibration objectives, or model ensembling. Through empirical studies with different OOD detection settings - far OOD, near OOD, and semantically coherent OOD - we establish a highly effective OOD detection approach.
Enabling machine learning-ready HPC ensembles with Merlin
With the growing complexity of computational and experimental facilities, many scientific researchers are turning to machine learning (ML) techniques to analyze large scale ensemble data. With complexities such as multi-component workflows, heterogeneous machine architectures, parallel file systems, and batch scheduling, care must be taken to facilitate this analysis in a high performance computing (HPC) environment. Here, we present Merlin, a workflow framework to enable large ML-friendly ensembles of scientific HPC simulations. By augmenting traditional HPC with distributed compute technologies, Merlin aims to lower the barrier for scientific subject matter experts to incorporate ML into their analysis. As a producer–consumer workflow model, Merlin enables multi-machine, cross-batch job, dynamically allocated yet persistent workflows capable of utilizing surge-compute resources. Key features of Merlin are a flexible HPC-centric interface, low per-task overhead, multi-tiered fault recovery, and a hierarchical sampling algorithm that allows for $\mathscr{O}$(N) task execution and $\mathscr{O}$(N ln N) task queuing to ensembles of millions of tasks. In addition to Merlin’s design, we test the algorithm’s performance in an HPC center and demonstrate the ability to enqueue 40 million simulations in 100 s, with a 30 millisecond per-task overhead that is independent of ensemble size. Finally, we describe some example applications that Merlin has enabled on leadership-class HPC resources, such as the ML-augmented optimization of nuclear fusion experiments and the calibration of infectious disease models to study the progression of and possible mitigation strategies for COVID-19.
Universal image representation based on a multimodal graph
A system for classifying a target image with segments having attributes is provided. The system generates a graph for the target image that includes vertices representing segments of the image and edges representing relationships between the connected vertices. For each vertex, the system generates a subgraph that includes the vertex as a home vertex and neighboring vertices representing segments of the target image within a neighborhood of the segment represented by the home vertex. The system applies an autoencoder to each subgraph to generate latent variables to represent the subgraph. The system applies a machine learning algorithm to a feature vector comprising a universal image representation of the target image that is derived from the generated latent variables of the subgraphs to generate a classification for the target image.