Search NASA⌕ Search

DOE OSTI · code-100450

Learning Universal Authorship Representations

Abstract

This code contains all the utilities required to reproduce the results of our EMNLP 2021 paper "Learning Universal Authorship Representations". It contains the utilities required for download the datasets, training our model, and performing all evaluations necessary for reproducing the results in the paper. Here's the abstract of our work: Determining whether two documents were composed by the same author, also known as authorship verification, has traditionally been tackled using statistical methods. Recently, authorship representations learned using neural networks have been found to outperform alternatives, particularly in large-scale settings involving hundreds of thousands of authors. But do such representations learned in a particular domain transfer to other domains? Or are these representations inherently entangled with domain-specific features? To study these questions, we conduct the first large-scale study of cross-domain transfer for authorship verification considering zero-shot transfers involving three disparate domains: Amazon reviews, fanfiction short stories, and Reddit comments. We find that although a surprising degree of transfer is possible between certain do- mains, it is not so successful between others. We examine properties of these domains that influence generalization and propose simple but effective methods to improve transfer.

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Rivera Soto, RafaelA, Miano, OliviaE, Ordonez, JuanitaA. 2021-10-05. Learning Universal Authorship Representations. https://doi.org/10.11578/dc.20230216.2

Cite the original work for its findings. Save a collection to share your selection of sources.