DOE OSTI · 1862329
Con Connections: Detecting Fraud from Abstracts using Topological Data Analysis
Abstract
In this paper we present a novel approach for identifying fraudulent papers from their titles and abstracts. The premise of the approach is that there are holes in the presentation of the approach and findings of fraudulent research papers. As an abstract is intended to highlight key features of the approach as well as important conclusions the authors seek to determine if the assumed existence of holes can be identified from analysis of abstracts alone. The data set considered is derived from papers sharing a single author with labels determined based on a formal linguistic analysis of the complete documents. To detect these logical and literary holes we utilize techniques from topological data analysis which summarizes data based on the presence of multi-dimensional, topological holes. We find that, in fact, topological features derived through a combination of techniques in natural language processing and time-series analysis allow for superior detection of the fraudulent papers than the natural language processing tools alone. Thus we conclude that the connections and holes present in the abstracts of research cons contributes to an ability to infer the scientific validity of the corresponding work.
Keep this discovery
Explore connections, maps & timelines
Tymochko, Sarah J., Chaput, Julien A., Doster, Timothy J., Purvine, Emilie AH, Warley, Jackson T., Emerson, Tegan H.. 2021-12-16. Con Connections: Detecting Fraud from Abstracts using Topological Data Analysis. https://doi.org/10.1109/icmla52953.2021.00069
Cite the original work for its findings. Save a collection to share your selection of sources.