Evaluation of deep neural network architectures for authorship obfuscation of Portuguese texts

Autor:	Antônio Marcos Rodrigues Franco, Ítalo Cunha, Leonardo B. Oliveira
Jazyk:	angličtina
Rok vydání:	2024
Předmět:	Authorship obfuscation Privacy Natural language processing Artificial intelligence Computational linguistics. Natural language processing P98-98.5
Zdroj:	Natural Language Processing Journal, Vol 9, Iss , Pp 100107- (2024)
Druh dokumentu:	article
ISSN:	2949-7191
DOI:	10.1016/j.nlp.2024.100107
Popis:	Preserving authorship anonymity is paramount to protect activists, freedom of expression, and critical journalism. Although there are several mechanisms to provide anonymity on the Internet, one can still identify anonymous authors through their writing style. With the advances in neural network and natural language processing research, the success of a classifier when identifying the author of a text is growing. On the other hand, new approaches that use recurrent neural networks for automatic generation of obfuscated texts have also arisen to fight anonymity adversaries. In this work, we evaluate two approaches that use neural networks to generate obfuscated texts. The first approach uses Generative Adversarial Networks to train an encoder–decoder to transform sentences from an input style into a target style. The second one trains an auto encoder with Gradient Reversal Layer to learn invariant representations. In our experiments, we compared the efficiency of both techniques when removing the stylistic attributes of a text and preserving its original semantics. Our evaluation on real texts clarifies each technique’s trade-offs for Portuguese texts and provides guidance on practical deployment.
Databáze:	Directory of Open Access Journals
Externí odkaz:	https://doaj.org/article/3b563a19b0154d19821ce267ccf7a865 Zobrazit plný text záznamu View record in DOAJ