Document Type

Conference Proceeding

Publication Title

FLP 2022 - 3rd Workshop on Figurative Language Processing, Proceedings of the Workshop

Abstract

We introduce EUREKA, an ensemble-based approach for performing automatic euphemism detection. We (1) identify and correct potentially mislabelled rows in the dataset, (2) curate an expanded corpus called EuphAug, (3) leverage model representations of Potentially Euphemistic Terms (PETs), and (4) explore using representations of semantically close sentences to aid in classification. Using our augmented dataset and kNN-based methods, EUREKA was able to achieve state-of-the-art results on the public leaderboard of the Euphemism Detection Shared Task, ranking first with a macro F1 score of 0.881.

First Page

111

Last Page

117

DOI

10.18653/v1/2022.flp-1.15

Publication Date

12-2022

Keywords

Computational linguistics

Comments

Archived thanks to ACL Anthology

License: CC by 4.0 DEED

Uploaded 30 November 2022

Share

COinS