Arlim, Novitasari and Siagian, Al Hafiz Akbar Maulana and Riyanto, Slamet and Rodiah, Rodiah and Kushadiani, Siti Kania and Hakim, Shidiq Al and Setiorini, Retno Asihanti and Apriani, Niken Fitria and Arianty, Rini and Susetianingtias, Diana Tri (2024) Dictionary-based extraction of hyperbole and swear words for sarcasm detection in Indonesian Tweets. International Journal of Information Technology, 17 (5). pp. 2671-2678. ISSN 2511-2104, 2511-2112
Full text not available from this repository. (Request a copy)Abstract
Detecting sarcasm in texts presents a formidable challenge due to the absence of clear signals, unlike in verbal conversations. Relying on hashtags for sarcasm detection may lack accuracy and standardization, while human annotation of sarcasm can be subjective. Moreover, detecting sarcasm in a low-resource language like Indonesian presents unique challenges. In this work, we utilize hyperbole, i.e., interjection (INJ), intensifier (INS), capital letters (CL), elongated words (EW), and punctuation marks (PUNC), as features for classifiers to detect sarcasm in Indonesian texts. We also propose using swear words (SW) as features to deal with this task. As our other proposal in this work, we consider a dictionary-based method to extract these features more effectively. To evaluate our work, we use a dataset containing Indonesian tweets collected from Drone Emprit. Experimental results show that incorporating hyperbolic features with SW improves sarcasm detection performance compared to using these features separately. Our proposed feature extraction method also produces better classification results than the baseline.
| Item Type: | Article |
|---|---|
| Uncontrolled Keywords: | Sarcasm, Hyperbole, Swear words, Indonesia, Dictionary |
| Subjects: | Computers, Control & Information Theory Language |
| Depositing User: | Defryan Aprisandani |
| Date Deposited: | 07 Sep 2026 04:40 |
| Last Modified: | 07 Sep 2026 04:40 |
| URI: | https://karya.brin.go.id/id/eprint/60171 |


Dimensions
Dimensions