Text annotation automation for hate speech detection using SVM-classifier based on feature extraction

S., Saifullah and N.H., Cahyana and Y., Fauziah and Aribowo, Agus Sasmito and F.A., Dwiyanto and R., Drezewski (2024) Text annotation automation for hate speech detection using SVM-classifier based on feature extraction. In: AIP Conf. Proc. 3167, 040003 (2024).

[img] Text
IoT-based monitoring system reduces oil and gas pipeline leaks Improving sustainability and safety.pdf
Restricted to Registered users only

Download (265kB)

Abstract

This article aims to develop a semi-supervised method for automatically annotating hate speech in social media using natural language processing (NLP) techniques. The approach is based on a Support Vector Machine (SVM) classifier that combines feature extraction algorithms, including ensemble meta-learners and meta-vectorizers. The system was trained on a dataset of 13,169 elements, and the results show that the accuracy of the model is highly dependent on the feature extraction method used. The optimal automatic annotation was achieved using TF-IDF feature extraction, resulting in an accuracy of 92.5%. The implications of this study are that automated hate speech annotation using NLP techniques can significantly improve the accuracy, reliability, and inclusiveness of identifying hate speech online. The results of this study suggest that SVM and TF-IDF are the most suitable methods for this task.

Item Type: Conference or Workshop Item (Paper)
Divisions: Faculty of Information and Communication Technology
Depositing User: NUR FARISAH JAFRIN
Date Deposited: 21 Jul 2026 08:25
Last Modified: 21 Jul 2026 08:25
URI: http://eprints.utem.edu.my/id/eprint/29836
Statistic Details: View Download Statistic

Actions (login required)

View Item View Item