>
Fa   |   Ar   |   En
   efficient arabic hate speech detection via llama-3: a prompting and instruction-tuning approach  
   
نویسنده abbass anas khudhur ,faili heshaam
منبع aut journal of modeling and simulation - 2025 - دوره : 57 - شماره : 2 - صفحه:127 -138
چکیده    Cyberspace generates user-generated content daily, but also gives freedom of expression, potentially spreading hate speech and endangering minorities. so, there's the issue of identifying hate speech as quickly as possible to stop it from being shared. this is particularly challenging for arabic given its rich morphology and lack of good quality linguistic resources. in this article, we examine zero-shot and few-shot prompting for detecting arab hate speech using the llama-3-8b language model while also refining performance via supervised fine-tuning utilizing a custom instruction-based dataset. in the zero-shot approach, the outputs from the model are unstructured textual outputs, so we take the unstructured responses and run them through a lightweight tf-idf + logistic regression classifier to classify the responses in one of the predefined hate speech categories. to obtain better classification, we construct the instruction-based training set by creating tweet embeddings from arabic-bert and using k-means clustering to enforce semantic/topical variety. next, we use the gpt-4o model to generate representative instructions from each cluster and create an instruction-based fine-tuning data set. we then fine-tune llama-3-8b using qlora, which also allows the model to be fine-tuned with a lower memory footprint. the experimental results presented in this paper show that zero-shot and few-shot prompting achieved relatively low f1-scores of 42.2% and 45.0%, respectively, and instruction-tuning fine-tuning achieves the overall performance of an f1-score of 90.1%, which exceeds stronger benchmarks like arabert. our results exemplify the potential impact of instruction tuning and qlora-based fine-tuning over prompting-based approaches in low-resource contexts like arabic.
کلیدواژه hate speech detection in arabic ,zero-shot prompting ,few-shot prompting ,supervised fine-tuning ,llama-3-8b
آدرس university of tehran, alborz campus, iran, university of tehran, school of electrical and computer engineering, college of engineering, iran
پست الکترونیکی hfaili@ut.ac.ir
 
     
   
Authors
  
 
 

Copyright 2023
Islamic World Science Citation Center
All Rights Reserved