Main Article Content

Abstract

The unregulated proliferation of hate speech, defined as any expression of beliefs that incite hatred, discrimination, or violence against individuals or social groups (Fortuna & Nunes, 2018), in social media has called into question the efficacy of removal-based content moderation and has driven the need to explore alternative approaches. AI-generated counterspeech could be a scalable and cost-effective intervention (Ashida & Komachi, 2022). Counterspeech is defined as a direct response to hate speech aimed at refuting or undermining it (Wachs et al., 2023). Research in this growing field is fragmented across computational linguistics and social sciences, lacking systematic methods to define and evaluate effective counterspeech, revealing a gap between AI's technical capabilities and evidenced social impact.
This review aims to examine evidence from these two fields to identify the differing conceptualizations of counterspeech, map the collaborative approaches used for counterspeech generation and evaluation, and examine the alignment between stated research goals (e.g., reducing hate speech in digital environments) and the operational measures employed to identify them.
Only a few studies evaluate counterspeech in real-world interactive settings or measure long-term impact. We identify some key misalignments, such as the reliance on text metrics and human or artificial annotators as proxies for complex social interventions, instead of integrating psychological and behavioral measures. We propose a linked approach to counterspeech that links theoretical frameworks, procedures, measurements and research goals from the perspective of both psychology and computational linguistics to identify promising directions for future interdisciplinary collaboration and more socially meaningful AI-assisted interventions.

Keywords

Counter-Speech Artificial Intelligence Large Language Models Human Augmentation

Article Details

How to Cite
Cosola, A., Paciello, M., Pollini, A., & D’Errico, F. (2026). Beyond automation: the collaborative role of AI in countering Hate Speech. A scoping review. Journal of E-Learning and Knowledge Society, 22(1), 75-90. https://doi.org/10.20368/1971-8829/1136378

References

  1. Ashida, M., & Komachi, M. (2022). Towards Automatic Generation of Messages Countering Online Hate Speech and Microaggressions. In K. Narang, A. Mostafazadeh Davani, L. Mathias, B. Vidgen, & Z. Talat (A c. Di), Proceedings of the Sixth Workshop on Online Abuse and Harms (WOAH) (pp. 11–23). Association for Computational Linguistics.
  2. Baez Santamaria, S., Gomez Adorno, H., & Markov, I. (2024). Contextualized Graph Representations for Generating Counter-Narratives against Hate Speech. In Y. Al-Onaizan, M. Bansal, & Y.-N. Chen (A c. Di), Findings of the Association for Computational Linguistics: EMNLP 2024 (pp. 7664–7674). Association for Computational Linguistics.
  3. Bandura, A., Barbaranelli, C., Caprara, G. V., & Pastorelli, C. (1996). Mechanisms of moral disengagement in the exercise of moral agency. Journal of Personality and Social Psychology, 71, 364–374.
  4. Bär, D., Maarouf, A., & Feuerriegel, S. (2024). Generative AI may backfire for counterspeech (arXiv:2411.14986). arXiv.
  5. Benesch, S. (2016). Counterspeech on Twitter: A Field Study.
  6. Bennie, M., Zhang, D., Xiao, B., Cao, J., Liu, C. X., Meng, J., & Tripp, A. (2025). PANDA -- Paired Anti-hate Narratives Dataset from Asia: Using an LLM-as-a-Judge to Create the First Chinese Counterspeech Dataset (arXiv:2501.00697). arXiv.
  7. Bilewicz, M., Tempska, P., Leliwa, G., Dowgiałło, M., Tańska, M., Urbaniak, R., & Wroczyński, M. (2021). Artificial intelligence against hate: Intervention reducing verbal aggression in the social network environment. Aggressive Behavior, 47(3), 260–266.
  8. Bonaldi, H., Damo, G., Ocampo, N. B., Cabrio, E., Villata, S., & Guerini, M. (2024). Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech Countering (arXiv:2410.03466). arXiv.
  9. Bonaldi, H., Dellantonio, S., Tekiroğlu, S. S., & Guerini, M. (2022). Human-Machine Collaboration Approaches to Build a Dialogue Dataset for Hate Speech Countering. In Y. Goldberg, Z. Kozareva, & Y. Zhang (A c. Di), Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (pp. 8031–8049). Association for Computational Linguistics.
  10. Cecconi, C., Poggi, I., & D’Errico, F. (2020). Schadenfreude: malicious joy in social media interactions. Frontiers in Psychology, 11, 558282.
  11. Cima, L., Miaschi, A., Trujillo, A., Avvenuti, M., Dell’Orletta, F., & Cresci, S. (2025). Contextualized Counterspeech: Strategies for Adaptation, Personalization, and Evaluation. WWW - Proc. ACM Web Conf., 5022–5033.
  12. Chung, Y.-L., Abercrombie, G., Enock, F., Bright, J., & Rieser, V. (2024). Understanding Counterspeech for Online Harm Mitigation. ArXiv.org.
  13. Chung, Y.-L., & Bright, J. (2024). On the Effectiveness of Adversarial Robustness for Abuse Mitigation with Counterspeech. In K. Duh, H. Gomez, & S. Bethard (A c. Di), Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) (pp. 6988–7002). Association for Computational Linguistics.
  14. Darley, J. M., & Latane, B. (1968). Bystander intervention in emergencies: Diffusion of responsibility. Journal of Personality and Social Psychology, 8(4, Pt.1), 377–383.
  15. Ding, X., Ping, K., Gunturi, U. S., Carik, B., Stil, S., Wilhelm, L. T., Daryanto, T., Hawdon, J., Lee, S. W., & Rho, E. H. (2025). Designing Human-AI Collaboration to Support Learning in Counterspeech Writing (arXiv:2410.03032). arXiv.
  16. Fanton, M., Bonaldi, H., Tekiroğlu, S. S., & Guerini, M. (2021). Human-in-the-Loop for Data Collection: A Multi-Target Counter Narrative Dataset to Fight Online Hate Speech. In C. Zong, F. Xia, W. Li, & R. Navigli (A c. Di), Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) (pp. 3226–3240). Association for Computational Linguistics.
  17. Fortuna, P., & Nunes, S. (2018). A Survey on Automatic Detection of Hate Speech in Text. ACM Computing Surveys, 51(4), 1–30.
  18. Garland, J., Ghazi‐Zahedi, K., Young, J. G., Hébert‐Dufresne, L., & Galesic, M. (2022). Impact and dynamics of hate and counter‐speech online. EPJ Data Science, 11(3):3.
  19. Hassan, S., & Alikhani, M. (2023). DisCGen: A Framework for Discourse-Informed Counterspeech Generation (arXiv:2311.18147). arXiv.
  20. Hengle, A., Padhi, A. K., Bandhakavi, A., & Chakraborty, T. (2025). CSEval: Towards Automated, Multi-Dimensional, and Reference-Free Counterspeech Evaluation using Auto-Calibrated LLMs. In L. Chiruzzo, A. Ritter, & L. Wang (A c. Di), Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) (pp. 5402–5419). Association for Computational Linguistics.
  21. Hengle, A., Padhi, A., Singh, S., Bandhakavi, A., Akhtar, M. S., & Chakraborty, T. (2024). Intent-conditioned and Non-toxic Counterspeech Generation using Multi-Task Instruction Tuning with RLAIF. In K. Duh, H. Gomez, & S. Bethard (A c. Di), Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) (pp. 6716–6733). Association for Computational Linguistics.
  22. Hong, L., Luo, P., Blanco, E., & Song, X. (2024). Outcome-Constrained Large Language Models for Countering Hate Speech (arXiv:2403.17146). arXiv.
  23. Horta Ribeiro, M., Hosseinmardi, H., West, R., & Watts, D. J. (2023). Deplatforming did not decrease Parler users’ activity on fringe social media. PNAS Nexus, 2(3).
  24. Kumar, A., Bandhakavi, A., & Chakraborty, T. (2025). Counterspeech the ultimate shield! Multi-Conditioned Counterspeech Generation through Attributed Prefix Learning (arXiv:2505.11958). arXiv.
  25. Mathew, B., Saha, P., Tharad, H., Rajgaria, S., Singhania, P., Maity, S. K., Goyal, P., & Mukherje, A. (2019). Thou shalt not hate: Countering Online Hate Speech (arXiv:1808.04409). arXiv.
  26. Mun, J., Allaway, E., Yerukola, A., Vianna, L., Leslie, S.-J., & Sap, M. (2023). Beyond Denouncing Hate: Strategies for Countering Implied Biases and Stereotypes in Language (arXiv:2311.00161). arXiv.
  27. Podolak, J., Łukasik, S., Balawender, P., Ossowski, J., Piotrowski, J., Bakowicz, K., & Sankowski, P. (2024). LLM generated responses to mitigate the impact of hate speech. In Y. Al-Onaizan, M. Bansal, & Y.-N. Chen (A c. Di), Findings of the Association for Computational Linguistics: EMNLP 2024 (pp. 15860–15876). Association for Computational Linguistics.
  28. Poggi, I., D’Errico, F., & Vincze, L. (2013). Comments by words, face and body. Journal on Multimodal User Interfaces, 7(1), 67-78.
  29. Prasannan, P., Kumaresan, P. K., Rajiakodi, S., Subalalitha, C. N., & Chakravarthi, B. R. (2025). Counter-speech generation for homophobic and transphobic social media content in Malayalam. Social Network Analysis and Mining, 15(1), 87.
  30. Sahoo, N. R., Beria, G. P., & Bhattacharyya, P. (2024). IndicCONAN: A Multilingual Dataset for Combating Hate Speech in Indian Context. Proceedings of the AAAI Conference on Artificial Intelligence, 38(20), 22313–22321.
  31. Sportelli, C., Cicirelli, P. G., Paciello, M., Corbelli, G., & D'Errico, F. (2025). “Let's Make the Difference!” Promoting Hate Counter‐Speech in Adolescence Through Empathy and Digital Intergroup Contact. Journal of Community & Applied Social Psychology, 35(1), e70028
  32. Tekiroğlu, S. S., Chung, Y.-L., & Guerini, M. (2020). Generating Counter Narratives against Online Hate Speech: Data and Strategies. In D. Jurafsky, J. Chai, N. Schluter, & J. Tetreault (A c. Di), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 1177–1190). Association for Computational Linguistics.
  33. Tricco, AC, Lillie, E, Zarin, W, O'Brien, KK, Colquhoun, H, Levac, D, Moher, D, Peters, MD, Horsley, T, Weeks, L, Hempel, S et al. PRISMA extension for scoping reviews (PRISMA-ScR): checklist and explanation. Ann Intern Med. 2018,169(7):467-473. doi: 10.7326/M18-0850
  34. Wachs, S., Valido, A., Espelage, D. L., Castellanos, M., Wettstein, A., & Bilz, L. (2023). The relation of classroom climate to adolescents’ countering hate speech via social skills: A positive youth development perspective. Journal of Adolescence, 95(6), 1127–1139.
  35. Wang, H., Pan, Y., Song, X., Zhao, X., Hu, M., & Zhou, B. (2024). F^2RL: Factuality and Faithfulness Reinforcement Learning Framework for Claim-Guided Evidence-Supported Counterspeech Generation. In Y. Al-Onaizan, M. Bansal, & Y.-N. Chen (A c. Di), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 4457–4470). Association for Computational Linguistics.
  36. Wang, M., Ma, S., Li, N., Zhang, P., Li, C., Gu, N., & Lu, T. (2026). Echoes of Norms: Investigating Counterspeech Bots’ Influence on Bystanders in Online Communities (arXiv:2603.03687). arXiv.
  37. Wu, C., Wang, Y., Zhang, Y., Wang, H., & Pang, Y. (2025). Confront hate with AI: How AI-generated counter speech helps against hate speech on social Media? Telematics and Informatics, 101, 102304.
  38. Zhu, W., & Bhat, S. (2021). Generate, Prune, Select: A Pipeline for Counterspeech Generation against Online Hate Speech (arXiv:2106.01625). arXiv.
  39. Zhu, L., Wang, X., & Wang, X. (2023, October 26). JudgeLM: Fine-tuned Large Language Models are Scalable Judges. arXiv.Org.
  40. Zubiaga, I., Soroa, A., & Agerri, R. (2024). A LLM-based Ranking Method for the Evaluation of Automatic Counter-Narrative Generation. In Y. Al-Onaizan, M. Bansal, & Y.-N. Chen (A c. Di), Findings of the Association for Computational Linguistics: EMNLP 2024 (pp. 9572–9585). Association for Computational Linguistics.