Multimodal Emotion Classification of Indonesian Memes on Platform X Using IndoBERT and YOLOv11

Authors

  • Pungkas Subarkah Universitas Amikom Purwokerto, Indonesia
  • Esti Widianti Universitas Amikom Purwokerto, Indonesia
  • Agus Pramono Universitas Amikom Purwokerto, Indonesia
Pages Icon

DOI:

https://doi.org/10.63158/journalisi.v8i4.1742

Keywords:

Multimodal Emotion Classification, Indonesian Memes, IndoBERT, YOLOv11, Weighted Late Fusion

Abstract

The widespread use of social media, particularly Platform X, has increased the popularity of memes as a multimodal communication medium that combines textual and visual elements to express emotions, opinions, and social reactions. This study proposes a multimodal emotion classification approach for Indonesian-language memes by integrating IndoBERT for textual analysis and YOLOv11 operating in image classification mode for visual analysis through a weighted late fusion strategy. An initial dataset of 2,810 Indonesian-language memes was collected through web scraping. After removing corrupted or unreadable images, 2,547 valid samples remained. Each meme was manually annotated into one of six emotion categories—Happiness, Disgust, Anger, Sadness, Fear, and Surprise—based on the annotators' judgment of the dominant emotion conveyed by the combination of text and image, following Ekman's Basic Emotion framework. The dataset was divided using a stratified 80:10:10 split into 2,037 training, 255 validation, and 255 testing samples. The validation set was used to determine the optimal fusion weight, while the held-out test set was reserved exclusively for final evaluation. The best-performing weighted late fusion model (α = 0.6) achieved 72.55% accuracy, 73.03% macro precision, 72.76% macro recall, and 72.66% macro F1-score on the test set. Within the constructed dataset, the proposed multimodal approach outperformed the evaluated IndoBERT-only and YOLOv11-only baselines, indicating that combining textual and visual information can improve emotion classification performance for Indonesian-language meme content.

Downloads

Download data is not yet available.

References

[1] R. Das and T. D. Singh, “Multimodal sentiment analysis: A survey of methods, trends, and challenges,” ACM Comput. Surv., vol. 55, no. 13s, Art. no. 270, Dec. 2023, doi: 10.1145/3586075.

[2] “Digital 2025: Global overview report,” DataReportal—Global Digital Insights, 2025. [Online]. Accessed: Jul. 24, 2026.

[3] M. Ihsan and R. S. Adnan, “Media sosial Twitter sebagai ruang publik virtual (Studi kasus penolakan Omnibus Law),” Syntax Lit.: J. Ilm. Indones., vol. 7, no. 3, pp. 3254–3267, Mar. 2022, doi: 10.36418/syntax-literate.v7i3.6612.

[4] U. Singh, K. Abhishek, and H. K. Azad, “A survey of cutting-edge multimodal sentiment analysis,” ACM Comput. Surv., vol. 56, no. 9, Art. no. 227, pp. 1–38, 2024, doi: 10.1145/3652149.

[5] C. Sharma et al., “SemEval-2020 Task 8: Memotion analysis—The visuo-lingual metaphor!,” in Proc. 14th Int. Workshop Semantic Evaluation (SemEval), 2020, pp. 759–773, doi: 10.18653/v1/2020.semeval-1.99.

[6] S. Pramanick, S. Sharma, D. Dimitrov, M. S. Akhtar, P. Nakov, and T. Chakraborty, “MOMENTA: A multimodal framework for detecting harmful memes and their targets,” in Findings Assoc. Comput. Linguistics: EMNLP, 2021, pp. 4439–4455, doi: 10.18653/v1/2021.findings-emnlp.379.

[7] D. Kiela et al., “The hateful memes challenge: Detecting hate speech in multimodal memes,” Adv. Neural Inf. Process. Syst., vol. 33, pp. 2611–2624, 2020.

[8] H. Uswatun and A. Ahmadi, “Meme politik di TikTok: Perspektif pragmatika humor sebagai kritik sosial digital,” J. Pendid. Bhs. Sastra Indones., vol. 2, no. 2, pp. 72–89, Jun. 2026, doi: 10.47134/jpbsi.v2i2.2625.

[9] A. Pandey and D. K. Vishwakarma, “Progress, achievements, and challenges in multimodal sentiment analysis using deep learning: A survey,” Appl. Soft Comput., vol. 152, Art. no. 111206, Feb. 2024, doi: 10.1016/j.asoc.2023.111206.

[10] T. Modi, E. Shah, S. Shah, J. Kanakia, and M. Tiwari, “Meme classification and offensive content detection using multimodal approach,” in Proc. 2024 OITS Int. Conf. Inf. Technol. (OCIT), 2024, pp. 635–640, doi: 10.1109/OCIT65031.2024.00116.

[11] B. Wilie et al., “IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding,” in Proc. 1st Conf. Asia-Pacific Chapter Assoc. Comput. Linguistics and 10th Int. Joint Conf. Natural Lang. Process. (AACL-IJCNLP), 2020, pp. 843–857, doi: 10.18653/v1/2020.aacl-main.85.

[12] Ultralytics, “Ultralytics YOLO11,” Ultralytics Docs. 2026, [Online]. Available: https://docs.ultralytics.com/models/yolo11/. Accessed: Jul. 22, 2026.

[13] J. H. Chowdhury and S. Ramanna, “MMLTC: A novel tolerance-based clustering framework for multimodal sentiment and harmful meme classification in multilingual settings,” Comput. Intell., vol. 42, no. 2, Art. no. e70219, Mar. 2026, doi: 10.1111/coin.70219.

[14] N. P. Bhapkar, “Memotion 2.0—Sentiment analysis and emotion classification of memes,” M.Sc. research project, Data Analytics, National College of Ireland, Dublin, Ireland, 2022.

[15] S. Patel, N. Shroff, and H. Shah, “Multimodal sentiment analysis using deep learning: A review,” Commun. Comput. Inf. Sci., vol. 2038, pp. 13–29, 2024, doi: 10.1007/978-3-031-59097-9_2.

[16] S. K. Baberwal, N. A. Shelke, and K. Anwar, “Systematic review of recent advances in multimodal sentiment analysis,” Discov. Comput., vol. 28, no. 1, Art. no. 270, Nov. 2025, doi: 10.1007/s10791-025-09727-7.

[17] M. Dhotay, M. Dharrao, S. Deokate, A. Bongale, and D. Dharrao, “Multimodal sentiment analysis: Emerging innovations, core challenges, and future directions,” Discov. Artif. Intell., vol. 6, no. 1, Art. no. 433, Mar. 2026, doi: 10.1007/s44163-026-01141-2.

[18] F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, “IndoLEM and IndoBERT: A benchmark dataset and pre-trained language model for Indonesian NLP,” in Proc. 28th Int. Conf. Comput. Linguistics (COLING), 2020, pp. 757–770, doi: 10.18653/v1/2020.coling-main.66.

[19] A. Conneau et al., “Unsupervised cross-lingual representation learning at scale,” in Proc. 58th Annu. Meeting Assoc. Comput. Linguistics (ACL), 2020, pp. 8440–8451, doi: 10.18653/v1/2020.acl-main.747.

[20] C. Shaw, P. M. LaCasse, and L. E. Champagne, “Exploring emotion classification of Indonesian tweets using large scale transfer learning via IndoBERT,” Soc. Netw. Anal. Min., vol. 15, no. 1, Art. no. 22, Mar. 2025, doi: 10.1007/s13278-025-01439-6.

[21] H. Jayadianti, W. Kaswidjanti, A. T. Utomo, S. Saifullah, F. A. Dwiyanto, and R. Drezewski, “Sentiment analysis of Indonesian reviews using fine-tuning IndoBERT and R-CNN,” ILKOM J. Ilm., vol. 14, no. 3, pp. 348–354, Dec. 2022, doi: 10.33096/ilkom.v14i3.1505.348-354.

[22] S. Li and W. Deng, “Deep facial expression recognition: A survey,” IEEE Trans. Affect. Comput., vol. 13, no. 3, pp. 1195–1215, Jul.–Sep. 2022, doi: 10.1109/TAFFC.2020.2981446.

[23] T. Baltrušaitis, C. Ahuja, and L.-P. Morency, “Multimodal machine learning: A survey and taxonomy,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 41, no. 2, pp. 423–443, Feb. 2019, doi: 10.1109/TPAMI.2018.2798607.

[24] A. Kumar, K. Sharma, and A. Sharma, “MEmoR: A multimodal emotion recognition using affective biomarkers for smart prediction of emotional health for people analytics in smart industries,” Image Vis. Comput., vol. 123, Art. no. 104483, Jul. 2022, doi: 10.1016/j.imavis.2022.104483.

[25] A. Nandi, F. Xhafa, L. Subirats, and S. Fort, “Reward-penalty weighted ensemble for emotion state classification from multi-modal data streams,” Int. J. Neural Syst., vol. 32, no. 12, Art. no. 2250049, Dec. 2022, doi: 10.1142/S0129065722500496.

[26] G. Chandrasekaran, T. N. Nguyen, and J. Hemanth D., “Multimodal sentimental analysis for social media applications: A comprehensive review,” Wiley Interdiscip. Rev. Data Min. Knowl. Discov., vol. 11, no. 5, Art. no. e1415, Sep. 2021, doi: 10.1002/WIDM.1415.

[27] K. Zhao, M. Zheng, Q. Li, and J. Liu, “Multimodal sentiment analysis—A comprehensive survey from a fusion methods perspective,” IEEE Access, vol. 13, pp. 64556–64583, 2025, doi: 10.1109/ACCESS.2025.3554665.

[28] R. K. Routhu and U. Baruah, “Sentiment analysis on memes: A review,” Expert Syst., vol. 42, no. 11, Art. no. e70133, Nov. 2025, doi: 10.1111/EXSY.70133.

[29] P. Ekman, “An argument for basic emotions,” Cogn. Emot., vol. 6, nos. 3–4, pp. 169–200, 1992, doi: 10.1080/02699939208411068.

[30] L. Zhu, Z. Zhu, C. Zhang, Y. Xu, and X. Kong, “Multimodal sentiment analysis based on fusion methods: A survey,” Inf. Fusion, vol. 95, pp. 306–325, Jul. 2023, doi: 10.1016/j.inffus.2023.02.028.

[31] M. L. Williams, P. Burnap, and L. Sloan, “Towards an ethical framework for publishing Twitter data in social research: Taking into account users’ views, online context and algorithmic estimation,” Sociology, vol. 51, no. 6, pp. 1149–1168, Dec. 2017, doi: 10.1177/0038038517708140.

[32] A. N. Markham, K. Tiidenberg, and A. Herman, “Ethics as methods: Doing ethics in the era of big data research—Introduction,” Soc. Media Soc., vol. 4, no. 3, Jul. 2018, doi: 10.1177/2056305118784502.

[33] S. Hazmoune and F. Bougamouza, “Using transformers for multimodal emotion recognition: Taxonomies and state-of-the-art review,” Eng. Appl. Artif. Intell., vol. 133, Art. no. 108339, Jul. 2024, doi: 10.1016/j.engappai.2024.108339.

[34] Y. Dui and H. Hu, “Social media public opinion detection using multimodal natural language processing and attention mechanisms,” IET Inf. Secur., vol. 2024, no. 1, Art. no. 8880804, 2024, doi: 10.1049/2024/8880804.

[35] E. Ferrara, O. Varol, C. Davis, F. Menczer, and A. Flammini, “The rise of social bots,” Commun. ACM, vol. 59, no. 7, pp. 96–104, Jul. 2016, doi: 10.1145/2818717.

[36] R. Smith, “An overview of the Tesseract OCR engine,” in Proc. 9th Int. Conf. Document Anal. Recognit. (ICDAR), Curitiba, Brazil, 2007, pp. 629–633, doi: 10.1109/ICDAR.2007.4376991.

[37] J. Cohen, “A coefficient of agreement for nominal scales,” Educ. Psychol. Meas., vol. 20, no. 1, pp. 37–46, 1960, doi: 10.1177/001316446002000104.

[38] Y. Jiao, J. Li, J. Wu, D. Hong, R. Gupta, and J. Shang, “SeNsER: Learning cross-building sensor metadata tagger,” in Findings Assoc. Comput. Linguistics: EMNLP, 2020, pp. 950–960, doi: 10.18653/v1/2020.findings-emnlp.85.

[39] M. Tan and Q. V. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in Proc. 36th Int. Conf. Mach. Learn. (ICML), vol. 97, 2019, pp. 6105–6114.

[40] R. Kohavi, “A study of cross-validation and bootstrap for accuracy estimation and model selection,” in Proc. 14th Int. Joint Conf. Artif. Intell. (IJCAI), Montreal, QC, Canada, 1995, pp. 1137–1143.

[41] P. K. Atrey, M. A. Hossain, A. El Saddik, and M. S. Kankanhalli, “Multimodal fusion for multimedia analysis: A survey,” Multimed. Syst., vol. 16, no. 6, pp. 345–379, Nov. 2010, doi: 10.1007/s00530-010-0182-0.

[42] H. He and E. A. Garcia, “Learning from imbalanced data,” IEEE Trans. Knowl. Data Eng., vol. 21, no. 9, pp. 1263–1284, Sep. 2009, doi: 10.1109/TKDE.2008.239.

[43] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On calibration of modern neural networks,” in Proc. 34th Int. Conf. Mach. Learn. (ICML), vol. 70, 2017, pp. 1321–1330.

Downloads

Published

2026-08-22

Issue

Section

Articles