LSTM-Based Causal Attribution Modeling of the 2025 Sumatra Flash Flood Discourse on YouTube
DOI:
https://doi.org/10.31961/eltikom.v10i1.2132Keywords:
causal attribution, disaster discourse, LSTM, social media analysis, YouTube commentsAbstract
Existing disaster sentiment analysis mainly focuses on emotional polarity classification, while often over-looking the causal reasoning that shapes public discourse on responsibility for disaster outcomes. This study proposes and assesses a Long Short-Term Memory (LSTM)-based causal attribution classification framework to examine YouTube comments related to the 2025 Sumatra flash flood. It compares LSTM performance with Sup-port Vector Machine (SVM) and Naïve Bayes baselines. A total of 17,503 publicly available comments were collected through the YouTube Data API v3 and processed into a final dataset of 12,299 comments. The com-ments were classified into two causal categories, human factor and nature/prayer factor, using lexicon-based scoring validated by three independent annotators (Cohen's κ = 0.81). The experimental results show that LSTM achieves 98.17% accuracy with strong stability (±0.25% standard deviation) under stratified five-fold cross-validation, substantially outperforming SVM (82.83%) and Naïve Bayes (75.04%). These findings indi-cate that sequence-based architectures can capture the contextual dependencies in causal attribution dis-course, offering a replicable framework for disaster risk communication monitoring systems.
Downloads
References
[1] Maruli, “TRAGEDI SUMATRA 2025 — ‘Banjir Neraka Sumatra Pecah Rekor: Korban Lewat 700 Jiwa, 1 Juta Warga Mengungsi, Negara Dikepung Krisis!,’” Gakorpan News, 2025. [Online]. Available: https://www.gakorpan.com/tragedi-sumatra-2025-banjir-neraka-sumatra-pecah-rekor-korban-lewat-700-jiwa-1-juta-warga-mengungsi-negara-dikepung-krisis
[2] Matius Alfons Hutajulu, “Korban Tewas Bencana Sumatera Capai 1.006 Orang, 217 Masih Hilang’” DetikNews, 2025. [Online]. Available: https://news.detik.com/berita/d-8258300/korban-tewas-bencana-sumatera-capai-1-006-orang-217-masih-hilang
[3] Agungnoe, “Bencana Banjir Bandang Sumatra, Pakar UGM Sebut Akibat Kerusakan Ekosistem Hutan di Hulu DAS,” Universitas Gadjah Mada, 2025. [Online]. Available: https://ugm.ac.id/id/berita/bencana-banjir-bandang-sumatra-pakar-ugm-sebut-akibat-kerusakan-ekosistem-hutan-di-hulu-das/
[4] J. Hladík, L. Herman, D. Snopková, and M. Konečný, “Spatio-temporal patterns of disaster impact and recovery in YouTube content,” Int. J. Digit. Earth, vol. 17, no. 1, pp. 1–23, 2024.
[5] J. Tian and R. Zhang, “Moral judgments influence emotional responses and comment lengths through the moderating role of linguistic style matching,” Sci. Rep., vol. 15, no. 1, p. 20972, 2025.
[6] S. Kwon and A. Park, “Examining thematic and emotional differences across Twitter, Reddit, and YouTube: The case of COVID-19 vaccine side effects,” Comput. Hum. Behav., vol. 144, p. 107734, 2023.
[7] K. L. Po, A. J. R. Sy, and R. D. G. Jamora, “Evaluating YouTube as a source of information on hemifacial spasm,” Clin. Park. Relat. Disord., vol. 12, p. 100311, 2025.
[8] Kompas.com, “Pakar ITB Ungkap Penyebab Banjir Sumatera, Termasuk Hilangnya Resapan,” Youtube, 2025. [Online]. Available: https://www.youtube.com/watch?v=tfJC0bn9gHs
[9] Narasi, “Meliput Banjir Sumatera: Air Masih Tinggi, Jalan Terputus | Mata Najwa,” Mata Najwa, 2025. [Online]. Available: https://www.youtube.com/watch?v=ouR11bOMzVM
[10] B. Ilyas and A. Sharifi, “A systematic review of social media-based sentiment analysis in disaster risk management,” Int. J. Disaster Risk Reduct., vol. 123, p. 105487, 2025.
[11] W. Maharani, H. Daud, N. Muhammad, and E. A. Kadir, “Leveraging Social Media Data for Forest Fires Sentiment Classification: A Data-Driven Method,” J. Inf. Syst. Eng. Bus. Intell., vol. 10, no. 3, pp. 392–407, 2024.
[12] P. M. Lavanya and E. Sasikala, “Auto capture on drug text detection in social media through NLP from the heterogeneous data,” Meas. Sens., vol. 24, p. 100550, 2022.
[13] N. Arlim et al., “Dictionary-based extraction of hyperbole and swear words for sarcasm detection in Indonesian Tweets,” Int. J. Inf. Technol., vol. 17, no. 5, pp. 2671–2678, 2024.
[14] A. F. Hidayatullah, R. A. Apong, D. T. C. Lai, and A. Qazi, “Pre-trained language model for code-mixed text in Indonesian, Javanese, and English using transformer,” Soc. Netw. Anal. Min., vol. 15, no. 1, p. 30, 2025.
[15] N. Khamidah, K. A. Notodiputro, and S. D. Oktarina, “Sentiment Analysis of Imbalanced Sarcastic Flood Disaster Texts Using Deep Learning Models,” Int. Res. J. Innov. Eng. Technol., vol. 08, no. 11, pp. 150–158, 2024.
[16] C. Mallikarjuna and S. Sivanesan, “Tweet question classification for enhancing Tweet Question Answering System,” Nat. Lang. Process. J., vol. 10, p. 100130, 2025.
[17] S. Christina and D. Ronaldo, “Studi Literatur Sistematis Terhadap Pengembangan Leksikon Sentiment,” J. ELTIKOM, vol. 4, no. 2, pp. 121–131, 2020.
[18] B. Ghojogh and A. Ghodsi, “Recurrent Neural Networks and Long Short-Term Memory Networks: Tutorial and Survey,” Apr. 22, 2023, arXiv: arXiv:2304.11461.
[19] Y. Yu, X. Si, C. Hu, and J. Zhang, “A Review of Recurrent Neural Networks: LSTM Cells and Network Architectures,” Neural Comput., vol. 31, no. 7, pp. 1235–1270, 2019.
[20] J. García Cabello and S. Carbó-García, “LSTM new gate for computing the efficiency on inputdata,” Knowl.-Based Syst., vol. 322, p. 113622, 2025.
[21] E. Kim, W. Luo, and H. Jho, “Perception and argumentation in the LK-99 superconductivity controversy: a sentiment and argument mining analysis,” Sci. Rep., vol. 15, no. 1, p. 13254, 2025.
[22] T. Abdurahmonov, M. N. T. Abiyyu, D. N. Khayat, and M. A. Heryanto, “Aspects-Based Sentiment Analysis of Extreme Weather on Twitter Using Long Short-Term Memory,” J. Inform. Web Eng., vol. 4, no. 2, pp. 430–443, 2025.
[23] Y. Ji, W. Tao, and C. Wan, “A Systematic Review of Attribution Theory Applied to Crisis Events in Communication Journals: Integration and Advancing Insights,” Journal. Mass Commun. Q., 2025.
[24] W. Kassa and R. Lavin, “Assessment of Blame and Responsibility Through Social Media in Disaster Recovery in the Case of #FlintWaterCrisis,” Front. Commun., vol. 3, p. 45, 2018.
[25] A. Yusima, “Digital Traces of Collective Trauma in the 2025 Aceh Climate Disaster,” Renai, vol. 12, no. 1, pp. 1–8, 2026.
[26] M. Siino, I. Tinnirello, and M. La Cascia, “Is text preprocessing still worth the time? A comparative survey on the influence of popular preprocessing methods on Transformers and traditional classifiers,” Inf. Syst., vol. 121, p. 102342, 2024.
[27] C. P. Chai, “Comparison of text preprocessing methods,” Nat. Lang. Eng., vol. 29, pp. 509–553, 2022.
[28] L. Zhu and D. Luo, “A Novel Efficient and Effective Preprocessing Algorithm for Text Classification,” J. Comput. Commun., vol. 11, no. 03, pp. 1–14, 2023.
[29] R. Patil, S. Boit, V. Gudivada, and J. Nandigam, “A Survey of Text Representation and Embedding Techniques in NLP,” IEEE Access, vol. 11, pp. 36120–36146, 2023.
[30] G. Erasmo Ndomba, M. Edmund Mswahili, and Y.-S. Jeong, “Tokenizers for African Languages,” IEEE Access, vol. 13, pp. 1046–1054, 2025.
[31] K. Madatov, S. Bekchanov, and J. Vičič, “Dataset of stopwords extracted from Uzbek texts,” Data Brief, vol. 43, p. 108351, 2022.
[32] H. T. Y. Achsan, H. Suhartanto, W. C. Wibowo, D. A. Dewi, and K. Ismed, “Automatic Extraction of Indonesian Stopwords,” Int. J. Adv. Comput. Sci. Appl., vol. 14, no. 2, 2023.
[33] Rianto, A. B. Mutiara, E. P. Wibowo, and P. I. Santosa, “Improving the accuracy of text classification using stemming method, a case of non-formal Indonesian conversation,” J. Big Data, vol. 8, no. 1, p. 26, 2021.
[34] T. W. G. Cuizon and H. S. Alar, “Lexicon-based Sentence Emotion Detection Utilizing Polarity-Intensity Unit Circle Mapping and Scoring Algorithm,” Procedia Comput. Sci., vol. 212, pp. 161–170, 2022.
[35] F. E. L. Otto and E. Raju, “Harbingers of decades of unnatural disasters,” Commun. Earth Environ., vol. 4, no. 1, p. 280, 2023.
[36] D. Marks and I. G. Baird, “The urban political ecology of worsening flooding in Phnom Penh, Cambodia: Neopatrimonialism, displacement, and uneven harm,” Int. J. Disaster Risk Reduct., vol. 118, p. 105229, 2025.
[37] I. Cero, J. Luo, and J. M. Falligant, “Lexicon-Based Sentiment Analysis in Behavioral Research,” Perspect. Behav. Sci., vol. 47, no. 1, pp. 283–310, 2024.
[38] B. Drury, H. Gonçalo Oliveira, and A. De Andrade Lopes, “A survey of the extraction and applications of causal relations,” Nat. Lang. Eng., vol. 28, no. 3, pp. 361–400, 2022.
[39] K. S. Tan, Y.-C. Yeh, P. S. Adusumilli, and W. D. Travis, “Quantifying Interrater Agreement and Reliability Between Thoracic Pathologists: Paradoxical Behavior of Cohen’s Kappa in the Presence of a High Prevalence of the Histopathologic Feature in Lung Cancer,” JTO Clin. Res. Rep., vol. 5, no. 1, p. 100618, 2024.
[40] Y. Ma, “Construction and Data Analysis of a New Media Content Popularity Prediction Model Based on Naive Bayes Algorithm,” Procedia Comput. Sci., vol. 261, pp. 294–302, 2025.
[41] B. A. Chandio, A. S. Imran, M. Bakhtyar, S. M. Daudpota, and J. Baber, “Attention-Based RU-BiLSTM Sentiment Analysis Model for Roman Urdu,” Appl. Sci., vol. 12, no. 7, p. 3641, 2022.
[42] J. Duan, P.-F. Zhang, R. Qiu, and Z. Huang, “Long short-term enhanced memory for sequential recommendation,” World Wide Web, vol. 26, no. 2, pp. 561–583, 2023.
[43] A. Iqbal, A. Shahid, M. Roman, M. T. Afzal, and U. U. Hassan, “Optimising window size of semantic of classification model for identification of in-text citations based on context and intent,” PLOS ONE, vol. 20, no. 3, p. e0309862, 2025.
[44] S. F. Taskiran, B. Turkoglu, E. Kaya, and T. Asuroglu, “A comprehensive evaluation of oversampling techniques for enhancing text classification performance,” Sci. Rep., vol. 15, no. 1, p. 21631, 2025.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Kunti Najma Jalia, Adi Suwondo, Hidayatus Sibyan

This work is licensed under a Creative Commons Attribution 4.0 International License.
All accepted papers will be published under a Creative Commons Attribution 4.0 International (CC BY 4.0) License. Authors retain copyright and grant the journal right of first publication. CC-BY Licenced means lets others to Share (copy and redistribute the material in any medium or format) and Adapt (remix, transform, and build upon the material for any purpose, even commercially).

