Peringkasan Teks Abstraktif Otomatis Menggunakan Model Seq2seq dengan LSTM Encoder-Decoder dan Attantion pada Data Set Amazon Fine Food Reviews
DOI:
https://doi.org/10.33830/saintek.v2i2.15475.2026Kata Kunci:
automatic teks summarization, bleu score, long short-term memory, natural language processing, sequence-to-sequenceAbstrak
Penelitian ini bertujuan untuk membangun model peringkasan teks otomatis dengan strategi abstraktif. Model dibangun menggunakan pendekatan sequence-to-sequence (seq2seq) dengan arsitektur encoder dan decoder berbasis Long Short-Term Memory (LSTM) yang efektif untuk menangani data teks yang memerlukan proses mengingat informasi penting dari awal urutan dan menghubungkannya pada bagian akhir urutan. Kemudian ditambahkan mekanisme attention untuk membantu decoder fokus pada bagian-bagian tertentu yang memerlukan perhatian. Model dilatih pada pasangan data teks ulasan dan teks ringkasan pada data set Amazon Fine Food Reviews kemudian dilatih menggunakan fungsi kerugian (loss) categorical crossentropy dan algoritma optimasi RMSProp. Dari pelatihan diperoleh akurasi sebesar 67.2% dan kerugian sebesar 1.6314. Metode evaluasi dilakukan dengan menggunakan metrik BLEU (Bilingual Evaluation Understudy) untuk membandingkan hasil prediksi keluaran model dengan referensi manusia. Dari hasil evaluasi, model mampu menghasilkan ringkasan dengan skor komulatif BLEU-1 sebesar 0.428571, BLEU-2 sebesar 0.654654, BLEU-3 sebesar 0.756080, dan BLEU-4 sebesar 0.809107. Hasil metrik BLEU menunjukkan model dapat menghasilkan prediksi ringkasan yang cukup koheren dan relevan.
Referensi
Akhter, B., & Mehra, M. (2022). A study of implementation of deep learning techniques for text summarization. International Journal of Innovative Research in Engineering and Management, 9(2), 18–28.
Bahdanau, D., Cho, K., & Bengio, Y. (2014). Neural Machine Translation by Jointly Learning to Align and Translate.
Bystrov, V., Naboka-Krell, V., Staszewska-Bystrova, A., & Winker, P. (2023). Analysing the Impact of Removing Infrequent Words on Topic Quality in LDA Models. ArXiv Preprint, 1–21.
Chaurasia, S., Dasgupta, D., & Regunathan, R. (2023). T5LSTM-RNN based Text Summarization Model for Behavioral Biology Literature. Procedia Computer Science, 218, 585–593. https://doi.org/10.1016/j.procs.2023.01.040
Febriana, A., Nur, Y., & Mariah, M. (2023). Pengaruh ulasan produk, kemudahan, dan keamanan terhadap keputusan pembelian produk secara online pada Shopee di Kota Makassar. Jurnal Malomo : Manajemen Dan Akuntansi, 1(3), 281–292.
Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. The MIT Press.
Luo, M., Xue, B., & Niu, B. (2024). A comprehensive survey for automatic text summarization: Techniques, approaches and perspectives. Neurocomputing, 603, 128280. https://doi.org/10.1016/j.neucom.2024.128280
Mastropaolo, A., Ciniselli, M., Di Penta, M., & Bavota, G. (2023). Evaluating Code Summarization Techniques: A New Metric and an Empirical Characterization.
Matrenin, P. V., Manusov, V. Z., Khalyasmaa, A. I., Antonenkov, D. V., Eroshenko, S. A., & Butusov, D. N. (2020). Improving Accuracy and Generalization Performance of Small-Size Recurrent Neural Networks Applied to Short-Term Load Forecasting. Mathematics, 8(12), 2169. https://doi.org/10.3390/math8122169
Muflikhah, L., Mahmudy, W. F., & Kurnianingtyas, D. (2023). Machine learning: Cetakan Pertama. UB Press.
Nambiar, S. K., Peter S, D., & Idicula, S. M. (2021). Attention based Abstractive Summarization of Malayalam Document. Procedia Computer Science, 189, 250–257. https://doi.org/10.1016/j.procs.2021.05.088
Niu, Z., Zhong, G., & Yu, H. (2021). A review on the attention mechanism of deep learning. Neurocomputing, 452, 48–62. https://doi.org/10.1016/j.neucom.2021.03.091
Paaß, G., & Giesselbach, S. (2023). Foundation Models for Natural Language Processing -- Pre-trained Language Models Integrating Media.
Papineni, K., Roukos, S., Ward, T., & Zhu, W.-J. (2002). BLEU: A method for automatic evaluation of machine translation. Proceedings of the 40th Annual Meeting on Association for Computational Linguistics - ACL ’02, 311–318. https://doi.org/10.3115/1073083.1073135
Setyawan, C., Benarkah, N., & Prasetyo, V. R. (2021). Automatic Text Summarization Berdasarkan Pendekatan Statistika pada Dokumen Berbahasa Indonesia. KELUWIH: Jurnal Sains Dan Teknologi, 2(1), 9–15. https://doi.org/10.24123/saintek.v2i1.4045
Shakil, H., Farooq, A., & Kalita, J. (2024). Abstractive text summarization: State of the art, challenges, and improvements. Neurocomputing, 603, 128255. https://doi.org/10.1016/j.neucom.2024.128255
Supriyono, S., Wibawa, A. P., Suyono, S., & Kurniawan, F. (2024). A survey of text summarization: Techniques, evaluation and challenges. Natural Language Processing Journal, 7, 100070. https://doi.org/10.1016/j.nlp.2024.100070
Tsai, C.-F., Chen, K., Hu, Y.-H., & Chen, W.-K. (2020). Improving text summarization of online hotel reviews with review helpfulness and sentiment. Tourism Management, 80, 104122. https://doi.org/10.1016/j.tourman.2020.104122
Yudistira, N., Alfiansih, L. M. D., Andriyani, N. I., Essayem, W., Maulida, N., Maghfiroh, N. A., & Nurdian, I. W. (2023). Prediksi deret waktu menggunakan deep learning. UB Press.
