IMPLEMENTATION OF THE TEXT SIMILARITY EVALUATION METHOD WITH A WEIGHTED APPROACH TO IMPROVE THE ACCURACY OF FAKE NEWS CLASSIFICATION

Iryna Ilina, Kyrylo Reuka

Abstract


The article examines how to enhance automated fake news detection accuracy using a novel semantic-aware text-similarity evaluation method. Traditional frequency-based models often fail to capture subtle semantic manipulations in the modern context of digital disinformation and hybrid warfare. This study addresses this gap by combining GloVe word embeddings with a specialized weighting approach. The goal of this study is to develop a highly effective classification methodology that integrates pre-trained GloVe embeddings with a norm-based weighted similarity metric to improve the identification accuracy of fake news. The achievement of this goal is quantitatively validated through measurable improvements in Accuracy, F1-score, Precision, Recall, and Log-loss compared to standard unweighted baselines. The tasks to be addressed include formulating a robust mathematical method to calculate word importance based on the L2-norm of semantic vectors, applying Manhattan distance for nuanced similarity scoring, and seamlessly integrating this feature into a comprehensive machine learning pipeline to ensure flexibility across various algorithms. The methods used in this study encompass extensive data preprocessing on a balanced dataset of 44,898 news records. Text vectorization is performed using 50-dimensional GloVe embeddings. Then, the semantic features are used to train and evaluate four supervised classifiers: Logistic Regression, Random Forest, XGBoost, and Support Vector Classifier (SVC) across incremental dataset sizes and under simulated noise conditions. Conclusions. The empirical results demonstrate that the proposed weighted similarity approach significantly improves fake news classification. Specifically, the method yields an average increase of 3–4% in Accuracy and up to a 4% improvement in F1-scores across all evaluated models. Furthermore, the proposed approach significantly reduces the impact of linguistic noise and maintains strong probabilistic confidence. The scientific novelty lies in replacing uniform semantic pooling with an embedding-norm-based weighting mechanism, offering a computationally efficient and highly accurate solution for real-world disinformation detec

Keywords


fake news classification; word importance; text similarity; machine learning; disinformation

References


Bahruz, E.T. Manipulation as a Form of infor-mation-psychological War. Revista Universidad y Sociedad, 2023, vol. 15, no. 5, pp. 143-150. Available at: https://rus.ucf.edu.cu/index.php/rus/article/view/4060/4628 (accessed 01.04.2025).

Bontridder, N. & Poullet, Y. The Role of Artificial Intelligence in Disinformation. Data & Policy, 2021, vol. 3, no. E32, DOI: 10.1017/dap.2021.20.

Lee, A., Chen, X. & Wood, I. Robust Detection of Fake News Using LSTM and GloVe. International Journal of Scientific Research and Management (IJSRM), 2022, vol. 10, no. 06, pp. 929-941. DOI: 10.18535/ijsrm/v10i6.ec06.

AlJamal, M., Alquran, R., Alsarhan, A., Aljaidi, M., Al-Jamal, W. & Alkoradees, A. Optimized Novel Text Embedding Approach for Fake News Detection on Twitter X: Integrating Social Context, Temporal Dynamics, and Enhanced Interpretability. International Journal of Computational Intelligence Systems, 2025, vol. 18, no. 1. DOI: 10.1007/s44196-024-00730-2.

Mishchenko, L. & Klymenko, I. Recognizing fake news based on natural language processing using the BM25 algorithm with fine-tuned parameters. Eastern-European Journal of Enterprise Technologies, 2023, vol. 6, no. 2 (126), pp. 33-40. DOI: 10.15587/1729-4061.2023.293513.

Shah, R., Magodia, S., Zala, Y. & Dabre, K. Fake news prediction using TF-IDF vectorizer. International Journal of Research and Analytical Reviews (IJRAR), 2021, vol. 8, no. 2, pp. 429-434. Available at: https://ijrar.org/papers/IJRAR21B1646.pdf (accessed 19.04.2025).

Toor, M.S., Shahbaz, H., Yasin, M., Ali, A., Fitriyani, N.L., Kim, C. & Syafrudin, M. An Optimized Weighted-Voting-Based Ensemble Learning Approach for Fake News Classification. Mathematics, 2025, vol. 13, no. 3, pp. 449-470. DOI: 10.3390/math13030449.

Beseiso M. & AI-Zahrani S. A Context-Enhanced Model for Fake News Detection. Engineering, Technology & Applied Science Research, 2025, vol. 15, no. 1, pp. 19128-19135. DOI: 10.48084/etasr.9192.

Khanam Z., Alwasel B. N., Sirafi H. & Rashid M. Fake News Detection Using Machine Learning Approaches. IOP Conference Series: Materials Science and Engineering, 2021. DOI: 10.1088/1757-899X/1099/1/012040.

Siddiqui, T., Hina, S., Asif, R., Ahmed, S. & Ahmed, M. An ensemble approach for the identification and classification of crime tweets in the English language. Computer Science and Information Technologies, 2023, vol. 4, no. 2, pp. 149-159. DOI: 10.11591/csit.v4i2.pp149-159.

Kumar, R. Fake News Detection using Passive Aggressive and TF-IDF Vectorizer. International Research Journal of Engineering and Technology (IRJET), 2020, vol. 07, no. 12, pp. 902-904. Available at: https://www.irjet.net/archives/V7/i12/IRJET-V7I12158.pdf (accessed 21.04.2025).

Mhatre, S. & Masurkar, A. A Hybrid Method for Fake News Detection using Cosine Similarity Scores. International Conference on Communication information and Computing Technology (ICCICT), 2021, pp. 1-6. DOI: 10.1109/ICCICT50803.2021.9510134

Naik, S. & Patil, A. Fake News Detection Using NLP. International Journal for Research in Applied Science and Engineering Technology, 2021, vol. 9, no. 12, pp. 2022-2031. DOI: 10.22214/ijraset.2021.39582.

De Boom, C., Van Canneyt, S., Demeester, T. & Dhoedt, B. Representation learning for very short texts using weighted word embedding aggregation. Pattern Recognition Letters, 2016, vol. 80, pp. 150-156.

DOI: 10.1016/j.patrec.2016.06.012.

Shtovba, S., Petrychko, M. & Petranova, M. A Similarity Metric of Categorical Distributions that Accounts for the Kinship of Different Categories. Visnyk of Vinnytsia Politechnical Institute, 2023, vol. 167, no. 2, pp. 49-57. DOI: 10.31649/1997-9266-2023-167-2-49-57.

Wang, J. & Dong, Y. Measurement of Text Similarity: A Survey. Information, 2020, vol. 11, no. 9, pp. 421. DOI: 10.3390/info11090421.

Rathi, R. N. & Mustafi, A. Rathi R. N. The importance of Term Weighting in semantic understanding of text: A review of techniques. Multimedia Tools and Applications, 2022, vol. 82, pp. 9761-9783. DOI: 10.1007/s11042-022-12538-3.

Seegmiller, B., Papanikolaou, D., Schmidt, L. & Seegmiller, B. Measuring Document Similarity with Word Embeddings. SSRN Electronic Journal, 2022. DOI: 10.2139/ssrn.4088443.

Luque, A., Carrasco, A., Martin, A. & de las Heras, A. The impact of class imbalance in classification performance metrics based on the binary confusion matrix. Pattern Recognition, 2019, vol. 91, pp. 216-231. DOI: 10.1016/j.patcog.2019.02.023.

Emmert-Streib, F., Moutari, S. & Dehmer, M. A comprehensive survey of error measures for evaluating binary decision making in data science. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 2019, vol. 9, no. 5, DOI: 10.1002/widm.1303.

Foody, G. M. Challenges in the real world use of classification accuracy metrics: From recall and precision to the Matthews correlation coefficient. PLOS ONE, 2023, vol. 18, no. 10. DOI: 10.1371/journal.pone.0291908.

Christen, P., Hand, D. J. & Kirielle, N. A review of the F-measure: Its History, Properties, Criticism, and Alternatives. ACM Computing Surveys, 2023. DOI: 10.1145/3606367.

Kravets, P., Pasichnyk, V. & Prodaniuk, M. Mathematical Model of Logistic Regression for Binary Classification. Part 1. Regression Models of Data Generalization. Vìsnik Nacìonalʹnogo unìversitetu "Lʹvìvsʹka polìtehnìka". Serìâ Ìnformacìjnì sistemi ta merežì, 2024, vol. 15, pp. 290-321. DOI: 10.23939/sisn2024.15.290. (In Ukrainian).

Reuka, K. News Detection Experiment: Weighted Text Similarity. GitHub, 2025. Available at: https://github.com/kreuka/fake-news-weighted-similarity




DOI: https://doi.org/10.32620/reks.2026.2.08

Refbacks

  • There are currently no refbacks.