Impact Factor
Call For Paper
Volume 12 Issue 07
July 2026
Author(s)
Abstract
The Exponential Growth Of Unstructured Text Data, Such As News Articles, Social Media Posts, And Online Discussions, Has Created The Need For Effective Methods Of Semantic Organisation. Text Clustering Plays A Vital Role In This Context By Grouping Similar Documents Without Labelled Data. However, The Inherent Challenges Of High Dimensionality And Sparsity In Textual Representations Hinder Clustering Performance. To Address This, Dimensionality Reduction Techniques Are Integrated With Clustering Algorithms. This Paper Presents A Comparative Study Of Two Approaches: Singular Value Decomposition (SVD) Combined With KMeans And Non-negative Matrix Factorisation (NMF) Combined With KMeans. The Experiments Were Conducted Using The 20 Newsgroups Dataset In MATLAB R2024b, With TF-IDF Employed As The Feature Extraction Technique. Results Demonstrate That SVD + KMeans Achieved Superior Clustering Accuracy With A Normalised Mutual Information (NMI) Score Of Approximately 0.55 On The Training Set And 0.50 On The Test Set, Whereas NMF + KMeans Attained Moderate Accuracy (NMI ≈ 0.45) But Offered More Interpretable, Topic-based Clusters. These Findings Confirm The Trade-off Between Accuracy And Interpretability, Suggesting That Method Selection Should Be Based On Specific Application Requirements. The Study Contributes By Providing A Reproducible MATLAB-based Framework And Offering Insights Into The Suitability Of Dimensionality Reduction Strategies For Large-scale Text Clustering.
Keywords
Paper ID
IJSARTV11I11104274
Publication Date
November 12, 2025
Research Area
Text Clustering / Machine Learning / Natural Language Processing (NLP)