Tuti Handayani, Sri Mardiyati
The increasing volume of mobile application reviews has created a need for automated sentiment classification capable of handling large-scale and imbalanced user-generated data. This study aimed to compare classical machine learning algorithms and evaluate the effectiveness of class-weighted learning in improving minority-class recognition in Indonesian mobile application reviews. The dataset was cleaned, deduplicated, and controlled for conflicting labels and data leakage, resulting in 357,327 unique reviews across Negative, Neutral, and Positive sentiment classes. Textual features were represented using Term Frequency-Inverse Document Frequency (TF-IDF) with 50,000 features, and the data were divided using an 80:20 stratified train-test split. Logistic Regression, Multinomial Naive Bayes, Linear Support Vector Machine (SVM), and Random Forest were evaluated alongside class-weighted Logistic Regression and Linear SVM. Logistic Regression achieved the highest accuracy of 83.47%, whereas Balanced Linear SVM achieved the highest Macro F1-score of 62.87% and improved the Neutral F1-score from 8.70% to 21.18%, while maintaining an accuracy of 79.09%. The findings demonstrate that the highest overall accuracy does not necessarily indicate the most balanced classifier under class imbalance. Balanced Linear SVM provided the most favorable trade-off between overall performance and minority-class recognition, highlighting the importance of class-sensitive evaluation and class-weighted learning for imbalanced sentiment classification.
Article Details
| Volume: | 4 |
| Issue: | 2 |
| Year: | 2026 |
| Published: | 2026-09-10 |
| Pages: | 95–105 |
| Section: | Articles |

This work is licensed under a Creative Commons Attribution 4.0 International License.
This work is licensed under a Creative Commons License.
Track citations and research impact