A Word2Vec-Based Machine Learning Framework for Binary Sentiment Classification | IJORET Volume 11 – Issue 4 | IJORET-V11I4P5

IJORET
International Journal of Research in Engineering Technology
ISSN 2455-1341 · Peer-Reviewed · Open Access
📚 Volume 11, Issue 4
📅 August 19, 2026
📄 Pages 36–41
🔖 ID: IJORET-V11I4P5

A Word2Vec-Based Machine Learning Framework for Binary Sentiment Classification

Author(s)

Mohit Kharbanda, Arshi Husain, Virendra P. Vishwakarma

Abstract

Automatically judging whether a piece of text expresses a positive or negative opinion– sentiment analysis– remains a core problem in natural language processing, and movie reviews are a widely used testbed for it. This paper builds a sentiment classification pipeline around Word2Vec embeddings and evaluates it on a corpus of 40,000 labelled movie reviews spanning both sentiment classes. Reviews are first cleaned, tokenized, filtered for stop-words, and lemmatized; the resulting tokens are then mapped into 100-dimensional Word2Vec vectors that serve as numerical input to the classifiers. Using a stratified 80:20 split (32,000 training instances and 8,000 held-out test instances), four classical supervised learners– Support Vector Machine (SVM), Random Forest (RF), K-Nearest Neighbors (KNN), and XGBoost– are trained and compared on accuracy, precision, recall, F1-score, and ROC-AUC. SVM comes out ahead of the other three, reaching 85.51% accuracy, 85.38% precision, 85.66% recall, an F1-score of 85.52%, and a ROC-AUC of 93.04%. XGBoost is the runner-up at 82.99% accuracy, with Random Forest close behind at 82.50% and KNN trailing at 79.66%. These results indicate that a Word2Vec-plus-SVM pipeline is a strong, computationally light baseline for binary movie-review sentiment classification and a reasonable starting point for future work on richer text representations.

Keywords

Movie Reviews, Sentiment Analysis, Word2Vec Embeddings, Natural Language Processing, Machine Learning, Support Vector Machine, Random Forest, K-Nearest Neighbors, XGBoost.

Conclusion

The aim of this paper was to develop a Word2Vec framework to classify movie reviews into positive and negative categories and test it using four different supervised classification algorithms on a set of 40,000 reviews in an 80:20 ratio. The performance of all the four classifiers demonstrated that the SVM was the best of the four classifiers with 85.51%, 85.52%, and 93.04% accuracy, F1 score, and ROC-AUC respectively.

References

[1] M. Giatsoglou, M. G. Vozalis, K. Diamantaras, A. Vakali,
G. Sarigiannidis, and K. C. Chatzisavvas, “Sentiment anal
ysis leveraging emotions and word embeddings,” Expert
Systems with Applications, vol. 69, pp. 214–224, 2016.
[2] H. J. Alantari, I. S. Currim, Y. Deng, and S. Singh, “An
empirical comparison of machine learning methods for
text-based sentiment analysis of online consumer reviews,”
International Journal of Research in Marketing, vol. 39,
pp. 1–19, 2021.
[3] V. Yadav, P. Verma, and V. Katiyar, “E-commerce product
reviews using aspect based Hindi sentiment analysis,” in
Proc. 2021 International Conference on Computer Com
munication and Informatics (ICCCI), Coimbatore, India,
Jan. 27–29, 2021.
[4] Z. Desai, K. Anklesaria, and H. Balasubramaniam, “Busi
ness Intelligence Visualization Using Deep Learning
Based Sentiment Analysis on Amazon Review Data,” in
Proc. 2021 12th International Conference on Computing
Communication and Networking Technologies (ICCCNT),
Kharagpur, India, July 6–8, 2021, pp. 1–7.
[5] K. K. Mohbey, “Sentiment analysis for product rating
using a deep learning approach,” in Proc. 2021 Interna
tional Conference on Artificial Intelligence and Smart
Systems (ICAIS), Coimbatore, India, Mar. 25–27, 2021,
pp. 121–126.
[6] M. S. Darokar, A. D. Raut, and V. M. Thakre, “Method
ological Review of Emotion Recognition for Social Me
dia: A Sentiment Analysis Approach,” in Proc. 2021
International Conference on Computing, Communication
and Green Engineering (CCGE), Pune, India, Sept. 23
25, 2021, pp. 1–5.
5
[7] M. Devika, C. Sunitha, and A. Ganesh, “Sentiment Anal
Linguistics (ACL), Baltimore, MD, USA, June 23–25,
ysis: A Comparative Study on Different Approaches,”
Procedia Computer Science, vol. 87, pp. 44–49, 2016.
[8] W. G. Mangold and D. J. Faulds, “Social media: The new
hybrid element of the promotion mix,” Business Horizons,
vol. 52, pp. 357–365, 2009.
[9] J. Foster, Ö. Çetinoglu, J. Wagner, J. Le Roux, S. Hogan,
J. Nivre, D. Hogan, and J. Van Genabith, “#hardtoparse:
POSTagging and Parsing the Twitterverse,” in Proc. AAAI
2011 Workshop on Analyzing Microtext, San Francisco,
CA, USA, Aug. 7–8, 2011, pp. 20–25.
[10] A. Yousefpour, R. Ibrahim, and H. N. A. Hamed,
“Ordinal-based and frequency-based integration of fea
ture selection methods for sentiment analysis,” Expert
Systems with Applications, vol. 75, pp. 80–93, 2017.
[11] R. Xia and C. Zong, “A POS-based ensemble model for
cross-domain sentiment classification,” in Proc. 5th In
ternational Joint Conference on Natural Language Pro
cessing, Chiang Mai, Thailand, Nov. 8–13, 2011, pp.
614–622.
[12] I. Goodfellow, Y. Bengio, and A. Courville, Deep Learn
ing. Cambridge, MA, USA: MIT Press, 2016.
[13] E. Cambria, S. Poria, R. Bajpai, and B. Schuller, “Sentic
Net 4: A semantic resource for sentiment analysis based
on conceptual primitives,” in Proc. 26th International
Conference on Computational Linguistics (COLING), Os
aka, Japan, Dec. 11–16, 2016, pp. 2666–2677.
[14] R.Jozefowicz, O. Vinyals, M. Schuster, N. Shazeer, and Y.
Wu, “Exploring the limits of language modeling,” arXiv
preprint arXiv:1602.02410, 2016.
[15] P. Vateekul and T. Koomsubha, “A study of sentiment
analysis using deep learning techniques on Thai Twitter
data,” in Proc. 13th International Joint Conference on
Computer Science and Software Engineering (JCSSE),
Khon Kaen, Thailand, July 13–15, 2016, pp. 1–6.
[16] S. Pal, S. Ghosh, and A. Nag, “Sentiment Analysis in
the Light of LSTM Recurrent Neural Networks,” Interna
tional Journal of Synthetic Emotions, vol. 9, pp. 33–39,
2018.
[17] A. Hassan and A. Mahmood, “Deep Learning approach
for sentiment analysis of short texts,” in Proc. 3rd Interna
tional Conference on Control, Automation and Robotics
(ICCAR), Nagoya, Japan, Apr. 24–26, 2017, pp. 705–710.
[18] M. Baroni, G. Dinu, and G. Kruszewski, “Don’t count,
predict! A systematic comparison of context-counting
vs. context-predicting semantic vectors,” in Proc. 52nd
Annual Meeting of the Association for Computational
2014, vol. 1, pp. 238–247.
[19] Y. Yuan and Y. Zhou, “Twitter sentiment analysis with
recursive neural networks,” in CS224D Course Project,
Stanford University, Stanford, CA, USA, 2015.
[20] G. Vinodhini and R. Chandrasekaran, “Sentiment analysis
and opinion mining: A survey,” International Journal, vol.
2, pp. 282–292, 2012.
[21] R. Singh and R. Kaur, “Sentiment Analysis on Social
Media and Online Review,” International Journal of Com
puter Applications, vol. 121, pp. 44–48, 2015.
[22] I. Hemalatha, G. Varma, and A. Govardhan, “Automated
Sentiment Analysis System Using Machine Learning Al
gorithms,” International Journal of Research in Computer
and Communication Technology, vol. 3, pp. 300–303,
2014.
[23] V. Kharde and P. Sonawane, “Sentiment analysis of
twitter data: A survey of techniques,” arXiv preprint
arXiv:1601.06971, 2016.

📋 How to Cite This Paper

Mohit Kharbanda, Arshi Husain, Virendra P. Vishwakarma (2026). A Word2Vec-Based Machine Learning Framework for Binary Sentiment Classification. International Journal of Research in Engineering Technology, 11(4), 36–41. ISSN: 2455-1341. DOI: https://doi.org/10.5281/zenodo.22011650
© 2026 International Journal of Research in Engineering Technology (IJORET). All rights reserved. · ijoret.com
Submit Your Research Paper