Threshold Text Classification with Kullback–Leibler Divergence Approach

Huynh, Hiep Xuan; Phan, Cang Anh; Tran, Tu Cam Thi; Nguyen, Hai Thanh; Truong, Dinh Quoc

doi:10.1007/978-981-19-6450-3_2

Hiep Xuan Huynh⁴,
Cang Anh Phan⁵,
Tu Cam Thi Tran⁵,
Hai Thanh Nguyen⁴ &
…
Dinh Quoc Truong⁴

Part of the book series: Studies in Computational Intelligence ((SCI,volume 1068))

452 Accesses

Abstract

Text classification based on thresholds belongs to the supervised learning method which assigns text material to predefined classes or categories based on different thresholds with divergence approach. These categories are identified by a set of documents trained by an automated algorithm. This work presents an approach of text classification using an automatic keyword extraction algorithm based on the Kullback–Leibler divergence approach. The proposed method is evaluated on 2000 documents in Vietnamese, covering ten topics, collected from various e-journals and news portal Web sites including vietnamnet.vn, vnexpress.net, and so on to generate a completely new set of keywords. Such keywords, then, are leveraged to categorize the topic of new text documents. The obtained results verifying the practicality of our approach are feasible as well as outperform the state-of-the-art method.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Subscribe and save

Springer+ Basic

EUR 32.99 /Month

Get 10 units per month
Download Article/Chapter or Ebook
1 Unit = 1 Article or 1 Chapter
Cancel anytime

Subscribe now

Buy Now

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 129.00; Price excludes VAT (USA)

Softcover Book: USD 119.99; Price excludes VAT (USA)

Hardcover Book: USD 169.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Text categorization based on a new classification by thresholds

Article 03 June 2021

Automatic text classification method based on Zipf’s law

Article 01 May 2015

Enhancing Decision Boundary Setting for Binary Text Classification

References

Liu, F., Pennell, D., Liu, F., & Liu, Y. (2009). Unsupervised approaches for automatic keyword extraction using meeting transcripts. In Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the ACL, NAACL’09, (June 2009) (pp. 620–628), Boulder, Colorado.
Google Scholar
Nguyen, C. T., Nguyen, T. K., Phan, X. H., Nguyen, L. M., & Ha, Q. T. (2006). Vietnamese word segmentation with CRFs and SVMs: an investigation. In Proceedings of the 20th Pacific Asia Conference on Language, Information and Computation (PACLIC 2006).
Google Scholar
Matsuo, Y., & Ishizuka, M. (2004). Keyword extraction from a single document using word co-occurrence statistical information. International Journal on AI Tools, 13(1), 157–169.
Article Google Scholar
Zaïane, O. R., & Antonie, M.-L. (2002). Classifying text documents by associating terms with text categories. Australian Computer Science and Communications, 24(2), 215–222.
Google Scholar
Truong, Q. D., Huynh, H. X., & Nguyen, C. N. (2016). An abstract-based approach for text classification. In Lecture Notes of the Institute for Computer Sciences, Social Informatics and Telecommunications Engineering (Vol 168, pp. 237–245). Springer. https://doi.org/10.1007/978-3-319-46909-6_22
Zang, T., Goetz, T., Johnson, D., & Oles, F. (2002). A decision tree-base symbolic rule induction system for text categorization. IBM Systems Journal, 428–437. https://doi.org/10.1147/sj.413.0428
Han, E.-H. (Sam), Karypis, G., & Kumar, V. (2001). Text categorization using weight adjusted k-nearest neighbor classification. Springer
Google Scholar
Dinh, Q. T., Le, H. P., Nguyen, T. M. H., Nguyen, C. T., & Mathias Rossignol, et al. (2008). Word segmentation of Vietnamese texts: A comparison of approaches. In 6th International Conference on Language Resources and Evaluation-LREC 2008, (May 2008), Marrakech, Morocco.
Google Scholar
Lai, K. P., Ho, J. C. S., & Lam, W. (2020). Cross-domain sentiment classification using topic attention and dual-task adversarial training. In Artificial Neural Networks and Machine Learning—ICANN 2020. ICANN 2020. Lecture Notes in Computer Science (Vol. 12397). Springer. https://doi.org/10.1007/978-3-030-61616-8_46
Wang, W., Guo, B., Shen, Y., et al. (2020). Twin labeled LDA: A supervised topic model for document classification. Applied Intelligence, 50, 4602–4615. https://doi.org/10.1007/s10489-020-01798-x
Article Google Scholar
Aggarwal, A. G. (2018). A multi-attribute online advertising budget allocation under uncertain preferences. Ingeniería Solidaria, 14(25), 1–10.
Article MathSciNet Google Scholar
Terrance, A. R., Shrivastava, S., Kumari, A., & Sivanandam, L. (2018). Competitive analysis of retail websites through search engine marketing. Ingeniería Solidaria, 14(25), 1–14.
Article Google Scholar
Lopez-Inga, M. E., & Guerrero-Huaranga, R. M. (2018). Cloud business intelligence and analytics model for SMES in the retail sector in Peru/Modelo de inteligencia de negocios y analitica en la nube para pymes del sector retail en Peru/Modelo de inteligencia de negocios e analitica em nuvem para pmes do setor varejista no Peru. Revista Ingenieria Solidaria, 14(24).
Google Scholar
Gupta, M., Solanki, V. K., & Singh, V. K. (2017). A novel framework to use association rule mining for classification of traffic accident severity. Ingeniería solidaria, 13(21), 37–44.
Article Google Scholar
Le, H. P., Nguyen, M. H. T., Roussanaly, A., & Ho, T.V. A. (2008). Hybrid approach to word segmentation of vietnamese texts. Language and Automata Theory and Applications, 240–249. https://doi.org/10.1007/978-3-540-88282-4_23
Pak, I., & The, P. L. (2018). Text segmentation techniques: A critical review. In: Innovative Computing, Optimization and Its Applications. Studies in Computational Intelligence (Vol. 741). Springer. https://doi.org/10.1007/978-3-319-66984-7
Weisstein, E. W. Graph arc. From MathWorld—A wolfram web resource. https://mathworld.wolfram.com/GraphArc.html
Wu, C.-H., Huang, C.-L., Su, C.-S., & Lee, K.-M. (2007). Speech retrieval using spoken keyword extraction and semantic verification. In TENCON 2007—2007 IEEE Region 10 Conference. https://doi.org/10.1109/TENCON.2007.4429138
http://xltiengviet.wikia.com/wiki/Danh_s%C3%A1ch_stop_word
Kullback, S. (1987). Letter to the editor: The Kullback-Leibler distance. The American Statistician, 41(4), 340–341. https://doi.org/10.1080/00031305.1987.10475510.JSTOR2684769
Article Google Scholar
Wallach, H. M. (2006). Topic modeling: Beyond bag-of-words. In Proceedings of the 23rd International Conference on Machine Learning, ICML (pp. 977–984), New York, NY, USA. ACM
Google Scholar
Thi Tran, T. C., Huynh, H. X., Tran, P. Q., & Truong, D. Q. (2019). Text classification based on keywords with different thresholds. In ACM International Conference Proceeding Series.
Google Scholar

Download references

Author information

Authors and Affiliations

College of Information and Communication Technology, Can Tho University, Can Tho, Vietnam
Hiep Xuan Huynh, Hai Thanh Nguyen & Dinh Quoc Truong
Vinh Long University of Technology Education, Vinh Long, Vietnam
Cang Anh Phan & Tu Cam Thi Tran

Authors

Hiep Xuan Huynh
View author publications
You can also search for this author in PubMed Google Scholar
Cang Anh Phan
View author publications
You can also search for this author in PubMed Google Scholar
Tu Cam Thi Tran
View author publications
You can also search for this author in PubMed Google Scholar
Hai Thanh Nguyen
View author publications
You can also search for this author in PubMed Google Scholar
Dinh Quoc Truong
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding author

Correspondence to Hiep Xuan Huynh .

Editor information

Editors and Affiliations

Hanoi University of Industry, Bac Tu Liem, Hanoi, Vietnam
Thi Dieu Linh Nguyen
Department of Computer Science, University of Huddersfield, Huddersfield, UK
Joan Lu

Rights and permissions

Reprints and permissions

Copyright information

About this chapter

Cite this chapter

Huynh, H.X., Phan, C.A., Tran, T.C.T., Nguyen, H.T., Truong, D.Q. (2023). Threshold Text Classification with Kullback–Leibler Divergence Approach. In: Nguyen, T.D.L., Lu, J. (eds) Machine Learning and Mechanics Based Soft Computing Applications. Studies in Computational Intelligence, vol 1068. Springer, Singapore. https://doi.org/10.1007/978-981-19-6450-3_2

Download citation

DOI: https://doi.org/10.1007/978-981-19-6450-3_2
Published: 28 February 2023
Publisher Name: Springer, Singapore
Print ISBN: 978-981-19-6449-7
Online ISBN: 978-981-19-6450-3
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics

Threshold Text Classification with Kullback–Leibler Divergence Approach

Abstract

Access this chapter

Subscribe and save

Buy Now

Similar content being viewed by others

Text categorization based on a new classification by thresholds

Automatic text classification method based on Zipf’s law

Enhancing Decision Boundary Setting for Binary Text Classification

References

Author information

Authors and Affiliations

Corresponding author

Editor information

Editors and Affiliations

Rights and permissions

Copyright information

About this chapter

Cite this chapter

Download citation

Publish with us

Subscribe and save

Buy Now

Navigation

Threshold Text Classification with Kullback–Leibler Divergence Approach

Abstract

Access this chapter

Subscribe and save

Buy Now

Similar content being viewed by others

Text categorization based on a new classification by thresholds

Automatic text classification method based on Zipf’s law

Enhancing Decision Boundary Setting for Binary Text Classification

References

Author information

Authors and Affiliations

Corresponding author

Editor information

Editors and Affiliations

Rights and permissions

Copyright information

About this chapter

Cite this chapter

Download citation

Share this chapter

Publish with us

Search

Navigation