Lecture Notes in Computer Science, 2009, Volume 5729/2009, 40-47, DOI: 10.1007/978-3-642-04208-9_9

Improving the Clustering of Blogosphere with a Self-term Enriching Technique

Fernando Perez-Tellez, David Pinto, John Cardiff and Paolo Rosso

View Related Documents

Abstract

The analysis of blogs is emerging as an exciting new area in the text processing field which attempts to harness and exploit the vast quantity of information being published by individuals. However, their particular characteristics (shortness, vocabulary size and nature, etc.) make it difficult to achieve good results using automated clustering techniques. Moreover, the fact that many blogs may be considered to be narrow domain means that exploiting external linguistic resources can have limited value. In this paper, we present a methodology to improve the performance of clustering techniques on blogs, which does not rely on external resources. Our results show that this technique can produce significant improvements in the quality of clusters produced.

Fulltext Preview

Image of the first page of the fulltext document