Adaptive document clustering based on query-based similarity

In information retrieval, cluster-based retrieval is a well-known attempt in resolving the problem of term mismatch. Clustering requires similarity information between the documents, which is difficult to calculate at a feasible time. The adaptive document clustering scheme has been investigated by...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	Information processing & management 2007-07, Vol.43 (4), p.887-901
Hauptverfasser:	Na, Seung-Hoon, Kang, In-Su, Lee, Jong-Hyeok
Format:	Artikel
Sprache:	eng
Schlagworte:	Adaptive document clustering Cluster analysis Cluster-based retrieval Clustering Exact sciences and technology Information and communication sciences Information retrieval Information retrieval systems Information retrieval systems. Information and document management system Information science. Documentation Language modeling approach Query-based similarity Sciences and techniques of general use Similarity measures Studies Term selection Trouble shooting
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	In information retrieval, cluster-based retrieval is a well-known attempt in resolving the problem of term mismatch. Clustering requires similarity information between the documents, which is difficult to calculate at a feasible time. The adaptive document clustering scheme has been investigated by researchers to resolve this problem. However, its theoretical viewpoint has not been fully discovered. In this regard, we provide a conceptual viewpoint of the adaptive document clustering based on query-based similarities, by regarding the user’s query as a concept. As a result, adaptive document clustering scheme can be viewed as an approximation of this similarity. Based on this idea, we derive three new query-based similarity measures in language modeling framework, and evaluate them in the context of cluster-based retrieval, comparing with K-means clustering and full document expansion. Evaluation result shows that retrievals based on query-based similarities significantly improve the baseline, while being comparable to other methods. This implies that the newly developed query-based similarities become feasible criterions for adaptive document clustering.
ISSN:	0306-4573 1873-5371
DOI:	10.1016/j.ipm.2006.08.008