Features Discovery for Web Classification Using Support Vector Machine

The ever fast-expanding web information resources pose a big challenge to internet users seeking the most relevant, latest and quality information. The sheer vast amount of web information has resulted in restructuring of the resources. Thus, an appropriate web classification method needs to be esta...

Ausführliche Beschreibung

Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Othman, M S, Yusuf, L M, Salim, J
Format: Tagungsbericht
Sprache:eng
Schlagworte:
Online-Zugang:Volltext bestellen
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
Beschreibung
Zusammenfassung:The ever fast-expanding web information resources pose a big challenge to internet users seeking the most relevant, latest and quality information. The sheer vast amount of web information has resulted in restructuring of the resources. Thus, an appropriate web classification method needs to be established in order for quality web information to be accessed. This paper intends to discuss the web document features that classify the web information resources. Six web document features have been identified which are text, meta tag and title (A), title and text (B), title (C), meta tag and title (D), meta tag (E) and text (F). The Support Vector Machine (SVM) method is used to classify the web document while four types of kernels namely: Radial Basis Function (RBF), linear, polynomial and sigmoid kernels was applied to test the accuracy of the classification. The studies show that the text, meta tag and title (A) features is the best features for classification of web document that employs the four kernels followed by the features on title and text (B) as well as the features on meta tag and title (C). The studies also found that the linear kernel is the best kernel in classifying the web document compared to the RBF, polynomial and sigmoid kernel.
DOI:10.1109/ICICCI.2010.16