Automatic Extraction of Web Page Text Information Based on Network Topology Coincidence Degree

In order to effectively solve the above problems, an automatic extraction method of web text information based on network topology coincidence degree is proposed. Search engine, web crawler, and hypertext tag are used to classify web text information, and then, dimensionality reduction is carried ou...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	Wireless communications and mobile computing 2022-03, Vol.2022, p.1-10
Hauptverfasser:	Shu, Zhinian, Li, Xiaorong
Format:	Artikel
Sprache:	eng
Schlagworte:	Data collection Hypertext Network topologies Optimization algorithms Search engines Similarity Websites
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	In order to effectively solve the above problems, an automatic extraction method of web text information based on network topology coincidence degree is proposed. Search engine, web crawler, and hypertext tag are used to classify web text information, and then, dimensionality reduction is carried out. After processing, the similarity of different features of web page text information is calculated, the similarity is sorted, and the similar text information is extracted according to the correlation based on segment estimation. The experimental results show that the designed method can simplify the complexity of the associated information of the data set and improve the amount of data collection and the success rate of information collection.
ISSN:	1530-8669 1530-8677
DOI:	10.1155/2022/9220661