Tagging Webcast Text in Baseball Videos by Video Segmentation and Text Alignment

Sports video annotation, an active research area in the field of multimedia content understanding, is an essential process in applications, such as summarization, highlight extraction, event detection, and retrieval. This paper considers the issue in relation to the annotation of baseball videos. Co...

Ausführliche Beschreibung

Gespeichert in:
Bibliographische Detailangaben
Veröffentlicht in:IEEE transactions on circuits and systems for video technology 2012-07, Vol.22 (7), p.999-1013
Hauptverfasser: CHIU, Chih-Yi, LIN, Po-Chih, LI, Sheng-Yang, TSAI, Tsung-Han, TSAI, Yu-Lung
Format: Artikel
Sprache:eng
Schlagworte:
Online-Zugang:Volltext bestellen
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
Beschreibung
Zusammenfassung:Sports video annotation, an active research area in the field of multimedia content understanding, is an essential process in applications, such as summarization, highlight extraction, event detection, and retrieval. This paper considers the issue in relation to the annotation of baseball videos. Conventional baseball video annotation frameworks are based primarily on video content analysis, such as scoreboard recognition and machine learning techniques, which require a substantial amount of human input to collect and organize training data. The performance of such frameworks might become unstable if they encounter audiovisual patterns not included in the training data. To address the issue, we propose a novel framework for baseball video annotation that aligns high-level webcast text with low-level video content. Several cues, which are derived from the video content and webcast text, are utilized for alignment by leveraging hierarchical agglomerative clustering and genetic algorithm optimization. In addition, we develop an unsupervised method to learn the pitch segment properties of baseball videos by Markov random walk, and thereby reduce the need for human intervention substantially. Our experiments demonstrate that the proposed framework yields a robust result against a variety of video content and enhances the automaticity in baseball video annotation.
ISSN:1051-8215
1558-2205
DOI:10.1109/TCSVT.2012.2189478