Dense-Captioning Events in Videos: SYSU Submission to ActivityNet Challenge 2020
This technical report presents a brief description of our submission to the dense video captioning task of ActivityNet Challenge 2020. Our approach follows a two-stage pipeline: first, we extract a set of temporal event proposals; then we propose a multi-event captioning model to capture the event-l...
Gespeichert in:
Hauptverfasser: | , , |
---|---|
Format: | Artikel |
Sprache: | eng |
Schlagworte: | |
Online-Zugang: | Volltext bestellen |
Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Zusammenfassung: | This technical report presents a brief description of our submission to the
dense video captioning task of ActivityNet Challenge 2020. Our approach follows
a two-stage pipeline: first, we extract a set of temporal event proposals; then
we propose a multi-event captioning model to capture the event-level temporal
relationships and effectively fuse the multi-modal information. Our approach
achieves a 9.28 METEOR score on the test set. |
---|---|
DOI: | 10.48550/arxiv.2006.11693 |