A Study of Using Synthetic Data for Effective Association Knowledge Learning

Association, aiming to link bounding boxes of the same identity in a video sequence, is a central component in multi-object tracking (MOT). To train association modules, e.g., parametric networks, real video data are usually used. However, annotating person tracks in consecutive video frames is expe...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	International journal of automation and computing 2023-04, Vol.20 (2), p.194-206
Hauptverfasser:	Liu, Yuchi, Wang, Zhongdao, Zhou, Xiangxin, Zheng, Liang
Format:	Artikel
Sprache:	eng
Schlagworte:	Adaptation Algorithms Artificial Intelligence Cameras Computer Science Datasets Domains Kalman filters Knowledge Learning Modules Movement Multiple target tracking Neural networks Performance evaluation Research Article Simulation Synthetic data Video data
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	Association, aiming to link bounding boxes of the same identity in a video sequence, is a central component in multi-object tracking (MOT). To train association modules, e.g., parametric networks, real video data are usually used. However, annotating person tracks in consecutive video frames is expensive, and such real data, due to its inflexibility, offer us limited opportunities to evaluate the system performance w.r.t. changing tracking scenarios. In this paper, we study whether 3D synthetic data can replace real-world videos for association training. Specifically, we introduce a large-scale synthetic data engine named MOTX, where the motion characteristics of cameras and objects are manually configured to be similar to those of real-world datasets. We show that, compared with real data, association knowledge obtained from synthetic data can achieve very similar performance on real-world test sets without domain adaption techniques. Our intriguing observation is credited to two factors. First and foremost, 3D engines can well simulate motion factors such as camera movement, camera view, and object movement so that the simulated videos can provide association modules with effective motion features. Second, the experimental results show that the appearance domain gap hardly harms the learning of association knowledge. In addition, the strong customization ability of MOTX allows us to quantitatively assess the impact of motion factors on MOT, which brings new insights to the community.
ISSN:	2731-538X 1476-8186 2153-182X 2731-5398 1751-8520 2153-1838
DOI:	10.1007/s11633-022-1380-x