PoNA: Pose-Guided Non-Local Attention for Human Pose Transfer

Human pose transfer, which aims at transferring the appearance of a given person to a target pose, is very challenging and important in many applications. Previous work ignores the guidance of pose features or only uses local attention mechanism, leading to implausible and blurry results. We propose...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	IEEE transactions on image processing 2020-01, Vol.29, p.9584-9599
Hauptverfasser:	Li, Kun, Zhang, Jinsong, Liu, Yebin, Lai, Yu-Kun, Dai, Qionghai
Format:	Artikel
Sprache:	eng
Schlagworte:	attention Computer Science Computer Science, Artificial Intelligence Engineering Engineering, Electrical & Electronic Feature extraction generative adversarial network (GAN) Generative adversarial networks Generators Human pose transfer Integrated circuits Science & Technology Shape Task analysis Technology Three-dimensional displays
Online-Zugang:	Volltext bestellen
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	Human pose transfer, which aims at transferring the appearance of a given person to a target pose, is very challenging and important in many applications. Previous work ignores the guidance of pose features or only uses local attention mechanism, leading to implausible and blurry results. We propose a new human pose transfer method using a generative adversarial network (GAN) with simplified cascaded blocks. In each block, we propose a pose-guided non-local attention (PoNA) mechanism with a long-range dependency scheme to select more important regions of image features to transfer. We also design pre-posed image-guided pose feature update and post-posed pose-guided image feature update to better utilize the pose and image features. Our network is simple, stable, and easy to train. Quantitative and qualitative results on Market-1501 and DeepFashion datasets show the efficacy and efficiency of our model. Compared with state-of-the-art methods, our model generates sharper and more realistic images with rich details, while having fewer parameters and faster speed. Furthermore, our generated images can help to alleviate data insufficiency for person re-identification.
ISSN:	1057-7149 1941-0042
DOI:	10.1109/TIP.2020.3029455