Autor: |
Yinglin Zheng, Ting Zhang, Jianmin Bao, Dong Chen, Ming Zeng |
Jazyk: |
angličtina |
Rok vydání: |
2024 |
Předmět: |
|
Zdroj: |
Graphical Models, Vol 135, Iss , Pp 101223- (2024) |
Druh dokumentu: |
article |
ISSN: |
1524-0703 |
DOI: |
10.1016/j.gmod.2024.101223 |
Popis: |
Instructional image editing has received a significant surge of attention recently. In this work, we are interested in the challenging problem of instructional image editing within the particular fashion realm, a domain with significant potential demand in both commercial and personal contexts. This specific domain presents heightened challenges owing to the stringent quality requirements. It necessitates not only the creation of vivid details in alignment with instructions, but also the preservation of precise attributes unrelated to the text guidance. Naive extensions of existing image editing methods produce noticeable artifacts. In order to achieve high-fidelity fashion editing, we propose a novel framework, leveraging the generative prior of a pre-trained human generator and performing edit in the latent space. In addition, we introduce a novel CLIP-based loss to better align the generated target with the instruction. Extensive experiments demonstrate that our approach outperforms prior works including GAN-based editing as well as diffusion-based editing by a large margin, showing impressive visual quality. |
Databáze: |
Directory of Open Access Journals |
Externí odkaz: |
|