CPT: Colorful Prompt Tuning for pre-trained vision-language models

Autor:	Yuan Yao, Ao Zhang, Zhengyan Zhang, Zhiyuan Liu, Tat-Seng Chua, Maosong Sun
Jazyk:	angličtina
Rok vydání:	2024
Předmět:	Vision-language pre-training models Prompt tuning Electronic computers. Computer science QA75.5-76.95
Zdroj:	AI Open, Vol 5, Iss , Pp 30-38 (2024)
Druh dokumentu:	article
ISSN:	2666-6510
DOI:	10.1016/j.aiopen.2024.01.004
Popis:	Vision-Language Pre-training (VLP) models have shown promising capabilities in grounding natural language in image data, facilitating a broad range of cross-modal tasks. However, we note that there exists a significant gap between the objective forms of model pre-training and fine-tuning, resulting in a need for large amounts of labeled data to stimulate the visual grounding capability of VLP models for downstream tasks. To address the challenge, we present Color-based Prompt Tuning (CPT), a novel paradigm for tuning VLP models, which reformulates visual grounding into a fill-in-the-blank problem with color-based co-referential markers in image and text, maximally mitigating the gap. In this way, CPT enables strong few-shot and even zero-shot visual grounding capabilities of VLP models. Comprehensive experimental results show that CPT achieves state-of-the-art performance on zero/few-shot visual grounding (e.g., 75.1 zero-shot accuracy in RefCOCO evaluation), outperforming fine-tuned and other prompt-tuned models by a large margin. Moreover, CPT can also be easily extended to achieve promising zero/few-shot performance on other vision-language tasks, such as visual relation detection, visual commonsense reasoning and visual question answering. We make the data and codes publicly available at https://github.com/thunlp/CPT.
Databáze:	Directory of Open Access Journals
Externí odkaz:	https://doaj.org/article/19ad385354e14dd7bf707d68a7194a53 Zobrazit plný text záznamu View record in DOAJ