Boosted Transformer for Image Captioning

Autor:	Jiangyun Li, Peng Yao, Longteng Guo, Weicun Zhang
Jazyk:	angličtina
Rok vydání:	2019
Předmět:	image captioning self-attention deep learning transformer Technology Engineering (General). Civil engineering (General) TA1-2040 Biology (General) QH301-705.5 Physics QC1-999 Chemistry QD1-999
Zdroj:	Applied Sciences, Vol 9, Iss 16, p 3260 (2019)
Druh dokumentu:	article
ISSN:	2076-3417
DOI:	10.3390/app9163260
Popis:	Image captioning attempts to generate a description given an image, usually taking Convolutional Neural Network as the encoder to extract the visual features and a sequence model, among which the self-attention mechanism has achieved advanced progress recently, as the decoder to generate descriptions. However, this predominant encoder-decoder architecture has some problems to be solved. On the encoder side, without the semantic concepts, the extracted visual features do not make full use of the image information. On the decoder side, the sequence self-attention only relies on word representations, lacking the guidance of visual information and easily influenced by the language prior. In this paper, we propose a novel boosted transformer model with two attention modules for the above-mentioned problems, i.e., “Concept-Guided Attention” (CGA) and “Vision-Guided Attention” (VGA). Our model utilizes CGA in the encoder, to obtain the boosted visual features by integrating the instance-level concepts into the visual features. In the decoder, we stack VGA, which uses the visual information as a bridge to model internal relationships among the sequences and can be an auxiliary module of sequence self-attention. Quantitative and qualitative results on the Microsoft COCO dataset demonstrate the better performance of our model than the state-of-the-art approaches.
Databáze:	Directory of Open Access Journals
Externí odkaz:	https://doaj.org/article/f3d9bebd9f14465c8bd0bcd597fce7a8 Zobrazit plný text záznamu View record in DOAJ