Document Type

Theses, Ph.D

Disciplines

1.2 COMPUTER AND INFORMATION SCIENCE, Computer Sciences

Abstract

Image captioning models enable us to automatically generate natural language image descriptions for previously unseen images. It combines the two fields of computer vision and natural language generation, allowing models to interpret the con tent of an image and communicate that knowledge through natural language text.

Research into image captioning has the potential benefit of reducing the gap in digital information availability between fully sighted individuals and those who are visually impaired. However, automatically generated captions often fail to provide the required level of detail and specificity to achieve this goal. Furthermore, current standard evaluation methods are insufficient at measuring some of the perceived shortcomings which need to be addressed to advance the field in a meaningful direction for visually impaired end-users.

In this thesis, a first step is taken towards elucidating a set of quality criteria for automatically generated captions, in the context of viewing image captioning as a tool for communicating clear and relevant information to real human end-users. The proposed criteria are based on the Gricean maxims of the Cooperative Principle– Quantity, Quality, Relation and Manner– and are interpreted here within the context of single-sentence automatically generated captions. Based on this theoretical framework, concrete techniques and architectural improvements are presented which advance the state of the art in a meaningful direction.

First, the issue with generic and ambiguous captions (a violation of the Quantity and Manner criteria) is addressed through a generally applicable, unsupervised fine tuning method that improves the specificity and diversity of the generated captions. The learning signal is provided by a natural language understanding component (in practice, an image retrieval model) through a contrastive margin loss function. The purpose of this fine-tuning method is to modify a pre-trained image captioning model by increasing the similarity of its generated captions to their corresponding images while simultaneously decreasing the similarity to another similar but distinct image. The experiments use an image retrieval model and an image captioning model that were first pre-trained using the same set of image-caption pairs training data. The subsequent fine-tuning phase requires only the images, thus enabling the possibility of leveraging additional unlabelled image data without the need for fur ther, expensive caption annotations. In the experiments, the same images (from the pre-training phase) were re-used during the fine-tuning phase, thus demonstrating the effectiveness of the training method even without access to additional image data. The quantitative results show that the improvements in caption specificity resulted in a noticeable increase in the model’s capacity to generate more diverse and novel captions, with a larger effective vocabulary.

Following this, the criteria of Quality (to only provide information for which there is adequate evidence) is addressed by disentangling the decision of what to describe from the decision of how to describe it. In a typical image captioning model, these two decisions are not separate; the decoder may mention any part of the image at any time during caption generation. In contrast, the method presented here requires that each caption describes a sequence of image regions in the exact order that they were presented to the model. This places a stricter requirement on the captioning model’s adherence to the visual evidence, and limits its ability to rely on memorised object co-occurrence rates. Along with the overall, original architectural idea, a practical improvement to the region attention timing is presented, in which a novel token is introduced to indicate the end-of-chunk of text describing each currently attended image region. A chunk timing precision of 86.55% and a recall of 97.92% were achieved, along with improvements over the previous state-of-the-art results on the standard set of image captioning metrics, and an effective vocabulary size with more than double the number of words compared to the previous state-of the-art model.

Building on this method, the criteria of Relation is approached by addressing the current issue of inflexible one-size-fits-all models. A component-based model is presented, which is capable of generating descriptions with a customisable amount of content that can be specified (and changed) during inference time. Experiments confirm that the model is capable of generating captions with a variable amount of detail (thus additionally addressing the criteria of Quantity). This preference based model achieves similar results on the standard metrics to a model which uses ground-truth region annotations by selecting the automatically detected regions with the closest match to the human-annotated regions mentioned in the ground-truth captions. This represents a step towards a more user-centred approach to automatic caption generation, which is an area of research in need of more attention.

Finally, the outcome of a user study is presented, where participants from the blind and low-vision community were asked about their opinions on the amount of content in captions from the preference-based model. The responses showed that participants generally preferred more content in the captions than is typical in the popular image captioning datasets. Furthermore, the outcome confirmed a statistically significant difference in preferences between individual participants, thus highlighting the need for models that can adapt to individual user preferences.

The discussion and methods presented in this thesis provide a starting point for categorising and addressing shortcomings in caption quality for the purpose of delivering single-sentence, informative image descriptions tailored to visually impaired human end-users; applied here in the context of the English language, In summary, the contributions to the field consist of: i) a proposed theoretical frame work for caption quality criteria based on the Gricean maxims of the Cooperative Principle; ii) a set of concrete algorithms which target existing shortcomings in the image captioning field with respect to each of the proposed criteria; and iii) a user study that confirms the importance of adapting to individual end-user preferences.

DOI

https://doi.org/10.21427/3amj-s550

Funder

Science Foundation Ireland

Creative Commons License

Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License
This work is licensed under a Creative Commons Attribution-NonCommercial-Share Alike 4.0 International License.


Share

COinS