ORCID
- Rajesh Jaiswal: 0000-0002-4530-7079
Abstract
Automatic Image Captioning (AIC) combines two distinct machine learning disciplines, Computer Vision (CV) and Natural Language Processing (NLP), forming a challenging research landscape. Much research has been carried out for AIC in the English language facilitated by the availability of English language datasets. However, this is not true for other languages. The aim of this work is to investigate existing AIC models' compatibility for the Italian language. The scope of this research will be aimed to reduced complexity, improved efficiency, and readily trainable models on smaller datasets while maintaining an acceptable level of caption quality. The popular Show and Tell encoder-decoder model was selected and trained on subsets of the MSCOCO-it dataset of 10k, 20k, and 30k with and without human validated captions, and then evaluated with BLEU-3/4, METEOR, ROUGE-L, and CIDEr. The model was also tested on unseen images. The evaluation results were not promising and were below 50%. We intend to investigate multimodal augmentation techniques for image-caption datasets. To enhance interpretability, explainable AI (XAI) methods such as attention visualization and saliency mapping are to be employed to reveal which image regions and linguistic features influenced caption generation. Furthermore, a gender-aware evaluation is planned to be researched and introduced, assessing whether generated captions reproduce gender bias present in training data. This dual focus on explainability and inclusivity strengthens trust in AIC systems while supporting fairer and more transparent applications.
Keywords
explainable AI, gender bias, image captioning
DOI Link
Publication Date
2026-02-16
Event
3rd International Conference on Human-Centred AI - Education and Practice, HCAI-ep 2026
Publication Title
HCAI-ep 2026 - Proceedings of the 2026 Conference on Human Centered Artificial Intelligence - Education and Practice
Publisher
Association for Computing Machinery (ACM)
ISBN
9798400721533
First Page
75
Last Page
78
Deposit Date
2026-05-11
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 International License.
Additional Links
Recommended Citation
De Amicis, Valentina; Jaiswal, Rajesh; and Perez-Tellez, Fernando, "On The Automatic Image Captioning Task In Italian: A Human-Centric Approach" (2026). Research Outputs: 2025-Present. 1.
https://arrow.tudublin.ie/dfrhro/1