基于ViLBERT与BiLSTM的图像描述算法

doi:10.15888/j.cnki.csa.008133

AIPUB归智期刊联盟

微信公众号

网站二维码

2025年4月24日 2:20 星期四

首页 > 过刊浏览>2021年第30卷第11期 >195-202. DOI:10.15888/j.cnki.csa.008133

PDF HTML阅读 XML下载导出引用引用提醒

基于ViLBERT与BiLSTM的图像描述算法
DOI:
                        10.15888/j.cnki.csa.008133
                    
CSTR:
                        
                    
作者:
                        许昊许昊
上海电力大学 计算机科学与技术学院, 上海 200090
在期刊界中查找
在百度中查找
在本站中查找
张凯张凯
上海电力大学 计算机科学与技术学院, 上海 200090
在期刊界中查找
在百度中查找
在本站中查找
田英杰田英杰
国家电网公司 上海电器科学研究院, 上海 200437
在期刊界中查找
在百度中查找
在本站中查找
种法广种法广
上海电力大学 计算机科学与技术学院, 上海 200090
在期刊界中查找
在百度中查找
在本站中查找
王子超王子超
上海电力大学 计算机科学与技术学院, 上海 200090
在期刊界中查找
在百度中查找
在本站中查找

                    
作者单位:
作者简介:
通讯作者:
中图分类号:
基金项目:国家自然科学基金(61872230, 61802248, 61802249); 上海高校青年教师培养资助计划(ZZsdl18006)

Image Caption Algorithm Based on ViLBERT and BiLSTM

Author:

XU Hao
XU Hao
College of Computer Science and Technology, Shanghai University of Electric Power, Shanghai 200090, China
在期刊界中查找
在百度中查找
在本站中查找
ZHANG Kai
ZHANG Kai
College of Computer Science and Technology, Shanghai University of Electric Power, Shanghai 200090, China
在期刊界中查找
在百度中查找
在本站中查找
TIAN Ying-Jie
TIAN Ying-Jie
Shanghai Electrical Research Institute, State Grid Corporation of China, Shanghai 200437, China
在期刊界中查找
在百度中查找
在本站中查找
CHONG Fa-Guang
CHONG Fa-Guang
College of Computer Science and Technology, Shanghai University of Electric Power, Shanghai 200090, China
在期刊界中查找
在百度中查找
在本站中查找
WANG Zi-Chao
WANG Zi-Chao
College of Computer Science and Technology, Shanghai University of Electric Power, Shanghai 200090, China
在期刊界中查找
在百度中查找
在本站中查找

Affiliation:

Fund Project:

摘要

图/表

访问统计

参考文献

相似文献

引证文献

资源附件

文章评论

摘要:

传统图像描述算法存在提取图像特征利用不足、缺少上下文信息学习和训练参数过多的问题, 提出基于ViLBERT和双层长短期记忆网络(BiLSTM)结合的图像描述算法. 使用ViLBERT作为编码器, ViLBERT模型能将图片特征和描述文本信息通过联合注意力的方式进行结合, 输出图像和文本的联合特征向量. 解码器使用结合注意力机制的BiLSTM来生成图像描述. 该算法在MSCOCO2014数据集进行训练和测试, 实验评价标准BLEU-4和BLEU得分分别达到36.9和125.2, 优于基于传统图像特征提取结合注意力机制图像描述算法. 通过生成文本描述对比可看出, 该算法生成的图像描述能够更细致地表述图片信息.

关键词:图像描述;ViLBERT;BiLSTM;注意力机制

Abstract:

Traditional image captioning has the problems of the under-utilization of extracted image features, the lack of context information learning and too many training parameters. This study proposes an image captioning algorithm based on Vision-and-Language BERT (ViLBERT) and Bidirectional Long Short-Term Memory network (BiLSTM). The ViLBERT model is used as an encoder, which can combine image features and descriptive text information through the co-attention mechanism and output the joint feature vector of image and text. The decoder uses a BiLSTM combined with attention mechanism to generate image caption. The algorithm is trained and tested on MSCOCO2014, and the scores of evaluation criteria BLEU-4 and BLEU are 36.9 and 125.2 respectively. This indicates that the proposed algorithm is better than the image captioning based on the traditional image feature extraction combined with the attention mechanism. The comparison of generated text descriptions demonstrates that the image caption generated by this algorithm can describe the image information in more detail.

Key words:image caption;Vision-and-Language BERT (ViLBERT);Bidirectional Long Short-Term Memory (BiLSTM);attention mechanism

引用本文

许昊,张凯,田英杰,种法广,王子超.基于ViLBERT与BiLSTM的图像描述算法.计算机系统应用,2021,30(11):195-202

复制

文章指标

点击次数:
下载次数:
HTML阅读次数:
引用次数:

历史

收稿日期:2020-12-29
最后修改日期:2021-02-03
录用日期:
在线发布日期: 2021-10-22
出版日期:

微信公众号

网站二维码

引用本文

分享

文章指标

历史

文章二维码

微信公众号

网站二维码

引用本文

分享

微信扫一扫：分享

文章指标

历史

文章二维码