基于YOLO的多模态特征差分注意融合行人检测
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:


Pedestrian Detection Based on Multimodal Feature Differential Attention Fusion and YOLO
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 增强出版
  • |
  • 文章评论
    摘要:

    针对可见光模态与热红外模态间的差异问题和如何充分利用多模态信息进行行人检测, 本文提出了一种基于YOLO的多模态特征差分注意融合行人检测方法. 该方法首先利用YOLOv3深度神经网络的特征提取主干分别提取多模态特征; 其次在对应多模态特征层之间嵌入模态特征差分注意模块充分挖掘模态间的差异信息, 并经过注意机制强化差异特征表示进而改善特征融合质量, 再将差异信息分别反馈到多模态特征提取主干中, 提升网络对多模态互补信息的学习融合能力; 然后对多模态特征进行分层融合得到融合后的多尺度特征; 最后在多尺度特征层上进行目标检测, 预测行人目标的概率和位置. 在KAIST和LLVIP公开多模态行人检测据集上的实验结果表明, 提出的多模态行人检测方法能有效解决模态间的差异问题, 实现多模态信息的充分利用, 具有较高的检测精度和速度, 具有实际应用价值.

    Abstract:

    In order to address the difference between visible light modality and thermal infrared modality and make full use of multimodal information to perform pedestrian detection, this study proposes a multimodal feature differential attention fusion pedestrian detection method based on YOLO. The method first uses the feature extraction backbone of the YOLOv3 deep neural network to extract multimodal features respectively. Second, the differential attention module of modal features is embedded between the corresponding multimodal feature layers to fully mine the difference information between modalities, and the difference feature representation is strengthened through the attention mechanism, so as to improve the quality of feature fusion. Then, the difference information is fed back to the multimodal feature extraction backbone to improve the network’s ability to learn and fuse multimodal complementary information. In addition, the multimodal features are fused in layers to obtain the multi-scale features. Finally, target detection is performed on the multi-scale feature layer to predict the probability and location of pedestrian targets. The experimental results on the public multimodal pedestrian detection datasets of KAIST and LLVIP show that the proposed multimodal pedestrian detection method can effectively address the difference between modalities and realize the full use of multimodal information. Furthermore, it has high detection accuracy and speed and is of practical application value.

    参考文献
    相似文献
    引证文献
引用本文

王钊,解文彬,文江.基于YOLO的多模态特征差分注意融合行人检测.计算机系统应用,2023,32(4):329-338

复制
分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2022-08-19
  • 最后修改日期:2022-09-22
  • 录用日期:
  • 在线发布日期: 2022-12-23
  • 出版日期:
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京海淀区中关村南四街4号 中科院软件园区 7号楼305房间,邮政编码:100190
电话:010-62661041 传真: Email:csa (a) iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号