结合动态自适应调制和结构关系学习的细粒度图像分类

doi:10.15888/j.cnki.csa.009571

AIPUB归智期刊联盟

微信公众号

网站二维码

2025年4月4日 3:55 星期五

首页 > 过刊浏览>2024年第33卷第8期 >166-175. DOI:10.15888/j.cnki.csa.009571

PDF HTML阅读 XML下载导出引用引用提醒

结合动态自适应调制和结构关系学习的细粒度图像分类
DOI:
                        10.15888/j.cnki.csa.009571
                    
CSTR:
                        32024.14.csa.009571
                    
作者:
                        王衍根王衍根
福州大学 计算机与大数据学院, 福州 350108
在期刊界中查找
在百度中查找
在本站中查找
陈飞陈飞
福州大学 计算机与大数据学院, 福州 350108
在期刊界中查找
在百度中查找
在本站中查找
陈权陈权
福州大学 计算机与大数据学院, 福州 350108
在期刊界中查找
在百度中查找
在本站中查找

                    
作者单位:
作者简介:
通讯作者:
中图分类号:
基金项目:国家自然科学基金(61771141); 福建省自然科学基金(2021J01620)

Fine Grained Image Classification Combining Dynamic Adaptive Modulation and Structural Relationship Learning

Author:

WANG Yan-Gen
WANG Yan-Gen
College of Computer and Data Science, Fuzhou University, Fuzhou 35108, China
在期刊界中查找
在百度中查找
在本站中查找
CHEN Fei
CHEN Fei
College of Computer and Data Science, Fuzhou University, Fuzhou 35108, China
在期刊界中查找
在百度中查找
在本站中查找
CHEN Quan
CHEN Quan
College of Computer and Data Science, Fuzhou University, Fuzhou 35108, China
在期刊界中查找
在百度中查找
在本站中查找

Affiliation:

Fund Project:

摘要

图/表

访问统计

参考文献

相似文献

引证文献

资源附件

文章评论

摘要:

由于细粒度图像类间差异小, 类内差异大的特点, 因此细粒度图像分类任务关键在于寻找类别间细微差异. 最近, 基于Vision Transformer的网络大多侧重挖掘图像最显著判别区域特征. 这存在两个问题: 首先, 网络忽略从其他判别区域挖掘分类线索, 容易混淆相似类别; 其次, 忽略了图像的结构关系, 导致提取的类别特征不准确. 为解决上述问题, 本文提出动态自适应调制和结构关系学习两个模块, 通过动态自适应调制模块迫使网络寻找多个判别区域, 再利用结构关系学习模块构建判别区域间结构关系; 最后利用图卷积网络融合语义信息和结构信息得出预测分类结果. 所提出的方法在CUB-200-2011数据集和NA-Birds数据集上测试准确率分别达到92.9%和93.0%, 优于现有最先进网络.

关键词:细粒度图像分类;Vision Transformer (ViT);动态自适应调制;结构关系学习;图卷积网络

Abstract:

Due to the small inter-class differences and large intra-class differences of fine-grained images, the key to fine-grained image classification tasks is to find subtle differences between categories. Recently, Vision Transformer-based networks mostly focus on mining the most prominent discriminative region features in images. There are two problems with this. Firstly, the network ignores mining classification clues from other discriminative regions, which can easily confuse similar categories. secondly, the structural relationships of images are ignored, resulting in inaccurate extraction of category features. To solve the above problems, this study proposes two modules: dynamic adaptive modulation and structural relationship learning. The dynamic adaptive modulation module forces the network to search for multiple discriminative regions, and then the structural relationship learning module is used to construct structural relationships between discriminative regions. Finally, the graph convolutional network is used to fuse semantic and structural information to obtain predicted classification results. The proposed method achieves testing accuracy of 92.9% and 93.0% on the CUB-200-2011 dataset and NA-Birds dataset, respectively, which is superior to existing state-of-the-art networks.

Key words:fine grained image classification;Vision Transformer (ViT);dynamic adaptive modulation;structural relationship learning;graph convolutional network (GCN)

引用本文

王衍根,陈飞,陈权.结合动态自适应调制和结构关系学习的细粒度图像分类.计算机系统应用,2024,33(8):166-175

复制

文章指标

点击次数:254
下载次数: 687
HTML阅读次数: 507
引用次数: 0

历史

收稿日期:2024-01-27
最后修改日期:2024-02-29
录用日期:
在线发布日期: 2024-06-28
出版日期:

微信公众号

网站二维码

引用本文

分享

文章指标

历史

文章二维码

微信公众号

网站二维码

引用本文

分享

微信扫一扫：分享

文章指标

历史

文章二维码