Unseen object categorization using multiple visual cues

详细信息查看全文

作者：B. Ramesh ; ^{bharath.ramesh03@u.nus.edu}Author Vitae ; C. XiangAuthor Vitae
关键词：Log-polar transform ; Object classification ; Structure-texture decomposition ; Shape extraction ; Bag-of-words model ; ETH-80 dataset
刊名：Neurocomputing
出版年：2017
出版时间：22 March 2017
年：2017
卷：230
期：Complete
页码：88-99
全文大小：1346 K
卷排序：230

文摘

In this paper, we propose an object categorization framework to extract different visual cues and tackle the problem of categorizing previously unseen objects under various viewpoints. Specifically, we decompose the input image into three visual cues: structure, texture and shape cues. Then, local features are extracted using the log-polar transform to achieve scale and rotation invariance. The local descriptors obtained from different visual cues are fused using the bag-of-words representation with some key contributions: (1) a keypoint detection scheme based on variational calculus is proposed for selecting sampling locations; (2) a codebook optimization scheme based on discrete entropy is proposed to choose the optimal codewords and at the same time increase the overall performance. We tested the proposed object classification framework on the ETH-80 dataset using the leave-one-object-out protocol to specifically tackle the problem of categorizing previously unseen objects under various viewpoints. On this popular dataset, the proposed object categorization system obtained a very high improvement in classification performance compared to state-of-the-art methods.

地址：北京市海淀区学院路29号邮编：100083

电话：办公室：(+86 10)66554848；文献借阅、咨询服务、科技查新：66554700