Spatio-Temporal VLAD Encoding for Human Action Recognition in Videos

详细信息查看全文

关键词：Action recognition ; Video classification ; Feature encoding ; Spatio ; temporal VLAD (ST ; VLAD)
刊名：Lecture Notes in Computer Science
出版年：2017
出版时间：2017
年：2017
卷：10132
期：1
页码：365-378
丛书名：MultiMedia Modeling
ISBN：978-3-319-51811-4
卷排序：10132

文摘

Encoding is one of the key factors for building an effective video representation. In the recent works, super vector-based encoding approaches are highlighted as one of the most powerful representation generators. Vector of Locally Aggregated Descriptors (VLAD) is one of the most widely used super vector methods. However, one of the limitations of VLAD encoding is the lack of spatial information captured from the data. This is critical, especially when dealing with video information. In this work, we propose Spatio-temporal VLAD (ST-VLAD), an extended encoding method which incorporates spatio-temporal information within the encoding process. This is carried out by proposing a video division and extracting specific information over the feature group of each video split. Experimental validation is performed using both hand-crafted and deep features. Our pipeline for action recognition with the proposed encoding method obtains state-of-the-art performance over three challenging datasets: HMDB51 (67.6%), UCF50 (97.8%) and UCF101 (91.5%).

地址：北京市海淀区学院路29号邮编：100083

电话：办公室：(+86 10)66554848；文献借阅、咨询服务、科技查新：66554700