Dynamic Pooling for Complex Event Recognition

Weixin Li, Qian Yu, Ajay Divakaran, Nuno Vasconcelos; Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2013, pp. 2728-2735

Abstract


The problem of adaptively selecting pooling regions for the classification of complex video events is considered. Complex events are defined as events composed of several characteristic behaviors, whose temporal configuration can change from sequence to sequence. A dynamic pooling operator is defined so as to enable a unified solution to the problems of event specific video segmentation, temporal structure modeling, and event detection. Video is decomposed into segments, and the segments most informative for detecting a given event are identified, so as to dynamically determine the pooling operator most suited for each sequence. This dynamic pooling is implemented by treating the locations of characteristic segments as hidden information, which is inferred, on a sequence-by-sequence basis, via a large-margin classification rule with latent variables. Although the feasible set of segment selections is combinatorial, it is shown that a globally optimal solution to the inference problem can be obtained efficiently, through the solution of a series of linear programs. Besides the coarselevel location of segments, a finer model of video structure is implemented by jointly pooling features of segmenttuples. Experimental evaluation demonstrates that the resulting event detector has state-of-the-art performance on challenging video datasets.

Related Material


[pdf]
[bibtex]
@InProceedings{Li_2013_ICCV,
author = {Li, Weixin and Yu, Qian and Divakaran, Ajay and Vasconcelos, Nuno},
title = {Dynamic Pooling for Complex Event Recognition},
booktitle = {Proceedings of the IEEE International Conference on Computer Vision (ICCV)},
month = {December},
year = {2013}
}