Multi-Modal Extreme Classification

Mittal, Anshul; Dahiya, Kunal; Malani, Shreya; Ramaswamy, Janani; Kuruvilla, Seba; Ajmera, Jitendra; Chang, Keng-hao; Agarwal, Sumeet; Kar, Purushottam; Varma, Manik

Multi-Modal Extreme Classification

Anshul Mittal, Kunal Dahiya, Shreya Malani, Janani Ramaswamy, Seba Kuruvilla, Jitendra Ajmera, Keng-hao Chang, Sumeet Agarwal, Purushottam Kar, Manik Varma; Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 12393-12402

Abstract

This paper develops the MUFIN technique for extreme classification (XC) tasks with millions of labels where datapoints and labels are endowed with visual and textual descriptors. Applications of MUFIN to product-to-product recommendation and bid query prediction over several millions of products are presented. Contemporary multi-modal methods frequently rely on purely embedding-based methods. On the other hand, XC methods utilize classifier architectures to offer superior accuracies than embedding-only methods but mostly focus on text-based categorization tasks. MUFIN bridges this gap by reformulating multi-modal categorization as an XC problem with several millions of labels. This presents the twin challenges of developing multi-modal architectures that can offer embeddings sufficiently expressive to allow accurate categorization over millions of labels; and training and inference routines that scale logarithmically in the number of labels. MUFIN develops an architecture based on cross-modal attention and trains it in a modular fashion using pre-training and positive and negative mining. A novel product-to-product recommendation dataset MM-AmazonTitles-300K containing over 300K products was curated from publicly available amazon.com listings with each product endowed with a title and multiple images. On the MM-AmazonTitles-300K and Polyvore datasets, and a dataset with over 4 million labels curated from click logs of the Bing search engine, MUFIN offered at least 3% higher accuracy than leading text-based, image-based and multi-modal techniques.

Related Material

[pdf]

[bibtex]

@InProceedings{Mittal_2022_CVPR, author = {Mittal, Anshul and Dahiya, Kunal and Malani, Shreya and Ramaswamy, Janani and Kuruvilla, Seba and Ajmera, Jitendra and Chang, Keng-hao and Agarwal, Sumeet and Kar, Purushottam and Varma, Manik}, title = {Multi-Modal Extreme Classification}, booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, month = {June}, year = {2022}, pages = {12393-12402} }