A

ELoPE: Fine-Grained Visual Classification with Efficient Localization, Pooling and Embedding

IEEE Winter Conference on Applications of Computer Vision, pp. 1236–1245

Abstract

The task of fine-grained visual classification (FGVC) deals with classification problems that display a small inter-class variance such as distinguishing between different bird species or car models. State-of-the-art approaches typically tackle this problem by integrating an elaborate attention mechanism or (part-) localization method into a standard convolutional neural network (CNN). Also in this work the aim is to enhance the performance of a backbone CNN such as ResNet by including three efficient and lightweight components specifically designed for FGVC. This is achieved by using global k-max pooling, a discriminative embedding layer trained by optimizing class means and an efficient localization module that estimates bounding boxes using only class labels for training. The resulting model achieves state-of-the-art recognition accuracies on multiple FGVC benchmark datasets.

Authors 2

  1. RWTH Aachen University

    Affiliation as printed

    Human Language Technology and Pattern Recognition Group, RWTH Aachen University, Aachen, Germany

    RWTH Aachen University,Human Language Technology and Pattern Recognition Group,Aachen,Germany,52062

  2. RWTH Aachen University

    Affiliation as printed

    Human Language Technology and Pattern Recognition Group, RWTH Aachen University, Aachen, Germany

    RWTH Aachen University,Human Language Technology and Pattern Recognition Group,Aachen,Germany,52062

Cited by 37 stored of 37

No patents citing this paper on Lens.org (checked 2026-10-06).

References 53