Efficient Feature Compression for the Object Tracking Task
Proceedings - International Conference on Image Processing, pp. 3505–3509
Abstract
In object tracking systems, often clients capture video, encode it and transmit it to a server that performs the actual machine task. In this paper we propose an alternative architecture, where we instead transmit features to the server. Specifically, we partition the Joint Detection and Embedding (JDE) person tracking network into client and server side sub-networks and code the intermediate tensors i.e. features. The features are compressed for transmission using a Deep Neural Network (DNN) we design and train specifically for carrying out the tracking task. The DNN uses trainable non-uniform quantizers, conditional probability estimators, hierarchical coding; concepts that have been used in the past for neural networks based image and video compression. Additionally, the DNN includes a novel parameterized dual-path layer that comprises of an autoencoder in one path and a convolution layer in the other. The tensor output by each path is added before being consumed by subsequent layers. The parameter value for this dual-path layer controls the output channel count and correspondingly the bitrate of transmitted bitstream. We demonstrate that our model improves coding efficiency by 43.67% over state-of-the-art Versatile Video Coding standard that codes the source video in pixel domain.
Authors 3
-
Robert Henzel Aachen
Affiliation as printed
RWTH Aachen University
-
Affiliation as printed
Sharp Laboratories of America
-
Affiliation as printed
Sharp Laboratories of America
Cited by 6 stored of 6
6 results
No patents citing this paper on Lens.org (checked 2026-10-06).
References 22
-
W6638667902details pending0citations
-
W6750227808details pending0citations
-
W6746023985details pending0citations
-
W4320930577details pending0citations