Multimodal Multi-User Surface Recognition With the Kernel Two-Sample Test
IEEE Transactions on Automation Science and Engineering, vol. 21, pp. 4432–4447
Abstract
Machine learning and deep learning have been used extensively to classify physical surfaces through images and time-series contact data. However, these methods rely on human expertise and entail the time-consuming processes of data and parameter tuning. To overcome these challenges, we propose an easily implemented framework that can directly handle heterogeneous data sources for classification tasks. Our data-versus-data approach automatically quantifies distinctive differences in distributions in a high-dimensional space via kernel two-sample testing between two sets extracted from multimodal data (e.g., images, sounds, haptic signals). We demonstrate the effectiveness of our technique by benchmarking against expertly engineered classifiers for visual-audio-haptic surface recognition due to the industrial relevance, difficulty, and competitive baselines of this application; ablation studies confirm the utility of key components of our pipeline. As shown in our open-source code, we achieve 97.2% accuracy on a standard multi-user dataset with 108 surface classes, outperforming the state-of-the-art machine-learning algorithm by 6% on a more difficult version of the task. The fact that our classifier obtains this performance with minimal data processing in the standard algorithm setting reinforces the powerful nature of kernel methods for learning to recognize complex patterns.Note to Practitioners—We demonstrate how to apply the kernel two-sample test to a surface-recognition task, discuss opportunities for improvement, and explain how to use this framework for other classification problems with similar properties. Automating surface recognition could benefit both surface inspection and robot manipulation. Our algorithm quantifies class similarity and therefore outputs an ordered list of similar surfaces. This technique is well suited for quality assurance and documentation of newly received materials or newly manufactured parts. More generally, our automated classification pipeline can handle heterogeneous data sources including images and high-frequency time-series measurements of vibrations, forces and other physical signals. As our approach circumvents the time-consuming process of feature engineering, both experts and non-experts can use it to achieve high-accuracy classification. It is particularly appealing for new problems without existing models and heuristics. In addition to strong theoretical properties, the algorithm is straightforward to use in practice since it requires only kernel evaluations. Its transparent architecture can provide fast insights into the given use case under different sensing combinations without costly optimization. Practitioners can also use our procedure to obtain the minimum data-acquisition time for independent time-series data from new sensor recordings.
Authors 4
-
Max Planck Institute for Intelligent Systems · University of Stuttgart
Affiliation as printed
Max Planck Institute for Intelligent Systems (MPI-IS), Stuttgart, Germany
Faculty of Engineering Design, Production Engineering and Automotive Engineering, University of Stuttgart, Stuttgart, Germany
-
RWTH Aachen University · Max Planck Institute for Intelligent Systems
Affiliation as printed
Institute for Data Science in Mechanical Engineering, RWTH Aachen University, Aachen, Germany
MPI-IS, Stuttgart, Germany
-
RWTH Aachen University · Max Planck Institute for Intelligent Systems
Affiliation as printed
Institute for Data Science in Mechanical Engineering, RWTH Aachen University, Aachen, Germany
MPI-IS, Stuttgart, Germany
-
Max Planck Institute for Intelligent Systems · University of Stuttgart
Affiliation as printed
Max Planck Institute for Intelligent Systems (MPI-IS), Stuttgart, Germany
Faculty of Engineering Design, Production Engineering and Automotive Engineering, University of Stuttgart, Stuttgart, Germany
Cited by 6 stored of 6
6 results
No patents citing this paper on Lens.org (checked 2026-10-06).
References 52
-
W6684191040details pending0citations
-
W6688325169details pending0citations
-
W3103934428details pending0citations
-
W6683633756details pending0citations
-
W1779010541details pending0citations
-
W1909952827details pending0citations
-
W2004162386details pending0citations
-
W2060081565details pending0citations
-
W2067755752details pending0citations
-
W2071356986details pending0citations
-
W2075654868details pending0citations
-
W2080317558details pending0citations
-
W2107780715details pending0citations
-
W2143934463details pending0citations
-
W2158662820details pending0citations
-
W2331974298details pending0citations
-
W2344756528details pending0citations