Skip to main navigation Skip to search Skip to main content

Multimodal Knowledge Distillation for Acoustic-Aware Object Detection

  • Saheli Hazra
  • , Nushrat Hussain
  • , Sudip Das
  • , Arindam Das
  • , Ujjwal Bhattachary
  • Indian Statistical Institute
  • Valeo

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Automatic object detection has traditionally relied on vision-based inputs, whereas human perception integrates both visual and auditory cues to interpret the environment effectively. Individuals with visual impairments, particularly those who are blind, often develop heightened auditory abilities that support navigation and spatial awareness. Motivated by the above fact, here we present our study of a multi-teacher single-student distillation framework capable of object detection from audio-based input only. However, its modality-specific experts are trained using RGB, thermal, and depth input information to guide its student network accepting only audio input from the environment. Unlike existing audio-visual distillation frameworks, our proposed model takes into consideration the following two issues: (i) the varying reliability of predictions made by the teacher network and (ii) the importance of focusing on regions of the key feature. In regard to the above, we propose here a confidence-driven distillation strategy that adaptively weighs supervision based on the prediction confidence, and introduce AAMCA (Audio-Guided Adaptive Multimodal Contrastive Attention) module. The proposed AAMCA module aligns audio features with their visual counterparts through attention-guided contrastive learning. This helps the audio based detector to obtain richer spatial representations and effectively mimic the performance of multimodal systems when the input to the detection network solely relies on the audio signal only during inference time. Extensive experiments demonstrate that our method consistently outperforms state-of-the-art approaches, highlighting the effectiveness of confidence-driven distillation and adaptive attention-based alignment strategy.

Original languageEnglish
Title of host publicationPattern Recognition - 28th International Conference, ICPR 2026, Proceedings
EditorsMaria De Marsico, Tin Kam Ho, Frederic Jurie, Cheng-Lin Liu, Daniel Lopresti, Ingela Nyström, Jean-Marc Ogier, Arun Ross, Liang Wang
PublisherSpringer Science and Business Media Deutschland GmbH
Pages397-412
Number of pages16
ISBN (Print)9783032319296
DOIs
Publication statusPublished - 2027
Externally publishedYes
Event28th International Conference on Pattern Recognition, ICPR 2026 - Lyon, France
Duration: 17 Aug 202622 Aug 2026

Publication series

NameLecture Notes in Computer Science
Volume16825 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference28th International Conference on Pattern Recognition, ICPR 2026
Country/TerritoryFrance
CityLyon
Period17/08/2622/08/26

Keywords

  • adaptive attention
  • Audio-based vehicle detection
  • audio-visual integration
  • contrastive learning
  • multimodal knowledge distillation

Fingerprint

Dive into the research topics of 'Multimodal Knowledge Distillation for Acoustic-Aware Object Detection'. Together they form a unique fingerprint.

Cite this