TY - GEN
T1 - Multimodal Knowledge Distillation for Acoustic-Aware Object Detection
AU - Hazra, Saheli
AU - Hussain, Nushrat
AU - Das, Sudip
AU - Das, Arindam
AU - Bhattachary, Ujjwal
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2027.
PY - 2027
Y1 - 2027
N2 - Automatic object detection has traditionally relied on vision-based inputs, whereas human perception integrates both visual and auditory cues to interpret the environment effectively. Individuals with visual impairments, particularly those who are blind, often develop heightened auditory abilities that support navigation and spatial awareness. Motivated by the above fact, here we present our study of a multi-teacher single-student distillation framework capable of object detection from audio-based input only. However, its modality-specific experts are trained using RGB, thermal, and depth input information to guide its student network accepting only audio input from the environment. Unlike existing audio-visual distillation frameworks, our proposed model takes into consideration the following two issues: (i) the varying reliability of predictions made by the teacher network and (ii) the importance of focusing on regions of the key feature. In regard to the above, we propose here a confidence-driven distillation strategy that adaptively weighs supervision based on the prediction confidence, and introduce AAMCA (Audio-Guided Adaptive Multimodal Contrastive Attention) module. The proposed AAMCA module aligns audio features with their visual counterparts through attention-guided contrastive learning. This helps the audio based detector to obtain richer spatial representations and effectively mimic the performance of multimodal systems when the input to the detection network solely relies on the audio signal only during inference time. Extensive experiments demonstrate that our method consistently outperforms state-of-the-art approaches, highlighting the effectiveness of confidence-driven distillation and adaptive attention-based alignment strategy.
AB - Automatic object detection has traditionally relied on vision-based inputs, whereas human perception integrates both visual and auditory cues to interpret the environment effectively. Individuals with visual impairments, particularly those who are blind, often develop heightened auditory abilities that support navigation and spatial awareness. Motivated by the above fact, here we present our study of a multi-teacher single-student distillation framework capable of object detection from audio-based input only. However, its modality-specific experts are trained using RGB, thermal, and depth input information to guide its student network accepting only audio input from the environment. Unlike existing audio-visual distillation frameworks, our proposed model takes into consideration the following two issues: (i) the varying reliability of predictions made by the teacher network and (ii) the importance of focusing on regions of the key feature. In regard to the above, we propose here a confidence-driven distillation strategy that adaptively weighs supervision based on the prediction confidence, and introduce AAMCA (Audio-Guided Adaptive Multimodal Contrastive Attention) module. The proposed AAMCA module aligns audio features with their visual counterparts through attention-guided contrastive learning. This helps the audio based detector to obtain richer spatial representations and effectively mimic the performance of multimodal systems when the input to the detection network solely relies on the audio signal only during inference time. Extensive experiments demonstrate that our method consistently outperforms state-of-the-art approaches, highlighting the effectiveness of confidence-driven distillation and adaptive attention-based alignment strategy.
KW - adaptive attention
KW - Audio-based vehicle detection
KW - audio-visual integration
KW - contrastive learning
KW - multimodal knowledge distillation
UR - https://www.scopus.com/pages/publications/105046996059
U2 - 10.1007/978-3-032-31930-2_27
DO - 10.1007/978-3-032-31930-2_27
M3 - Conference contribution
AN - SCOPUS:105046996059
SN - 9783032319296
T3 - Lecture Notes in Computer Science
SP - 397
EP - 412
BT - Pattern Recognition - 28th International Conference, ICPR 2026, Proceedings
A2 - De Marsico, Maria
A2 - Ho, Tin Kam
A2 - Jurie, Frederic
A2 - Liu, Cheng-Lin
A2 - Lopresti, Daniel
A2 - Nyström, Ingela
A2 - Ogier, Jean-Marc
A2 - Ross, Arun
A2 - Wang, Liang
PB - Springer Science and Business Media Deutschland GmbH
T2 - 28th International Conference on Pattern Recognition, ICPR 2026
Y2 - 17 August 2026 through 22 August 2026
ER -