Perception-Inspired Bayesian Causal Fusion for Audiovisual Source Localization
Preprint 2026
Abstract
Multimodal fusion promises more accurate perception but only when the modalities share a common cause. When they do not, the second modality carries no information about the target, and fusing it can only corrupt the estimate. We cast this whether-to-fuse decision as Bayesian causal inference, following the optimal-observer model of human multisensory perception, and implement it as a plug-and-play layer above frozen audio and visual models for sound event localization and detection. The model infers a common-cause posterior over visible candidates, then gates precision-weighted fusion accordingly. Fusing unconditionally more than doubles the direction error, whereas the causal gate improves on-screen localization while limiting off-screen degradation, without any joint network retraining.
