IEEE Transactions on Multimedia 2024

Cross Modality Bias in Visual Question Answering: A Causal View with Possible Worlds VQA

Figure: Cross Modality Bias in Visual Question Answering: A Causal View with Possible Worlds VQA

Ali Vosoughi, Shijian Deng, Songyang Zhang, Yapeng Tian, Chenliang Xu, Jiebo Luo

First causal method to reduce vision and language bias in VQA at once, with large gains on VQA-CP v2 and roughly doubled accuracy on numerical questions.

First causal method to reduce vision and language bias in VQA at once, with large gains on VQA-CP v2 and roughly doubled accuracy on numerical questions.

All publications