CASIA OpenIR  > 数字内容技术与服务研究中心  > 听觉模型与认知计算
Cascaded Mutual Modulation for Visual Reasoning
Yao, Yiqun; Xu, Jiaming; Wang, Feng; Xu, Bo
Conference NameIn Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP2018)
Conference Date2018/10/31-2018/11/04
Conference PlaceBrussels, Belgium

Visual reasoning is a special visual question answering problem that is multi-step and compositional by nature, and also requires intensive text-vision interactions. We propose CMM: Cascaded Mutual Modulation as a novel end-to-end visual reasoning model. CMM includes a multi-step comprehension process for both question and image. In each step, we use a Feature-wise Linear Modulation (FiLM) technique to enable textual/visual pipeline to mutually control each other. Experiments show that CMM significantly outperforms most related models, and reach state-of-the-arts on two visual reasoning benchmarks: CLEVR and NLVR, collected from both synthetic and natural languages. Ablation studies confirm that both our multi-step framework and our visual-guided language modulation are critical to the task. Our code is available at

KeywordVisual Reasoning Natural Language Understanding Question Answering
Document Type会议论文
Corresponding AuthorXu, Jiaming
Recommended Citation
GB/T 7714
Yao, Yiqun,Xu, Jiaming,Wang, Feng,et al. Cascaded Mutual Modulation for Visual Reasoning[C],2018:975-980.
Files in This Item:
File Name/Size DocType Version Access License
028-非正式-2018-EMNLP-C(480KB)会议论文 开放获取CC BY-NC-SAView Download
Related Services
Recommend this item
Usage statistics
Export to Endnote
Google Scholar
Similar articles in Google Scholar
[Yao, Yiqun]'s Articles
[Xu, Jiaming]'s Articles
[Wang, Feng]'s Articles
Baidu academic
Similar articles in Baidu academic
[Yao, Yiqun]'s Articles
[Xu, Jiaming]'s Articles
[Wang, Feng]'s Articles
Bing Scholar
Similar articles in Bing Scholar
[Yao, Yiqun]'s Articles
[Xu, Jiaming]'s Articles
[Wang, Feng]'s Articles
Terms of Use
No data!
Social Bookmark/Share
File name: 028-非正式-2018-EMNLP-Cascaded Mutual Modulation for Visual Reasoning.pdf
Format: Adobe PDF
All comments (0)
No comment.

Items in the repository are protected by copyright, with all rights reserved, unless otherwise indicated.