motivation
Perception and reasoning run as separate steps
LLM-based multimodal reasoning either aligns image features with the language space or turns the image into text for the LLM. Both keep visual perception and language reasoning apart. Chain-of-thought drifts away from the question's intent. Direct prediction skips the knowledge the question needs.
People solve such questions by abduction. They guess candidate answers from what they see, justify each one with visual cues and background knowledge, and keep the explanation that fits best. ViLA models this process across the VLM and the LLM.






