EACL 2026 Findings

ViLA visual–linguistic abductive reasoning

Visual–Linguistic Abductive Reasoning with LLMs for Knowledge-based Visual Question Answering

Jieun Kim·Yujin Jeong·Sung-Bae Cho

Yonsei University

ViLA overview: a VLM proposes answer candidates and visual hypotheses, an LLM proposes linguistic hypotheses, they form an abductive chain, and fuzzy scoring from both models selects the final hypotheses that the VLM uses to answer.
ViLA hypothesizes plausible answers, generates visual and linguistic premises, and keeps the most coherent ones through fuzzy scoring before the VLM gives its final answer.

motivation

Perception and reasoning run as separate steps

LLM-based multimodal reasoning either aligns image features with the language space or turns the image into text for the LLM. Both keep visual perception and language reasoning apart. Chain-of-thought drifts away from the question's intent. Direct prediction skips the knowledge the question needs.

People solve such questions by abduction. They guess candidate answers from what they see, justify each one with visual cues and background knowledge, and keep the explanation that fits best. ViLA models this process across the VLM and the LLM.

For 'What is required to be on the ground to do this sport?', chain-of-thought answers 'suitable environment' and direct prediction answers 'Skis', both wrong; the abductive chain scores premises for Skis and Snow and answers Snow.

abstract

Hypothesize, justify, then select

Recent attempts to leverage large language models (LLMs) for reasoning and pre-trained knowledge in multi-modal reasoning focus on two main approaches: aligning image features with linguistic space, and converting images into textual cues to exploit the implicit reasoning capabilities of LLMs. Although they integrate visual information into the reasoning pipeline, they often treat visual perception and language reasoning as separate processes, limiting the potential for fully unified multi-modal reasoning. In this paper, we propose a novel method, Visual–Linguistic Abductive Reasoning (ViLA), inspired by human abductive reasoning processes. ViLA hypothesizes a plausible answer, generates the corresponding visual and textual premises, and employs fuzzy scoring to select the most coherent combination, thus deriving the final inference. This process integrates visual and linguistic modalities into interpretable abductive reasoning chains, enabling unified multi-modal reasoning. Without fine-tuning LLMs or retrieving external knowledge, ViLA improves performance by 2.31% on AOKVQA, 1.7% on OKVQA, and 1.7% on GQA over previous state-of-the-art models, while also improving interpretability and stability.

method

An abductive chain across the VLM and the LLM

Two cats lying on a desk, one of them on a keyboard.

QWhere are the cats sitting on top of?

o₁desk
  • Cats are on a desk.
    Sa 0.9Sr 0.50.86kept
  • Desk is a common place for cats to sit.
    Sa 0.9Sr 0.10.82kept

H₁ ∧ H₂ ⊨ desk

o₂keyboard
  • Cats are on keyboard.
    Sa 0.9Sr 0.50.86kept
  • Cats are on computer keyboard.
    Sa 0.9Sr 0.50.86kept

H₃ ∧ H₄ ⊨ keyboard

Sa adequacy to the image (VLM)Sr relevance to the question (LLM)bar: validity μV, line at t = 0.7

VLM with Hfinal → keyboard ✓ LLaVA → desk ✗ CoT → desk ✗

Example and scores from Figure 6 of the paper. Validity is computed with wv = 0.9 and wl = 0.1.

Why keep several candidates

The VLM's first beam is right 58% of the time, but the correct answer is within the top 2 beams 72% of the time and within the top 10, 89%. The model often finds the right answer and then discards it at the final selection. ViLA keeps the top-k beam outputs as answer candidates for abductive reasoning.

Top-k accuracy rises from 58% at k=1 to 72% at k=2 and 89% at k=10, while Hit@k falls from 58 to 2.

Fuzzy scoring by prompt

Both models score every premise with the same multiple-choice prompt. The VLM judges whether it is reasonable to the image (adequacy μA), and the LLM whether it is relevant to the question (relevance μR). The options map to very low 0.1, low 0.5, medium 0.7, high 0.9 and very high 0.95. The value 0.5 counts as low because both models give moderate scores even for weak entailment.

Validity μV = wv·μA + wl·μR, a convex combination with wv + wl = 1. Premises with μV ≥ t become Hfinal.

Choose the most appropriate score to indicate
whether the information provided is {reasonable
to the image / relevant to the question}.
Return only the letter corresponding to your
choice from the options below:
(a) 0.1  (b) 0.5  (c) 0.7  (d) 0.9  (e) 0.95
(f) unknown

results

Higher accuracy without fine-tuning or retrieval

Evaluated on OK-VQA, A-OKVQA and GQA. Unless noted, the VLM is LLaVA-NeXT (Mistral-7B) and the LLM is Mistral-7B, with 2 answer candidates, 2 premises, t = 0.7 and wv = 0.9.

69.98
A-OKVQA accuracy, best in the comparison
68.91
OK-VQA accuracy, best in the comparison
+4.85
GQA accuracy over LLaVA-NeXT

OK-VQA and A-OKVQA

MethodOK-VQAA-OKVQA
Direct answering
MAVEx41.37–
UnifER42.13–
ViLBERT35.2030.60
LXMERT36.9130.70
KRISP38.4033.70
PICa48.0043.37
Prophet61.1058.20
LION57.3360.87
QACap68.2066.30
Structured reasoning
ViperGPT51.9039.50
VisProg*41.8452.12
VCTP56.2053.20
Base VLMs
BLIP232.3238.94
InstructBLIP61.1662.23
LLaVA-NeXT55.6166.48
  + CoT (7B)13.3434.10
  + CoT (GPT-4)30.5652.71
Base VLMs + ViLA
BLIP2 + ViLA (7B)34.0539.37
InstructBLIP + ViLA (7B)61.5162.31
LLaVA-NeXT + ViLA (7B)57.9268.18
LLaVA-NeXT-13B + ViLA (GPT-4)68.9169.98

Accuracy (%). GPT-4 refers to gpt-4o-mini.

GQA

MethodKnowledge sourceGQA
LXMERTLXMERT60.0
BLIP2Vicuna-13B41.0
InstructBLIPVicuna-7B49.2
InstructBLIPVicuna-13B49.5
ViperGPTGPT-348.1
VisProgCLIP, GPT-350.5
LLaVA-NeXTMistral-7B58.74
  + CoTMistral-7B32.1
  + ViLAMistral-7B63.59

Accuracy (%) on GQA test-dev.

Turning the image into text for chain-of-thought hurts: LLaVA-NeXT drops from 58.74 to 32.1 on GQA with CoT. ViLA keeps each modality as its own premises and gains 4.85 instead.

Reasoning depth and question type

On GQA, ViLA beats LLaVA at every reasoning depth, most at one step (+18.1). It improves both binary and open questions. Its predicted answer distribution is also closer to the ground truth, with a lower chi-square distance (Dist).

GQA accuracy of LLaVA and ViLA by reasoning step (gains +18.1, +7.9, +0.9, +2.3, +0.4) and by question type (Binary +1.9, Open +7.3, Acc +4.9, Dist -0.1).

GQA question categories

ViLA improves every structural and semantic category. The largest gains are in Query among structural types and in Global and Category among semantic types.

Structural reasoning
ModelChoCompLogQryVerAvg
LLaVA84.1567.2377.9371.2281.3576.38
+ ViLA85.5668.2579.8778.5583.7579.20
Semantic reasoning
ModelAttrCatGlobObjRelAvg
LLaVA68.0137.9521.6684.3251.5352.69
+ ViLA71.3351.7264.3386.8954.9765.85

Rationale quality

On the A-OKVQA validation set, ViLA's rationales are closer to the ground-truth rationales than chain-of-thought rationales from the same LLM and from GPT-4o-mini. Similarity is the average cosine similarity of CLIP (ViT-B/16) text embeddings. CoT* generates the rationale before the answer.

MethodLLMSentence similarity
CoTMistral-7B38.7
CoT*Mistral-7B40.9
CoTGPT-4o-mini39.02
ViLAMistral-7B44.8

Qualitative examples

For each answer candidate, ViLA builds premises and scores them with adequacy Sa and relevance Sr; the premises shown passed the fuzzy evaluation. With these premises, ViLA answers keyboard and water, while LLaVA and CoT answer desk and pond.

Two examples. Cats: premises for desk and keyboard with their scores; LLaVA and CoT answer desk, ViLA answers keyboard. Zebras: premises for pond and water; LLaVA and CoT answer pond, ViLA answers water.

Ablations

Accuracy stays at 66.41 for VLM weights 0.1 to 0.8 and jumps to 68.18 at wv = 0.9. Visual adequacy matters more than textual relevance when judging premises.

Accuracy is flat at about 66.4% for VLM weight 0.1 to 0.8 and rises to about 68.2% at 0.9.

On A-OKVQA validation, 3 premises with 1 candidate or 2 of each work best. Too many premises or candidates add ambiguous evidence and lower accuracy.

Accuracy heatmap over number of answer candidates (1 to 5) and premises (1 to 5); the best cells are 68.18 at 1 candidate and 3 premises and 68.03 at 2 and 2.

Premises from each model help in different ways. Using only the LLM's or only the VLM's premises gives 65.33; combining both gives 68.18, a gain of 2.85 points.

LLM premisesVLM premisesAcc. (%)
✓65.33
✓65.33
✓✓68.18

cite

BibTeX

@inproceedings{kim-etal-2026-visual,
  title     = {Visual{--}Linguistic Abductive Reasoning with {LLM}s for
               Knowledge-based Visual Question Answering},
  author    = {Kim, Jieun and Jeong, Yujin and Cho, Sung-Bae},
  booktitle = {Findings of the Association for Computational Linguistics:
               EACL 2026},
  year      = {2026},
  url       = {https://aclanthology.org/2026.findings-eacl.343/}
}