ICCV 2025

FuzzyCD fuzzy contrastive decoding

Fuzzy Contrastive Decoding to Alleviate Object Hallucination in Large Vision-Language Models

Jieun Kim·Jinmyeong Kim·Yoonji Kim·Sung-Bae Cho

Yonsei University

FuzzyCD overview: original and filter-augmented images pass through the LVLM; top-1 log-probabilities are fuzzified, fuzzy rules produce rule weights, and amplified logits are used in contrastive decoding.
Phase I reads the model's confidence on the original and filter-augmented images through Takagi-Sugeno fuzzy rules. Phase II uses the rule-weighted amplified logits in contrastive decoding.

motivation

The distorted input does not always hallucinate

Contrastive decoding subtracts the logits of a distorted input x′ from the original logits. This assumes x′ amplifies hallucination. When it does not, the subtraction creates new hallucinations instead of removing them.

In this example, a filtered image makes the model more confident in the correct color. Subtracting its logits pushes a wrong color to the top. FuzzyCD first checks how confident the model is on each input, then decides how much to subtract.

Q: What color is this?

Logits from Figure 5 of the paper. Step 3 applies Eq. 2 with α = 1.

abstract

Fuzzy rules decide how much to contrast

Large vision-language models (LVLMs) often exhibit object hallucination, a phenomenon where models generate descriptions of non-existent objects within images. Prior methods have sought to mitigate this issue by adjusting model logits to reduce linguistic bias, but they often lack precise control over visual uncertainty, sometimes exacerbating hallucinations instead of mitigating them. To address this limitation, we propose a novel decoding strategy called fuzzy contrastive decoding (FuzzyCD) that uses Takagi-Sugeno fuzzy inference to refine hallucination control. FuzzyCD adaptively assigns weights to high-hallucination logits while mitigating unnecessary linguistic bias. Specifically, it transforms the log-probabilities of top-1 tokens from both standard and hallucination logits into a confidence linguistic fuzzy set. Through Takagi-Sugeno fuzzy inference, it dynamically adjusts hallucination logits to prevent the model from over-relying on spurious linguistic patterns. Experimental results on object hallucination datasets demonstrate that hallucination is mitigated by 11%p compared to conventional LVLMs. In-depth analyses highlight the effectiveness of FuzzyCD in enhancing the reliability of vision-language models.

method

From confidence to contrast, step by step

01 Input
02 Fuzzify
03 Rule weights
04 Rule outputs
05 Aggregate
06 Contrast

Why log-probability

To choose the confidence score, a simple classifier trained on VizWiz detects hallucination in LLaVA-1.5 from five logit statistics. Probability and log-probability reach the highest AUC in every setting.

VizWizPOPE
Criterionvalrandompopularadversarial
Logits-mean0.580.240.260.31
Logits-min0.560.250.270.31
Logits-max0.700.690.720.72
Probability0.750.790.780.76
Log-probability0.750.790.780.76

Hallucination detection AUC, LLaVA-1.5. Highlighted: the score FuzzyCD uses.

Density plots of logits mean, logits max, probability and log-probability for correct and incorrect answers.

results

Best accuracy in every POPE setting

Evaluated on POPE, MME and ROPE against VCD, ICD, VDD and OPERA, with LLaVA-1.5 and InstructBLIP (Vicuna-7B).

6 / 6
POPE settings with the highest accuracy
+5.43
avg. POPE accuracy gain, LLaVA-1.5
641.66
MME total, highest among methods

POPE (MSCOCO)

LLaVA-1.5InstructBLIP
MethodAcc.Prec.Rec.F1Acc.Prec.Rec.F1

Bold: best per column. Highlighted: FuzzyCD.

MME, LLaVA-1.5

Largest gains on object-level tasks: Existence +19.33, Count +56.66.

ObjectAttribute
MethodExistenceCountPositionColorTotal
Default175.67106.67114.00160.00556.34
+ VCD184.66138.33128.67153.00604.66
+ VDD190.00138.33126.67165.00620.00
+ FuzzyCD195.00163.33123.33160.00641.66

ROPE, LLaVA-1.5

Best on all six splits. Single-object Hom. 31.88 → 57.88.

Multi-objectSingle-object
MethodWildHom.Het.WildHom.Het.
LLaVA-1.513.9631.883.9813.9631.883.98
+ VCD13.5728.686.3724.7648.339.84
+ ICD14.3532.635.6922.9245.4710.24
+ VCC17.8040.777.4027.4952.1211.06
+ OPERA13.2037.143.8213.2037.143.82
+ FuzzyCD21.1246.447.4529.0157.8811.30

Across model families and scales

POPE accuracy improves on all five LVLMs from 7B to 72B, most on 7B to 13B models.

Accuracy of base decoding and FuzzyCD across five LVLMs: LLaVA 1.5 7B +5.12, InstructBLIP 7B +5.41, LLaVA-NeXT 13B +4.46, QwenVL 32B +0.72, QwenVL 72B +1.05.

Choice of image filter

POPE accuracy stays within 86.5 to 89.3 across five filters. Sobel is best.

FilterAcc. (%)
Bilateral86.86
Median86.53
Sobel89.33
Sharpen (default)88.76
Gaussian87.83

Qualitative example

Asked “What is this a picture of?”, the model is confident on both the original and the filtered image (High 0.99 and 0.91). Rule 9 gets the highest weight, which signals no hallucination, so the filtered logits are not amplified and FuzzyCD answers sugar. Standard contrastive decoding suppresses Sugar and Domino, the top candidates on both inputs, and answers k.

Rule weights for the Domino Sugar example: Rule 9 is highest; FuzzyCD multiplies the amplified logits by zero and answers sugar, while visual contrastive decoding answers k.

cite

BibTeX

@inproceedings{kim2025fuzzycd,
  title     = {Fuzzy Contrastive Decoding to Alleviate Object Hallucination
               in Large Vision-Language Models},
  author    = {Kim, Jieun and Kim, Jinmyeong and Kim, Yoonji and Cho, Sung-Bae},
  booktitle = {Proceedings of the IEEE/CVF International Conference
               on Computer Vision (ICCV)},
  pages     = {20572--20581},
  year      = {2025},
  doi       = {10.1109/ICCV51701.2025.01913}
}