Skip to content

About

Visual Grounding Score Guided Decoding for Hallucination Mitigation in Medical VLMs

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

VGS Logo VGS-Decoding: Visual Grounding Score Guided Decoding for Hallucination Mitigation in Medical VLMs

Typing SVG

Project Page arXiv VQA-RAD License Visitors

🌐 Project Page Β |Β  πŸ“„ Paper Β |Β  πŸ’» Code

Govinda Kolli*, Adinath Madhavrao Dukre*, Yifan Lu, Ziyun Zou, Dwarikanath Mahapatra, Behzad Bozorgtabar, Imran Razzak

* Equal first authors

GenMI Lab

πŸ”₯ News

  • [21 Sep 2026] πŸš€ Inference code for LLaVA-Med and MedGemma on VQA-RAD, the project page and an interactive demo setup are released.
  • [19 Mar 2026] β›³ Our preprint is live on arXiv. Check it out for details.

Overview

Medical vision-language models (VLMs) can produce fluent answers that are not supported by the image, often because the language prior outweighs the visual evidence. We propose Visual Grounding Score (VGS) guided decoding, a training-free inference strategy that measures, for every candidate token, how much its probability depends on the image. At each step VGS compares the next-token distribution under the original image with the distribution under a Gaussian-plus-Poisson perturbed copy. Visually grounded tokens are amplified and tokens that remain likely without reliable visual evidence are suppressed, with no change to the base model.

VGS-Decoding Framework

framework

πŸ† Main Results

Table 1. Comparison of decoding methods across three medical VLM backbones and three medical VQA benchmarks. Open: token-level recall on open-ended questions; Closed: accuracy on closed-ended questions; Overall: question-count-weighted mixed score (all in %). Bold: best within each backbone and metric.

BackboneMethodVQA-RADSLAKEMIMIC-Diff-VQA
OpenClosedOverallOpenClosedOverallOpenClosedOverall
LLaVA-MedGreedy34.4568.9253.6440.8162.2549.2228.0448.3935.39
VCD30.8561.2047.7139.5060.5647.7625.9846.4233.33
DoLA32.7658.9647.3442.5461.9750.1628.7547.9435.68
OPERA33.2261.6949.0531.2558.5941.9721.1846.0230.14
VGS (Ours)38.9072.9157.7541.1175.2154.4830.8455.7939.86
CheXagentGreedy22.0270.9249.2444.1469.3054.0044.0682.0757.79
VCD21.7368.5347.7843.0166.2052.1038.8879.1453.42
DoLA20.7368.9247.5542.9569.0153.1739.7881.9455.02
OPERA20.5069.3247.6738.1969.3050.3936.1982.0352.75
VGS (Ours)23.4871.3150.1043.7570.1454.1043.9982.5357.91
MedGemmaGreedy49.5061.7556.3254.7473.5662.1225.9773.5543.16
VCD50.2957.7754.4554.0066.3558.8429.3868.7643.61
DoLA51.9172.5163.3858.7179.5766.8932.5676.8248.55
OPERA48.9065.7458.2758.6176.9265.7928.3575.5345.40
VGS (Ours)53.0573.5864.4854.2585.8266.6334.9582.9152.28

Open, closed and overall scores across decoding methods and backbones
Performance of Greedy, VCD, DoLA, OPERA and VGS across LLaVA-Med, CheXagent and MedGemma. See Table 1 for exact scores.

Qualitative comparison on an abdominal CT question
Qualitative example on abdominal CT: baselines name the wrong organ, while VGS answers small bowel.

Token-level analysis: mean VGS per token category (VQA-RAD)

Mean VGS per token category

πŸ“– Contents

⛏️ Installation

Note

Requirements: Python 3.10, PyTorch 2.7.1, Transformers 4.53.0 and a CUDA GPU. MedGemma additionally needs bfloat16 support. Our experiments ran on NVIDIA A100 40 GB GPUs.

  1. Clone the repository and navigate to the project folder
git clone https://github.com/genmilab/VGS-Decoding.git
cd VGS-Decoding
  1. Set up the environment and install in editable mode
conda create -n vgs python=3.10 -y
conda activate vgs
pip install torch==2.7.1 torchvision==0.22.1 --index-url https://download.pytorch.org/whl/cu128
pip install -e .

Tip

We recommend a separate environment per backbone (e.g. vgs-llavamed and vgs-medgemma). The CUDA 12.8 wheels need a compatible NVIDIA driver; see PyTorch previous versions for other builds.

πŸ”„ Upgrade to the latest code base
git pull
pip install -e .

🧩 Models and Weights

Model CLI name Checkpoint Precision
LLaVA-Med-v1.5 (7B) llava-med chaoyinshe/llava-med-v1.5-mistral-7b-hf float16
MedGemma (4B) medgemma google/medgemma-4b-it bfloat16

Weights are downloaded automatically to the Hugging Face cache (~/.cache/huggingface) on the first run. MedGemma is gated: accept its terms on Hugging Face, then log in once.

huggingface-cli login
πŸ“¦ Pre-download weights to a local folder (clusters / offline machines)
huggingface-cli download chaoyinshe/llava-med-v1.5-mistral-7b-hf \
  --revision 627be53734c667cbb1669608dac747a4485a22d7 \
  --local-dir checkpoints/llava-med-v1.5-mistral-7b-hf

huggingface-cli download google/medgemma-4b-it \
  --revision 290cda5eeccbee130f987c4ad74a59ae6f196408 \
  --local-dir checkpoints/medgemma-4b-it

Suggested layout:

VGS-Decoding/
β”œβ”€β”€ checkpoints/
β”‚   β”œβ”€β”€ llava-med-v1.5-mistral-7b-hf/
β”‚   └── medgemma-4b-it/
└── data/
    └── vqa-rad/        # VQA-RAD test split as an Arrow file

Then pass --model-path checkpoints/<name> (see Multi-GPU and Offline Runs).

Warning

The LLaVA-Med checkpoint is a Hugging Face-format conversion; weights from the original LLaVA repository will not load. Both revisions are pinned in vgs_decoding.py. Never put access tokens in source files, commands or commits.

⚑ Quick Start

CLI Inference

Run VGS decoding on VQA-RAD directly from the command line:

# LLaVA-Med + VGS
python vgs_llavamed_vqarad.py \
  --device cuda:0 --limit 3 \
  --output-dir outputs/llava-med-vgs

# MedGemma + VGS
python vgs_medgemma_vqarad.py \
  --device cuda:0 --limit 3 \
  --output-dir outputs/medgemma-vgs

Remove --limit to process the full test split. The installed console scripts vgs-llavamed-vqarad, vgs-medgemma-vqarad and vgs-vqarad --model <name> are equivalent.

Optional arguments:

Argument Description Default
--alpha Strength of VGS reweighting (0 = greedy) 1.0
--min-factor Lower bound on the multiplicative factor 0.01
--gaussian-std Gaussian noise std on RGB values in [0, 1] 0.07
--poisson-scale Poisson rate scale 70.0
--max-new-tokens Maximum generated tokens 64
--save-token-trace Save per-token probabilities and VGS values off

Baseline Inference

Set --alpha 0 to obtain the greedy baseline with the identical prompt, precision and stopping rule:

python vgs_llavamed_vqarad.py --alpha 0 --output-dir outputs/llava-med-greedy
python vgs_medgemma_vqarad.py --alpha 0 --output-dir outputs/medgemma-greedy

Note

VCD, DoLA and OPERA baselines, the CheXagent backbone, and the SLAKE and MIMIC-Diff-VQA pipelines reported in the paper are not part of this release. Please use the official implementations of those methods.

Script Inference

Answer a question about any image programmatically with the shared decoder in vgs_decoding.py:

import torch
from PIL import Image
from vgs_decoding import DecodingConfig, decode, load_model, perturb_image

bundle = load_model("medgemma", torch.device("cuda:0"))   # or "llava-med"
config = DecodingConfig(alpha=1.0)                         # alpha=0 -> greedy

image = Image.open("./path/to/image.png").convert("RGB")
question = "Is there evidence of pneumothorax?"
distorted = perturb_image(image, config, seed=config.seed)

result = decode(
    bundle.model,
    bundle.prepare(image, question),
    bundle.prepare(distorted, question),
    config,
    bundle.stop_token_ids(),
)
print(bundle.processor.tokenizer.decode(result.token_ids[0], skip_special_tokens=True).strip())
πŸ’‘ Inspect the token-level VGS trace.
for step in result.steps:
    print(step)   # chosen token, clean probability and VGS value

πŸ‘‰ Only batch size 1 and the two listed backbones are supported. Other architectures need a new adapter in vgs_decoding.py.

Gradio Web Interface

Build and launch the demo locally with Docker:

docker build -f hf_demo/Dockerfile -t vgs-decoding-demo .
docker run --rm --gpus all -p 127.0.0.1:7860:7860 -e HF_TOKEN vgs-decoding-demo

Open http://127.0.0.1:7860, choose MedGemma or LLaVA-Med, load a VQA-RAD example or upload a de-identified image, adjust alpha, and inspect the answer with its token-level VGS trace. See hf_demo/README.md for a non-Docker setup.

πŸ› οΈ Advanced Usage

Parameter Settings

  • alpha (β‰₯ 0): Strength of VGS reweighting

    • 0: plain greedy decoding
    • Higher = stronger promotion of visually grounded tokens
    • Default: 1.0
  • gaussian-std / poisson-scale: Perturbation applied to the image copy

    • Defaults: 0.07 / 70.0
  • min-factor (0, 1]: Floor on the reweighting factor, so no token is removed entirely

    • Default: 0.01

The rule applied at every step, with $p$ and $q$ the clean and perturbed next-token distributions:

$$ \mathrm{VGS}(y) = \frac{p(y) - q(y)}{p(y) + q(y) + \epsilon}, \qquad P_{\text{final}}(y) \propto p(y)\cdot\max\bigl(1 + \alpha,\mathrm{VGS}(y),\ m\bigr) $$

Full details are in docs/decoding.md.

Multi-GPU and Offline Runs

Each command loads one model on one device. Split the dataset across GPUs with --start / --limit; noise is seeded by seed + dataset_index, so partitions match a single run.

python vgs_medgemma_vqarad.py --device cuda:0 --start 0   --limit 226 --output-dir outputs/mg-part0 &
python vgs_medgemma_vqarad.py --device cuda:1 --start 226             --output-dir outputs/mg-part1 &

Run from local weights and data without network access:

HF_HUB_OFFLINE=1 python vgs_medgemma_vqarad.py \
  --model-path checkpoints/medgemma-4b-it \
  --dataset-arrow data/vqa-rad/vqa-rad-test.arrow \
  --local-files-only --output-dir outputs/medgemma-offline

Outputs

outputs/<run-name>/
β”œβ”€β”€ predictions.jsonl   # question ID, question, answer, model, seed, token count, stop reason
└── run.json            # settings, checkpoint revision, dataset fingerprint, versions, status

Tip

Each output directory must be new; existing runs are never overwritten. To resume after a failure, use a new directory with --start at the next unfinished index and the same seed.

πŸ—‚οΈ Dataset

  • VQA-RAD: radiology visual question answering (test split by default). Downloaded automatically.

The paper additionally evaluates on SLAKE and MIMIC-Diff-VQA; runners for these are not included in this release.

Note

Only the image, question and question ID are read. Reference answers are never passed to the model or saved. Models and data remain subject to their own licenses.

πŸ“Š Evaluation

This release generates answers only; no accuracy, recall or hallucination metrics are computed. The paper reports open-question recall, closed-question accuracy and a question-weighted overall score. Complete tables for all backbones and datasets are on the project page.

πŸ“ Citation

If you find our paper and code useful in your research, please cite using this BibTeX:

@misc{kolli2026vgs,
  title         = {VGS-Decoding: Visual Grounding Score Guided Decoding for
                   Hallucination Mitigation in Medical VLMs},
  author        = {Govinda Kolli and Adinath Madhavrao Dukre and
                   Behzad Bozorgtabar and Dwarikanath Mahapatra and Imran Razzak},
  year          = {2026},
  eprint        = {2603.20314},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  doi           = {10.48550/arXiv.2603.20314},
  url           = {https://arxiv.org/abs/2603.20314}
}

πŸ“š Acknowledgments

This project builds upon the following open-source works:

We thank the authors for their valuable contributions to the medical AI community.

πŸ“¨ Contact

For questions or collaboration, please open an issue or reach out to the code contributors: Govinda Kolli and Adinath Madhavrao Dukre.

πŸ“œ License

This project is licensed under the MIT License. See the LICENSE file for details.

🧰 Intended Use

VGS-Decoding is intended for research on hallucination and visual grounding in medical vision-language models.

Key Applications

  • πŸ”¬ Research Utility: study how strongly generated tokens depend on visual evidence, and compare decoding strategies.
  • πŸ§ͺ Model Analysis: inspect token-level VGS traces to find answers that rely on language priors.
  • πŸŽ“ Education: illustrate grounding and hallucination behaviour of medical VLMs.

Important

VGS is a perturbation-sensitivity signal, not a clinical confidence score. Outputs must not inform any clinical decision.


Limitations and Recommendations
  1. Perturbation Choice: the Gaussian-plus-Poisson perturbation is empirical and not a validated acquisition model for every imaging modality.
  2. Scope: experiments concern research benchmarks with short answers; performance on other tasks or modalities is untested.
  3. Compute: each generated token requires two forward passes (clean and perturbed image).
  4. Reproducibility: identical seeds do not guarantee bitwise identity across GPUs, package versions or model revisions.
Ethical Considerations
  • Patient Privacy: all input images must be fully de-identified and compliant with HIPAA, GDPR or equivalent local regulations.
  • Responsible Use: outputs may contain inaccuracies and should be interpreted with caution.
  • Accountability: responsibility for verification lies with the end user.
Disclaimer

This code is intended solely for research and educational purposes. It is not approved by the FDA, CE or any other regulatory authority for clinical use. For medical diagnosis or treatment, please consult a licensed healthcare professional.

About

Visual Grounding Score Guided Decoding for Hallucination Mitigation in Medical VLMs

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages