π Project Page Β |Β π Paper Β |Β π» Code
Govinda Kolli*, Adinath Madhavrao Dukre*, Yifan Lu, Ziyun Zou, Dwarikanath Mahapatra, Behzad Bozorgtabar, Imran Razzak
* Equal first authors
- [21 Sep 2026] π Inference code for LLaVA-Med and MedGemma on VQA-RAD, the project page and an interactive demo setup are released.
- [19 Mar 2026] β³ Our preprint is live on arXiv. Check it out for details.
Medical vision-language models (VLMs) can produce fluent answers that are not supported by the image, often because the language prior outweighs the visual evidence. We propose Visual Grounding Score (VGS) guided decoding, a training-free inference strategy that measures, for every candidate token, how much its probability depends on the image. At each step VGS compares the next-token distribution under the original image with the distribution under a Gaussian-plus-Poisson perturbed copy. Visually grounded tokens are amplified and tokens that remain likely without reliable visual evidence are suppressed, with no change to the base model.
Table 1. Comparison of decoding methods across three medical VLM backbones and three medical VQA benchmarks. Open: token-level recall on open-ended questions; Closed: accuracy on closed-ended questions; Overall: question-count-weighted mixed score (all in %). Bold: best within each backbone and metric.
| Backbone | Method | VQA-RAD | SLAKE | MIMIC-Diff-VQA | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Open | Closed | Overall | Open | Closed | Overall | Open | Closed | Overall | ||
| LLaVA-Med | Greedy | 34.45 | 68.92 | 53.64 | 40.81 | 62.25 | 49.22 | 28.04 | 48.39 | 35.39 |
| VCD | 30.85 | 61.20 | 47.71 | 39.50 | 60.56 | 47.76 | 25.98 | 46.42 | 33.33 | |
| DoLA | 32.76 | 58.96 | 47.34 | 42.54 | 61.97 | 50.16 | 28.75 | 47.94 | 35.68 | |
| OPERA | 33.22 | 61.69 | 49.05 | 31.25 | 58.59 | 41.97 | 21.18 | 46.02 | 30.14 | |
| VGS (Ours) | 38.90 | 72.91 | 57.75 | 41.11 | 75.21 | 54.48 | 30.84 | 55.79 | 39.86 | |
| CheXagent | Greedy | 22.02 | 70.92 | 49.24 | 44.14 | 69.30 | 54.00 | 44.06 | 82.07 | 57.79 |
| VCD | 21.73 | 68.53 | 47.78 | 43.01 | 66.20 | 52.10 | 38.88 | 79.14 | 53.42 | |
| DoLA | 20.73 | 68.92 | 47.55 | 42.95 | 69.01 | 53.17 | 39.78 | 81.94 | 55.02 | |
| OPERA | 20.50 | 69.32 | 47.67 | 38.19 | 69.30 | 50.39 | 36.19 | 82.03 | 52.75 | |
| VGS (Ours) | 23.48 | 71.31 | 50.10 | 43.75 | 70.14 | 54.10 | 43.99 | 82.53 | 57.91 | |
| MedGemma | Greedy | 49.50 | 61.75 | 56.32 | 54.74 | 73.56 | 62.12 | 25.97 | 73.55 | 43.16 |
| VCD | 50.29 | 57.77 | 54.45 | 54.00 | 66.35 | 58.84 | 29.38 | 68.76 | 43.61 | |
| DoLA | 51.91 | 72.51 | 63.38 | 58.71 | 79.57 | 66.89 | 32.56 | 76.82 | 48.55 | |
| OPERA | 48.90 | 65.74 | 58.27 | 58.61 | 76.92 | 65.79 | 28.35 | 75.53 | 45.40 | |
| VGS (Ours) | 53.05 | 73.58 | 64.48 | 54.25 | 85.82 | 66.63 | 34.95 | 82.91 | 52.28 | |
Performance of Greedy, VCD, DoLA, OPERA and VGS across LLaVA-Med, CheXagent and MedGemma. See Table 1 for exact scores.
Qualitative example on abdominal CT: baselines name the wrong organ, while VGS answers small bowel.
- π Main Results
- βοΈ Installation
- π§© Models and Weights
- β‘ Quick Start
- π οΈ Advanced Usage
- ποΈ Dataset
- π Evaluation
- π Citation
- π Acknowledgments
- π¨ Contact
- π License
- π§° Intended Use
Note
Requirements: Python 3.10, PyTorch 2.7.1, Transformers 4.53.0 and a CUDA GPU. MedGemma additionally needs bfloat16 support. Our experiments ran on NVIDIA A100 40 GB GPUs.
- Clone the repository and navigate to the project folder
git clone https://github.com/genmilab/VGS-Decoding.git
cd VGS-Decoding- Set up the environment and install in editable mode
conda create -n vgs python=3.10 -y
conda activate vgs
pip install torch==2.7.1 torchvision==0.22.1 --index-url https://download.pytorch.org/whl/cu128
pip install -e .Tip
We recommend a separate environment per backbone (e.g. vgs-llavamed and vgs-medgemma). The CUDA 12.8 wheels need a compatible NVIDIA driver; see PyTorch previous versions for other builds.
π Upgrade to the latest code base
git pull
pip install -e .| Model | CLI name | Checkpoint | Precision |
|---|---|---|---|
| LLaVA-Med-v1.5 (7B) | llava-med |
chaoyinshe/llava-med-v1.5-mistral-7b-hf | float16 |
| MedGemma (4B) | medgemma |
google/medgemma-4b-it | bfloat16 |
Weights are downloaded automatically to the Hugging Face cache (~/.cache/huggingface) on the first run. MedGemma is gated: accept its terms on Hugging Face, then log in once.
huggingface-cli loginπ¦ Pre-download weights to a local folder (clusters / offline machines)
huggingface-cli download chaoyinshe/llava-med-v1.5-mistral-7b-hf \
--revision 627be53734c667cbb1669608dac747a4485a22d7 \
--local-dir checkpoints/llava-med-v1.5-mistral-7b-hf
huggingface-cli download google/medgemma-4b-it \
--revision 290cda5eeccbee130f987c4ad74a59ae6f196408 \
--local-dir checkpoints/medgemma-4b-itSuggested layout:
VGS-Decoding/
βββ checkpoints/
β βββ llava-med-v1.5-mistral-7b-hf/
β βββ medgemma-4b-it/
βββ data/
βββ vqa-rad/ # VQA-RAD test split as an Arrow file
Then pass --model-path checkpoints/<name> (see Multi-GPU and Offline Runs).
Warning
The LLaVA-Med checkpoint is a Hugging Face-format conversion; weights from the original LLaVA repository will not load. Both revisions are pinned in vgs_decoding.py. Never put access tokens in source files, commands or commits.
Run VGS decoding on VQA-RAD directly from the command line:
# LLaVA-Med + VGS
python vgs_llavamed_vqarad.py \
--device cuda:0 --limit 3 \
--output-dir outputs/llava-med-vgs
# MedGemma + VGS
python vgs_medgemma_vqarad.py \
--device cuda:0 --limit 3 \
--output-dir outputs/medgemma-vgsRemove --limit to process the full test split. The installed console scripts vgs-llavamed-vqarad, vgs-medgemma-vqarad and vgs-vqarad --model <name> are equivalent.
Optional arguments:
| Argument | Description | Default |
|---|---|---|
--alpha |
Strength of VGS reweighting (0 = greedy) |
1.0 |
--min-factor |
Lower bound on the multiplicative factor | 0.01 |
--gaussian-std |
Gaussian noise std on RGB values in [0, 1] | 0.07 |
--poisson-scale |
Poisson rate scale | 70.0 |
--max-new-tokens |
Maximum generated tokens | 64 |
--save-token-trace |
Save per-token probabilities and VGS values | off |
Set --alpha 0 to obtain the greedy baseline with the identical prompt, precision and stopping rule:
python vgs_llavamed_vqarad.py --alpha 0 --output-dir outputs/llava-med-greedy
python vgs_medgemma_vqarad.py --alpha 0 --output-dir outputs/medgemma-greedyNote
VCD, DoLA and OPERA baselines, the CheXagent backbone, and the SLAKE and MIMIC-Diff-VQA pipelines reported in the paper are not part of this release. Please use the official implementations of those methods.
Answer a question about any image programmatically with the shared decoder in vgs_decoding.py:
import torch
from PIL import Image
from vgs_decoding import DecodingConfig, decode, load_model, perturb_image
bundle = load_model("medgemma", torch.device("cuda:0")) # or "llava-med"
config = DecodingConfig(alpha=1.0) # alpha=0 -> greedy
image = Image.open("./path/to/image.png").convert("RGB")
question = "Is there evidence of pneumothorax?"
distorted = perturb_image(image, config, seed=config.seed)
result = decode(
bundle.model,
bundle.prepare(image, question),
bundle.prepare(distorted, question),
config,
bundle.stop_token_ids(),
)
print(bundle.processor.tokenizer.decode(result.token_ids[0], skip_special_tokens=True).strip())π‘ Inspect the token-level VGS trace.
for step in result.steps:
print(step) # chosen token, clean probability and VGS valueπ Only batch size 1 and the two listed backbones are supported. Other architectures need a new adapter in
vgs_decoding.py.
Build and launch the demo locally with Docker:
docker build -f hf_demo/Dockerfile -t vgs-decoding-demo .
docker run --rm --gpus all -p 127.0.0.1:7860:7860 -e HF_TOKEN vgs-decoding-demoOpen http://127.0.0.1:7860, choose MedGemma or LLaVA-Med, load a VQA-RAD example or upload a de-identified image, adjust alpha, and inspect the answer with its token-level VGS trace. See hf_demo/README.md for a non-Docker setup.
-
alpha(β₯ 0): Strength of VGS reweighting0: plain greedy decoding- Higher = stronger promotion of visually grounded tokens
- Default: 1.0
-
gaussian-std/poisson-scale: Perturbation applied to the image copy- Defaults: 0.07 / 70.0
-
min-factor(0, 1]: Floor on the reweighting factor, so no token is removed entirely- Default: 0.01
The rule applied at every step, with
Full details are in docs/decoding.md.
Each command loads one model on one device. Split the dataset across GPUs with --start / --limit; noise is seeded by seed + dataset_index, so partitions match a single run.
python vgs_medgemma_vqarad.py --device cuda:0 --start 0 --limit 226 --output-dir outputs/mg-part0 &
python vgs_medgemma_vqarad.py --device cuda:1 --start 226 --output-dir outputs/mg-part1 &Run from local weights and data without network access:
HF_HUB_OFFLINE=1 python vgs_medgemma_vqarad.py \
--model-path checkpoints/medgemma-4b-it \
--dataset-arrow data/vqa-rad/vqa-rad-test.arrow \
--local-files-only --output-dir outputs/medgemma-offlineoutputs/<run-name>/
βββ predictions.jsonl # question ID, question, answer, model, seed, token count, stop reason
βββ run.json # settings, checkpoint revision, dataset fingerprint, versions, status
Tip
Each output directory must be new; existing runs are never overwritten. To resume after a failure, use a new directory with --start at the next unfinished index and the same seed.
- VQA-RAD: radiology visual question answering (test split by default). Downloaded automatically.
The paper additionally evaluates on SLAKE and MIMIC-Diff-VQA; runners for these are not included in this release.
Note
Only the image, question and question ID are read. Reference answers are never passed to the model or saved. Models and data remain subject to their own licenses.
This release generates answers only; no accuracy, recall or hallucination metrics are computed. The paper reports open-question recall, closed-question accuracy and a question-weighted overall score. Complete tables for all backbones and datasets are on the project page.
If you find our paper and code useful in your research, please cite using this BibTeX:
@misc{kolli2026vgs,
title = {VGS-Decoding: Visual Grounding Score Guided Decoding for
Hallucination Mitigation in Medical VLMs},
author = {Govinda Kolli and Adinath Madhavrao Dukre and
Behzad Bozorgtabar and Dwarikanath Mahapatra and Imran Razzak},
year = {2026},
eprint = {2603.20314},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
doi = {10.48550/arXiv.2603.20314},
url = {https://arxiv.org/abs/2603.20314}
}This project builds upon the following open-source works:
- LLaVA-Med: a biomedical vision-language assistant.
- MedGemma: medical vision-language models from Google.
- VQA-RAD: a radiology visual question answering dataset.
- Hugging Face Transformers: model loading and generation utilities.
We thank the authors for their valuable contributions to the medical AI community.
For questions or collaboration, please open an issue or reach out to the code contributors: Govinda Kolli and Adinath Madhavrao Dukre.
This project is licensed under the MIT License. See the LICENSE file for details.
VGS-Decoding is intended for research on hallucination and visual grounding in medical vision-language models.
- π¬ Research Utility: study how strongly generated tokens depend on visual evidence, and compare decoding strategies.
- π§ͺ Model Analysis: inspect token-level VGS traces to find answers that rely on language priors.
- π Education: illustrate grounding and hallucination behaviour of medical VLMs.
Important
VGS is a perturbation-sensitivity signal, not a clinical confidence score. Outputs must not inform any clinical decision.
Limitations and Recommendations
- Perturbation Choice: the Gaussian-plus-Poisson perturbation is empirical and not a validated acquisition model for every imaging modality.
- Scope: experiments concern research benchmarks with short answers; performance on other tasks or modalities is untested.
- Compute: each generated token requires two forward passes (clean and perturbed image).
- Reproducibility: identical seeds do not guarantee bitwise identity across GPUs, package versions or model revisions.
Ethical Considerations
- Patient Privacy: all input images must be fully de-identified and compliant with HIPAA, GDPR or equivalent local regulations.
- Responsible Use: outputs may contain inaccuracies and should be interpreted with caution.
- Accountability: responsibility for verification lies with the end user.
Disclaimer
This code is intended solely for research and educational purposes. It is not approved by the FDA, CE or any other regulatory authority for clinical use. For medical diagnosis or treatment, please consult a licensed healthcare professional.

