Stories about Qwen2.5-VL-72B
1 related stories
Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects
AI InsightThis study systematically tests two multimodal LLMs as reviewers on ICLR 2026 submissions, finding sensitivity to author identity and figures, with limited error detection. Compared to prior work on LLM-generated reviews, it reveals specific limitations in identity bias and error spotting.Key TakeawayShift from evaluating LLM-generated reviews to auditing their critical review capacity.Why It MattersLLM reviewers may introduce systematic bias and miss errors, directly affecting review quality and fairness.Who's Affected- AI ResearchersGet empirical data on LLM review bias and error detection, shaping future research.
- Academic PublishersNeed caution with LLM-assisted review and consider bias calibration.
- Paper AuthorsAuthor identity or figure presentation may affect LLM review outcomes.
What's NextWatch for debiasing methods for LLM reviewers and adoption by real conferences.Importance 82/100