2 papers
cs.CV2026
Learning to Look Again: Loss-Gap Supervision for Free-form Crop Routing in Vision-Language Models
Jinchang Zhu, Rong Fu, Yi Ding +3
Vision-language models (VLMs) fail many detail-centric questions for a concrete reason: the answer is visible in the image, yet lost after the image is compressed into a low-resolu…
cs.AI2026
Where vs What: Decomposing Structural and Content Failures in LLM-Generated Structured Outputs
Yiwei Zhang, Chengke Wu, Li Wang +1
Structured outputs such as JSON and tables are central to modern LLM-based systems, yet generation failures are evaluated monolithically, conflating two distinct error modes: place…