1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CL2025
PyVision: Agentic Vision with Dynamic Tooling
Shitian Zhao, Haoquan Zhang, Shaoheng Lin +4
LLMs are increasingly deployed as agents, systems capable of planning, reasoning, and dynamically calling external tools. However, in visual reasoning, prior approaches largely rem…
cs.CV2025
Improving Autoregressive Image Generation through Coarse-to-Fine Token Prediction
Ziyao Guo, Kaipeng Zhang, Michael Qizhe Shieh
Autoregressive models have shown remarkable success in image generation by adapting sequential prediction techniques from language modeling. However, applying these approaches to i…
cs.CV2024★ 1 cited
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
Fanqing Meng, Jin Wang, Chuanhao Li +9
The capability to process multiple images is crucial for Large Vision-Language Models (LVLMs) to develop a more thorough and nuanced understanding of a scene. Recent multi-image LV…