Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation
Zekai Zhang, Jiahao Li, Jie Zhang +18
While text-to-image (T2I) models have achieved remarkable progress, they struggle with real-world requests that are often underspecified, implicit, or dependent on up-to-date knowl…
cs.CV2026
Gold Points Sniper: Self-guided Visual Reasoning in VLM for Fine-grained Action Understanding
Haodi Liu, Xinhang Yang, Kunda Yan +3
Robots operating in everyday environments must understand fine-grained human actions, intentions, and contextual cues from broad views where people occupy only small regions, a cap…