1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CV2025
START: Spatial and Textual Learning for Chart Understanding
Zhuoming Liu, Xiaofeng Gao, Feiyang Niu +3
Chart understanding is crucial for deploying multimodal large language models (MLLMs) in real-world scenarios such as analyzing scientific papers and technical reports. Unlike natu…
cs.CV2025
ProxT2I: Efficient Reward-Guided Text-to-Image Generation via Proximal Diffusion
Zhenghan Fang, Jian Zheng, Qiaozi Gao +2
Diffusion models have emerged as a dominant paradigm for generative modeling across a wide range of domains, including prompt-conditional generation. The vast majority of samplers,…
cs.AI2024★ 1 cited
TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft
Qian Long, Zhi Li, Ran Gong +3
Collaboration is a cornerstone of society. In the real world, human teammates make use of multi-sensory data to tackle challenging tasks in ever-changing environments. It is essent…