collaborators

7 papers

cs.CV2026

Self-Evolving Code-with-Image Reasoning

Tianze Yang, Liang Wu, Ruitong Sun +6

Multimodal models increasingly reach for tools when solving visual tasks (crop, zoom, rotate, brighten), a paradigm known as thinking-with-images. The central challenge is one of p…

cs.AI2026

Trust the Right Teacher: Quality-Aware Self-Distillation for GUI Grounding

Jingyuan Huang, Zuming Huang, Yucheng Shi +4

Graphical user interface (GUI) grounding requires vision-language models (VLMs) to identify small target elements in high-resolution screenshots and predict precise screen coordina…

cs.CV2026

RSTR: Reducing SpatioTemporal Redundancy in Diffusion Transformers

Ruitong Sun, Tianze Yang, Wei Niu +1

Diffusion Transformers (DiTs) have achieved remarkable success in image generation, yet their deployment is hindered by high computational costs. We identify two sources of redunda…

cs.CV2026

Self-Improving Small Object Grounding in LVLMs

Tianze Yang, Yucheng Shi, Ruitong Sun +2

Can internal attention patterns in Large Vision Language Models (LVLMs) identify reliable small-object boxes without fine-tuning? In this work, we provide an affirmative answer. At…

cs.AI2026

TRON: Targeted Rule-Verifiable Online Environments for Visual Reasoning RL

Tianze Yang, Yucheng Shi, Ruitong Sun +3

Reinforcement learning (RL) for visual reasoning needs scalable, verifiable, and controllable training signals. Existing visual RL post-training trains on static curated datasets,…

cs.CV2026

Common Inpainted Objects In-N-Out of Context

Tianze Yang, Tyson Jordan, Ruitong Sun +2

We present Common Inpainted Objects In-N-Out of Context (COinCO), a novel dataset addressing the scarcity of out-of-context examples in existing vision datasets. By systematically…