activity
20212026
most citedGRiT: A Generative Region-to-text Transformer for Object Understanding

30 citations · 37 across the 23 of their papers we have counts for

collaborators

47 papers

cs.CV2026

Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation

Tianyi Xiong, Zhengyuan Yang, Xiaofei Wang +10

Large vision-language models have shown strong progress in UI-to-code generation, yet their test-time self-evolution remains unstable. We first identify a fundamental obstacle, ter…

cs.CV2026

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors

Yilin Wang, Xiangxi Zheng, Dongxing Mao +6

Understanding long videos with multimodal large language models (MLLMs) requires selecting a compact set of frames from thousands of candidates, yet identifying the right frames se…

cs.AI2026

Planning with the Views

Kangrui Wang, Linjie Li, Zhengyuan Yang +7

Can VLMs predict how each camera move changes the view, and plan many such moves ahead? We call this capability view planning, requiring (1)understanding how a single action transf…

cs.CL2026

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement

Jiajie Jin, Yuyang Hu, Kai Qiu +15

Scientific progress depends on a repeated loop of exploration, experimentation, and abstraction. Researchers test candidate directions, interpret the evidence, and carry the result…

cs.CV2026

3D-CoS: A New 3D Reconstruction Paradigm Based on VLM Code Synthesis

Yuhao Wang, Puyi Wang, Linjie Li +3

Most recent 3D reconstruction and editing systems operate on implicit and explicit representations such as NeRF, point clouds, or meshes. While these representations enable high-fi…

cs.CV2026

TextGround4M: A Prompt-Aligned Dataset for Layout-Aware Text Rendering

Dongxing Mao, Yilin Wang, Linjie Li +2

Despite recent advances in text-to-image generation, models still struggle to accurately render prompt-specified text with correct spatial layout -- especially in multi-span, struc…