activity
20242026
most citedGSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts

1 citations · 3 across the 40 of their papers we have counts for

collaborators

43 papers

cs.CL2026

PaperGym: Rubric-Centered Evolution for Research-Plan Generation

Yuhan Wang, Zhengxi Lu, Yuchen Yan +6

Research planning is the decisive capability of AI scientists. Yet a research plan admits no verifiable answer, so reinforcement learning lacks the environment it requires: tasks p…

cs.CL2026

TTPO: Test-Time Policy Optimization

Aozhe Wang, Zhengxi Lu, Jianze Wang +8

Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large l…

cs.CL2026

BrowserForge: Scaling Web Episode via Parallel Browser Sandboxes

Fei Tang, Huawen Shen, Zhiqiong Lu +7

Web agents that act from rendered pixels avoid the fragility and heavy token cost of reading a page's HTML or accessibility tree, but training them depends on large amounts of high…

cs.CV2026

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning

Hongxing Li, Xiufeng Huang, Dingming Li +11

Fine-grained visual reasoning remains challenging for vision-language models, especially when small but critical visual cues are buried in high-resolution images. Existing approach…

cs.CV2026

InstructSAM: Segment Any Instance with Any Instructions

Yuqian Yuan, Wentong Li, Zhaocheng Li +6

In this paper, we introduce InstructSAM, a unified and streamlined framework designed for multi-instance segmentation under arbitrary instructions. We formulates instruction-driven…

cs.CV2026

CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark

Wei Wang, Yuqian Yuan, Tianwei Lin +4

Spatial intelligence requires multimodal large language models (MLLMs) to move beyond single-view perception and reason consistently about objects, visibility, geometry, and intera…