activity
20232026
most citedA Survey of Reinforcement Learning for Large Reasoning Models

2 citations · 7 across the 32 of their papers we have counts for

collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV2026

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

Yuandong Pu, Le Zhuo, Sayak Paul +11

Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not…

cs.CV2026

WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations

Xiaojie Xu, Zhengyuan Lin, Runyi Li +3

Learning to generate or reconstruct explorable worlds requires video paired with more than RGB: camera motion, scene geometry, temporal correspondence and, for interactive models,…

cs.CV2026

Towards Physics-Faithful Generation of Scientific Diagrams

Minghui Zhang, Jinxin Shi, Yifan Chang +12

Text-to-image generation has reached photorealistic quality, yet state-of-the-art systems remain unreliable at producing scientific diagrams, whose value depends not on appearance…

cs.CV2026

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

Siyu Yan, Zhuoran Yan, Haiying Xu +10

Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these vi…

cs.CV2026

Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning

Jiayi Lei, Yuandong Pu, Xingyu Han +8

Text-to-image (T2I) generation models have achieved remarkable progress in producing visually realistic images from natural language prompts. Yet it remains unclear whether their s…

cs.CV2026

Faithful, Enriched, and Precise: Benchmarking Natural-Science Illustration Generation by T2I models

Yifan Chang, Jiaxin Ai, Jianwen Sun +9

Scientific illustrations are essential tools for communicating research findings, especially in natural science, where they visualize complex concepts and processes. As Text-to-Ima…