46 papers
Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors
Yilin Wang, Xiangxi Zheng, Dongxing Mao +6
Understanding long videos with multimodal large language models (MLLMs) requires selecting a compact set of frames from thousands of candidates, yet identifying the right frames se…
Planning with the Views
Kangrui Wang, Linjie Li, Zhengyuan Yang +7
Can VLMs predict how each camera move changes the view, and plan many such moves ahead? We call this capability view planning, requiring (1)understanding how a single action transf…
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
Jiajie Jin, Yuyang Hu, Kai Qiu +15
Scientific progress depends on a repeated loop of exploration, experimentation, and abstraction. Researchers test candidate directions, interpret the evidence, and carry the result…
3D-CoS: A New 3D Reconstruction Paradigm Based on VLM Code Synthesis
Yuhao Wang, Puyi Wang, Linjie Li +3
Most recent 3D reconstruction and editing systems operate on implicit and explicit representations such as NeRF, point clouds, or meshes. While these representations enable high-fi…
Residual Decoder Adapter: ID-Preserving Tokenizer Adaption for Autoregressive Text Rendering
Dongxing Mao, Jinpeng Wang, Jiahao Tang +6
Visual Autoregressive (AR) models generate images by predicting discrete tokens that are decoded by a visual tokenizer. Despite demonstrating strong overall image generation abilit…
SceneCode: Executable World Programs for Editable Indoor Scenes with Articulated Objects
Puyi Wang, Yuhao Wang, Linjie Li +4
Indoor scene synthesis underpins embodied AI, robotic manipulation, and simulation-based policy evaluation, where a useful scene must specify not only what the environment looks li…