2 papers
cs.CV2026
EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal
Feier Wu, Wanke Xia, Xu He +8
Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing meth…
cs.CV2026
VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents
Zhengbo Zhang, Jinbo Su, Zhaowen Zhou +14
The rapid advancement of Multimodal Large Language Models (MLLMs) has enabled browsing agents to acquire and reason over multimodal information in the real world. But existing benc…