3 papers
cs.CV2026
TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions
Ce Chen, Yi Ren, Yuanming Li +5
Traditional Shot Boundary Detection (SBD) inherently struggles with complex transitions by formulating the task around isolated cut points, frequently yielding corrupted video shot…
cs.CV2026
Generate Your Talking Avatar from Video Reference
Zujin Guo, Zhenhui Ye, Yi Ren +4
Existing talking avatar methods typically adopt an image-to-video pipeline conditioned on a static reference image within the same scene as the target generation. This restricted,…
cs.AI2026
HippoCamp: Benchmarking Contextual Agents on Personal Computers
Zhe Yang, Shulin Tian, Kairui Hu +9
We present HippoCamp, a new benchmark designed to evaluate agents' capabilities on multimodal file management. Unlike existing agent benchmarks that focus on tasks like web interac…