2 papers
cs.CV2026
LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding
Xiaodong Wang, Langling Huang, Zhirong Wu +4
The development of multimodal large language models (MLLMs) has advanced general video understanding. However, existing video evaluation benchmarks primarily focus on non-interacti…
cs.CV2024
TIGER: Text-Instructed 3D Gaussian Retrieval and Coherent Editing
Teng Xu, Jiamin Chen, Peng Chen +3
Editing objects within a scene is a critical functionality required across a broad spectrum of applications in computer vision and graphics. As 3D Gaussian Splatting (3DGS) emerges…