2 papers
cs.CV2026
ImageAttributionBench: How Far Are We from Generalizable Attribution?
Tingshu Mou, Zhipeng Wei, Chao Gong +2
The rapid advancement of generative AI has enabled the creation of highly realistic and diverse synthetic images, posing critical challenges for image provenance and misinformation…
cs.CV2026
ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models
Tingshu Mou, Jiabo He, Renying Wang +5
Recent advances in Multi-modal Large Language Models (MLLMs) target 3D spatial intelligence, yet the progress has been largely driven by post-training on curated benchmarks, leavin…