3 papers
cs.CV2026
RUTA: Principled Visual Token Allocation via Rate-Utility Optimization
Jian Zou, Xiaoyu Xu, Zhihua Wang +3
High-resolution images and long videos provide vision-language models with rich context for multimodal reasoning and fine-grained perception, but the resulting long visual token se…
cs.CV2026
Learning Stable Canonical Worlds for Novel View Synthesis and Beyond
Xiaoyu Xu, Jian Zou, Sheyang Tang +3
Feed-forward Gaussian splatting (FFGS) facilitates real-time novel view synthesis, yet current methods often remain tied to view-dependent predictions. As more input views are adde…
cs.CV2026
Aesthetic Camera Viewpoint Suggestion with 3D Aesthetic Field
Sheyang Tang, Armin Shafiee Sarvestani, Jialu Xu +2
The aesthetic quality of a scene depends strongly on camera viewpoint. Existing approaches for aesthetic viewpoint suggestion are either single-view adjustments, predicting limited…