7 papers · 1 filter
When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware
Hao Dou, Ruiwen Tian
Fewer visual tokens do not guarantee lower end-to-end latency. We evaluate break-even with a reproducible protocol that accounts for decision overhead, shared work, and the operato…
DreamX-World 1.0: A General-Purpose Interactive World Model
DreamX Team, Yancheng Bai, Rui Chen +20
DreamX-World 1.0 is a general-purpose interactive text/image-to-video world model for controllable long-horizon generation. It supports camera navigation, revisits to previously ob…
RISE: Reliable Improvement in Self-Evolving Vision-Language Models
Chaoran Xu, Yingmao Miao, Pengfei Zhang +3
Vision-language models (VLMs) have achieved strong multimodal reasoning capabilities, but further improving them still relies heavily on large-scale human-constructed supervision f…
Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models
Meiqi Wu, Zhixin Cai, Fufangchen Zhao +13
Video--based world models have emerged along two dominant paradigms: video generation and 3D reconstruction. However, existing evaluation benchmarks either focus narrowly on visual…
DiffStyle3D: Consistent 3D Gaussian Stylization via Attention Optimization
Yitong Yang, Xuexin Liu, Yinglin Wang +4
3D style transfer enables the creation of visually expressive 3D content, enriching the visual appearance of 3D scenes and objects. However, existing VGG- and CLIP-based methods st…
PositionIC: Unified Position and Identity Consistency for Image Customization
Junjie Hu, Tianyang Han, Kai Ma +6
Recent subject-driven image customization excels in fidelity, yet fine-grained instance-level spatial control remains an elusive challenge, hindering real-world applications. This…