8 papers
VC-Tooler: Learning Compositional and Adaptive Visual Tool Use
Yizheng Wu, Jiashen Hua, Bing Deng +1
Agentic multimodal reasoning extends passive image understanding by allowing VLMs to actively acquire and refine visual evidence through visual tool interactions. Effective visual…
Illuminating Visual Identity in Universal Multimodal Embeddings
Jiawei Cao, Junyi Feng, Jiashen Hua +5
Universal Multimodal Embeddings (UMEs) aim to unify various modalities and tasks into a shared representation space. In recent years, this field has witnessed substantial progress…
EchoStyle: Unlocking High-Fidelity Video Stylization with Reverse Data Synthesis
Huaqiu Li, Jiahao Wang, Sijia Cai +4
While image stylization has been studied extensively, video stylization remains a critical and largely unsolved challenge in the field of intelligent content creation. Existing met…
AnyID: Ultra-Fidelity Universal Identity-Preserving Video Generation from Any Visual References
Jiahao Wang, Hualian Sheng, Sijia Cai +5
Identity-preserving video generation offers powerful tools for creative expression, allowing users to customize videos featuring their beloved characters. However, prevailing metho…
Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
Zhiyu Pan, Yizheng Wu, Jiashen Hua +5
Reasoning has emerged as a key capability of large language models. In linguistic tasks, this capability can be enhanced by self-improving techniques that refine reasoning paths fo…
EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer
Yuxiao Yang, Hualian Sheng, Sijia Cai +6
Video generation models have advanced significantly, yet they still struggle to synthesize complex human movements due to the high degrees of freedom in human articulation. This li…