1 citations · 1 across the 4 of their papers we have counts for
5 papers
LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation
Ethan Chern, Zhulin Hu, Bohao Tang +4
Real-time video generation via diffusion is essential for building general-purpose multimodal interactive AI systems. However, the simultaneous denoising of all video frames with b…
Visual Programmability: A Guide for Code-as-Thought in Chart Understanding
Bohao Tang, Yan Ma, Fei Zhang +6
Chart understanding presents a critical test to the reasoning capabilities of Vision-Language Models (VLMs). Prior approaches face critical limitations: some rely on external tools…
Interaction as Intelligence: Deep Research With Human-AI Partnership
Lyumanshan Ye, Xiaojie Cai, Xinkai Wang +23
This paper introduces "Interaction as Intelligence" research series, presenting a reconceptualization of human-AI relationships in deep research tasks. Traditional approaches treat…
Thinking with Generated Images
Ethan Chern, Zhulin Hu, Steffi Chern +5
We present Thinking with Generated Images, a novel paradigm that fundamentally transforms how large multimodal models (LMMs) engage with visual reasoning by enabling them to native…
PC Agent: While You Sleep, AI Works -- A Cognitive Journey into Digital World
Yanheng He, Jiahe Jin, Shijie Xia +5
Imagine a world where AI can handle your work while you sleep - organizing your research materials, drafting a report, or creating a presentation you need for tomorrow. However, wh…