activity
20212026
most citedIs Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection

132 citations · 610 across the 214 of their papers we have counts for

collaborators
Showing 2025 · cs.CVShow all

46 papers · 2 filters

cs.CV2025

ShowUI-: Flow-based Generative Models as GUI Dexterous Hands

Siyuan Hu, Kevin Qinghong Lin, Mike Zheng Shou

Building intelligent agents capable of dexterous manipulation is essential for achieving human-like automation in both robotics and digital environments. However, existing GUI agen…

cs.CV2025

Factorized Learning for Temporally Grounded Video-Language Models

Wenzheng Zeng, Difei Gao, Mike Zheng Shou +1

Recent video-language models have shown great potential for video understanding, but still struggle with accurate temporal grounding for event-level perception. We observe that two…

cs.CV2025

Mitty: Diffusion-based Human-to-Robot Video Generation

Yiren Song, Cheng Liu, Weijia Mao +1

Learning directly from human demonstration videos is a key milestone toward scalable and generalizable robot learning. Yet existing methods rely on intermediate representations suc…

cs.CV2025

OmniPSD: Layered PSD Generation with Diffusion Transformer

Cheng Liu, Yiren Song, Haofan Wang +1

Recent advances in diffusion models have greatly improved image generation and editing, yet generating or reconstructing layered PSD files with transparent alpha channels remains h…

cs.CV2025

X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale

Pei Yang, Hai Ci, Yiren Song +1

The advancement of embodied AI has unlocked significant potential for intelligent humanoid robots. However, progress in both Vision-Language-Action (VLA) models and world models is…

cs.CV2025

DiffSeg30k: A Multi-Turn Diffusion Editing Benchmark for Localized AIGC Detection

Hai Ci, Ziheng Peng, Pei Yang +2

Diffusion-based editing enables realistic modification of local image regions, making AI-generated content harder to detect. Existing AIGC detection benchmarks focus on classifying…