5 papers · 1 filter
X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction
Xiaoming Ren, Ru Zhen, Chao Li +11
Inspired by the development of OpenClaw, there is a growing demand for mobile-based personal agents capable of handling complex and intuitive interactions. In this technical report…
Click-to-Ask: An AI Live Streaming Assistant with Offline Copywriting and Online Interactive QA
Ruizhi Yu, Keyang Zhong, Peng Liu +5
Live streaming commerce has become a prominent form of broadcasting in the modern era. To facilitate more efficient and convenient product promotions for streamers, we present Clic…
LoopAnimate: Loopable Salient Object Animation
Fanyi Wang, Peng Liu, Haotian Hu +6
Research on diffusion model-based video generation has advanced rapidly. However, limitations in object fidelity and generation length hinder its practical applications. Additional…
Zero-shot High-fidelity and Pose-controllable Character Animation
Bingwen Zhu, Fanyi Wang, Tianyi Lu +7
Image-to-video (I2V) generation aims to create a video sequence from a single image, which requires high temporal coherence and visual fidelity. However, existing approaches suffer…
Lightweight high-resolution Subject Matting in the Real World
Peng Liu, Fanyi Wang, Jingwen Su +2
Existing saliency object detection (SOD) methods struggle to satisfy fast inference and accurate results simultaneously in high resolution scenes. They are limited by the quality o…