Publications (30)
Learning Interaction-aware 3D Gaussian Splatting for One-shot Hand Avatars
Xuan Huang, Hanhui Li, Wanquan Liu +4
In this paper, we propose to create animatable avatars for interacting hands with 3D Gaussian Splatting (GS) and single-image inputs. Existing GS-based methods designed for single…
Revealing Directions for Text-guided 3D Face Editing
Zhuo Chen, Yichao Yan, Sehngqi Liu +5
3D face editing is a significant task in multimedia, aimed at the manipulation of 3D face models across various control signals. The success of 3D-aware GAN provides expressive 3D…
Nuanced Emotion Recognition Based on a Segment-based MLLM Framework Leveraging Qwen3-Omni for AH Detection
Liang Tang, Hongda Li, Jiayu Zhang +5
Emotion recognition in videos is a pivotal task in affective computing, where identifying subtle psychological states such as Ambivalence and Hesitancy holds significant value for…
WildGHand: Learning Anti-Perturbation Gaussian Hand Avatars from Monocular In-the-Wild Videos
Hanhui Li, Xuan Huang, Wanquan Liu +5
Despite recent progress in 3D hand reconstruction from monocular videos, most existing methods rely on data captured in well-controlled environments and therefore degrade in real-w…
Beyond endoscopy for over with ramification 1: Poisson summation
Yuhao Cheng
At the beginning of this century, Langlands introduced a strategy known as \emph{Beyond Endoscopy} to attack the principle of functoriality. AltuÄ studied over $\m…
Multimodal Latent Diffusion Model for Complex Sewing Pattern Generation
Shengqi Liu, Yuhao Cheng, Zhuo Chen +6
Generating sewing patterns in garment design is receiving increasing attention due to its CG-friendly and flexible-editing nature. Previous sewing pattern generation methods have b…
ProPhy: Progressive Physical Alignment for Dynamic World Simulation
Zijun Wang, Panwen Hu, Jing Wang +7
Recent advances in video generation have shown remarkable potential for constructing world simulators. However, current models still struggle to produce physically consistent resul…
GarmentAligner: Text-to-Garment Generation via Retrieval-augmented Multi-level Corrections
Shiyue Zhang, Zheng Chong, Xujie Zhang +4
General text-to-image models bring revolutionary innovation to the fields of arts, design, and media. However, when applied to garment generation, even the state-of-the-art text-to…
MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts
Jingnan Gao, Zhe Wang, Xianze Fang +7
Recent advances in language and vision have demonstrated that scaling up model capacity consistently improves performance across diverse tasks. In 3D visual geometry reconstruction…
Evolving in Tasks: Empowering the Multi-modality Large Language Model as the Computer Use Agent
Yuhao Cheng, Liang Tang, Shuxian Li +5
Computer use agents represent an emerging area in artificial intelligence, aiming to operate computers autonomously to fulfill user tasks, attracting significant attention from bot…
Beyond endoscopy for over with ramification 2: bounds towards the Ramanujan conjecture
Yuhao Cheng
We continue generalizing AltuÄ's work on over in the unramified setting for \emph{Beyond Endoscopy} to the ramified case where ramification occurs at…
ConsistentID: Portrait Generation with Multimodal Fine-Grained Identity Preserving
Jiehui Huang, Xiao Dong, Wenhui Song +9
Diffusion-based technologies have made significant strides, particularly in personalized and customized facialgeneration. However, existing methods face challenges in achieving hig…
AutoStudio: Crafting Consistent Subjects in Multi-turn Interactive Image Generation
Junhao Cheng, Xi Lu, Hanhui Li +5
As cutting-edge Text-to-Image (T2I) generation models already excel at producing remarkable single images, an even more challenging task, i.e., multi-turn interactive image generat…
LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation
Wenhui Song, Hanhui Li, Jiehui Huang +5
In this paper, we present LaVieID, a novel \underline{l}ocal \underline{a}utoregressive \underline{vi}d\underline{e}o diffusion framework designed to tackle the challenging \underl…
TheaterGen: Character Management with LLM for Consistent Multi-turn Image Generation
Junhao Cheng, Baiqiao Yin, Kaixin Cai +9
Recent advances in diffusion models can generate high-quality and stunning images from text. However, multi-turn image generation, which is of high demand in real-world scenarios,…
Rethink Predicting the Optical Flow with the Kinetics Perspective
Yuhao Cheng, Siru Zhang, Yiqiang Yan
Optical flow estimation is one of the fundamental tasks in low-level computer vision, which describes the pixel-wise displacement and can be used in many other tasks. From the appa…
BAMI: Training-Free Bias Mitigation in GUI Grounding
Borui Zhang, Bo Zhang, Bo Wang +6
GUI grounding is a critical capability for enabling GUI agents to execute tasks such as clicking and dragging. However, in complex scenarios like the ScreenSpot-Pro benchmark, exis…
Beyond endoscopy for the symmetric square representation: The simple trace formula case
Yuhao Cheng
At the beginning of this century, Langlands introduced a strategy known as \emph{Beyond Endoscopy} to attack the principle of functoriality. AltuÄ studied over $\m…
Beyond endoscopy for over with ramification 4: contribution of non-elliptic parts
Yuhao Cheng
We continue our work on over in the ramified setting for \emph{Beyond Endoscopy}. We establish asymptotic formulas for each term of the trace formula w…
Monocular Identity-Conditioned Facial Reflectance Reconstruction
Xingyu Ren, Jiankang Deng, Yuhao Cheng +5
Recent 3D face reconstruction methods have made remarkable advancements, yet there remain huge challenges in monocular high-quality facial reflectance reconstruction. Existing meth…
GANHead: Towards Generative Animatable Neural Head Avatars
Sijing Wu, Yichao Yan, Yunhao Li +5
To bring digital avatars into people's lives, it is highly demanded to efficiently generate complete, realistic, and animatable head avatars. This task is challenging, and it is di…
Simple and Robust Loss Design for Multi-Label Learning with Missing Labels
Youcai Zhang, Yuhao Cheng, Xinyu Huang +4
Multi-label learning in the presence of missing labels (MLML) is a challenging problem. Existing methods mainly focus on the design of network structures or training schemes, which…
Topo4D: Topology-Preserving Gaussian Splatting for High-Fidelity 4D Head Capture
Xuanchen Li, Yuhao Cheng, Xingyu Ren +4
4D head capture aims to generate dynamic topological meshes and corresponding texture maps from videos, which is widely utilized in movies and games for its ability to simulate fac…
EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation
Yongxin Wang, Meng Cao, Haokun Lin +5
Multimodal large language models (MLLMs) have achieved remarkable progress on various visual question answering and reasoning tasks leveraging instruction fine-tuning specific data…
Head3D: Complete 3D Head Generation via Tri-plane Feature Distillation
Yuhao Cheng, Yichao Yan, Wenhan Zhu +3
Head generation with diverse identities is an important task in computer vision and computer graphics, widely used in multimedia applications. However, current full head generation…
Semantic Role Labeling with Associated Memory Network
Chaoyu Guan, Yuhao Cheng, Hai Zhao
Semantic role labeling (SRL) is a task to recognize all the predicate-argument pairs of a sentence, which has been in a performance improvement bottleneck after a series of latest…
Beyond endoscopy for over with ramification 3: contribution of the elliptic part
Yuhao Cheng
We continue to work on \emph{Beyond Endoscopy} for over with ramification at (where ), generalizing the fina…
Towards High-fidelity 3D Talking Avatar with Personalized Dynamic Texture
Xuanchen Li, Jianyu Wang, Yuhao Cheng +5
Significant progress has been made for speech-driven 3D face animation, but most works focus on learning the motion of mesh/geometry, ignoring the impact of dynamic texture. In thi…
SingingBot: An Avatar-Driven System for Robotic Face Singing Performance
Zhuoxiong Xu, Xuanchen Li, Yuhao Cheng +3
Equipping robotic faces with singing capabilities is crucial for empathetic Human-Robot Interaction. However, existing robotic face driving research primarily focuses on conversati…
LSCD: A Large-Scale Screen Content Dataset for Video Compression
Yuhao Cheng, Siru Zhang, Yiqiang Yan +2
Multimedia compression allows us to watch videos, see pictures and hear sounds within a limited bandwidth, which helps the flourish of the internet. During the past decades, multim…