activity
20242026
most citedAsk Questions with Double Hints: Visual Question Generation with Answer-awareness and Region-reference

4 citations · 4 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

EmbodiedHead: Real-Time Listening and Speaking Avatar for Conversational Agents

Yu Zhang, Kaiyuan Shen, Yang Li

We present EmbodiedHead, a speech-driven talking-head framework that equips LLMs with real-time visual avatars for conversation. A practical embodied avatar must achieve real-time…

cs.CV2025

OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions

Wendong Bu, Kaihang Pan, Yuze Lin +6

Large language models (LLMs) have unified diverse linguistic tasks within a single framework, yet such unification remains unexplored in human motion generation. Existing methods a…

cs.CV2025

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL

Kaihang Pan, Wendong Bu, Yuruo Wu +7

Recent studies extend the autoregression paradigm to text-to-image generation, achieving performance comparable to diffusion models. However, our new PairComp benchmark -- featurin…

cs.CV2025

Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning

Kaihang Pan, Yang Wu, Wendong Bu +9

Recent endeavors in Multimodal Large Language Models (MLLMs) aim to unify visual comprehension and generation. However, these two capabilities remain largely independent, as if the…

cs.CV20244 cited

Ask Questions with Double Hints: Visual Question Generation with Answer-awareness and Region-reference

Kai Shen, Lingfei Wu, Siliang Tang +4

The visual question generation (VQG) task aims to generate human-like questions from an image and potentially other side information (e.g. answer type). Previous works on VQG fall…

cs.CV2024

T2S-GPT: Dynamic Vector Quantization for Autoregressive Sign Language Production from Text

Aoxiong Yin, Haoyuan Li, Kai Shen +2

In this work, we propose a two-stage sign language production (SLP) paradigm that first encodes sign language sequences into discrete codes and then autoregressively generates sign…