activity
20232026
most citedGAIA: Zero-shot Talking Avatar Generation

2 citations · 2 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL2026

Training-Free Token-Level Steering for LLM Personalized Co-Writing

Wenhao Mao, Chengbin Hou, Weixiao Wang +4

While Large Language Models (LLMs) show great promise for personalization, they often lack specialized domain knowledge. Conventional solutions like fine-tuning struggle with high…

cs.CL2026

RE-TRAC: REcursive TRAjectory Compression for Deep Search Agents

Jialiang Zhu, Gongrui Zhang, Xiaolong Ma +17

LLM-based deep research agents are largely built on the ReAct framework. This linear design makes it difficult to revisit earlier states, branch into alternative search directions,…

cs.CL2025

InfoAgent: Advancing Autonomous Information-Seeking Agents

Gongrui Zhang, Jialiang Zhu, Ruiqi Yang +15

Building Large Language Model agents that expand their capabilities by interacting with external tools represents a new frontier in AI research and applications. In this paper, we…

cs.CV2025

Phi-Ground Tech Report: Advancing Perception in GUI Grounding

Miaosen Zhang, Ziqiang Xu, Jialiang Zhu +8

With the development of multimodal reasoning models, Computer Use Agents (CUAs), akin to Jarvis from \textit{"Iron Man"}, are becoming a reality. GUI grounding is a core component…

cs.CV2023

Multiple View Geometry Transformers for 3D Human Pose Estimation

Ziwei Liao, Jialiang Zhu, Chunyu Wang +2

In this work, we aim to improve the 3D reasoning ability of Transformers in multi-view 3D human pose estimation. Recent works have focused on end-to-end learning-based transformer…

cs.CV2023

GAIA: Zero-shot Talking Avatar Generation

Tianyu He, Junliang Guo, Runyi Yu +10

Zero-shot talking avatar generation aims at synthesizing natural talking videos from speech and a single portrait image. Previous methods have relied on domain-specific heuristics…