activity
20242026
most citedGeneralist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms

1 citations · 1 across the 11 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

Iron: Intent-Aligned and Retrospective Dual Learning Framework for Enhancing Generalist Virtual Agents

Jiahe Ying, Wendong Bu, Kaihang Pan +6

Achieving virtual agents capable of automating tasks across diverse digital environments remains a pivotal challenge in Embodied AI. While Multimodal Large Language Models (MLLMs)…

cs.CV2025

OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions

Wendong Bu, Kaihang Pan, Yuze Lin +6

Large language models (LLMs) have unified diverse linguistic tasks within a single framework, yet such unification remains unexplored in human motion generation. Existing methods a…

cs.CV2025

WiseEdit: Benchmarking Cognition- and Creativity-Informed Image Editing

Kaihang Pan, Weile Chen, Haiyi Qiu +6

Recent image editing models boast next-level intelligent capabilities, facilitating cognition- and creativity-informed image editing. Yet, existing benchmarks provide too narrow a…

cs.CV2025

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities

Wendong Bu, Yang Wu, Qifan Yu +10

As multimodal large language models (MLLMs) advance, MLLM-based virtual agents have demonstrated remarkable performance. However, existing benchmarks face significant limitations,…

cs.CV2025

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL

Kaihang Pan, Wendong Bu, Yuruo Wu +7

Recent studies extend the autoregression paradigm to text-to-image generation, achieving performance comparable to diffusion models. However, our new PairComp benchmark -- featurin…

cs.CV2025

Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning

Kaihang Pan, Yang Wu, Wendong Bu +9

Recent endeavors in Multimodal Large Language Models (MLLMs) aim to unify visual comprehension and generation. However, these two capabilities remain largely independent, as if the…