activity
20232026
most citedGPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation

10 citations · 14 across the 15 of their papers we have counts for

collaborators

18 papers

cs.CV2026

Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation

Tianyi Xiong, Zhengyuan Yang, Xiaofei Wang +10

Large vision-language models have shown strong progress in UI-to-code generation, yet their test-time self-evolution remains unstable. We first identify a fundamental obstacle, ter…

cs.CV2025

TechImage-Bench: Rubric-Based Evaluation for Technical Image Generation

Minheng Ni, Zhengyuan Yang, Yaowen Zhang +9

We study technical image generation, where a model must synthesize information-dense, scientifically precise illustrations from detailed descriptions rather than merely produce vis…

cs.CL2025

SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models

Cheng-Han Chiang, Xiaofei Wang, Linjie Li +7

Current large language models (LLMs) and spoken language models (SLMs) begin thinking and taking actions only after the user has finished their turn. This prevents the model from i…

cs.CV2025

EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn Editing

Tianyu Chen, Yasi Zhang, Zhi Zhang +13

Instruction-based image editing has advanced rapidly, yet reliable and interpretable evaluation remains a bottleneck. Current protocols either (i) depend on paired reference images…

cs.CL2025

STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models

Cheng-Han Chiang, Xiaofei Wang, Linjie Li +7

Spoken Language Models (SLMs) are designed to take speech inputs and produce spoken responses. However, current SLMs lack the ability to perform an internal, unspoken thinking proc…

cs.CV2025

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Xiyao Wang, Zhengyuan Yang, Chao Feng +10

Reinforcement learning (RL) has shown great effectiveness for fine-tuning large language models (LLMs) using tasks that are challenging yet easily verifiable, such as math reasonin…