7 papers
LlamaSeg: Image Segmentation via Autoregressive Mask Generation
Jiru Deng, Tengjin Weng, Tianyu Yang +3
We present \textbf{LlamaSeg}, a visual autoregressive framework that unifies multiple image segmentation tasks via natural language instructions. By reformulating segmentation as v…
CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration
Hairui Zhu, Yiying Yang, Tengjin Weng +5
Complex image creation and editing often require more than a single generation or editing model. A user request may involve synthesizing images, localizing objects, segmenting regi…
EmoTrans: A Benchmark for Understanding, Reasoning, and Predicting Emotion Transitions in Multimodal LLMs
He Hu, Tengjin Weng, Zebang Cheng +5
Recent multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and generation, and are increasingly used in applications such as social ro…
GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning
Jingyi Wang, Lei Zhu, Tengjin Weng +8
Reinforcement Learning with Verifiable Rewards (RLVR) has advanced the reasoning capabilities of Large Language Models (LLMs) by leveraging direct outcome verification instead of l…
OddGridBench: Exposing the Lack of Fine-Grained Visual Discrepancy Sensitivity in Multimodal Large Language Models
Tengjin Weng, Wenhao Jiang, Jingyi Wang +3
Multimodal large language models (MLLMs) have achieved remarkable performance across a wide range of vision language tasks. However, their ability in low-level visual perception, p…
BEAP-Agent: Backtrackable Execution and Adaptive Planning for GUI Agents
Ziyu Lu, Tengjin Weng, Yiying Yang +3
GUI agents are designed to automate repetitive tasks and enhance productivity. However, existing GUI agents struggle to recover once they follow an incorrect exploration path, ofte…