5 papers
ServImage: An Image Generation and Editing Benchmark from Real-world Commercial Imaging Services
Fengxian Ji, Jingpu Yang, Zirui Song +5
Recent image generation and editing models demonstrate robust adherence to instructions and high visual quality on academic benchmarks. However, their performance on paid, real-wor…
PrefIx: Understand and Adapt to User Preference in Human-Agent Interaction
Jialin Li, Zhenhao Chen, Hanjun Luo +1
LLM-based agents can complete tasks correctly yet still frustrate users through poor interaction patterns, such as excessive confirmations, opaque reasoning, or misaligned pacing.…
ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models
Zirui Song, Guangxian Ouyang, Mingzhe Li +10
Large Vision-Language Models (LVLMs) have recently advanced robotic manipulation by leveraging vision for scene perception and language for instruction following. However, existing…
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
Zirui Song, Qian Jiang, Mingxuan Cui +9
The rise of Large Audio Language Models (LAMs) brings both potential and risks, as their audio outputs may contain harmful or unethical content. However, current research lacks a s…
MMAC-Copilot: Multi-modal Agent Collaboration Operating Copilot
Zirui Song, Yaohang Li, Meng Fang +6
Large language model agents that interact with PC applications often face limitations due to their singular mode of interaction with real-world environments, leading to restricted…