15 papers
BabyVision: Visual Reasoning Beyond Language
Liang Chen, Weichu Xie, Yiyan Liang +27
While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile…
Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
Haozhe Zhao, Shuzheng Si, Liang Chen +4
Large vision-language models (LVLMs) have achieved impressive results in various vision-language tasks. However, despite showing promising performance, LVLMs suffer from hallucinat…
SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?
Kean Shi, Zihang Li, Tianyi Ma +13
Computer-Using Agents (CUAs) are rapidly extending large language models (LLMs) beyond text-based reasoning toward action execution in more complex environments, such as web browse…
RoadmapBench: Evaluating Long-Horizon Agentic Software Development Across Version Upgrades
Xinbo Xu, Ruihan Yang, Haiyang Shen +13
Coding agents are increasingly deployed in real software development, where a single version iteration requires months of coordinated work across many files. However, most existing…
Improving MLLM Training Efficiency via Stage-Aware Sparsity
Kean Shi, Liang Chen, Haozhe Zhao +1
Multimodal Large Language Models (MLLMs) have demonstrated outstanding performance across a variety of domains. However, training MLLMs is often inefficient, as much of the computa…
LongR: Unleashing Long-Context Reasoning via Reinforcement Learning with Dense Utility Rewards
Bowen Ping, Zijun Chen, Yiyao Yu +3
Reinforcement Learning has emerged as a key driver for LLM reasoning. This capability is equally pivotal in long-context scenarios--such as long-dialogue understanding and structur…