3 papers
cs.CV2026
BabyVision: Visual Reasoning Beyond Language
Liang Chen, Weichu Xie, Yiyan Liang +27
While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile…
cs.AI2026
EvoCode-Bench: Evaluating Coding Agents in Multi-Turn Iterative Interactions
Haiyang Shen, Xuanzhong Chen, Wendong Xu +3
Coding agents are increasingly used as iterative development partners, but most benchmarks still evaluate one specification followed by one final assessment. This leaves out a basi…
cs.SE2026
RoadmapBench: Evaluating Long-Horizon Agentic Software Development Across Version Upgrades
Xinbo Xu, Ruihan Yang, Haiyang Shen +13
Coding agents are increasingly deployed in real software development, where a single version iteration requires months of coordinated work across many files. However, most existing…