Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Long-Horizon Embodied Decision-Making via Multimodal Memory Compression
Bingxuan Li, Rui Yang, Cheng Qian +6
Agents are increasingly expected to act not only as task executors, but also as decision-makers on behalf of human users. This shift requires agents to accumulate evidence over lon…
cs.CV2025
DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
Chengke Zou, Xingang Guo, Rui Yang +3
The rapid advancements in Vision-Language Models (VLMs) have shown great potential in tackling mathematical reasoning tasks that involve visual context. Unlike humans who can relia…