17 citations · 17 across the 3 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding
Jianghao Yin, Qingbin Li, Kun Sun +10
While Multimodal Large Language Models (MLLMs) excel at single-image understanding, they exhibit significantly degraded performance in multi-image reasoning scenarios. Multi-image…
cs.CV2026
3D Dynamics-Aware Manipulation: Endowing Manipulation Policies with 3D Foresight
Yuxin He, Ruihao Zhang, Xianzu Wu +3
The incorporation of world modeling into manipulation policy learning has pushed the boundary of manipulation performance. However, existing efforts simply model the 2D visual dyna…
cs.CV2024
Human Stone Toolmaking Action Grammar (HSTAG): A Challenging Benchmark for Fine-grained Motor Behavior Recognition
Cheng Liu, Xuyang Yan, Zekun Zhang +5
Action recognition has witnessed the development of a growing number of novel algorithms and datasets in the past decade. However, the majority of public benchmarks were constructe…