2 citations · 2 across the 1 of their papers we have counts for
2 papers
cs.RO2026
TEMPO: Learning Temporal Context for Dynamic Robot Manipulation
Zhenyang Feng, Jimin Heo, Erik B. Sudderth +1
Vision-language-action (VLA) models have achieved impressive performance in quasi-static manipulation, but struggle in dynamic manipulation tasks because they operate on a single o…
cs.HC2025★ 2 cited
"It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language Models
Kapil Garg, Xinru Tang, Jimin Heo +4
Vision-Language Models (VLMs) are increasingly used by blind and low-vision (BLV) people to identify and understand products in their everyday lives, such as food, personal care it…