Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
Mahtab Bigverdi, Zelun Luo, Cheng-Yu Hsieh +4
Multimodal language models (MLMs) still face challenges in fundamental visual perception tasks where specialized models excel. Tasks requiring reasoning about 3D structures benefit…
cs.CV2024
Few-Shot Classification of Interactive Activities of Daily Living (InteractADL)
Zane Durante, Robathan Harries, Edward Vendrow +5
Understanding Activities of Daily Living (ADLs) is a crucial step for different applications including assistive robots, smart homes, and healthcare. However, to date, few benchmar…