2 papers
cs.CV2026
Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos
Hao Zheng, Jinyi Huang, Tiantian Zheng +2
Assembly action understanding is a key enabler for effective human-robot collaborative assembly, yet it remains challenging due to subtle motions and fine-grained hand-object inter…
cs.CV2026
Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model
Jiaxin Liu, Xun Xu, Zhenhao Zhang +5
Vision-Language-Action (VLA) models have become an important paradigm of embodied AI. However, existing VLA models typically assume well-lit and stable indoor settings, while real-…