2 papers
cs.CV2026
Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos
Hao Zheng, Jinyi Huang, Tiantian Zheng +2
Assembly action understanding is a key enabler for effective human-robot collaborative assembly, yet it remains challenging due to subtle motions and fine-grained hand-object inter…
cs.CV2026
Polyphony: Diffusion-based Dual-Hand Action Segmentation with Alternating Vision Transformer and Semantic Conditioning
Hao Zheng, Hu Wang, Tiantian Zheng +2
Dual-hand action segmentation, densely predicting actions for both hands from untrimmed videos, is essential for understanding complex bimanual activities. However, it poses severa…