2 papers
cs.CV2026
InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos
Chi Zhang, Haoyang Shi, Yueyi Liu +4
Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by multimodal assistants,avatar…
cs.CV2026
Progressive Reasoning with Primitive Correction for Compositional Zero-Shot Learning
Ziyi Chen, Haoyan Shi, Sunhan Xu +1
Compositional Zero-Shot Learning (CZSL) aims to combine known attributes and objects as primitives for recognizing previously unseen attribute-object pairs. Prior works either pred…