1 paper
Fufangchen Zhao, Songbai Tan, Xuerui Qiu +7
Existing video large language models (VLLMs) primarily leverage prompt agnostic visual encoders, which extract untargeted facial representations without awareness of the queried in…