1 paper
Jiaze Li, Hao Yin, Wenhui Tan +7
Self-reflection mechanisms that rely on purely text-based rethinking processes perform well in most multimodal tasks. However, when directly applied to long-form video understandin…