1 paper
Yulin Fei, Yuhui Gao, Xingyuan Xian +3
With the rise of multimodal large language models, accurately extracting and understanding textual information from video content, referred to as video based optical character reco…