Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
LinkedOut: Linking World Knowledge Representation Out of Video LLM for Next-Generation Video Recommendation
Haichao Zhang, Yao Lu, Lichen Wang +4
Video Large Language Models (VLLMs) unlock world-knowledge-aware video understanding through pretraining on internet-scale data and have already shown promise on tasks such as movi…
cs.CV2025
PP-DocBee: Improving Multimodal Document Understanding Through a Bag of Tricks
Feng Ni, Kui Huang, Yao Lu +4
With the rapid advancement of digitalization, various document images are being applied more extensively in production and daily life, and there is an increasingly urgent need for…