2 papers
cs.CV2025
HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices
HyperAI Team, Yuchen Liu, Kaiyang Han +26
Current multimodal large lanauge models possess strong perceptual and reasoning capabilities, however high computational and memory requirements make them difficult to deploy direc…
cs.CV2025
Multi-Level LVLM Guidance for Untrimmed Video Action Recognition
Liyang Peng, Sihan Zhu, Yunjie Guo
Action recognition and localization in complex, untrimmed videos remain a formidable challenge in computer vision, largely due to the limitations of existing methods in capturing f…