4 citations · 5 across the 3 of their papers we have counts for
3 papers
cs.CV2025★ 4 cited
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Weiyun Wang, Zhangwei Gao, Lixin Gu +72
We introduce InternVL 3.5, a new family of open-source multimodal models that significantly advances versatility, reasoning capability, and inference efficiency along the InternVL…
cs.CV2025
Multimodal Long Video Modeling Based on Temporal Dynamic Context
Haoran Hao, Jiaming Han, Yiyuan Zhang +1
Recent advances in Large Language Models (LLMs) have led to significant breakthroughs in video understanding. However, existing models still struggle with long video processing due…
cs.LG2024★ 1 cited
Learning for Long-Horizon Planning via Neuro-Symbolic Abductive Imitation
Jie-Jing Shao, Hao-Ran Hao, Xiao-Wen Yang +1
Recent learning-to-imitation methods have shown promising results in planning via imitating within the observation-action space. However, their ability in open environments remains…