1 citations · 1 across the 1 of their papers we have counts for
5 papers
BabyVision: Visual Reasoning Beyond Language
Liang Chen, Weichu Xie, Yiyan Liang +27
While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile…
Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols
Xianchao Zeng, Xinyu Zhou, Youcheng Li +5
Vision-Language-Action (VLA) models have recently achieved remarkable progress in robotic manipulation, yet they remain limited in failure diagnosis and learning from failures. Add…
Kimi K2.5: Visual Agentic Intelligence
Kimi Team, Tongtong Bai, Yifan Bai +339
We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that…
DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry
Zhenyang Cai, Jiaming Zhang, Junjie Zhao +21
Reliable interpretation of multimodal data in dentistry is essential for automated oral healthcare, yet current multimodal large language models (MLLMs) struggle to capture fine-gr…
Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation
Nan Xiang, Tianyi Liang, Haiwen Huang +6
Text-to-3D (T23D) generation has transformed digital content creation, yet remains bottlenecked by blind trial-and-error prompting processes that yield unpredictable results. While…