4 papers
Towards Physically Executable 3D Gaussian for Embodied Navigation
Bingchen Miao, Rong Wei, Zhiqi Ge +8
3D Gaussian Splatting (3DGS), a 3D representation method with photorealistic real-time rendering capabilities, is regarded as an effective tool for narrowing the sim-to-real gap. H…
On Path to Multimodal Generalist: General-Level and General-Bench
Hao Fei, Yuan Zhou, Juncheng Li +29
The Multimodal Large Language Model (MLLM) is currently experiencing rapid growth, driven by the advanced capabilities of LLMs. Unlike earlier specialists, existing MLLMs are evolv…
Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining
Zhiqi Ge, Juncheng Li, Xinglei Pang +7
Digital agents are increasingly employed to automate tasks in interactive digital environments such as web pages, software applications, and operating systems. While text-based age…
Unified Generative and Discriminative Training for Multi-modal Large Language Models
Wei Chow, Juncheng Li, Qifan Yu +7
In recent times, Vision-Language Models (VLMs) have been trained under two predominant paradigms. Generative training has enabled Multimodal Large Language Models (MLLMs) to tackle…