81 citations · 82 across the 2 of their papers we have counts for
5 papers
Agentic Robot: A Brain-Inspired Framework for Vision-Language-Action Models in Embodied Agents
Zhejian Yang, Yongchao Chen, Xueyang Zhou +8
Long-horizon robotic manipulation poses significant challenges for autonomous systems, requiring extended reasoning, precise execution, and robust error recovery across complex seq…
Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
Yuting Li, Lai Wei, Kaipeng Zheng +6
Despite the rapid progress of multimodal large language models (MLLMs), they have largely overlooked the importance of visual processing. In a simple yet revealing experiment, we i…
First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training
Lai Wei, Yuting Li, Chen Wang +4
Improving Multi-modal Large Language Models (MLLMs) in the post-training stage typically relies on supervised fine-tuning (SFT) or reinforcement learning (RL), which require expens…
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
Lai Wei, Yuting Li, Kaipeng Zheng +5
Recent advancements in large language models (LLMs) have demonstrated impressive chain-of-thought reasoning capabilities, with reinforcement learning (RL) playing a crucial role in…
HTR-VT: Handwritten Text Recognition with Vision Transformer
Yuting Li, Dexiong Chen, Tinglong Tang +1
We explore the application of Vision Transformer (ViT) for handwritten text recognition. The limited availability of labeled data in this domain poses challenges for achieving high…