8 papers
TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding
Lianyu Hu, Xiaoyu Ma, Zeqin Liao +1
Chain-of-thought (CoT) reasoning has proven effective for enhancing problem-solving in large language models. However, when applied to multimodal LLMs (MLLMs), existing CoT approac…
Balancing Multimodal Learning through Label Space Reshaping
Xiaoyu Ma, Weijie Zhang, Yuanhao Gao +3
Multimodal learning often suffers from modality imbalance, where modalities that converge faster dominate optimization while others remain undertrained. Existing approaches typical…
Resource-Efficient Reinforcement for Reasoning Large Language Models via Dynamic One-Shot Policy Refinement
Yunjian Zhang, Sudong Wang, Yang Li +5
Large language models (LLMs) have exhibited remarkable performance on complex reasoning tasks, with reinforcement learning under verifiable rewards (RLVR) emerging as a principled…
BrainMem: Brain-Inspired Evolving Memory for Embodied Agent Task Planning
Xiaoyu Ma, Lianyu Hu, Wenbing Tang +4
Embodied task planning requires agents to execute long-horizon, goal-directed actions in complex 3D environments, where success depends on both immediate perception and accumulated…
BLURR: A Boosted Low-Resource Inference for Vision-Language-Action Models
Xiaoyu Ma, Zhengqing Yuan, Zheyuan Zhang +3
Vision-language-action (VLA) models enable impressive zero shot manipulation, but their inference stacks are often too heavy for responsive web demos or high frequency robot contro…
Revisit Modality Imbalance at the Decision Layer
Xiaoyu Ma, Hao Chen
Multimodal learning integrates information from different modalities to enhance model performance, yet it often suffers from modality imbalance, where dominant modalities overshado…