collaborators

8 papers

cs.CV2026

TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding

Lianyu Hu, Xiaoyu Ma, Zeqin Liao +1

Chain-of-thought (CoT) reasoning has proven effective for enhancing problem-solving in large language models. However, when applied to multimodal LLMs (MLLMs), existing CoT approac…

cs.LG2026

Balancing Multimodal Learning through Label Space Reshaping

Xiaoyu Ma, Weijie Zhang, Yuanhao Gao +3

Multimodal learning often suffers from modality imbalance, where modalities that converge faster dominate optimization while others remain undertrained. Existing approaches typical…

cs.AI2026

Resource-Efficient Reinforcement for Reasoning Large Language Models via Dynamic One-Shot Policy Refinement

Yunjian Zhang, Sudong Wang, Yang Li +5

Large language models (LLMs) have exhibited remarkable performance on complex reasoning tasks, with reinforcement learning under verifiable rewards (RLVR) emerging as a principled…

cs.RO2026

BrainMem: Brain-Inspired Evolving Memory for Embodied Agent Task Planning

Xiaoyu Ma, Lianyu Hu, Wenbing Tang +4

Embodied task planning requires agents to execute long-horizon, goal-directed actions in complex 3D environments, where success depends on both immediate perception and accumulated…

cs.RO2025

BLURR: A Boosted Low-Resource Inference for Vision-Language-Action Models

Xiaoyu Ma, Zhengqing Yuan, Zheyuan Zhang +3

Vision-language-action (VLA) models enable impressive zero shot manipulation, but their inference stacks are often too heavy for responsive web demos or high frequency robot contro…

cs.LG2025

Revisit Modality Imbalance at the Decision Layer

Xiaoyu Ma, Hao Chen

Multimodal learning integrates information from different modalities to enhance model performance, yet it often suffers from modality imbalance, where dominant modalities overshado…