activity
20242026
collaborators

9 papers

cs.CV2026

PixVL: Self-Supervised Training of Pixel-Level MLLMs via a Unified Mask--Text Consistency Cycle

Yicheng Xiao, Haoxuan Ma, Caorui Li +7

Recent studies develop pixel-level multimodal large language models (MLLMs) that support both Region Segmentation and Region Understanding, extending multimodal interaction from wh…

cs.AI2026

Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents

Yutao Sun, Yanting Miao, Hao-Xuan Ma +8

Improving vision-language models (VLMs) on visual reasoning typically requires retraining or hand-designed prompts and tools. We present Dynamo, a training-free framework that adap…

cs.CV2026

REKEY: Metadata-Grounded Visual-Key Regeneration for Contamination-Resilient VQA Evaluation

Tengjie Lin, Yutao Sun, Jingwei Ni +7

Static visual question answering (VQA) benchmarks age quickly: Once the items leak into training corpora, scores can reflect memorization rather than genuine visual ability, thus o…

cs.LG2026

Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation

Guo Yu, Wenlin Liu, Yulan Hu +3

On-policy distillation (OPD) has recently become a prominent post-training recipe by combining two desirable ingredients: on-policy student-generated trajectories and dense token-l…

cs.CV2026

Adversarial Orthogonal Disentanglement for LVLM Hallucination Mitigation

Ruoxi Cheng, Haoxuan Ma, Zhengfei Hai +6

Large Vision-Language Models (LVLMs) have advanced multimodal understanding, yet their reliability is limited by hallucination, where generated content conflicts with visual facts.…

cs.AI2026

MMR-Bench: A Comprehensive Benchmark for Multimodal LLM Routing

Haoxuan Ma, Guannan Lai, Han-Jia Ye

Multimodal large language models (MLLMs) have advanced rapidly, yet heterogeneity in architecture, alignment strategies, and efficiency means that no single model is uniformly supe…