collaborators

10 papers

cs.CV2026

PixVL: Self-Supervised Training of Pixel-Level MLLMs via a Unified Mask--Text Consistency Cycle

Yicheng Xiao, Haoxuan Ma, Caorui Li +7

Recent studies develop pixel-level multimodal large language models (MLLMs) that support both Region Segmentation and Region Understanding, extending multimodal interaction from wh…

cs.AI2026

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally

Shaowen Wang, Yuke Zheng, Tansheng Zhu +4

Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform frequency schedule and scaling a…

cs.CL2026

Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents

Shijue Huang, Hangyu Guo, Guanting Dong +8

Multimodal deep search requires an agent to solve open-world problems by chaining search, tool use, and visual reasoning over evolving textual and visual context. Two bottlenecks l…

cs.AI2026

TaskGround: Structured Executable Task Inference for Full-Scene Household Reasoning

ZhiYuan Feng, Yu Deng, Ruichuan An +11

In real home deployments, household agents must often operate from a complete household scene and a situated household request, rather than from a clean task specification. Such re…

cs.CL2026

Checkup2Action: A Multimodal Clinical Check-up Report Dataset for Patient-Oriented Action Card Generation

Sike Xiang, Shuang Chen, Kevin Qinghong Lin +4

Routine clinical check-up reports combine laboratory measurements, physiological assessments, imaging findings and visually structured information, but rarely tell patients what to…

cs.AI2026

Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence?

Qianshan Wei, Yishan Yang, Siyi Wang +12

Multimodal Large Language Models (MLLMs) are evolving from passive observers into active agents, solving problems through Visual Expansion (invoking visual tools) and Knowledge Exp…