7 papers
Distributionally Robust Set Representation Learning Under Inference-Time Element Corruption
Yankai Chen, Hanrong Zhang, Bowei He +2
Standard Set Representation Learning methods typically excel on curated data but often overlook the challenge of inference-time element corruption. This refers to scenarios where d…
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments
Qiuyue Wang, Mingsheng Li, Jian Guan +37
Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generali…
Revisiting Multimodal Positional Encoding in Vision-Language Models
Jie Huang, Xuejing Liu, Sibo Song +4
Multimodal position encoding is essential for vision-language models, yet there has been little systematic investigation into multimodal position encoding. We conduct a comprehensi…
SSP-SAM: SAM with Semantic-Spatial Prompt for Referring Expression Segmentation
Wei Tang, Xuejing Liu, Yanpeng Sun +1
The Segment Anything Model (SAM) excels at general image segmentation but has limited ability to understand natural language, which restricts its direct application in Referring Ex…
Qwen3-VL Technical Report
Shuai Bai, Yuxuan Cai, Ruizhe Chen +61
We introduce Qwen3-VL, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks. It natively…
Qwen2.5-VL Technical Report
Shuai Bai, Keqin Chen, Xuejing Liu +24
We introduce Qwen2.5-VL, the latest flagship model of Qwen vision-language series, which demonstrates significant advancements in both foundational capabilities and innovative func…