activity
20242026
collaborators

9 papers

astro-ph.SR2026

The Production of Electron-Capture Elements in Thermonuclear Supernovae: Theory vs. Observations

S. Shiber, P. Hoeflich, T. Mera +9

Type Ia supernovae (SNe Ia) explosively destroy carbon-oxygen white dwarfs (WDs) in multiple stellar systems. They produce approximately 50% of the iron-group elements in the Unive…

cs.CV2026

VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling

Yuqi Zhang, Cheng Chen, Yuyu Guo +6

Vision Language Models (VLMs) face significant challenges with ultra-long, interleaved image-text sequences due to the quadratic complexity of self-attention. Current solutions eit…

cs.CV2026

Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification

Wujian Peng, Lingchen Meng, Yuxuan Cai +7

Unified Multimodal Modeling aims to integrate visual understanding and generation within a single system. However, existing approaches typically rely on two disparate visual tokeni…

cs.RO2026

Unify Robot Actions in Camera Frame

Sicheng Xie, Lingchen Meng, Zijie Diao +9

Cross-embodiment robot learning requires a unified action representation with consistent semantics across robot platforms. Existing representations suffer from platform-specific in…

cs.CV2026

INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning

Wujian Peng, Lingchen Meng, Yitong Chen +7

Large Multimodal Models (LMMs) have made significant breakthroughs with the advancement of instruction tuning. However, while existing models can understand images and videos at a…

cs.CV2025

Qwen3-VL Technical Report

Shuai Bai, Yuxuan Cai, Ruizhe Chen +61

We introduce Qwen3-VL, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks. It natively…