collaborators

15 papers

cs.CL2026

Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning

Xinyu Luo, Hui Liu, Yihua Shao +3

The paper introduces Conditional Retrieval Alignment (CoRA), a gradient‑free method that turns a frozen encoder into a task‑conditioned retriever for on‑device in‑context learning,…

cs.CV2026

GrainGS: Gradient-Decoupled Gaussian Splatting for Efficient Dynamic Novel View Synthesis

Jiahao He, Yihua Shao, Zhengkai Zhao +6

Dynamic scene reconstruction with 3D Gaussian Splatting requires a balance between fine-grained motion modeling, structural stability, and compact representation. Existing per-prim…

cs.CV2026

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding

Shida Gao, Feng Xue, Xiangfeng Wang +8

Multimodal large language models (MLLMs) are rapidly expanding from general video understanding to finer-grained understanding such as spatio-temporal video grounding (STVG) and re…

cs.CV2026

3DSceneEditor: Controllable 3D Scene Editing with Gaussian Splatting

Ziyang Yan, Yihua Shao, Minwen Liao +7

The creation of 3D scenes has traditionally been both labor-intensive and costly, requiring designers to meticulously configure 3D assets and environments. Recent advancements in g…

cs.CV2026

OralGPT-Plus: Learning to Use Visual Tools via Reinforcement Learning for Panoramic X-ray Analysis

Yuxuan Fan, Jing Hao, Hong Chen +5

Panoramic dental radiographs require fine-grained spatial reasoning, bilateral symmetry understanding, and multi-step diagnostic verification, yet existing vision-language models o…

cs.CV2026

Nüwa: Mending the Spatial Integrity Torn by VLM Token Pruning

Yihong Huang, Fei Ma, Yihua Shao +4

Vision token pruning has proven to be an effective acceleration technique for the efficient Vision Language Model (VLM). However, existing pruning methods demonstrate excellent per…