18 papers
LiST: Local-Simplex Test-Time LoRA Fusion
Yihua Shao, Jia Li, Siyu Chen +12
Task-specific LoRA adapters offer a modular way to specialize large language and vision-language models. However, existing adapter composition methods are mostly static and cannot…
HUG-VIS: A Multimodal Benchmark for Human-centered Understanding and Generation in Visual Intelligence
Fei Ma, Zebang Cheng, Minghui Li +11
Visual intelligence seeks to perceive, interpret, and synthesize the visual world and is central to modern computer vision. Human-centered visual intelligence is especially demandi…
Human-Centric Intelligence in the Era of Foundation Models: A Survey
Yang Chen, Tianqi Wang, Xiaorui Jiang +13
Human-centric intelligence is evolving in the foundation-model era, with growing emphasis on scale, transferability, and general-purpose modeling. Yet it has not fully integrated w…
Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning
Xinyu Luo, Hui Liu, Yihua Shao +3
On-device in-context learning (ICL) relies on pre-inference retrieval to select demonstrations for useful context before downstream model inference. This retrieval must exploit tas…
GrainGS: Gradient-Decoupled Gaussian Splatting for Efficient Dynamic Novel View Synthesis
Jiahao He, Yihua Shao, Zhengkai Zhao +6
Dynamic scene reconstruction with 3D Gaussian Splatting requires a balance between fine-grained motion modeling, structural stability, and compact representation. Existing per-prim…
OralGPT-Plus: Learning to Use Visual Tools via Reinforcement Learning for Panoramic X-ray Analysis
Yuxuan Fan, Jing Hao, Hong Chen +5
Panoramic dental radiographs require fine-grained spatial reasoning, bilateral symmetry understanding, and multi-step diagnostic verification, yet existing vision-language models o…