3 papers
cs.CV2026
Learning to Look Again: Loss-Gap Supervision for Free-form Crop Routing in Vision-Language Models
Jinchang Zhu, Rong Fu, Yi Ding +3
Vision-language models (VLMs) fail many detail-centric questions for a concrete reason: the answer is visible in the image, yet lost after the image is compressed into a low-resolu…
cs.MM2026
Adaptive Hierarchical Representation Alliance for Multimodal Learning
Chunlei Meng, Pengbin Feng, Jacqueline J. Pang +5
Multimodal models often align language, vision, and audio in a single final-layer latent space, implicitly assuming that task-relevant evidence emerges at the same semantic depth a…
cs.CL2026
FCPRAG: Fusion-Controller Parametric Retrieval-Augmented Generation for Stable Multi-Passage LoRA Injection
Jinchang Zhu, Jindong Li, Yi Ding +5
Parametric retrieval-augmented generation (PRAG) injects retrieved evidence into a large language model (LLM) through passage-specific LoRA adapters, reducing reliance on long in-c…