6 papers · 1 filter
IIR-VLM: In-Context Instance-level Recognition for Large Vision-Language Models
Liang Shi, Wei Li, Kevin M Beussman +2
Instance-level recognition (ILR) concerns distinguishing individual instances from one another, with person re-identification as a prominent example. Despite the impressive visual…
Restore-R1: Efficient Image Restoration Agents via Reinforcement Learning with Multimodal LLM Perceptual Feedback
Jianglin Lu, Yuanwei Wu, Ziyi Zhao +4
Complex image restoration aims to recover high-quality images from inputs affected by multiple degradations such as blur, noise, rain, and compression artifacts. Recent restoration…
SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
Yiyang Huang, Liang Shi, Yitian Zhang +2
Large Vision-Language Models (LVLMs) excel in diverse cross-modal tasks. However, object hallucination, where models produce plausible but inaccurate object descriptions, remains a…
ExpertGen: Training-Free Expert Guidance for Controllable Text-to-Face Generation
Liang Shi, Yun Fu
Recent advances in diffusion models have significantly improved text-to-face generation, but achieving fine-grained control over facial features remains a challenge. Existing metho…
LightAvatar: Efficient Head Avatar as Dynamic Neural Light Field
Huan Wang, Feitong Tan, Ziqian Bai +9
Recent works have shown that neural radiance fields (NeRFs) on top of parametric models have reached SOTA quality to build photorealistic head avatars from a monocular video. Howev…
Training-free Editioning of Text-to-Image Models
Jinqi Wang, Yunfei Fu, Zhangcan Ding +3
Inspired by the software industry's practice of offering different editions or versions of a product tailored to specific user groups or use cases, we propose a novel task, namely,…