From the 1 of 29 linked papers with an AI index.
29 papers
Test-Time Hallucination Control in Large Vision-Language Models
Mehran Tamjidi, Hamidreza Dastmalchi, Ali Cheraghian +3
Object Hallucination in large vision-language models (LVLMs), where models generate non-factual content about input images, remains a critical barrier to their reliability in real-…
Beyond Global Editing: Per-Instance Disentangled Subspaces for Training-Free Hallucination Mitigation in LVLMs
Ali Cheraghian, Hamidreza Dastmalchi, Hamed Barzamini +4
Recent advances in large vision-language models (LVLMs) have enabled powerful multimodal reasoning by integrating visual encoders with large language models (LLMs). However, their…
HRDiT: Training-Free High-Resolution Image Generation with Off-the-Shelf Diffusion Transformer Models
Yu Xue, Haoxuan Qu, Zhuoling Li +5
Training-free text-to-high-resolution image generation has recently attracted growing research attention. However, existing studies on this task primarily focus on adapting off-the…
Introspective Attention Modulation for Safe Text-to-Image Generation
Basim Azam, Hossein Rahmani, Naveed Akhtar
The paper proposes a method that monitors and adjusts the attention mechanisms of text‑to‑image diffusion models at inference time to prevent the generation of unsafe content while…
Training-free Controllable Human Motion Generation under Heterogeneous Constraints
Xiaofei Hui, Bo Yan, Haoxuan Qu +2
Training-free controllable motion generation has attracted growing interest for enabling flexible constraint enforcement without constraint-specific training. However, existing tra…
TraMP-LLaMA: Generative Interpretability with Decoupled Instruction Tuning for Facial Expression Quality Assessment
Shuchao Duan, Alan Whone, Hossein Rahmani +2
Existing facial expression quality assessment (FEQA) methods typically produce only a severity score, without explicitly communicating the observable facial motion evidence that su…