1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.CV2025
Causally-Grounded Dual-Path Attention Intervention for Object Hallucination Mitigation in LVLMs
Liu Yu, Zhonghao Chen, Ping Kuang +4
Object hallucination remains a critical challenge in Large Vision-Language Models (LVLMs), where models generate content inconsistent with visual inputs. Existing language-decoder…
cs.AI2025★ 1 cited
Ming-Omni: A Unified Multimodal Model for Perception and Generation
Inclusion AI, Biao Gong, Cheng Zou +55
We propose Ming-Omni, a unified multimodal model capable of processing images, text, audio, and video, while demonstrating strong proficiency in both speech and image generation. M…
cs.CV2024
Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning
Harold Haodong Chen, Harry Yang, Ser-Nam Lim
Recent advances in video generation have outpaced progress in video editing, which remains constrained by several limiting factors, namely: (a) the task's dependency on supervision…