3 papers
cs.CV2025
Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding
Meng Luo, Shengqiong Wu, Liqiang Jing +12
Recent advancements in large video models (LVMs) have significantly enhance video understanding. However, these models continue to suffer from hallucinations, producing content tha…
cs.CV2025
BoxSeg: Quality-Aware and Peer-Assisted Learning for Box-supervised Instance Segmentation
Jinxiang Lai, Wenlong Wu, Jiawei Zhan +5
Box-supervised instance segmentation methods aim to achieve instance segmentation with only box annotations. Recent methods have demonstrated the effectiveness of acquiring high-qu…
cs.CV2025
Spider: Any-to-Many Multimodal LLM
Jinxiang Lai, Jie Zhang, Jun Liu +3
Multimodal LLMs (MLLMs) have emerged as an extension of Large Language Models (LLMs), enabling the integration of various modalities. However, Any-to-Any MLLMs are limited to gener…