most citedExploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation

1 citations · 1 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

MLCR: Multi-Level Cue Refinement for Long-Term Multimodal Action Quality Assessment

Qiqi Li, Pengfei Wang, Hongyu Chen +1

Long-term multimodal action quality assessment (AQA) evaluates action execution in several-minute audiovisual sequences by mining discriminative quality cues for score prediction.…

cs.CV2026

PhysVideoGenerator: Towards Physically Aware Video Generation via Latent Physics Guidance

Siddarth Nilol Kundur Satish, Devesh Jaiswal, Hongyu Chen +1

Current video generation models produce high-quality aesthetic videos but often struggle to learn representations of real-world physics dynamics, resulting in artifacts such as unn…

cs.CV2025

UML-CoT: Structured Reasoning and Planning with Unified Modeling Language for Robotic Room Cleaning

Hongyu Chen, Guangrun Wang

Chain-of-Thought (CoT) prompting improves reasoning in large language models (LLMs), but its reliance on unstructured text limits interpretability and executability in embodied tas…

cs.CV2025

Position: Reasoning After Perception Means Reasoning Without Vision

Hongcheng Gao, Zihao Huang, Jingyi Tang +12

A common belief in multimodal research is that the perceptual weaknesses of vision--language models can be compensated by stronger language reasoning (e.g., chain-of-thought, in-co…

cs.CV20251 cited

Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation

Hongcheng Gao, Jiashu Qu, Jingyi Tang +6

The hallucination of large multimodal models (LMMs), providing responses that appear correct but are actually incorrect, limits their reliability and applicability. This paper aims…