2 papers
cs.CV2026
Semantic and Visual Evidence for Efficient Long-Video Reasoning: A Solution for the HD-EPIC VQA Challenge
Yinsong Xu, Wei Jing, Liuxin Zhang +2
Understanding long-form egocentric videos remains challenging for multimodal large language models (MLLMs) due to limited context length and insufficient grounding of fine-grained…
cs.CV2026
ConsDreamer: Advancing Multi-View Consistency for Zero-Shot Text-to-3D Generation
Yuan Zhou, Shilong Jin, Litao Hua +3
Recent advances in zero-shot text-to-3D generation have revolutionized 3D content creation by enabling direct synthesis from textual descriptions. While state-of-the-art methods le…