3 papers
cs.AI2025
Memory-QA: Answering Recall Questions Based on Multimodal Memories
Hongda Jiang, Xinyuan Zhang, Siddhant Garg +10
We introduce Memory-QA, a novel real-world task that involves answering recall questions about visual content from previously stored multimodal memories. This task poses unique cha…
cs.CV2025
OmniCam: Unified Multimodal Video Generation via Camera Control
Xiaoda Yang, Jiayang Xu, Kaixuan Luan +9
Camera control, which achieves diverse visual effects by changing camera position and pose, has attracted widespread attention. However, existing methods face challenges such as co…
cs.CV2025
Astrea: A MOE-based Visual Understanding Model with Progressive Alignment
Xiaoda Yang, JunYu Lu, Hongshun Qiu +12
Vision-Language Models (VLMs) based on Mixture-of-Experts (MoE) architectures have emerged as a pivotal paradigm in multimodal understanding, offering a powerful framework for inte…