4 papers · 1 filter
MUSE: A Unified Agentic Harness for MLLMs
Jianglin Lu, Hailing Wang, Xu Ma +4
Despite rapid progress, multimodal large language models (MLLMs) still fail on tasks that humans solve effortlessly, such as navigating a grid maze from a screenshot or selecting t…
D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition
Yiyang Huang, Yizhou Wang, Yun Fu
Video large language models (Vid-LLMs), which excel in diverse video-language tasks, can be effectively constructed by adapting image-pretrained vision-language models (VLMs). Howe…
Towards Zero-shot 3D Anomaly Localization
Yizhou Wang, Kuan-Chuan Peng, Yun Fu
3D anomaly detection and localization is of great significance for industrial inspection. Prior 3D anomaly detection and localization methods focus on the setting that the testing…
Don't Judge by the Look: Towards Motion Coherent Video Representation
Yitian Zhang, Yue Bai, Huan Wang +2
Current training pipelines in object recognition neglect Hue Jittering when doing data augmentation as it not only brings appearance changes that are detrimental to classification,…