5 papers
EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking
Enjun Du, Siyi Liu, Zirong Chen +8
Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to retain, an attribute to modify, and context to ignore. Yet exis…
VDE Bench: Evaluating The Capability of Image Editing Models to Modify Visual Documents
Hongzhu Yi, Yujia Yang, Yuanxiang Wang +18
In recent years, image editing models have made significant progress, enabling users to manipulate visual content in a flexible and interactive manner through natural language inst…
Omni IIE Bench: Benchmarking the Practical Capabilities of Image Editing Models
Yujia Yang, Yuanxiang Wang, Zhenyu Guan +11
While Instruction-based Image Editing (IIE) has achieved significant progress, existing benchmarks pursue task breadth via mixed evaluations. This paradigm obscures a critical fail…
FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
Yuxuan Cai, Xiaozhuan Liang, Xinghua Wang +7
As large language models (LLMs) become increasingly powerful, the sequential nature of autoregressive generation creates a fundamental throughput bottleneck that limits the practic…
ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts
Yuying Ge, Yixiao Ge, Chen Li +15
Real-world user-generated short videos, especially those distributed on platforms such as WeChat Channel and TikTok, dominate the mobile internet. However, current large multimodal…