8 papers
Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods
Xingsong Ye, Yongkun Du, Jiaxin Zhang +5
WordArt (artistic text) features highly customized fonts, textures, and layouts, making WordArt-oriented scene TExt Recognition (WATER) substantially more challenging than general…
MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation
Haojie Zhang, Di Wu, Bingyan Liu +5
While video foundation models excel at single-shot generation, real-world cinematic storytelling inherently relies on complex multi-shot sequencing. Further progress is constrained…
Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models
Yuancheng Wei, Linli Yao, Lei Li +4
Multimodal reward models have advanced substantially in text and image domains, yet progress in video understanding reward modeling remains severely limited by the lack of robust e…
MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation
Shichao Kan, Xuyang Zhang, Haojie Zhang +7
Evaluating image captions without references remains challenging because global embedding similarity often misses fine-grained mismatches such as hallucinated objects, missing attr…
DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning
Yuancheng Wei, Haojie Zhang, Linli Yao +7
Image Difference Captioning (IDC) generates natural language descriptions that precisely identify differences between two images, serving as a key benchmark for fine-grained change…
Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
Yongyi Su, Haojie Zhang, Shijie Li +11
Multimodal large language models (MLLMs) have advanced rapidly in recent years. However, existing approaches for vision tasks often rely on indirect representations, such as genera…