17 citations · 19 across the 6 of their papers we have counts for
7 papers
OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning
Kaihang Pan, Qi Tian, Jianwei Zhang +11
While proprietary systems such as Seedance-2.0 have achieved remarkable success in omni-capable video generation, open-source alternatives significantly lag behind. Most academic m…
Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis
Shuang Chen, Quanxin Shou, Hangting Chen +16
Unified multimodal models provide a natural and promising architecture for understanding diverse and complex real-world knowledge while generating high-quality images. However, the…
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation
Shi-Xue Zhang, Hongfa Wang, Duojun Huang +3
Video captions play a crucial role in text-to-video generation tasks, as their quality directly influences the semantic coherence and visual fidelity of the generated videos. Altho…
Transformer-based Reasoning for Learning Evolutionary Chain of Events on Temporal Knowledge Graph
Zhiyu Fang, Shuai-Long Lei, Xiaobin Zhu +4
Temporal Knowledge Graph (TKG) reasoning often involves completing missing factual elements along the timeline. Although existing methods can learn good embeddings for each factual…
Scene Text Recognition with Single-Point Decoding Network
Lei Chen, Haibo Qin, Shi-Xue Zhang +2
In recent years, attention-based scene text recognition methods have been very popular and attracted the interest of many researchers. Attention-based methods can adaptively focus…
Adaptive Boundary Proposal Network for Arbitrary Shape Text Detection
Shi-Xue Zhang, Xiaobin Zhu, Chun Yang +2
Arbitrary shape text detection is a challenging task due to the high complexity and variety of scene texts. In this work, we propose a novel adaptive boundary proposal network for…