1 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.CV2026
Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning
Zhengjian Yao, Yongzhi Li, Xinyuan Gao +3
We present "Narrative Weaver", a novel framework that addresses a fundamental challenge in generative AI: achieving multi-modal controllable, long-range, and consistent visual cont…
cs.CV2024★ 1 cited
POINTS1.5: Building a Vision-Language Model towards Real World Applications
Yuan Liu, Le Tian, Xiao Zhou +4
Vision-language models have made significant strides recently, demonstrating superior performance across a range of tasks, e.g. optical character recognition and complex diagram an…
cs.CL2023★ 1 cited
Ziya-Visual: Bilingual Large Vision-Language Model via Multi-Task Instruction Tuning
Junyu Lu, Dixiang Zhang, Xiaojun Wu +5
Recent advancements enlarge the capabilities of large language models (LLMs) in zero-shot image-to-text generation and understanding by integrating multi-modal inputs. However, suc…