2 citations · 2 across the 1 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025★ 2 cited
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Haoji Zhang, Xin Gu, Jiawen Li +7
The video reasoning ability of multimodal large language models (MLLMs) is crucial for downstream tasks like video question answering and temporal grounding. While recent approache…
cs.CV2020
ReLaText: Exploiting Visual Relationships for Arbitrary-Shaped Scene Text Detection with Graph Convolutional Networks
Chixiang Ma, Lei Sun, Zhuoyao Zhong +1
We introduce a new arbitrary-shaped text detection approach named ReLaText by formulating text detection as a visual relationship detection problem. To demonstrate the effectivenes…