Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
InterSketch: An Interleaved Reasoning Model with Self-correcting Visual Sketch and Stepwise Reward
Zhiwei Ning, Wenwen Tong, Xiangli Kong +12
While vision-language models (VLMs) have exhibited multi-turn visual reasoning capabilities, their reasoning trajectories remain relatively shallow and are dominated by a text-cent…
cs.CV2026
GeoSym127K: Scalable Symbolically-verifiable Synthesis for Multimodal Geometric Reasoning
Jinhao Jing, Zheng Ma, Jinwei Liang +9
Large Multimodal Models (LMMs) often struggle with geometric reasoning due to visual hallucinations and a lack of mathematically precise Chain-of-Thought (CoT) data. To address thi…
cs.CV2025
Towards Fine-Grained Recognition with Large Visual Language Models: Benchmark and Optimization Strategies
Cong Pang, Hongtao Yu, Zixuan Chen +2
Large Vision Language Models (LVLMs) have made remarkable progress, enabling sophisticated vision-language interaction and dialogue applications. However, existing benchmarks prima…