Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
DV-SFT: Direct Vision Supervision for Fine-Grained Visual Understanding
Jianfei Zhao, Feng Zhang, Xin Sun +3
Multimodal large language models are typically trained end-to-end to predict ground-truth answers, yet supervision signals are applied exclusively to text tokens. Visual tokens, th…
cs.CV2026
OneVision: An End-to-End Generative Framework for Multi-view E-commerce Vision Search
Zexin Zheng, Huangyu Dai, Lingtao Mao +8
Traditional vision search, similar to search and recommendation systems, follows the multi-stage cascading architecture (MCA) paradigm to balance efficiency and conversion. Specifi…
cs.CV2025
UniDGF: A Unified Detection-to-Generation Framework for Hierarchical Object Visual Recognition
Xinyu Nan, Lingtao Mao, Huangyu Dai +8
Achieving visual semantic understanding requires a unified framework that simultaneously handles object detection, category prediction, and attribute recognition. However, current…