10 papers
Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker
Zongjian Wu, Lei Zhang
Referring expression comprehension (REC) aims to localize a target object within an image based on a given expression. Although recent advances in vision-language models have led t…
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
Shihao Wang, Shilong Liu, Yuanguo Kuang +10
Vision-language models (VLMs) commonly formulate visual grounding and detection as a coordinate-token generation problem, serializing each 2D box into multiple 1D tokens that are l…
Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild
Mao Zheng, Zheng Li, Tao Chen +10
Hy-MT2 is a family of fast-thinking multilingual translation models designed for complex real-world scenarios. It includes three model sizes: 1.8B, 7B, and 30B-A3B (MoE), all of wh…
Weighted Reverse Convolution for Feature Upsampling
Wentong Li, Zhiyuan Qi, Zichen Zhao +2
Pre-trained vision foundation models (VFMs) provide strong semantic representations, yet their patch-level features are inherently coarse, limiting their effectiveness on tasks req…
DEL: Digit Entropy Loss for Numerical Learning of Large Language Models
Zhaohui Zheng, Chenhang He, Shihao Wang +3
Number prediction stands as a fundamental capability of large language models (LLMs) in mathematical problem-solving and code generation. The widely adopted maximum likelihood esti…
Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models
Yi Zhong, Haotong Qin, Xindong Zhang +2
Low-bit post-training quantization (PTQ) is a pivotal technique for deploying Vision-Language Models (VLMs) on resource-constrained devices. However, existing PTQ methods often deg…