1 citations · 1 across the 3 of their papers we have counts for
12 papers
Rethinking Token Reduction for Large Vision-Language Models
Yi Wang, Haofei Zhang, Qihan Huang +7
Large Vision-Language Models (LVLMs) excel in visual understanding and reasoning, but the excessive visual tokens lead to high inference costs. Although recent token reduction meth…
Syn-GRPO: Self-Evolving Data Synthesis for MLLM Perception Reasoning
Qihan Huang, Haofei Zhang, Rong Wei +4
RL (reinforcement learning) methods (e.g., GRPO) for MLLM (Multimodal LLM) perception ability has attracted wide research interest owing to its remarkable generalization ability. N…
HuggingR: A Progressive Reasoning Framework for Discovering Optimal Model Companions
Shaoyin Ma, Chenggong Hu, Huiqiong Wang +3
Building effective LLM agents increasingly requires selecting appropriate AI models as tools from large open repositories (e.g., HuggingFace with > 2M models) based on natural lang…
ITFormer: Bridging Time Series and Natural Language for Multi-Modal QA with Large-Scale Multitask Dataset
Yilin Wang, Peixuan Lei, Jie Song +6
Time-series data are critical in diverse applications, such as industrial monitoring, medical diagnostics, and climate research. However, effectively integrating these high-dimensi…
Diffusion Model Quantization: A Review
Qian Zeng, Chenggong Hu, Mingli Song +1
Recent success of large text-to-image models has empirically underscored the exceptional performance of diffusion models in generative tasks. To facilitate their efficient deployme…
Sampling-Aware Quantization for Diffusion Models
Qian Zeng, Jie Song, Yuanyu Wan +2
Diffusion models have recently emerged as the dominant approach in visual generation tasks. However, the lengthy denoising chains and the computationally intensive noise estimation…