3 papers
cs.LG2026
How Much Information Can a Vision Token Hold? A Scaling Law for Recognition Limits in VLMs
Shuxin Zhuang, Zi Liang, Runsheng Yu +4
Recent vision-centric approaches have made significant strides in long-context modeling. Represented by DeepSeek-OCR, these models encode rendered text into continuous vision token…
cs.LG2025
Particle Dynamics for Latent-Variable Energy-Based Models
Shiqin Tang, Shuxin Zhuang, Rong Feng +3
Latent-variable energy-based models (LVEBMs) assign a single normalized energy to joint pairs of observed data and latent variables, offering expressive generative modeling while c…
cs.LG2024
Tailed Low-Rank Matrix Factorization for Similarity Matrix Completion
Changyi Ma, Runsheng Yu, Xiao Chen +1
Similarity matrix serves as a fundamental tool at the core of numerous downstream machine-learning tasks. However, missing data is inevitable and often results in an inaccurate sim…