2 papers
cs.CV2026
CodeShrink: Adaptive Visual Compression for Efficient Multimodal Code Understanding
Wenxin Tang, Jingyu Xiao, Zhenyu Liu +6
Rendering source code as images offers a promising way to reduce the input costs of Multimodal Large Language Models (MLLMs). Adjusting image resolution can trade visual token cost…
cs.CL2025
Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models
Yang Zhang, Yu Yu, Bo Tang +8
With the rapid development of Large Language Models (LLMs), aligning these models with human preferences and values is critical to ensuring ethical and safe applications. However,…