activity
20242026
collaborators

5 papers

cs.CL2026

Positional Encoding via Token-Aware Phase Attention

Yu Wang, Sheng Shen, Rémi Munos +2

We prove under practical assumptions that Rotary Positional Embedding (RoPE) introduces an intrinsic distance-dependent bias in attention scores that limits RoPE's ability to model…

cs.SE2026

Learning to Solve and Verify: A Self-Play Framework for Code and Test Generation

Zi Lin, Sheng Shen, Ilia Kulikov +3

Recent advances in large language models (LLMs) have improved their performance on coding benchmarks. However, improvement is plateauing due to the exhaustion of readily available…

cs.CV2026

Faster and Better Alignment for Flow Matching Models via Step-aware Advantages

Zhixiong Yue, Zixuan Ni, Feiyang Ye +4

Recent advances in flow matching models, particularly with reinforcement learning (RL), have significantly enhanced human preference alignment in few-step text-to-image generators.…

cs.CL2025

Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study

Menglong Cui, Pengzhi Gao, Wei Liu +2

Large language models (LLMs) have shown continuously improving multilingual capabilities, and even small-scale open-source models have demonstrated rapid performance enhancement. I…

cs.CV2024

BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions

Anas Awadalla, Le Xue, Manli Shu +13

We introduce BLIP3-KALE, a dataset of 218 million image-text pairs that bridges the gap between descriptive synthetic captions and factual web-scale alt-text. KALE augments synthet…