Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Fractional Rotation, Full Potential? Investigating Performance and Convergence of Partial RoPE
Mohammad Aflah Khan, Krishna P. Gummadi, Manish Gupta +1
Rotary Positional Embedding (RoPE) is a common choice in transformer architectures for encoding relative positional information. Although earlier work has examined omitting RoPE in…
cs.LG2025
Rethinking Memorization Measures and their Implications in Large Language Models
Bishwamittra Ghosh, Soumi Das, Qinyuan Wu +4
Concerned with privacy threats, memorization in LLMs is often seen as undesirable, specifically for learning. In this paper, we study whether memorization can be avoided when optim…