3 papers
cs.CL2026
The Attribution-Compression Frontier in Retrieval-Augmented Generation
Deepanshu Mody
Context compression reduces generator input in retrieval-augmented generation, but answer quality alone does not characterize citation attribution. We measure citation attribution…
cs.LG2026
Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal +1
Activation steering controls model behavior by editing internal activations at inference time. We study its input-side dual: optimizing a fluent prompt so that a chosen internal la…
cs.CL2026
Joint Optimization for Greedy Longest-match Tokenization
Adhiraj Singh, Deepanshu Mody, Ghina Al Shdaifat +4
Recent work has shown that subword vocabularies can be trained to optimize compression for a specific inference rule rather than relying on greedy heuristics such as Byte Pair Enco…