1 paper · 1 filter
Liu Qing, Ou Wu, Yi Du
Token selection is pivotal for effective LLM post-training. However, existing methods mostly rely on local heuristics and rarely formulate token selection as a principled valuation…