Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Beyond Tokens: Semantic-Aware Speculative Decoding for Efficient Inference by Probing Internal States
Ximing Dong, Shaowei Wang, Dayi Lin +2
Large Language Models (LLMs) achieve strong performance across many tasks but suffer from high inference latency due to autoregressive decoding. The issue is exacerbated in Large R…
cs.CL2024
A Framework for Real-time Safeguarding the Text Generation of Large Language Model
Ximing Dong, Dayi Lin, Shaowei Wang +1
Large Language Models (LLMs) have significantly advanced natural language processing (NLP) tasks but also pose ethical and societal risks due to their propensity to generate harmfu…