1 paper · 1 filter
Yu Wu, Shuo Wu, Ye Tao +2
Standard LLM cascades improve efficiency by deferring difficult queries from weak to strong models. However, these systems are typically static: when faced with repeated or semanti…