2 papers
cs.AI2025
TripScore: Benchmarking and rewarding real-world travel planning with fine-grained evaluation
Yincen Qu, Huan Xiao, Feng Li +4
Travel planning is a valuable yet complex task that poses significant challenges even for advanced large language models (LLMs). While recent benchmarks have advanced in evaluating…
cs.CL2024
Deploying Multi-task Online Server with Large Language Model
Yincen Qu, Chao Ma, Xiangying Dai +3
In the industry, numerous tasks are deployed online. Traditional approaches often tackle each task separately by its own network, which leads to excessive costs for developing and…