3 papers
cs.IR2026
SODA: Semantic-Oriented Distributional Alignment for Generative Recommendation
Ziqi Xue, Dingxian Wang, Yimeng Bai +7
Generative recommendation has emerged as a scalable alternative to traditional retrieve-and-rank pipelines by operating in a compact token space. However, existing methods mainly r…
cs.AI2025
UpBench: A Dynamically Evolving Real-World Labor-Market Agentic Benchmark Framework Built for Human-Centric AI
Darvin Yi, Teng Liu, Mattie Terzolo +4
As large language model (LLM) agents increasingly undertake digital work, reliable frameworks are needed to evaluate their real-world competence, adaptability, and capacity for hum…
cs.LG2024
Evaluating Cost-Accuracy Trade-offs in Multimodal Search Relevance Judgements
Silvia Terragni, Hoang Cuong, Joachim Daiber +2
Large Language Models (LLMs) have demonstrated potential as effective search relevance evaluators. However, there is a lack of comprehensive guidance on which models consistently p…