43 citations · 80 across the 39 of their papers we have counts for
41 papers
Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs
Jingtan Wang, Arun Verma, Xiaoqiang Lin +4
How to divide a fixed annotation budget between supervised fine-tuning (SFT) and reinforcement learning (RL) during LLM post-training remains an open problem. Existing work charact…
CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia
Bryan Chen Zhengyu Tan, Weihua Zheng, Thong T. Doan +30
Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common use case: users seeking pract…
SagaQA: A Multi-hop Reasoning Benchmark for Long-form Narrative Understanding in TV Series
Galann Pennec, Zhengyuan Liu, Nicholas Asher +2
We introduce SagaQA, a long-form video benchmark for multi-hop reasoning over full-length TV series. Existing video reasoning benchmarks often emphasize local understanding of adja…
ARBOR: Online Process Rewards via a Reusable Rubric Buffer for Search Agents
Zheng Liu, Longxiang Zhang, Xintong Wang +8
LLM-based search agents are trained predominantly with outcome-only reward, leaving the search process itself unsupervised. This signal degenerates on outcome-homogeneous groups wh…
How Hard Can It Be? Hardness-Aware Multi-Objective Unlearning
Jiangwei Chen, Xinyuan Niu, Rachael Hwee Ling Sim +3
Machine unlearning aims to remove the influence of specific forget training data due to privacy, copyright or bias concerns while maintaining the model performance on the remaining…
SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures
Nedjma Ousidhoum, Junho Myung, Carla Perez-Almendros +27
We present our shared task on evaluating the adaptability of LLMs and NLP systems across multiple languages and cultures. The task data consist of an extended version of our manual…