5 papers
Marginal Advantage Accumulation for Memory-Driven Agent Self-Evolution
Mingyu Yang, Keye Zheng, Congchao Cheng +4
In batch-style trace distillation, the same memory operation may receive contradictory feedback across different batches. Existing methods lack a cross-batch, operation-level evide…
LISA: Linear-Indexed Sparse Attention for Efficient Long-Context Reasoning
Yu Zhao, Zekun Zhang, Fan Jiang +6
Recent advances in long chain-of-thought reasoning models such as DeepSeek-R1 have led to increasingly longer inference context lengths under the test-time scaling paradigm. Howeve…
Marco-MoE: Open Multilingual Mixture-of-Expert Language Models with Efficient Upcycling
Fan Jiang, Yu Zhao, Chenyang Lyu +5
We present Marco-MoE, a suite of fully open multilingual sparse Mixture-of-Experts (MoE) models. Marco-MoE features a highly sparse design in which only around 5\% of the total par…
CulturALL: Benchmarking Multilingual and Multicultural Competence of LLMs on Grounded Tasks
Peiqin Lin, Chenyang Lyu, Wenjiang Luo +22
Large language models (LLMs) are now deployed worldwide, inspiring a surge of benchmarks that measure their multilingual and multicultural abilities. However, these benchmarks prio…
Language Bias in Information Retrieval: The Nature of the Beast and Mitigation Methods
Jinrui Yang, Fan Jiang, Timothy Baldwin
Language fairness in multilingual information retrieval (MLIR) systems is crucial for ensuring equitable access to information across diverse languages. This paper sheds light on t…