activity
20242026
collaborators

5 papers

cs.CL2026

Towards Compositional Generalization of LLMs via Skill Taxonomy Guided Data Synthesis

Yifan Wei, Li Du, Xiaoyan Yu +2

Large Language Models (LLMs) and agent-based systems often struggle with compositional generalization due to a data bottleneck in which complex skill combinations follow a long-tai…

stat.ML2025

GeoERM: Geometry-Aware Multi-Task Representation Learning on Riemannian Manifolds

Aoran Chen, Yang Feng

Multi-Task Learning (MTL) seeks to boost statistical power and learning efficiency by discovering structure shared across related tasks. State-of-the-art MTL representation methods…

cs.CL2025

MoCE: Adaptive Mixture of Contextualization Experts for Byte-based Neural Machine Translation

Langlin Huang, Mengyu Bu, Yang Feng

Byte-based machine translation systems have shown significant potential in massively multilingual settings. Unicode encoding, which maps each character to specific byte(s), elimina…

cs.LG2024

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration

Zhuofan Wen, Shangtong Gui, Yang Feng

Inference acceleration of large language models (LLMs) has been put forward in many application scenarios and speculative decoding has shown its advantage in addressing inference a…

cs.AI2024

Hokoff: Real Game Dataset from Honor of Kings and its Offline Reinforcement Learning Benchmarks

Yun Qu, Boyuan Wang, Jianzhun Shao +15

The advancement of Offline Reinforcement Learning (RL) and Offline Multi-Agent Reinforcement Learning (MARL) critically depends on the availability of high-quality, pre-collected o…