collaborators

6 papers

cs.CR2025

CryptoX : Compositional Reasoning Evaluation of Large Language Models

Jiajun Shi, Chaoren Wei, Liqun Yang +7

The compositional reasoning capacity has long been regarded as critical to the generalization and intelligence emergence of large language models LLMs. However, despite numerous re…

cs.CL2024

TEGEE: Task dEfinition Guided Expert Ensembling for Generalizable and Few-shot Learning

Xingwei Qu, Yiming Liang, Yucheng Wang +10

Large Language Models (LLMs) exhibit the ability to perform in-context learning (ICL), where they acquire new tasks directly from examples provided in demonstrations. This process…

cs.AI2024

Read to Play (R2-Play): Decision Transformer with Multimodal Game Instruction

Yonggang Jin, Ge Zhang, Hao Zhao +7

Developing a generalist agent is a longstanding objective in artificial intelligence. Previous efforts utilizing extensive offline datasets from various tasks demonstrate remarkabl…

cs.CL2024

The Fine Line: Navigating Large Language Model Pretraining with Down-streaming Capability Analysis

Chen Yang, Junzhuo Li, Xinyao Niu +11

Uncovering early-stage metrics that reflect final model performance is one core principle for large-scale pretraining. The existing scaling law demonstrates the power-law correlati…

cs.SD2024

MuPT: A Generative Symbolic Music Pretrained Transformer

Xingwei Qu, Yuelin Bai, Yinghao Ma +25

In this paper, we explore the application of Large Language Models (LLMs) to the pre-training of music. While the prevalent use of MIDI in music modeling is well-established, our f…

cs.CL2024

StructLM: Towards Building Generalist Models for Structured Knowledge Grounding

Alex Zhuang, Ge Zhang, Tianyu Zheng +7

Structured data sources, such as tables, graphs, and databases, are ubiquitous knowledge sources. Despite the demonstrated capabilities of large language models (LLMs) on plain tex…