activity
20242026
collaborators

6 papers

cs.LG2026

On the Residual Scaling of Looped Transformers: Stability and Transferability

Shaowen Wang, Bingrui Li, Ge Zhang +3

Looped (weight-tied) Transformers apply a shared residual block times (, same at each step), increasing effective depth without adding p…

cs.CL2025

II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language Models

Ziqiang Liu, Feiteng Fang, Xi Feng +23

The rapid advancements in the development of multimodal large language models (MLLMs) have consistently led to new breakthroughs on various benchmarks. In response, numerous challe…

cs.CL2024

COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning

Yuelin Bai, Xinrun Du, Yiming Liang +19

Remarkable progress on English instruction tuning has facilitated the efficacy and reliability of large language models (LLMs). However, there remains a noticeable gap in instructi…

cs.CL2024

Can MLLMs Understand the Deep Implication Behind Chinese Images?

Chenhao Zhang, Xi Feng, Yuelin Bai +18

As the capabilities of Multimodal Large Language Models (MLLMs) continue to improve, the need for higher-order capability evaluation of MLLMs is increasing. However, there is a lac…

cs.CV2024

LIME: Less Is More for MLLM Evaluation

King Zhu, Qianbo Zang, Shian Jia +18

Multimodal Large Language Models (MLLMs) are evaluated on various benchmarks, such as image captioning, visual question answering, and reasoning. However, many of these benchmarks…

cs.CL2024

PsychoGAT: A Novel Psychological Measurement Paradigm through Interactive Fiction Games with LLM Agents

Qisen Yang, Zekun Wang, Honghui Chen +6

Psychological measurement is essential for mental health, self-understanding, and personal development. Traditional methods, such as self-report scales and psychologist interviews,…