most citedTencentLLMEval: A Hierarchical Evaluation of Real-World Capabilities for Human-Aligned LLMs

2 citations · 3 across the 2 of their papers we have counts for

collaborators

7 papers

cs.CL2024

Fact-and-Reflection (FaR) Improves Confidence Calibration of Large Language Models

Xinran Zhao, Hongming Zhang, Xiaoman Pan +4

For a LLM to be trustworthy, its confidence level should be well-calibrated with its actual performance. While it is now common sense that LLM performances are greatly impacted by…

cs.LG2024

Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment

Rui Yang, Xiaoman Pan, Feng Luo +4

We consider the problem of multi-objective alignment of foundation models with human preferences, which is a critical step towards helpful and harmless AI systems. However, it is g…

cs.CL20241 cited

InFoBench: Evaluating Instruction Following Ability in Large Language Models

Yiwei Qin, Kaiqiang Song, Yebowen Hu +7

This paper introduces the Decomposed Requirements Following Ratio (DRFR), a new metric for evaluating Large Language Models' (LLMs) ability to follow instructions. Addressing a gap…

cs.CL2023

Zebra: Extending Context Window with Layerwise Grouped Local-Global Attention

Kaiqiang Song, Xiaoyang Wang, Sangwoo Cho +2

This paper introduces a novel approach to enhance the capabilities of Large Language Models (LLMs) in processing and understanding extensive text sequences, a critical aspect in ap…

cs.CL20232 cited

TencentLLMEval: A Hierarchical Evaluation of Real-World Capabilities for Human-Aligned LLMs

Shuyi Xie, Wenlin Yao, Yong Dai +11

Large language models (LLMs) have shown impressive capabilities across various natural language tasks. However, evaluating their alignment with human preferences remains a challeng…

cs.CL2023

MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning

Fuxiao Liu, Xiaoyang Wang, Wenlin Yao +5

With the rapid development of large language models (LLMs) and their integration into large multimodal models (LMMs), there has been impressive progress in zero-shot completion of…