1 citations · 3 across the 5 of their papers we have counts for
5 papers
Halu-J: Critique-Based Hallucination Judge
Binjie Wang, Steffi Chern, Ethan Chern +1
Large language models (LLMs) frequently generate non-factual content, known as hallucinations. Existing retrieval-augmented-based hallucination detection approaches typically addre…
ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
Ethan Chern, Jiadi Su, Yan Ma +1
Previous open-source large multimodal models (LMMs) have faced several limitations: (1) they often lack native integration, requiring adapters to align visual representations with…
BeHonest: Benchmarking Honesty in Large Language Models
Steffi Chern, Zhulin Hu, Yuqing Yang +5
Previous works on Large Language Models (LLMs) have mainly focused on evaluating their helpfulness or harmlessness. However, honesty, another crucial alignment criterion, has recei…
Reformatted Alignment
Run-Ze Fan, Xuefeng Li, Haoyang Zou +5
The quality of finetuning data is crucial for aligning large language models (LLMs) with human values. Current methods to improve data quality are either labor-intensive or prone t…
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
Steffi Chern, Ethan Chern, Graham Neubig +1
Despite the utility of Large Language Models (LLMs) across a wide range of tasks and scenarios, developing a method for reliably evaluating LLMs across varied contexts continues to…