activity
20142024
most citedGPT-4V(ision) as a Generalist Evaluator for Vision-Language Tasks

13 citations · 48 across the 20 of their papers we have counts for

collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2024

RAG-QA Arena: Evaluating Domain Robustness for Long-form Retrieval Augmented Question Answering

Rujun Han, Yuhao Zhang, Peng Qi +6

Question answering based on retrieval augmented generation (RAG-QA) is an important research topic in NLP and has a wide range of real-world applications. However, most existing da…

cs.CL2024

Towards Safety and Helpfulness Balanced Responses via Controllable Large Language Models

Yi-Lin Tuan, Xilun Chen, Eric Michael Smith +5

As large language models (LLMs) become easily accessible nowadays, the trade-off between safety and helpfulness can significantly impact user experience. A model that prioritizes s…

cs.CL2024

Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts

Michael Saxon, Yiran Luo, Sharon Levy +3

Benchmarks of the multilingual capabilities of text-to-image (T2I) models compare generated images prompted in a test language to an expected image distribution over a concept set.…

cs.CL2023

ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language Models

Alex Mei, Sharon Levy, William Yang Wang

As large language models are integrated into society, robustness toward a suite of prompts is increasingly important to maintain reliability in a high-variance environment.Robustne…

cs.CL20233 cited

Zero-Shot Detection of Machine-Generated Codes

Xianjun Yang, Kexun Zhang, Haifeng Chen +3

This work proposes a training-free approach for the detection of LLMs-generated codes, mitigating the risks associated with their indiscriminate usage. To the best of our knowledge…

cs.CL20231 cited

Multilingual Conceptual Coverage in Text-to-Image Models

Michael Saxon, William Yang Wang

We propose "Conceptual Coverage Across Languages" (CoCo-CroLa), a technique for benchmarking the degree to which any generative text-to-image system provides multilingual parity to…