Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
TRUEBench: Can LLM Response Meet Real-world Constraints as Productivity Assistant?
Jiho Park, Jongyoon Song, Minjin Choi +3
Large language models (LLMs) are increasingly integral as productivity assistants, but existing benchmarks fall short in rigorously evaluating their real-world instruction-followin…
cs.CL2025
DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture
Arijit Maji, Raghvendra Kumar, Akash Ghosh +6
We introduce DRISHTIKON, a first-of-its-kind multimodal and multilingual benchmark centered exclusively on Indian culture, designed to evaluate the cultural understanding of genera…