Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
SamasÄmayik: A Parallel Dataset for Hindi-Sanskrit Machine Translation
N J Karthika, Keerthana Suryanarayanan, Jahanvi Purohit +3
We release SamasÄmayik, a novel, meticulously curated, large-scale Hindi-Sanskrit corpus, comprising 92,196 parallel sentences. Unlike most data available in Sanskrit, which focus…
cs.CL2025
Consistency Is the Key: Detecting Hallucinations in LLM Generated Text By Checking Inconsistencies About Key Facts
Raavi Gupta, Pranav Hari Panicker, Sumit Bhatia +1
Large language models (LLMs), despite their remarkable text generation capabilities, often hallucinate and generate text that is factually incorrect and not grounded in real-world…
cs.CL2024
SMART: Submodular Data Mixture Strategy for Instruction Tuning
H S V N S Kowndinya Renduchintala, Sumit Bhatia, Ganesh Ramakrishnan
Instruction Tuning involves finetuning a language model on a collection of instruction-formatted datasets in order to enhance the generalizability of the model to unseen tasks. Stu…