2 papers
cs.CL2026
Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty
Sravanthi Machcha, Sushrita Yerra, Sahil Gupta +4
Current evaluation of large language models (LLMs) overwhelmingly prioritizes accuracy; however, in real-world and safety-critical applications, the ability to abstain when uncerta…
cs.CL2024
LongLaMP: A Benchmark for Personalized Long-form Text Generation
Ishita Kumar, Snigdha Viswanathan, Sushrita Yerra +11
Long-text generation is seemingly ubiquitous in real-world applications of large language models such as generating an email or writing a review. Despite the fundamental importance…