Showing 2025Show all
3 papers · 1 filter
cs.CL2025
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
Sai Krishna Mendu, Harish Yenala, Aditi Gulati +2
Large language models (LLMs) have become integral to various real-world applications, leveraging massive, web-sourced datasets like Common Crawl, C4, and FineWeb for pretraining. W…
cs.CL2025
SCULPT: Systematic Tuning of Long Prompts
Shanu Kumar, Akhila Yesantarao Venkata, Shubhanshu Khandelwal +3
Prompt optimization is essential for effective utilization of large language models (LLMs) across diverse tasks. While existing optimization methods are effective in optimizing sho…
cs.CL2025
READ: Reinforcement-based Adversarial Learning for Text Classification with Limited Labeled Data
Rohit Sharma, Shanu Kumar, Avinash Kumar
Pre-trained transformer models such as BERT have shown massive gains across many text classification tasks. However, these models usually need enormous labeled data to achieve impr…