2 papers
cs.LG2026
Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection
Vijeta Deshpande, Tootiya Giyahchi, Veena Padmanabhan +2
Safety detection models require examples of HHH (Helpful, Harmless, Honest)-violating outputs for robust generalization, however such examples are scarce. Activation Steering (AS)…
cs.AI2026
Rethinking Evaluation for LLM Hallucination Detection: A Desiderata, A New RAG-based Benchmark, New Insights
Wenbo Chen, Veena Padmanabhan, Tootiya Giyahchi +2
Hallucination, broadly referring to unfaithful, fabricated, or inconsistent content generated by LLMs, has wide-ranging implications. Therefore, a large body of effort has been dev…