7 papers
Deployable Per-Instance Multi-Layer Activation Steering for Large Language Models
Muhammad Faishal Adly Nelwan, Alfan Farizki Wicaksono
Activation steering edits the behaviour of a frozen language model by adding a learned vector to its residual stream, and current practice fixes the injection layers globally per t…
Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection
Muhammad Alif Al Hakim, Alfan Farizki Wicaksono, Fajri Koto
Quantization is widely adopted to reduce the computational cost of large language models (LLMs); however, its implications for fairness and safety, particularly in dynamic quantiza…
Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages
Lyzander Marciano Andrylie, Inaya Rahmanisa, Mahardika Krisna Ihsani +3
Understanding the multilingual mechanisms of large language models (LLMs) provides insight into how they process different languages, yet this remains challenging. Existing studies…
Unveiling the Influence of Amplifying Language-Specific Neurons
Inaya Rahmanisa, Lyzander Marciano Andrylie, Mahardika Krisna Ihsani +3
Language-specific neurons in LLMs that strongly correlate with individual languages have been shown to influence model behavior by deactivating them. However, their role in amplifi…
IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages
Muhammad Falensi Azmi, Muhammad Dehan Al Kautsar, Alfan Farizki Wicaksono +1
Although region-specific large language models (LLMs) are increasingly developed, their safety remains underexplored, particularly in culturally diverse settings like Indonesia, wh…
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
Naila Shafirni Hidayat, Muhammad Dehan Al Kautsar, Alfan Farizki Wicaksono +1
The performance of large language models (LLMs) continues to improve, as reflected in rising scores on standard benchmarks. However, the lack of transparency around training data r…