4 papers
When LRP Diverges from Leave-One-Out in Transformers
Weiqiu You, Siqi Zeng, Yao-Hung Hubert Tsai +2
Leave-One-Out (LOO) provides an intuitive measure of feature importance but is computationally prohibitive. While Layer-Wise Relevance Propagation (LRP) offers a potentially effici…
Who is In Charge? Dissecting Role Conflicts in Instruction Following
Siqi Zeng
Large language models should follow hierarchical instructions where system prompts override user inputs, yet recent work shows they often ignore this rule while strongly obeying so…
MergeBench: A Benchmark for Merging Domain-Specialized LLMs
Yifei He, Siqi Zeng, Yuzheng Hu +3
Model merging provides a scalable alternative to multi-task training by combining specialized finetuned models through parameter arithmetic, enabling efficient deployment without t…
Learning Structured Representations with Hyperbolic Embeddings
Aditya Sinha, Siqi Zeng, Makoto Yamada +1
Most real-world datasets consist of a natural hierarchy between classes or an inherent label structure that is either already available or can be constructed cheaply. However, most…