1 paper
Haojun Zhang, Yi Zou, Min Chen +7
Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause d…