1 paper
Yupeng Chen, Junchi Yu, Aoxi Liu +3
Recent advances in end-to-end trained omni-models have substantially improved audio capabilities by strengthening text-audio modality alignment. However, whether such alignment ina…