2 papers
cs.LG2026
Judge Model for Large-scale Multimodality Benchmarks
Min-Han Shih, Yu-Hsin Wu, Yu-Wei Chen
We propose a dedicated multimodal Judge Model designed to provide reliable, explainable evaluation across a diverse suite of tasks. Our benchmark spans text, audio, image, and vide…
cs.CL2024
SpeechCaps: Advancing Instruction-Based Universal Speech Models with Multi-Talker Speaking Style Captioning
Chien-yu Huang, Min-Han Shih, Ke-Han Lu +2
Instruction-based speech processing is becoming popular. Studies show that training with multiple tasks boosts performance, but collecting diverse, large-scale tasks and datasets i…