2 papers
cs.CV2025
ESCA: Contextualizing Embodied Agents via Scene-Graph Generation
Jiani Huang, Amish Sethi, Matthew Kuo +6
Multi-modal large language models (MLLMs) are making rapid progress toward general-purpose embodied agents. However, existing MLLMs do not reliably capture fine-grained links betwe…
cs.CL2025
Optimal Multi-Task Learning at Regularization Horizon for Speech Translation Task
JungHo Jung, Junhyun Lee
End-to-end speech-to-text translation typically suffers from the scarcity of paired speech-text data. One way to overcome this shortcoming is to utilize the bitext data from the Ma…