2 papers
cs.LG2026
Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization
Joesph An, Joseph An, Phillip Keung +3
Audio language models process input audio into rich frame-level representations, but the standard approach to temporal localization generates timestamps as sequences of text tokens…
cs.CL2024
Summarization-Based Document IDs for Generative Retrieval with Language Models
Haoxin Li, Daniel Cheng, Phillip Keung +2
Generative retrieval (Wang et al., 2022; Tay et al., 2022) is a popular approach for end-to-end document retrieval that directly generates document identifiers given an input query…