2 papers
cs.LG2026
MMTM: Tri-Modal Topic Modeling for Long-Form Video via Similarity-Gated Fusion
Ali Abusaleh, Bhuvanesh Verma, Alexander Mehler
We introduce MMTM, a modular pipeline for topic discovery in long-form video that integrates speech recognition, audio and visual embeddings, and BERTopic clustering through a dete…
cs.CL2026
Large Language Models Decide Early and Explain Later
Ayan Datta, Zhixue Zhao, Bhuvanesh Verma +3
Large Language Models often achieve strong performance by generating long intermediate chain-of-thought reasoning. However, it remains unclear when a model's final answer is actual…