2 papers
cs.CL2025
Visual-Aware Speech Recognition for Noisy Scenarios
Lakshmipathi Balaji, Karan Singla
Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automa…
cs.CL2024
News Reporter: A Multi-lingual LLM Framework for Broadcast T.V News
Tarun Jain, Yufei Gao, Sridhar Vanga +1
Large Language Models (LLMs) have fast become an essential tools to many conversational chatbots due to their ability to provide coherent answers for varied queries. Datasets used…