4 papers
Automatic Teaching Platform on Vision Language Retrieval Augmented Generation
Ruslan Gokhman, Jialu Li, Youshan Zhang
Automating teaching presents unique challenges, as replicating human interaction and adaptability is complex. Automated systems cannot often provide nuanced, real-time feedback tha…
SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing
Varun Biyyala, Bharat Chanderprakash Kathuria, Jialu Li +1
Video editing models have advanced significantly, but evaluating their performance remains challenging. Traditional metrics, such as CLIP text and image scores, often fall short: t…
Leapfrog Latent Consistency Model (LLCM) for Medical Images Generation
Lakshmikar R. Polamreddy, Kalyan Roy, Sheng-Han Yueh +4
The scarcity of accessible medical image data poses a significant obstacle in effectively training deep learning models for medical diagnosis, as hospitals refrain from sharing the…
SparrowVQE: Visual Question Explanation for Course Content Understanding
Jialu Li, Manish Kumar Thota, Ruslan Gokhman +2
Visual Question Answering (VQA) research seeks to create AI systems to answer natural language questions in images, yet VQA methods often yield overly simplistic and short answers.…