13 papers
Omni2LoRA: Coherence-Preserving Parametric Memory for Efficient Omni Language Models
Puneet Mathur, Manan Suri, Dinesh Manocha
Omnimodal language models (OLMs) enable unified audio-visual understanding, but processing long joint token sequences makes inference computationally prohibitive. While recent toke…
CyberForge: Verified Vulnerability Injection at Repository Level for Cybersecurity Agent Training
Amine Lbath, Manan Suri, Aurelien Delaitre +4
Despite recent advances, frontier large language model (LLM) agents remain limited in discovering and patching complex vulnerabilities in real-world software. Generally available a…
Frames2LoRA: Parametric Video Internalization for Vision-Language Models
Manan Suri, Sarvesh Baskar, Dinesh Manocha
Processing video in vision-language models is expensive: each frame occupies hundreds of tokens, and inference cost scales with every frame and every repeated query. We introduce F…
Platform architecture determines whether recommendation algorithms can shape information quality on social media
Mohammad Hammas Saeed, David A. Broniatowski, Joseph Simons +3
Social media platforms shape public discourse through two fundamental design choices that naturally co-occur in any field investigation: platform architecture, which defines what t…
DIAGRAMS: A Review Framework for Reasoning-Level Attribution in Diagram QA
Anirudh Iyengar Kaniyar Narayana Iyengar, Tampu Ravi Kumar, Manan Suri +4
Diagram question answering (Diagram QA) requires reasoning-level attribution that links each question-answer pair to all visual regions needed to derive the answer, rather than onl…
DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams
Anirudh Iyengar Kaniyar Narayana Iyengar, Tampu Ravi Kumar, Gaurav Najpande +4
Diagram question answering (DQA) requires models to interpret structured visual representations such as charts, maps, infographics, circuit schematics, and scientific diagrams. Rec…