activity
20242026
collaborators

6 papers

cs.CV2026

Towards Artwork Explanation in Large-scale Vision Language Models

Kazuki Hayashi, Yusuke Sakai, Hidetaka Kamigaito +2

Large-scale Vision-Language Models (LVLMs) output text from images and instructions, demonstrating capabilities in text generation and comprehension. However, it has not been clari…

cs.CL2025

Considering Length Diversity in Retrieval-Augmented Summarization

Juseon-Do, Jaesung Hwang, Jingun Kwon +2

This study investigates retrieval-augmented summarization by specifically examining the impact of exemplar summary lengths under length constraints, not covered by previous work. W…

cs.CL2025

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models

Utkarsh Tiwari, Aryan Seth, Adi Mukherjee +3

We introduce DebateBench, a novel dataset consisting of an extensive collection of transcripts and metadata from some of the world's most prestigious competitive debates. The datas…

cs.CL2025

ARISE: Iterative Rule Induction and Synthetic Data Generation for Text Classification

Yashwanth M., Vaibhav Singh, Ayush Maheshwari +2

We propose ARISE, a framework that iteratively induces rules and generates synthetic data for text classification. We combine synthetic data generation and automatic rule induction…

cs.CL2024

GPT-4o System Card

OpenAI, :, Aaron Hurst +416

GPT-4o is an autoregressive omni model that accepts as input any combination of text, audio, image, and video, and generates any combination of text, audio, and image outputs. It's…

cs.CL2024

Losing Visual Needles in Image Haystacks: Vision Language Models are Easily Distracted in Short and Long Contexts

Aditya Sharma, Michael Saxon, William Yang Wang

We present LoCoVQA, a dynamic benchmark generator for evaluating long-context extractive reasoning in vision language models (VLMs). LoCoVQA augments test examples for mathematical…