activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation

Shuowei Li, Yuming Zhao, Parth Bhalerao +1

Text-to-video (T2V) generation has rapidly progressed in visual fidelity, yet its ability to faithfully represent multiple cultures within a single prompt remains underexplored. We…

cs.CV2026

ChartE: A Comprehensive Benchmark for End-to-End Chart Editing

Shuo Li, Jiajun Sun, Zhekai Wang +9

Charts are a fundamental visualization format for structured data analysis. Enabling end-to-end chart editing according to user intent is of great practical value, yet remains chal…

cs.CV2025

The Role of Entropy in Visual Grounding: Analysis and Optimization

Shuo Li, Jiajun Sun, Zhihao Zhang +10

Recent advances in fine-tuning multimodal large language models (MLLMs) using reinforcement learning have achieved remarkable progress, particularly with the introduction of variou…

cs.CV2025

Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbations

Shuo Li, Jiajun Sun, Guodong Zheng +10

Recently, multimodal large language models (MLLMs) have demonstrated remarkable performance in visual-language tasks. However, the authenticity of the responses generated by MLLMs…

cs.CV2024

Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs

Shuo Li, Tao Ji, Xiaoran Fan +10

In the study of LLMs, sycophancy represents a prevalent hallucination that poses significant challenges to these models. Specifically, LLMs often fail to adhere to original correct…