most citedPAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs

1 citations · 1 across the 4 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Generative World Renderer at the Speed of Play

Guixu Lin, Zheng-Hui Huang, Siqi Yang +3

Generative world renderer AlayaRenderer receives structured world states exported from physics engines and synthesizes RGB frames. Unlike models that generate frames from text/cont…

cs.CV20261 cited

PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs

Bowen Sun, Yujun Cai, Ming-Hsuan Yang +2

Video LLMs suffer from temporal inconsistency: small shifts in frame timing can flip attention and suppress relevant frames. We trace this instability to the common extension of Ro…

cs.CV2025

MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs

Haonan Ge, Yiwei Wang, Ming-Hsuan Yang +1

Large Vision-Language Models (LVLMs) have shown strong performance across multimodal tasks. However, they often produce hallucinations -- text that is inconsistent with visual inpu…

cs.CV2025

Efficiently Disentangling CLIP for Multi-Object Perception

Samyak Rawlekar, Yujun Cai, Yiwei Wang +2

Vision-language models like CLIP excel at recognizing the single, prominent object in a scene. However, they struggle in complex scenes containing multiple objects. We identify a f…

cs.CV2025

Text Speaks Louder than Vision: ASCII Art Reveals Textual Biases in Vision-Language Models

Zhaochen Wang, Bryan Hooi, Yiwei Wang +3

Vision-language models (VLMs) have advanced rapidly in processing multimodal information, but their ability to reconcile conflicting signals across modalities remains underexplored…

cs.CV2025

Lost in Edits? A -Compass for AIGC Provenance

Wenhao You, Bryan Hooi, Yiwei Wang +5

Recent advancements in diffusion models have driven the growth of text-guided image editing tools, enabling precise and iterative modifications of synthesized content. However, as…