3 papers
cs.CV2025
PixCLIP: Achieving Fine-grained Visual Language Understanding via Any-granularity Pixel-Text Alignment Learning
Yicheng Xiao, Yu Chen, Haoxuan Ma +5
While the Contrastive Language-Image Pretraining(CLIP) model has achieved remarkable success in a variety of downstream vison language understanding tasks, enhancing its capability…
cs.CV2025
Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
Jianfeng Cai, Wengang Zhou, Zongmeng Zhang +3
Multimodal large language models (MLLMs) have achieved remarkable progress in video understanding.However, hallucination, where the model generates plausible yet incorrect outputs,…
cs.CL2025
BriLLM: Brain-inspired Large Language Model
Hai Zhao, Hongqiu Wu, Dongjie Yang +2
We introduce BriLLM, a brain-inspired large language model that fundamentally redefines the foundations of machine learning through its implementation of Signal Fully-connected flo…