7 papers
Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR
Yulong Zhang, Tianyi Liang, Xinyue Huang +5
Optical Character Recognition (OCR) is fundamental to Vision-Language Models (VLMs) and high-quality data generation for LLM training. Yet, despite progress in average OCR accuracy…
Beyond Fixed Anchors: Precisely Erasing Concepts with Sibling Exclusive Counterparts
Tong Zhang, Ru Zhang, Jianyi Liu +2
Existing concept erasure methods for text-to-image diffusion models commonly rely on fixed anchor strategies, which often lead to critical issues such as concept re-emergence and e…
NCV: A Node-Wise Consistency Verification Approach for Low-Cost Structured Error Localization in LLM Reasoning
Yulong Zhang, Li Wang, Wei Du +7
Verifying multi-step reasoning in large language models is difficult due to imprecise error localization and high token costs. Existing methods either assess entire reasoning chain…
FedRS-Bench: Realistic Federated Learning Datasets and Benchmarks in Remote Sensing
Haodong Zhao, Peng Peng, Chiyu Chen +2
Remote sensing (RS) images are usually produced at an unprecedented scale, yet they are geographically and institutionally distributed, making centralized model training challengin…
InsightVision: A Comprehensive, Multi-Level Chinese-based Benchmark for Evaluating Implicit Visual Semantics in Large Vision Language Models
Xiaofei Yin, Yijie Hong, Ya Guo +4
In the evolving landscape of multimodal language models, understanding the nuanced meanings conveyed through visual cues - such as satire, insult, or critique - remains a significa…
U-GIFT: Uncertainty-Guided Firewall for Toxic Speech in Few-Shot Scenario
Jiaxin Song, Xinyu Wang, Yihao Wang +4
With the widespread use of social media, user-generated content has surged on online platforms. When such content includes hateful, abusive, offensive, or cyberbullying behavior, i…