collaborators

12 papers

cs.CV2026

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models

Yuansheng Gao, Jinman Zhao, Tong Zhang +5

Although Video Large Multimodal Models have achieved strong performance in video understanding, they still suffer from hallucination. Existing inference-time intervention methods u…

cs.CV2026

GT-SVJ: Generative-Transformer-Based Self-Supervised Video Judge For Efficient Video Reward Modeling

Shivanshu Shekhar, Uttaran Bhattacharya, Raghavendra Addanki +3

Aligning video generative models with human preferences remains challenging: current approaches rely on Vision-Language Models (VLMs) for reward modeling, but these models struggle…

cs.CV2026

DICE: Disentangling Artist Style from Content via Contrastive Subspace Decomposition in Diffusion Models

Tong Zhang, Ru Zhang, Jianyi Liu

The recent proliferation of diffusion models has made style mimicry effortless, enabling users to imitate unique artistic styles without authorization. In deployed platforms, this…

cs.IR2026

Farewell to Item IDs: Unlocking the Scaling Potential of Large Ranking Models via Semantic Tokens

Zhen Zhao, Tong Zhang, Jie Xu +5

Recent studies on scaling up ranking models have achieved substantial improvement for recommendation systems and search engines. However, most large-scale ranking systems rely on i…

cs.IT2026

Edge Collaborative Gaussian Splatting with Integrated Rendering and Communication

Yujie Wan, Chenxuan Liu, Shuai Wang +5

Gaussian splatting (GS) struggles with degraded rendering quality on low-cost devices. To address this issue, we present edge collaborative GS (ECO-GS), where each user can switch…

cs.CL2025

Test-time Scaling of LLMs: A Survey from A Subproblem Structure Perspective

Zhuoyi Yang, Xu Guo, Tong Zhang +2

With this paper, we survey techniques for improving the predictive accuracy of pretrained large language models by allocating additional compute at inference time. In categorizing…