12 papers
Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models
Yuansheng Gao, Jinman Zhao, Tong Zhang +5
Although Video Large Multimodal Models have achieved strong performance in video understanding, they still suffer from hallucination. Existing inference-time intervention methods u…
GT-SVJ: Generative-Transformer-Based Self-Supervised Video Judge For Efficient Video Reward Modeling
Shivanshu Shekhar, Uttaran Bhattacharya, Raghavendra Addanki +3
Aligning video generative models with human preferences remains challenging: current approaches rely on Vision-Language Models (VLMs) for reward modeling, but these models struggle…
DICE: Disentangling Artist Style from Content via Contrastive Subspace Decomposition in Diffusion Models
Tong Zhang, Ru Zhang, Jianyi Liu
The recent proliferation of diffusion models has made style mimicry effortless, enabling users to imitate unique artistic styles without authorization. In deployed platforms, this…
Farewell to Item IDs: Unlocking the Scaling Potential of Large Ranking Models via Semantic Tokens
Zhen Zhao, Tong Zhang, Jie Xu +5
Recent studies on scaling up ranking models have achieved substantial improvement for recommendation systems and search engines. However, most large-scale ranking systems rely on i…
Edge Collaborative Gaussian Splatting with Integrated Rendering and Communication
Yujie Wan, Chenxuan Liu, Shuai Wang +5
Gaussian splatting (GS) struggles with degraded rendering quality on low-cost devices. To address this issue, we present edge collaborative GS (ECO-GS), where each user can switch…
Test-time Scaling of LLMs: A Survey from A Subproblem Structure Perspective
Zhuoyi Yang, Xu Guo, Tong Zhang +2
With this paper, we survey techniques for improving the predictive accuracy of pretrained large language models by allocating additional compute at inference time. In categorizing…