works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.CV2026

Let RGB Be the Language of Vision

Timing Yang, Jinrui Yang, Xinlong Li +11

The paper proposes a unified vision framework that encodes all visual signals—including images, masks, and depth maps—as RGB images, turning diverse tasks into a common RGB-to-RGB…

cs.CV2026

InterSketch: An Interleaved Reasoning Model with Self-correcting Visual Sketch and Stepwise Reward

Zhiwei Ning, Wenwen Tong, Xiangli Kong +12

While vision-language models (VLMs) have exhibited multi-turn visual reasoning capabilities, their reasoning trajectories remain relatively shallow and are dominated by a text-cent…

cs.CV2026

Group Editing: Edit Multiple Images in One Go

Yue Ma, Xinyu Wang, Qianli Ma +9

In this paper, we tackle the problem of performing consistent and unified modifications across a set of related images. This task is particularly challenging because these images m…

cs.CV2026

Bridging Cognitive Gap: Hierarchical Description Learning for Artistic Image Aesthetics Assessment

Henglin Liu, Nisha Huang, Chang Liu +6

The aesthetic quality assessment task is crucial for developing a human-aligned quantitative evaluation system for AIGC. However, its inherently complex nature, spanning visual per…

q-fin.CP2025

Advancing Financial Engineering with Foundation Models: Progress, Applications, and Challenges

Liyuan Chen, Shuoling Liu, Jiangpeng Yan +8

The advent of foundation models (FMs), large-scale pre-trained models with strong generalization capabilities, has opened new frontiers for financial engineering. While general-pur…

cs.CV2025

Linear Differential Vision Transformer: Learning Visual Contrasts via Pairwise Differentials

Yifan Pu, Jixuan Ying, Qixiu Li +7

Vision Transformers (ViTs) have become a universal backbone for both image recognition and image generation. Yet their Multi-Head Self-Attention (MHSA) layer still performs a quadr…