activity
20242026
collaborators

6 papers

cs.CL2026

Self-Supervised Skill Optimization

Siran Peng, Cuiyu Yang, Tianyu Fu +9

Agent skills provide frozen large language model (LLM) agents with reusable procedural guidance, and recent work shows that such skills can be optimized with ground-truth (GT) feed…

cs.CV2026

Dive Into the Implicit Biases of Low-rank Vision-language Alignment

Mingjia Shi, Shuo Wang, Xiaobo Wang +7

Vision-language alignment, the stage that bridges pretrained vision encoders and large language models, is widely treated as a form of pretraining requiring full-parameter updates.…

cs.CV2026

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation

Wei Dong, Tianyu Fu, Zhe Yu +9

As web agents increasingly demonstrate capabilities in automated task execution, the development of robust evaluation frameworks for assessing their navigation and task completion…

cs.MM2025

Mano Technical Report

Tianyu Fu, Anyang Su, Chenxu Zhao +20

Graphical user interfaces (GUIs) are the primary medium for human-computer interaction, yet automating GUI interactions remains challenging due to the complexity of visual elements…

cs.CV2025

Salvaging the Overlooked: Leveraging Class-Aware Contrastive Learning for Multi-Class Anomaly Detection

Lei Fan, Junjie Huang, Donglin Di +4

For anomaly detection (AD), early approaches often train separate models for individual classes, yielding high performance but posing challenges in scalability and resource managem…

cs.LG2024

Boundary-Guided Learning for Gene Expression Prediction in Spatial Transcriptomics

Mingcheng Qu, Yuncong Wu, Donglin Di +4

Spatial transcriptomics (ST) has emerged as an advanced technology that provides spatial context to gene expression. Recently, deep learning-based methods have shown the capability…