Publications (5)
Towards Fine-Grained Recognition with Large Visual Language Models: Benchmark and Optimization Strategies
Cong Pang, Hongtao Yu, Zixuan Chen +2
Large Vision Language Models (LVLMs) have made remarkable progress, enabling sophisticated vision-language interaction and dialogue applications. However, existing benchmarks prima…
ISE: An Execution-Grounded Recipe for Multi-Turn OS-Agent Trajectories
Siyuan Luo, Nairong Zheng, Lin Zhou +6
The paper introduces ISE, a three-stage pipeline for creating a large dataset of multi‑turn OS‑agent interactions that include structured intents, simulated dialogues, and real too…
ICA: Information-Aware Credit Assignment for Visually Grounded Long-Horizon Information-Seeking Agents
Cong Pang, Xuyu Feng, Yujie Yi +7
Despite the strong performance achieved by reinforcement learning-trained information-seeking agents, learning in open-ended web environments remains severely constrained by low si…
Learned Smartphone ISP on Mobile GPUs with Deep Learning, Mobile AI & AIM 2022 Challenge: Report
Andrey Ignatov, Radu Timofte, Shuai Liu +35
The role of mobile cameras increased dramatically over the past few years, leading to more and more research in automatic image quality enhancement and RAW photo processing. In thi…
Lepton flavor violating Higgs couplings and single production of the Higgs boson via e γcollision
Chong-Xing Yue, Cong Pang, Yu-Chen Guo
Taking into account the constraints on the lepton flavor violation (LFV) couplings of the Standard Model (SM) Higgs boson H with leptons from low energy experiments and the recent…