From the 1 of 10 linked papers with an AI index.
7 papers · 1 filter
CoVStream: Edge-Cloud Collaboration for Understanding of Long Video Streams
Xu Liu, Guikun Chen, Zihao Yan +2
The paper introduces CoVStream, an edge‑cloud system that compresses raw video into compact visual features and captions on the device, sends them to the cloud for graph‑based reas…
SinkTrack: Attention Sink based Context Anchoring for Large Language Models
Xu Liu, Guikun Chen, Wenguan Wang
Large language models (LLMs) suffer from hallucination and context forgetting. Prior studies suggest that attention drift is a primary cause of these problems, where LLMs' focus sh…
A Survey on 3D Gaussian Splatting
Guikun Chen, Wenguan Wang
3D Gaussian splatting (GS) has emerged as a transformative technique in radiance fields. Unlike mainstream implicit neural models, 3D GS uses millions of learnable 3D Gaussians for…
DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
Zongxin Yang, Guikun Chen, Xiaodi Li +2
Recent LLM-driven visual agents mainly focus on solving image-based tasks, which limits their ability to understand dynamic scenes, making it far from real-life applications like g…
Hydra-SGG: Hybrid Relation Assignment for One-stage Scene Graph Generation
Minghan Chen, Guikun Chen, Wenguan Wang +1
DETR introduces a simplified one-stage framework for scene graph generation (SGG) but faces challenges of sparse supervision and false negative samples. The former occurs because e…
Scene Graph Generation with Role-Playing Large Language Models
Guikun Chen, Jin Li, Wenguan Wang
Current approaches for open-vocabulary scene graph generation (OVSGG) use vision-language models such as CLIP and follow a standard zero-shot pipeline -- computing similarity betwe…