activity
20242026
most citedDeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

884 citations · 884 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

DeepTaxon: An Interpretable Retrieval-Augmented Multimodal Framework for Unified Species Identification and Discovery

Jiawei Wang, Ming Lei, Yaning Yang +8

Identifying species in biology among tens of thousands of visually similar taxa while discovering unknown species in open-world environments remains a fundamental challenge in biod…

cs.CV2026

LongCat-Next: Lexicalizing Modalities as Discrete Tokens

Meituan LongCat Team, Bin Xiao, Chao Wang +86

The prevailing Next-Token Prediction (NTP) paradigm has driven the success of large language models through discrete autoregressive modeling. However, contemporary multimodal syste…

cs.CV2025

Length Matters: Length-Aware Transformer for Temporal Sentence Grounding

Yifan Wang, Ziyi Liu, Xiaolong Sun +2

Temporal sentence grounding (TSG) is a highly challenging task aiming to localize the temporal segment within an untrimmed video corresponding to a given natural language descripti…

cs.CV2025

Seed1.5-VL Technical Report

Dong Guo, Faming Wu, Feida Zhu +194

We present Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning. Seed1.5-VL is composed with a 532M-parameter v…

cs.CV2025

Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Liping Yuan, Jiawei Wang, Haomiao Sun +2

We introduce Tarsier2, a state-of-the-art large vision-language model (LVLM) designed for generating detailed and accurate video descriptions, while also exhibiting superior genera…

cs.CV2024

Tarsier: Recipes for Training and Evaluating Large Video Description Models

Jiawei Wang, Liping Yuan, Yuchen Zhang +1

Generating fine-grained video descriptions is a fundamental challenge in video understanding. In this work, we introduce Tarsier, a family of large-scale video-language models desi…