◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Yifeng Dai

4 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author1
  • last author1

Across the 2 of 4 papers where every author was matched, so the position is known.

fields
  • cs.CV3
  • cs.RO1
ORCID 0000-0002-5037-3695

identity via Semantic Scholar / OpenAlex

collaborators

4 papers

cs.RO2025

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy

Tianyi Zhang, Haonan Duan, Haoran Hao +3

Vision-Language-Action (VLA) models frequently encounter challenges in generalizing to real-world environments due to inherent discrepancies between observation and action spaces.…

cs.CV2025

Spatial Frequency Modulation for Semantic Segmentation

Linwei Chen, Ying Fu, Lin Gu +2

High spatial frequency information, including fine details like textures, significantly contributes to the accuracy of semantic segmentation. However, according to the Nyquist-Shan…

cs.CV2025

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Gen Luo, Ganlin Yang, Ziyang Gong +15

The remarkable progress of Multimodal Large Language Models (MLLMs) has attracted increasing attention to extend them to physical entities like legged robot. This typically require…

cs.CV2024

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models

Chenyu Yang, Xuan Dong, Xizhou Zhu +7

Large Vision-Language Models (VLMs) have been extended to understand both images and videos. Visual token compression is leveraged to reduce the considerable token length of visual…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.