NewEvery arXiv paper, its researchers & institutions — mapped.
papers

Publications (41)

cs.RO2026

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation

Junjin Xiao, Dongyang Li, Yandan Yang +9

cs.CV2024

Reshaping the Online Data Buffering and Organizing Mechanism for Continual Test-Time Adaptation

Zhilin Zhu, Xiaopeng Hong, Zhiheng Ma +4

cs.CV2026

Trajectory-Diversity-Driven Robust Vision-and-Language Navigation

Jiangyang Li, Cong Wan, SongLin Dong +4

cs.SD2025

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization

Detao Bai, Zhiheng Ma, Xihan Wei +1

cs.CV2023

Can SAM Count Anything? An Empirical Study on SAM Counting

Zhiheng Ma, Xiaopeng Hong, Qinnan Shangguan

cs.CV2026

ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment

Yuzhi Chen, Ronghan Chen, Dongjie Huo +11

cs.CV2019

Bayesian Loss for Crowd Count Estimation with Point Supervision

Zhiheng Ma, Xing Wei, Xiaopeng Hong +1

cs.CV2026

Dance Across Shifts: Forward-Facilitation Continual Test-Time Adaptation through Dynamic Style Bridging

Zhilin Zhu, Yabin Wang, Zhiheng Ma +3

cs.CV2023

Remind of the Past: Incremental Learning with Analogical Prompts

Zhiheng Ma, Xiaopeng Hong, Beinan Liu +3

cs.AI2026

Beyond World-Frame Action Heads: Motion-Centric Action Frames for Vision-Language-Action Models

Huoren Yang, Jianchao Zhao, Hu Yusong +7

cs.CV2024

Linguistic Profiling of Deepfakes: An Open Database for Next-Generation Deepfake Detection

Yabin Wang, Zhiwu Huang, Zhiheng Ma +1

cs.CV2026

ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents

Dongjie Huo, Haoyun Liu, Guoqing Liu +10

cs.CV2025

Curriculum Dataset Distillation

Zhiheng Ma, Anjia Cao, Funing Yang +2

cs.CV2025

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training

Anjia Cao, Xing Wei, Zhiheng Ma

cs.LG2026

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams

Cong Wan, Zeyu Guo, Zijian Cai +6

cs.CV2026

P2L-CA: An Effective Parameter Tuning Framework for Rehearsal-Free Multi-Label Class-Incremental Learning

Songlin Dong, Jiangyang Li, Chenhao Ding +4

cs.CV2024

ComprehendEdit: A Comprehensive Dataset and Evaluation Framework for Multimodal Knowledge Editing

Yaohui Ma, Xiaopeng Hong, Shizhou Zhang +4

cs.CV2024

Gramformer: Learning Crowd Counting via Graph-Modulated Transformer

Hui Lin, Zhiheng Ma, Xiaopeng Hong +2

cs.CV2026

ProSR: Process-Shaped Spatial Reasoning for Reliable Chain-of-Thought in VLMs

Jiangyang Li, Cong Wan, Changjie Wu +8

cs.CV2025

Few-shot Online Anomaly Detection and Segmentation

Shenxing Wei, Xing Wei, Zhiheng Ma +3

cs.CV2022

Boosting Crowd Counting via Multifaceted Attention

Hui Lin, Zhiheng Ma, Rongrong Ji +2

cs.CV2026

ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning

Yandan Yang, Shuang Zeng, Tong Lin +11

cs.CV2022

Isolation and Impartial Aggregation: A Paradigm of Incremental Learning without Interference

Yabin Wang, Zhiheng Ma, Zhiwu Huang +3

cs.RO2023

Towards Practical Multi-Robot Hybrid Tasks Allocation for Autonomous Cleaning

Yabin Wang, Xiaopeng Hong, Zhiheng Ma +3

cs.CV2026

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

Ronghan Chen, Yandan Yang, Zuojin Tang +18

cs.CV2026

Continuous Expert Assembly: Instance-Conditioned Low-Rank Residuals for All-in-One Image Restoration

Haisen He, Xiangyu Zou, SongLin Dong +3

cs.CV2024

Prompt Customization for Continual Learning

Yong Dai, Xiaopeng Hong, Yabin Wang +3

cs.CV2021

Direct Measure Matching for Crowd Counting

Hui Lin, Xiaopeng Hong, Zhiheng Ma +4

cs.CV2024

Multi-modal Crowd Counting via Modal Emulation

Chenhao Wang, Xiaopeng Hong, Zhiheng Ma +3

cs.CV2025

CKAA: Cross-subspace Knowledge Alignment and Aggregation for Robust Continual Learning

Lingfeng He, De Cheng, Zhiheng Ma +4

cs.CV2026

HumanOmni-Speaker: Identifying Who said What and When

Detao Bai, Zhiheng Ma, Xihan Wei

cs.CV2021

Anomaly Detection via Self-organizing Map

Ning Li, Kaitao Jiang, Zhiheng Ma +3

cs.CV2026

OmniEncoder: See, Hear, and Feel Continuous Motion Like Humans With One Encoder

Detao Bai, Shimin Yao, Weixuan Chen +4

cs.RO2026

Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models

Haoyun Liu, Jianzhuang Zhao, Xinyuan Chang +11

cs.RO2026

DLAM: Distributional Latent Actions with Temporal Constraints

Zuojin Tang, Feifan Luo, Haoyun Liu +10

The paper introduces DLAM, a distributional latent-action model that encodes video transitions as diagonal Gaussians with temporal constraints, improving reconstruction consistency…

#latent action modeling#vision-language-action#temporal constraints#distributional dynamics
cs.CV2026

ReMoT: Reinforcement Learning with Motion Contrast Triplets

Cong Wan, Zeyu Guo, Jiangyang Li +5

cs.CV2022

Semi-supervised Crowd Counting via Density Agency

Hui Lin, Zhiheng Ma, Xiaopeng Hong +2

cs.RO2026

Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs

Jianchao Zhao, Huoren Yang, Yusong Hu +6

cs.CV2025

DD-Ranking: Rethinking the Evaluation of Dataset Distillation

Zekai Li, Xinhao Zhong, Samir Khaki +49

cs.RO2026

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models

Zuojin Tang, Haoyun Liu, Xinyuan Chang +11

cs.CV2024

Semi-supervised Counting via Pixel-by-pixel Density Distribution Modelling

Hui Lin, Zhiheng Ma, Rongrong Ji +4