Publications (8)
ReMoT: Reinforcement Learning with Motion Contrast Triplets
Cong Wan, Zeyu Guo, Jiangyang Li +5
We present ReMoT, a unified training paradigm to systematically address the fundamental shortcomings of VLMs in spatio-temporal consistency -- a critical failure point in navigatio…
Eigenvalues of Autocovariance Matrix: A Practical Method to Identify the Koopman Eigenfrequencies
Yicun Zhen, Bertrand Chapron, Etienne Memin +1
To infer eigenvalues of the infinite-dimensional Koopman operator, we study the leading eigenvalues of the autocovariance matrix associated with a given observable of a dynamical s…
Joint News, Attention Spillover,and Stock Returns
Li Guo, Lin Peng, Yubo Tao +1
Analyzing a comprehensive news dataset, we document that joint news coverage triggers attention contagion, causing temporarily inflated valuations for affected stocks. Tracing SEC…
DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
Cong Wan, Zeyu Guo, Zijian Cai +6
Raw multimodal streams are abundant but noisy, redundant, and unaligned with any particular training objective. Turning them into supervision today means either brittle heuristics…
Selection of Supervised Learning-based Sparse Matrix Reordering Algorithms
Tao Tang, Youfu Jiang, Yingbo Cui +4
Sparse matrix ordering is a vital optimization technique often employed for solving large-scale sparse matrices. Its goal is to minimize the matrix bandwidth by reorganizing its ro…
Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Models
Boyang Guo, Liang Li, Lin Peng +3
Prompt learning has emerged as an efficient alternative to fine-tuning pre-trained vision-language models (VLMs). Despite its promise, current methods still struggle to maintain ta…
Dexonomy: Synthesizing All Dexterous Grasp Types in a Grasp Taxonomy
Jiayi Chen, Yubin Ke, Lin Peng +1
Generalizable dexterous grasping with suitable grasp types is a fundamental skill for intelligent robots. Developing such skills requires a large-scale and high-quality dataset tha…
CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models
Lin Peng, Cong Wan, Zeyu Guo +2
The paper introduces CoRe, a framework that improves vision-language models' ability to perform fine-grained cross‑image comparative reasoning by providing a large triplet‑based da…