6 papers
Point Tracking Improves World Action Models
Jiarui Guan, Wenshuai Zhao, Yue Pei +3
Robot policy learning benefits from world-action models that capture environment dynamics, but pixel-level prediction entangles dynamics with nuisance factors such as lighting and…
Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision Language Models
Zhengtao Zou, Ya Gao, Jiarui Guan +2
Large Vision-Language Models (LVLMs) typically process visual inputs as a prefix to the language decoder. As the model autoregressively generates text, this initial visual informat…
Latent-Compressed Variational Autoencoder for Video Diffusion Models
Jiarui Guan, Wenshuai Zhao, Zhengtao Zou +2
Video variational autoencoders (VAEs) used in latent diffusion models typically require a sufficiently large number of latent channels to ensure high-quality video reconstruction.…
MorphingDB: A Task-Centric AI-Native DBMS for Model Management and Inference
Wu Sai, Xia Ruichen, Yang Dingyu +9
The increasing demand for deep neural inference within database environments has driven the emergence of AI-native DBMSs. However, existing solutions either rely on model-centric d…
OmniD: Generalizable Robot Manipulation Policy via Image-Based BEV Representation
Jilei Mao, Jiarui Guan, Yingjuan Tang +7
The visuomotor policy can easily overfit to its training datasets, such as fixed camera positions and backgrounds. This overfitting makes the policy perform well in the in-distribu…
Designing Human-AI System for Legal Research: A Case Study of Precedent Search in Chinese Law
Jiarui Guan, Ruishi Zou, Jiajun Zhang +4
Recent advancements in AI technology have seen researchers and industry professionals actively exploring the application of AI tools in legal workflows. Despite this prevailing tre…