7 papers
EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization
Lihuang Fang, Yuchen Zou, kebing Jin +1
Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognition (MER) tasks and lifted MER to a new level that is complex emotion un…
FLM-Occ: Feed-forward Likelihood Maximization for Efficient Indoor Occupancy Prediction
Guangcheng Chen, Lihuang Fang, Huaqi Tao +3
Recent indoor occupancy prediction methods adopt Gaussian primitives as a sparse 3D representation for computational efficiency. However, their training relies on voxel classificat…
FlowPalm: Optical Flow Driven Non-Rigid Deformation for Geometrically Diverse Palmprint Generation
Yuchen Zou, Huikai Shao, Lihuang Fang +2
Recently, synthetic palmprints have been increasingly used as substitutes for real data to train recognition models. To be effective, such synthetic data must reflect the diversity…
Uncertainty-Aware Hierarchical Re-Localization in OpenStreetMap via Semantic Alignment
Yuchen Zou, Xiao Hu, Lihuang Fang +1
Monocular re-localization enables robots to estimate camera poses from visual observations. However, many existing methods rely on dense maps or large reference image databases, wh…
CogStereo: Neural Stereo Matching with Implicit Spatial Cognition Embedding
Lihuang Fang, Xiao Hu, Yuchen Zou +1
Deep stereo matching has advanced significantly on benchmark datasets through fine-tuning but falls short of the zero-shot generalization seen in foundation models in other vision…
NiteDR: Nighttime Image De-Raining with Cross-View Sensor Cooperative Learning for Dynamic Driving Scenes
Cidan Shi, Lihuang Fang, Han Wu +3
In real-world environments, outdoor imaging systems are often affected by disturbances such as rain degradation. Especially, in nighttime driving scenes, insufficient and uneven li…