2 citations · 2 across the 6 of their papers we have counts for
6 papers
PRISM: Pareto-Efficient Retrieval over Intent-Aware Structured Memory for Long-Horizon Agents
Jingyi Peng, Zhongwei Wan, Weiting Liu +1
Long-horizon language agents accumulate conversation history far faster than any fixed context window can hold, making memory management critical to both answer accuracy and servin…
FORGE: Fine-grained Multimodal Evaluation for Manufacturing Scenarios
Xiangru Jian, Hao Xu, Wei Pang +13
The manufacturing sector is increasingly adopting Multimodal Large Language Models (MLLMs) to transition from simple perception to autonomous execution, yet current evaluations fai…
Tighnari: Multi-modal Plant Species Prediction Based on Hierarchical Cross-Attention Using Graph-Based and Vision Backbone-Extracted Features
Haixu Liu, Penghao Jiang, Zerui Tao +2
Predicting plant species composition in specific spatiotemporal contexts plays an important role in biodiversity management and conservation, as well as in improving species identi…
Multi-Modal Video Feature Extraction for Popularity Prediction
Haixu Liu, Wenning Wang, Haoxiang Zheng +4
This work aims to predict the popularity of short videos using the videos themselves and their related features. Popularity is measured by four key engagement metrics: view count,…
nnY-Net: Swin-NeXt with Cross-Attention for 3D Medical Images Segmentation
Haixu Liu, Zerui Tao, Wenzhen Dong +1
This paper provides a novel 3D medical image segmentation model structure called nnY-Net. This name comes from the fact that our model adds a cross-attention module at the bottom o…
Unveiling and Controlling Anomalous Attention Distribution in Transformers
Ruiqing Yan, Xingbo Du, Haoyu Deng +7
With the advent of large models based on the Transformer architecture, researchers have observed an anomalous phenomenon in the Attention mechanism--there is a very high attention…