Publications (13)
Can Intelligent User Interfaces Engage in Philosophical Discussions? A Longitudinal Study of Philosophers' Evolving Perceptions
Yibo Meng, Lyumanshan Ye, Eve He +5
This study investigates the evolving attitudes of philosophy scholars towards the participation of generative AI based Intelligent User Interfaces (IUIs) in philosophical discourse…
Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis
Zhipeng Xu, Zulong Chen, Qing Liu +6
Key Information Extraction (KIE) converts visually rich documents into structured data, but practical deployment remains challenging: strong performance often relies on costly on-s…
Learning High-order Structural and Attribute information by Knowledge Graph Attention Networks for Enhancing Knowledge Graph Embedding
Wenqiang Liu, Hongyun Cai, Xu Cheng +3
The goal of representation learning of knowledge graph is to encode both entities and relations into a low-dimensional embedding spaces. Many recent works have demonstrated the ben…
1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World
Qiao Xu, Yipeng Yu, Chengxiao Feng +1
The paper introduces 1D-Bench, a benchmark for generating executable React UI code from design renderings and imperfect intermediate representations, supporting iterative component…
Deep Research of Deep Research: From Transformer to Agent, From AI to AI for Science
Yipeng Yu
With the advancement of large language models (LLMs) in their knowledge base and reasoning capabilities, their interactive modalities have evolved from pure text to multimodality a…
LLMAtKGE: Large Language Models as Explainable Attackers against Knowledge Graph Embeddings
Ting Li, Yang Yang, Yipeng Yu +3
Adversarial attacks on knowledge graph embeddings (KGE) aim to disrupt the model's ability of link prediction by removing or inserting triples. A recent black-box method has attemp…
Resolving Multi-Condition Confusion for Finetuning-Free Personalized Image Generation
Qihan Huang, Siming Fu, Jinlong Liu +3
Personalized text-to-image generation methods can generate customized images based on the reference images, which have garnered wide research interest. Recent methods propose a fin…
MHSNet:An MoE-based Hierarchical Semantic Representation Network for Accurate Duplicate Resume Detection with Large Language Model
Yu Li, Zulong Chen, Wenjian Xu +4
To maintain the company's talent pool, recruiters need to continuously search for resumes from third-party websites (e.g., LinkedIn, Indeed). However, fetched resumes are often inc…
SEAL: Structure and Element Aware Learning to Improve Long Structured Document Retrieval
Xinhao Huang, Zhibo Ren, Yipeng Yu +3
In long structured document retrieval, existing methods typically fine-tune pre-trained language models (PLMs) using contrastive learning on datasets lacking explicit structural in…
UniVid: Pyramid Diffusion Model for High Quality Video Generation
Xinyu Xiao, Binbin Yang, Tingtian Li +2
Diffusion-based text-to-video generation (T2V) or image-to-video (I2V) generation have emerged as a prominent research focus. However, there exists a challenge in integrating the t…
Colin: A Multimodal Human-AI Co-Creation Storytelling System To Support Children's Multi-Level Narrative Skills
Lyumanshan Ye, Jiandong Jiang, Yuhan Liu +7
Children develop narrative skills by understanding and actively building connections between elements, image text matching, and consequences. However, it is challenging for childre…
Deep Music Retrieval for Fine-Grained Videos by Exploiting Cross-Modal-Encoded Voice-Overs
Tingtian Li, Zixun Sun, Haoruo Zhang +5
Recently, the witness of the rapidly growing popularity of short videos on different Internet platforms has intensified the need for a background music (BGM) retrieval system. Howe…
RecGPT-Mobile: On-Device Large Language Models for User Intent Understanding in Taobao Feed Recommendation
Bin Zhang, Weipeng Huang, Dimin Wang +9
Predicting a user's next search query from recent interaction behaviors is a critical problem in modern e-commerce systems, particularly in scenarios where user intent evolves rapi…