papers

Publications (13)

cs.HC2025

Can Intelligent User Interfaces Engage in Philosophical Discussions? A Longitudinal Study of Philosophers' Evolving Perceptions

Yibo Meng, Lyumanshan Ye, Eve He +5

This study investigates the evolving attitudes of philosophy scholars towards the participation of generative AI based Intelligent User Interfaces (IUIs) in philosophical discourse…

cs.CV2026

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis

Zhipeng Xu, Zulong Chen, Qing Liu +6

Key Information Extraction (KIE) converts visually rich documents into structured data, but practical deployment remains challenging: strong performance often relies on costly on-s…

cs.AI2019

Learning High-order Structural and Attribute information by Knowledge Graph Attention Networks for Enhancing Knowledge Graph Embedding

Wenqiang Liu, Hongyun Cai, Xu Cheng +3

The goal of representation learning of knowledge graph is to encode both entities and relations into a low-dimensional embedding spaces. Many recent works have demonstrated the ben…

cs.SE2026

1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World

Qiao Xu, Yipeng Yu, Chengxiao Feng +1

The paper introduces 1D-Bench, a benchmark for generating executable React UI code from design renderings and imperfect intermediate representations, supporting iterative component…

#design-to-code#ui code generation#benchmark#iterative editing
cs.AI2026

Deep Research of Deep Research: From Transformer to Agent, From AI to AI for Science

Yipeng Yu

With the advancement of large language models (LLMs) in their knowledge base and reasoning capabilities, their interactive modalities have evolved from pure text to multimodality a…

cs.CL2025

LLMAtKGE: Large Language Models as Explainable Attackers against Knowledge Graph Embeddings

Ting Li, Yang Yang, Yipeng Yu +3

Adversarial attacks on knowledge graph embeddings (KGE) aim to disrupt the model's ability of link prediction by removing or inserting triples. A recent black-box method has attemp…

cs.CV2024

Resolving Multi-Condition Confusion for Finetuning-Free Personalized Image Generation

Qihan Huang, Siming Fu, Jinlong Liu +3

Personalized text-to-image generation methods can generate customized images based on the reference images, which have garnered wide research interest. Recent methods propose a fin…

cs.AI2025

MHSNet:An MoE-based Hierarchical Semantic Representation Network for Accurate Duplicate Resume Detection with Large Language Model

Yu Li, Zulong Chen, Wenjian Xu +4

To maintain the company's talent pool, recruiters need to continuously search for resumes from third-party websites (e.g., LinkedIn, Indeed). However, fetched resumes are often inc…

cs.IR2025

SEAL: Structure and Element Aware Learning to Improve Long Structured Document Retrieval

Xinhao Huang, Zhibo Ren, Yipeng Yu +3

In long structured document retrieval, existing methods typically fine-tune pre-trained language models (PLMs) using contrastive learning on datasets lacking explicit structural in…

cs.CV2026

UniVid: Pyramid Diffusion Model for High Quality Video Generation

Xinyu Xiao, Binbin Yang, Tingtian Li +2

Diffusion-based text-to-video generation (T2V) or image-to-video (I2V) generation have emerged as a prominent research focus. However, there exists a challenge in integrating the t…

cs.HC2025

Colin: A Multimodal Human-AI Co-Creation Storytelling System To Support Children's Multi-Level Narrative Skills

Lyumanshan Ye, Jiandong Jiang, Yuhan Liu +7

Children develop narrative skills by understanding and actively building connections between elements, image text matching, and consequences. However, it is challenging for childre…

cs.IR2021

Deep Music Retrieval for Fine-Grained Videos by Exploiting Cross-Modal-Encoded Voice-Overs

Tingtian Li, Zixun Sun, Haoruo Zhang +5

Recently, the witness of the rapidly growing popularity of short videos on different Internet platforms has intensified the need for a background music (BGM) retrieval system. Howe…

cs.IR2026

RecGPT-Mobile: On-Device Large Language Models for User Intent Understanding in Taobao Feed Recommendation

Bin Zhang, Weipeng Huang, Dimin Wang +9

Predicting a user's next search query from recent interaction behaviors is a critical problem in modern e-commerce systems, particularly in scenarios where user intent evolves rapi…