Publications (136)
S3Mem: Structured Spatiotemporal Scene-Event Memory for Long-Horizon Interactive Question Answering
Encheng Su, Jianyu Wu, Jinouwen Zhang +9
Long-horizon memory question answering often requires sparse evidence from heterogeneous histories, including events, object states, visual observations, temporal relations, and ca…
CSGuard: Toward Forgery-Resistant Watermarking in Diffusion Models via Compressed Sensing Constraint
Jiewei Lai, Lan Zhang, Chen Tang +4
Latent-based diffusion model watermarking embeds watermarks into generated images' latent space to enable content attribution, offering a training-free solution for intellectual pr…
SizeGS: Size-aware Compression of 3D Gaussian Splatting via Mixed Integer Programming
Shuzhao Xie, Jiahang Liu, Weixiang Zhang +7
Recent advances in 3D Gaussian Splatting (3DGS) have greatly improved 3D reconstruction. However, its substantial data size poses a significant challenge for transmission and stora…
Train & Constrain: Phonologically Informed Tongue-Twister Generation from Topics and Paraphrases
Tyler Loakman, Chen Tang, Chenghua Lin
Previous work in phonologically and phonetically grounded language generation has mainly focused on domains such as puns and poetry. In this article, we present new work on the gen…
GlanceSeg: Real-time microaneurysm lesion segmentation with gaze-map-guided foundation model for early detection of diabetic retinopathy
Hongyang Jiang, Mengdi Gao, Zirong Liu +5
Early-stage diabetic retinopathy (DR) presents challenges in clinical diagnosis due to inconspicuous and minute microangioma lesions, resulting in limited research in this area. Ad…
DERMARK: A Dynamic, Efficient and Robust Multi-bit Watermark for Large Language Models
Qihao Lin, Chen Tang, Lan zhang +2
As large language models (LLMs) grow more powerful, concerns over copyright infringement of LLM-generated texts have intensified. LLM watermarking has been proposed to trace unauth…
Domain Knowledge Driven Pseudo Labels for Interpretable Goal-Conditioned Interactive Trajectory Prediction
Lingfeng Sun, Chen Tang, Yaru Niu +5
Motion forecasting in highly interactive scenarios is a challenging problem in autonomous driving. In such scenarios, we need to accurately predict the joint behavior of interactin…
GAQAT: gradient-adaptive quantization-aware training for domain generalization
Jiacheng Jiang, Yuan Meng, Chen Tang +4
Research on loss surface geometry, such as Sharpness-Aware Minimization (SAM), shows that flatter minima improve generalization. Recent studies further reveal that flatter minima c…
Enhancing Biomedical Lay Summarisation with External Knowledge Graphs
Tomas Goldsack, Zhihao Zhang, Chen Tang +2
Previous approaches for automatic lay summarisation are exclusively reliant on the source article that, given it is written for a technical audience (e.g., researchers), is unlikel…
Breaking the Curse of Knowledge: Towards Effective Multimodal Recommendation using Knowledge Soft Integration
Kai Ouyang, Chen Tang, Zenghao Chai +4
A critical challenge in contemporary recommendation systems lies in effectively leveraging multimodal content to enhance recommendation personalization. Although various solutions…
HGNET: A Hierarchical Feature Guided Network for Occupancy Flow Field Prediction
Zhan Chen, Chen Tang, Lu Xiong
Predicting the motion of multiple traffic participants has always been one of the most challenging tasks in autonomous driving. The recently proposed occupancy flow field predictio…
Functional data analysis: An application to COVID-19 data in the United States
Chen Tang, Tiandong Wang, Panpan Zhang
The COVID-19 pandemic so far has caused huge negative impacts on different areas all over the world, and the United States (US) is one of the most affected countries. In this paper…
ComposableNav: Instruction-Following Navigation in Dynamic Environments via Composable Diffusion
Zichao Hu, Chen Tang, Michael J. Munje +6
This paper considers the problem of enabling robots to navigate dynamic environments while following instructions. The challenge lies in the combinatorial nature of instruction spe…
SciReasoner: Laying the Scientific Reasoning Ground Across Disciplines
Yizhou Wang, Chen Tang, Han Deng +29
We present a scientific reasoning foundation model that aligns natural language with heterogeneous scientific representations. The model is pretrained on a 206B-token corpus spanni…
Click-aware Structure Transfer with Sample Weight Assignment for Post-Click Conversion Rate Estimation
Kai Ouyang, Wenhao Zheng, Chen Tang +2
Post-click Conversion Rate (CVR) prediction task plays an essential role in industrial applications, such as recommendation and advertising. Conventional CVR methods typically suff…
Data-Aware Gradient Compression for FL in Communication-Constrained Mobile Computing
Rongwei Lu, Yutong Jiang, Yinan Mao +4
Federated Learning (FL) in mobile environments faces significant communication bottlenecks. Gradient compression has proven as an effective solution to this issue, offering substan…
Permit: Permission-Aware Representation Intervention for Controlled Generation in Large Language Models
Pengcheng Sun, Lan Zhang, Zhaopeng Zhang +2
Large language models (LLMs) are increasingly deployed in enterprise settings where they handle sensitive documents and user context, raising acute concerns over security and contr…
NGEP: A Graph-based Event Planning Framework for Story Generation
Chen Tang, Zhihao Zhang, Tyler Loakman +2
To improve the performance of long text generation, recent studies have leveraged automatically planned event structures (i.e. storylines) to guide story generation. Such prior wor…
One QuantLLM for ALL: Fine-tuning Quantized LLMs Once for Efficient Deployments
Ke Yi, Yuhui Xu, Heng Chang +4
Large Language Models (LLMs) have advanced rapidly but face significant memory demands. While quantization has shown promise for LLMs, current methods typically require lengthy tra…
Untraceable DeepFakes via Traceable Fingerprint Elimination
Jiewei Lai, Lan Zhang, Chen Tang +3
Recent advancements in DeepFakes attribution technologies have significantly enhanced forensic capabilities, enabling the extraction of traces left by generative models (GMs) in im…
Zero-shot Deep Reinforcement Learning Driving Policy Transfer for Autonomous Vehicles based on Robust Control
Zhuo Xu, Chen Tang, Masayoshi Tomizuka
Although deep reinforcement learning (deep RL) methods have lots of strengths that are favorable if applied to autonomous driving, real deep RL applications in autonomous driving h…
Investigating the Impact of Quantization on Adversarial Robustness
Qun Li, Yuan Meng, Chen Tang +2
Quantization is a promising technique for reducing the bit-width of deep models to improve their runtime performance and storage efficiency, and thus becomes a fundamental step for…
Terminology-aware Medical Dialogue Generation
Chen Tang, Hongbo Zhang, Tyler Loakman +2
Medical dialogue generation aims to generate responses according to a history of dialogue turns between doctors and patients. Unlike open-domain dialogue generation, this requires…
TwistList: Resources and Baselines for Tongue Twister Generation
Tyler Loakman, Chen Tang, Chenghua Lin
Previous work in phonetically-grounded language generation has mainly focused on domains such as lyrics and poetry. In this paper, we present work on the generation of tongue twist…
Clustering and Forecasting Multiple Functional Time Series
Chen Tang, Han Lin Shang, Yanrong Yang
Modelling and forecasting homogeneous age-specific mortality rates of multiple countries could lead to improvements in long-term forecasting. Data fed into joint models are often g…
Hecaton: Training Large Language Models with Scalable Chiplet Systems
Zongle Huang, Shupei Fan, Chen Tang +3
Large Language Models (LLMs) have achieved remarkable success in various fields, but their training and finetuning require massive computation and memory, necessitating parallelism…
CoopReflect: Towards Natural Language Communication for Cooperative Autonomous Driving via Multi-Agent Learning
Jiaxun Cui, Chen Tang, Jarrett Holtz +4
Past work has demonstrated that autonomous vehicles can drive more safely if they communicate with each other. However, this communication is usually not human-understandable. Usin…
DRE: An Effective Dual-Refined Method for Integrating Small and Large Language Models in Open-Domain Dialogue Evaluation
Kun Zhao, Bohao Yang, Chen Tang +4
Large Language Models (LLMs) excel at many tasks but struggle with ambiguous scenarios where multiple valid responses exist, often yielding unreliable results. Conversely, Small La…
DAPD: Dual-Anchored Policy Distillation
Jianyu Wu, Yizhou Wang, Encheng Su +2
On-policy (self) distillation (OPSD) is increasingly adopted for language-model post-training. It strengthens the teacher with privileged information but can induce a privilege ill…
Broadening of the drumhead mode spectrum due to in-plane thermal fluctuations of two-dimensional trapped ion crystals in a Penning trap
Athreya Shankar, Chen Tang, Matthew Affolter +5
Two-dimensional crystals of ions stored in Penning traps are a leading platform for quantum simulation and sensing experiments. For small amplitudes, the out-of-plane motion of suc…
WOMD-Reasoning: A Large-Scale Dataset for Interaction Reasoning in Driving
Yiheng Li, Cunxin Fan, Chongjian Ge +9
Language models uncover unprecedented abilities in analyzing driving scenarios, owing to their limitless knowledge accumulated from text-based pre-training. Naturally, they should…
JAQ: Joint Efficient Architecture Design and Low-Bit Quantization with Hardware-Software Co-Exploration
Mingzi Wang, Yuan Meng, Chen Tang +9
The co-design of neural network architectures, quantization precisions, and hardware accelerators offers a promising approach to achieving an optimal balance between performance an…
Quantifying Agent Interaction in Multi-agent Reinforcement Learning for Cost-efficient Generalization
Yuxin Chen, Chen Tang, Ran Tian +4
Generalization poses a significant challenge in Multi-agent Reinforcement Learning (MARL). The extent to which an agent is influenced by unseen co-players depends on the agent's po…
Active Exploration in Iterative Gaussian Process Regression for Uncertainty Modeling in Autonomous Racing
Tommaso Benciolini, Chen Tang, Marion Leibold +3
Autonomous racing creates challenging control problems, but Model Predictive Control (MPC) has made promising steps toward solving both the minimum lap-time problem and head-to-hea…
EtriCA: Event-Triggered Context-Aware Story Generation Augmented by Cross Attention
Chen Tang, Chenghua Lin, Henglin Huang +2
One of the key challenges of automatic story generation is how to generate a long narrative that can maintain fluency, relevance, and coherence. Despite recent progress, current st…
Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning
Chen Tang, Yizhou Wang, Jianyu Wu +26
Structure-property relationships are foundational to biology, chemistry and materials science, where function, reactivity and physical response emerge from spatial, chemical and pe…
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
Yijun Liu, Yuan Meng, Fang Wu +7
Large language models (LLMs) have exhibited exciting progress in multiple scenarios, while the huge computational demands hinder their deployments in lots of real-world application…
SpecFuse: Ensembling Large Language Models via Next-Segment Prediction
Bo Lv, Nayu Liu, Chen Tang +3
Ensembles of generative large language models (LLMs) are a promising way to compensate for individual model limitations, integrating the strengths of different LLMs. Existing LLM e…
ZPD Teaching Strategies for Deep Reinforcement Learning from Demonstrations
Daniel Seita, David Chan, Roshan Rao +3
Learning from demonstrations is a popular tool for accelerating and reducing the exploration requirements of reinforcement learning. When providing expert demonstrations to human s…
Pre-training on Synthetic Driving Data for Trajectory Prediction
Yiheng Li, Seth Z. Zhao, Chenfeng Xu +5
Accumulating substantial volumes of real-world driving data proves pivotal in the realm of trajectory forecasting for autonomous driving. Given the heavy reliance of current trajec…
STAR: Skeleton-aware Text-based 4D Avatar Generation with In-Network Motion Retargeting
Zenghao Chai, Chen Tang, Yongkang Wong +1
The creation of 4D avatars (i.e., animated 3D avatars) from text description typically uses text-to-image (T2I) diffusion models to synthesize 3D avatars in the canonical space and…
Recent Advances in Neural Text Generation: A Task-Agnostic Survey
Chen Tang, Frank Guerin, Chenghua Lin
In recent years, considerable research has been dedicated to the application of neural models in the field of natural language generation (NLG). The primary objective is to generat…
MEReQ: Max-Ent Residual-Q Inverse RL for Sample-Efficient Alignment from Intervention
Yuxin Chen, Chen Tang, Jianglan Wei +6
Aligning robot behavior with human preferences is crucial for deploying embodied AI agents in human-centered environments. A promising solution is interactive imitation learning fr…
TIC-VLA: A Think-in-Control Vision-Language-Action Model for Robot Navigation in Dynamic Environments
Zhiyu Huang, Yun Zhang, Johnson Liu +3
Robots in dynamic, human-centric environments must follow language instructions while maintaining real-time reactive control. Vision-language-action (VLA) models offer a promising…
3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding
Xiaoye Wang, Chen Tang, Xiangyu Yue +1
This paper addresses the challenge of training a single network to jointly perform multiple dense prediction tasks, such as segmentation and depth estimation, i.e., multi-task lear…
Dealing with the Unknown: Pessimistic Offline Reinforcement Learning
Jinning Li, Chen Tang, Masayoshi Tomizuka +1
Reinforcement Learning (RL) has been shown effective in domains where the agent can learn policies by actively interacting with its operating environment. However, if we change the…
Enhancing Dialogue Generation via Dynamic Graph Knowledge Aggregation
Chen Tang, Hongbo Zhang, Tyler Loakman +2
Incorporating external graph knowledge into neural chatbot models has been proven effective for enhancing dialogue generation. However, in conventional graph neural networks (GNNs)…
BeTAIL: Behavior Transformer Adversarial Imitation Learning from Human Racing Gameplay
Catherine Weaver, Chen Tang, Ce Hao +3
Imitation learning learns a policy from demonstrations without requiring hand-designed reward functions. In many robotic tasks, such as autonomous racing, imitated policies must mo…
ACL Anthology Helper: A Tool to Retrieve and Manage Literature from ACL Anthology
Chen Tang, Frank Guerin, Chenghua Lin
The ACL Anthology is an online repository that serves as a comprehensive collection of publications in the field of natural language processing (NLP) and computational linguistics…
HiEdit: Lifelong Model Editing with Hierarchical Reinforcement Learning
Yangfan Wang, Tianyang Sun, Chen Tang +3
Lifelong model editing (LME) aims to sequentially rectify outdated or inaccurate knowledge in deployed LLMs while minimizing side effects on unrelated inputs. However, existing app…
SciIF: Benchmarking Scientific Instruction Following Towards Rigorous Scientific Intelligence
Encheng Su, Jianyu Wu, Chen Tang +9
As large language models (LLMs) transition from general knowledge retrieval to complex scientific discovery, their evaluation standards must also incorporate the rigorous norms of…
MesonGS: Post-training Compression of 3D Gaussians via Efficient Attribute Transformation
Shuzhao Xie, Weixiang Zhang, Chen Tang +4
3D Gaussian Splatting demonstrates excellent quality and speed in novel view synthesis. Nevertheless, the huge file size of the 3D Gaussians presents challenges for transmission an…
Exploring Social Posterior Collapse in Variational Autoencoder for Interaction Modeling
Chen Tang, Wei Zhan, Masayoshi Tomizuka
Multi-agent behavior modeling and trajectory forecasting are crucial for the safe navigation of autonomous agents in interactive scenarios. Variational Autoencoder (VAE) has been w…
Emphasising Structured Information: Integrating Abstract Meaning Representation into LLMs for Enhanced Open-Domain Dialogue Evaluation
Bohao Yang, Kun Zhao, Dong Liu +3
Automatic open-domain dialogue evaluation has attracted increasing attention, yet remains challenging due to the complexity of assessing response appropriateness. Traditional evalu…
LECTOR: Joint Optimization of Scientific Reasoning Graphs and Introduction Generation
Jiabei Xiao, Yizhou Wang, Chen Tang +3
AI Scientists have shown promising progress across multiple stages of the research pipeline, among which automatic scientific paper writing remains a formidable challenge. The Intr…
MimiCAT: Mimic with Correspondence-Aware Cascade-Transformer for Category-Free 3D Pose Transfer
Zenghao Chai, Chen Tang, Yongkang Wong +2
3D pose transfer aims to transfer the pose-style of a source mesh to a target character while preserving both the target's geometry and the source's pose characteristic. Existing m…
Coverage-Driven Adaptive Keyframe Selection for Video Understanding
Junyang Zhang, Puhan Luo, Chen Tang +2
Recent advances in large vision-language models (LVLMs) have enabled long-video understanding and analysis. However, processing the large number of frames in a video incurs substan…
Effective Distillation of Table-based Reasoning Ability from LLMs
Bohao Yang, Chen Tang, Kun Zhao +2
Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of natural language processing tasks. However, their enormous parameter size and extremely…
TMPQ-DM: Joint Timestep Reduction and Quantization Precision Selection for Efficient Diffusion Models
Haojun Sun, Chen Tang, Zhi Wang +4
Diffusion models have emerged as preeminent contenders in the realm of generative models. Distinguished by their distinctive sequential generative processes, characterized by hundr…
Outracing Human Racers with Model-based Planning and Control for Time-trial Racing
Ce Hao, Chen Tang, Eric Bergkvist +4
Autonomous racing has become a popular sub-topic of autonomous driving in recent years. The goal of autonomous racing research is to develop software to control the vehicle at its…
VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
Lingxiao Li, Yifan Wang, Xinyan Gao +3
Chain-of-Thought (CoT) prompting has proven remarkably effective for eliciting complex reasoning in large language models (LLMs). Yet, its potential in multimodal large language mo…
ElasticViT: Conflict-aware Supernet Training for Deploying Fast Vision Transformer on Diverse Mobile Devices
Chen Tang, Li Lyna Zhang, Huiqiang Jiang +6
Neural Architecture Search (NAS) has shown promising performance in the automatic design of vision transformers (ViT) exceeding 1G FLOPs. However, designing lightweight and low-lat…
Retraining-free Model Quantization via One-Shot Weight-Coupling Learning
Chen Tang, Yuan Meng, Jiacheng Jiang +5
Quantization is of significance for compressing the over-parameterized deep neural models and deploying them on resource-limited devices. Fixed-precision quantization suffers from…
Improving Biomedical Abstractive Summarisation with Knowledge Aggregation from Citation Papers
Chen Tang, Shun Wang, Tomas Goldsack +1
Abstracts derived from biomedical literature possess distinct domain-specific characteristics, including specialised writing styles and biomedical terminologies, which necessitate…
Test-time Sparsity for Extreme Fast Action Diffusion
Kangye Ji, Yuan Meng, Jianbo Zhou +3
Action diffusion excels at high-fidelity action generation but incurs heavy computational costs owing to its iterative denoising nature. Despite current technologies showing promis…
TwinIR: Coordinated Invisible Dual-Point Attacks on Online HD Map Construction
Haibo Hu, Jianghuai Deng, Chen Tang +3
Online HD map construction is critical to prediction and planning in autonomous driving. We find that existing physical attacks against online map construction are limited by a cro…
Mixed-Precision Neural Network Quantization via Learned Layer-wise Importance
Chen Tang, Kai Ouyang, Zhi Wang +4
The exponentially large discrete search space in mixed-precision quantization (MPQ) makes it hard to determine the optimal bit-width for each layer. Previous works usually resort t…
Learning Online Belief Prediction for Efficient POMDP Planning in Autonomous Driving
Zhiyu Huang, Chen Tang, Chen Lv +2
Effective decision-making in autonomous driving relies on accurate inference of other traffic agents' future behaviors. To achieve this, we propose an online belief-update-based be…
PreTraM: Self-Supervised Pre-training via Connecting Trajectory and Map
Chenfeng Xu, Tian Li, Chen Tang +5
Deep learning has recently achieved significant progress in trajectory forecasting. However, the scarcity of trajectory data inhibits the data-hungry deep-learning models from lear…
Deep Reinforcement Learning for Robotics: A Survey of Real-World Successes
Chen Tang, Ben Abbatematteo, Jiaheng Hu +3
Reinforcement learning (RL), particularly its combination with deep neural networks referred to as deep RL (DRL), has shown tremendous promise across a wide range of applications,…
Simple Recipe Works: Vision-Language-Action Models are Natural Continual Learners with Reinforcement Learning
Jiaheng Hu, Jay Shim, Chen Tang +4
Continual Reinforcement Learning (CRL) for Vision-Language-Action (VLA) models is a promising direction toward self-improving embodied agents that can adapt in openended, evolving…
MesonGS++: Post-training Compression of 3D Gaussian Splatting with Hyperparameter Searching
Shuzhao Xie, Junchen Ge, Weixiang Zhang +10
3D Gaussian Splatting (3DGS) achieves high-quality novel view synthesis with real-time rendering, but its storage cost remains prohibitive for practical deployment. Existing post-t…
SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model Acceleration
Ye Li, Yuan Meng, Zewen Sun +7
Vision-Language-Action (VLA) models have attracted increasing attention for their strong control capabilities. However, their high computational cost and low execution frequency hi…
Editing Driver Character: Socially-Controllable Behavior Generation for Interactive Traffic Simulation
Wei-Jer Chang, Chen Tang, Chenran Li +3
Traffic simulation plays a crucial role in evaluating and improving autonomous driving planning systems. After being deployed on public roads, autonomous vehicles need to interact…
Double-Iterative Gaussian Process Regression for Modeling Error Compensation in Autonomous Racing
Shaoshu Su, Ce Hao, Catherine Weaver +3
Autonomous racing control is a challenging research problem as vehicles are pushed to their limits of handling to achieve an optimal lap time; therefore, vehicles exhibit highly no…
Efficient and Accurate Image Provenance Analysis: A Scalable Pipeline for Large-scale Images
Jiewei Lai, Lan Zhang, Chen Tang +1
The rapid proliferation of modified images on social networks that are driven by widely accessible editing tools demands robust forensic tools for digital governance. Image provena…
Residual Q-Learning: Offline and Online Policy Customization without Value
Chenran Li, Chen Tang, Haruki Nishimura +3
Imitation Learning (IL) is a widely used framework for learning imitative behavior from demonstrations. It is especially appealing for solving complex real-world tasks where handcr…
SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation
Michael J. Munje, Chen Tang, Shuijing Liu +6
Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-Language Models (VLMs) exhibit…
Improving Medical Dialogue Generation with Abstract Meaning Representations
Bohao Yang, Chen Tang, Chenghua Lin
Medical Dialogue Generation serves a critical role in telemedicine by facilitating the dissemination of medical expertise to patients. Existing studies focus on incorporating textu…
AutoAgent: Evolving Cognition and Elastic Memory Orchestration for Adaptive Agents
Xiaoxing Wang, Ning Liao, Shikun Wei +2
Autonomous agent frameworks still struggle to reconcile long-term experiential learning with real-time, context-sensitive decision-making. In practice, this gap appears as static c…
High Thermoelectric Cooling Performance of Junction Thermoelectric Transistors
Chen Tang, Bohang Nan, Xiaodong Liu +1
To achieve high performance thermoelectric materials and devices, thermoelectric transistors, which integrate thermoelectric effects with transistor technology, represent a promisi…
Dynamics of Baxter-Wu model
Chen Tang, Konstantinos Sfairopoulos, Wanzhou Zhang +1
Using Monte Carlo simulations, we investigate the dynamical properties of the Baxter-Wu (BW) model under linear quenches. For the linear cooling process, the scaling behavior of th…
RTF-Q: Efficient Unsupervised Domain Adaptation with Retraining-free Quantization
Nanyang Du, Chen Tang, Yuxiao Jiang +2
Performing unsupervised domain adaptation on resource-constrained edge devices is challenging. Existing research typically adopts architecture optimization (e.g., designing slimmab…
Hierarchical Planning Through Goal-Conditioned Offline Reinforcement Learning
Jinning Li, Chen Tang, Masayoshi Tomizuka +1
Offline Reinforcement learning (RL) has shown potent in many safe-critical tasks in robotics where exploration is risky and expensive. However, it still struggles to acquire skills…
Market Efficiency and the Heterogeneous Impact of Financial Liberalization: Evidence from the Shanghai-Hong Kong Stock Connect
Jiaqi Liu, Chen Tang
This paper examines how the Shanghai-Hong Kong Stock Connect (SHHK Stock Connect) affects the A-H share price premium and whether the policy effect depends on pre-existing market e…
Residual-MPPI: Online Policy Customization for Continuous Control
Pengcheng Wang, Chenran Li, Catherine Weaver +4
Policies developed through Reinforcement Learning (RL) and Imitation Learning (IL) have shown great potential in continuous control tasks, but real-world applications often require…
Point2Primitive: CAD Reconstruction from Point Cloud by Direct Primitive Prediction
Xinzhu Ma, Cheng Wang, Chen Tang +5
Recovering CAD models from point clouds requires reconstructing their topology and sketch-based extrusion primitives. A dominant paradigm for representing sketches involves implici…
FlashMesh: Faster and Better Autoregressive Mesh Synthesis via Structured Speculation
Tingrui Shen, Yiheng Zhang, Chen Tang +6
Autoregressive models can generate high-quality 3D meshes by sequentially producing vertices and faces, but their token-by-token decoding results in slow inference, limiting practi…
AgriEval: A Comprehensive Chinese Agricultural Benchmark for Large Language Models
Lian Yan, Haotian Wang, Chen Tang +5
In the agricultural domain, the deployment of large language models (LLMs) is hindered by the lack of training data and evaluation benchmarks. To mitigate this issue, we propose Ag…
Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers
Lei Chen, Yuan Meng, Chen Tang +5
Recent advancements in diffusion models, particularly the architectural transformation from UNet-based models to Diffusion Transformers (DiTs), significantly improve the quality an…
Accelerating Parallel Diffusion Model Serving with Residual Compression
Jiajun Luo, Yicheng Xiao, Jianru Xu +5
Diffusion models produce realistic images and videos but require substantial computational resources, necessitating multi-accelerator parallelism for real-time deployment. However,…
SLIDE: A Framework Integrating Small and Large Language Models for Open-Domain Dialogues Evaluation
Kun Zhao, Bohao Yang, Chen Tang +2
The long-standing one-to-many problem of gold standard responses in open-domain dialogue systems presents challenges for automatic evaluation metrics. Though prior works have demon…
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale
Yicheng Zou, Dongsheng Zhu, Lin Zhu +174
We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancem…
Improving Chinese Story Generation via Awareness of Syntactic Dependencies and Semantics
Henglin Huang, Chen Tang, Tyler Loakman +2
Story generation aims to generate a long narrative conditioned on a given input. In spite of the success of prior works with the application of pre-trained models, current neural m…
KA2L: A Knowledge-Aware Active Learning Framework for LLMs
Haoxuan Yin, Bojian Liu, Chen Tang +3
Fine-tuning large language models (LLMs) with high-quality knowledge has been shown to enhance their performance effectively. However, there is a paucity of research on the depth o…
Adaptive Probabilistic Vehicle Trajectory Prediction Through Physically Feasible Bayesian Recurrent Neural Network
Chen Tang, Jianyu Chen, Masayoshi Tomizuka
Probabilistic vehicle trajectory prediction is essential for robust safety of autonomous driving. Current methods for long-term trajectory prediction cannot guarantee the physical…
KCS: Diversify Multi-hop Question Generation with Knowledge Composition Sampling
Yangfan Wang, Jie Liu, Chen Tang +2
Multi-hop question answering faces substantial challenges due to data sparsity, which increases the likelihood of language models learning spurious patterns. To address this issue,…
A Survey of Scientific Large Language Models: From Data Foundations to Agent Frontiers
Ming Hu, Chenglong Ma, Wei Li +117
Scientific Large Language Models (Sci-LLMs) are transforming how knowledge is represented, integrated, and applied in scientific research, yet their progress is shaped by the compl…
TerraTransfer: Learning End-to-End Driving Policies Without Expert Demonstrations
Zikang Xiong, Weixin Li, Zhouchonghao Wu +6
The paper proposes a method to train end-to-end autonomous driving policies without expert demonstrations by pretraining a policy via self‑play in a fast vectorized simulator and t…
Optimizing Diffusion Models for Joint Trajectory Prediction and Controllable Generation
Yixiao Wang, Chen Tang, Lingfeng Sun +8
Diffusion models are promising for joint trajectory prediction and controllable generation in autonomous driving, but they face challenges of inefficient inference steps and high c…