Publications (91)
RetClean: Retrieval-Based Data Cleaning Using Foundation Models and Data Lakes
Zan Ahmad Naeem, Mohammad Shahmeer Ahmad, Mohamed Eltabakh +2
Can foundation models (such as ChatGPT) clean your data? In this proposal, we demonstrate that indeed ChatGPT can assist in data cleaning by suggesting corrections for specific cel…
ReLAM: Learning Anticipation Model for Rewarding Visual Robotic Manipulation
Nan Tang, Jing-Cheng Pang, Guanlin Li +2
Reward design remains a critical bottleneck in visual reinforcement learning (RL) for robotic manipulation. In simulated environments, rewards are conventionally designed based on…
Observation via spin Seebeck effect of macroscopic magnetic transport from emergent magnetic monopoles
Nan Tang, Stephan Glamsch, Aisha Aqeel +7
Magnetic monopoles, elusive in high-energy physics, have been realised as emergent quasiparticles in solid-state systems, where their unique properties hold promise for novel spint…
Phonon spectrum of PrZrO and PrIrO as an evidence of coupling of the lattice with electronic and magnetic degrees of freedom
Yuanyuan Xu, Huiyuan Man, Nan Tang +5
Magnetic materials with pyrochlore crystal structure form exotic magnetic states due to the high lattice frustration. In this work we follow the effects of coupling of the lattice…
SCAR: A Characterization Scheme for Multi-Modal Dataset
Ri Su, Zhao Chen, Caleb Chen Cao +2
Foundation models exhibit remarkable generalization across diverse tasks, largely driven by the characteristics of their training data. Recent data-centric methods like pruning and…
Hole doping in compositionally complex correlated oxide enables tunable exchange biasing
Alessandro R. Mazza, Elizabeth Skoropata, Jason Lapano +8
Magnetic interfaces and the phenomena arising from them drive both the design of modern spintronics and fundamental research. Recently, it was revealed that through designing magne…
Importance of dynamic lattice effects for crystal field excitations in quantum spin ice candidate PrZrO
Yuanyuan Xu, Huiyuan Man, Nan Tang +5
PrZrO is a pyrochlore quantum spin-ice candidate. Using Raman scattering spectroscopy we probe crystal electric field excitations of Pr, and demonstrate the impo…
DocSage: An Information Structuring Agent for Multi-Doc Multi-Entity Question Answering
Teng Lin, Yizhang Zhu, Zhengxuan Zhang +2
Multi-document Multi-entity Question Answering inherently demands models to track implicit logic between multiple entities across scattered documents. However, existing Large Langu…
Will LLMs be Professional at Fund Investment? DeepFund: A Live Arena Perspective
Changlun Li, Yao Shi, Yuyu Luo +1
Large Language Models (LLMs) have demonstrated impressive capabilities across various domains, but their effectiveness in financial decision-making remains inadequately evaluated.…
Magnetism in Metastable and Annealed Compositionally Complex Alloys
Nan Tang, Lizabeth Quigley, Walker L. Boldman +6
Compositionally complex materials (CCMs) present a potential paradigm shift in the design of magnetic materials. These alloys exhibit long-range structural order coupled with limit…
HAIChart: Human and AI Paired Visualization System
Yupeng Xie, Yuyu Luo, Guoliang Li +1
The growing importance of data visualization in business intelligence and data science emphasizes the need for tools that can efficiently generate meaningful visualizations from la…
AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent Framework
Meihao Fan, Ju Fan, Nan Tang +3
Answering natural language (NL) questions about tables, known as Tabular Question Answering (TQA), is crucial because it allows users to quickly and efficiently extract meaningful…
Controlling magnetic configuration in soft-hard bilayers probed by polarized neutron reflectometry
Nan Tang, Jung-Wei Liao, Siu-Tat Chui +4
Hard/soft magnetic bilayer thin films have been widely used in data storage technologies and permanent magnet applications. The magnetic configuration and response to temperatures…
Skyrmion-Excited Spin Wave Fractal Network
Nan Tang, W. L. N. C. Liyanage, Sergio A. Montoya +9
Magnetic skyrmions exhibit unique, technologically relevant pseudo-particle behaviors which arise from their topological protection, including well-defined, three-dimensional dynam…
Review-Then-Refine: A Dynamic Framework for Multi-Hop Question Answering with Temporal Adaptability
Xiangsen Chen, Xuming Hu, Nan Tang
Retrieve-augmented generation (RAG) frameworks have emerged as a promising solution to multi-hop question answering(QA) tasks since it enables large language models (LLMs) to incor…
PASTA: Table-Operations Aware Fact Verification via Sentence-Table Cloze Pre-training
Zihui Gu, Ju Fan, Nan Tang +3
Fact verification has attracted a lot of research attention recently, e.g., in journalism, marketing, and policymaking, as misinformation and disinformation online can sway one's o…
A Risk Decomposition Framework for Pre-Hoc Fine-Tuning Prediction
Yuxiang Luo, Chen Wang, Nan Tang
The high cost of fine-tuning LLMs poses a significant economic barrier; pre-hoc performance prediction offers a critical solution to substantially reduce this expense. However, the…
On Summarizing Graph Streams
Nan Tang, Qing Chen, Prasenjit Mitra
Graph streams, which refer to the graph with edges being updated sequentially in a form of a stream, have wide applications such as cyber security, social networks and transportati…
ChatPipe: Orchestrating Data Preparation Program by Optimizing Human-ChatGPT Interactions
Sibei Chen, Hanbing Liu, Weiting Jin +5
Orchestrating a high-quality data preparation program is essential for successful machine learning (ML), but it is known to be time and effort consuming. Despite the impressive cap…
A Unified Model for Cardinality Estimation by Learning from Data and Queries via Sum-Product Networks
Jiawei Liu, Ju Fan, Tongyu Liu +5
Cardinality estimation is a fundamental component in database systems, crucial for generating efficient execution plans. Despite advancements in learning-based cardinality estimati…
LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning
Xiaotian Lin, Yanlin Qi, Yizhang Zhu +4
Instruction tuning has emerged as a critical paradigm for improving the capabilities and alignment of large language models (LLMs). However, existing iterative model-aware data sel…
BWArea Model: Learning World Model, Inverse Dynamics, and Policy for Controllable Language Generation
Chengxing Jia, Pengyuan Wang, Ziniu Li +4
Large language models (LLMs) have catalyzed a paradigm shift in natural language processing, yet their limited controllability poses a significant challenge for downstream applicat…
RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation
Nan Tang, Ju Fan, Fangyi Li +5
Can AI help automate human-easy but computer-hard data preparation tasks that burden data scientists, practitioners, and crowd workers? We answer this question by presenting RPT, a…
SRAG: Structured Retrieval-Augmented Generation for Multi-Entity Question Answering over Wikipedia Graph
Teng Lin, Yizhang Zhu, Yuyu Luo +1
Multi-entity question answering (MEQA) poses significant challenges for large language models (LLMs), which often struggle to consolidate scattered information across multiple docu…
InteractComp: Evaluating Search Agents With Ambiguous Queries
Mingyi Deng, Lijun Huang, Yani Fan +23
Language agents have demonstrated remarkable potential in web search and information retrieval. However, many search-agent benchmarks assume that user queries are complete and unam…
EllieSQL: Cost-Efficient Text-to-SQL with Complexity-Aware Routing
Yizhang Zhu, Runzhi Jiang, Boyan Li +2
Text-to-SQL automatically translates natural language queries to SQL, allowing non-technical users to retrieve data from databases without specialized SQL knowledge. Despite the su…
Empowering Language Models with Active Inquiry for Deeper Understanding
Jing-Cheng Pang, Heng-Bo Fan, Pengyuan Wang +6
The rise of large language models (LLMs) has revolutionized the way that we interact with artificial intelligence systems through natural language. However, LLMs often misinterpret…
Reuse and Adaptation for Entity Resolution through Transfer Learning
Saravanan Thirumuruganathan, Shameem A Puthiya Parambath, Mourad Ouzzani +2
Entity resolution (ER) is one of the fundamental problems in data integration, where machine learning (ML) based classifiers often provide the state-of-the-art results. Considerabl…
Data Agents: Levels, State of the Art, and Open Problems
Yuyu Luo, Guoliang Li, Ju Fan +1
Data agents are an emerging paradigm that leverages large language models (LLMs) and tool-using agents to automate data management, preparation, and analysis tasks. However, the te…
Time Travel is Cheating: Going Live with DeepFund for Real-Time Fund Investment Benchmarking
Changlun Li, Yao Shi, Chen Wang +7
Large Language Models (LLMs) have demonstrated notable capabilities across financial tasks, including financial report summarization, earnings call transcript analysis, and asset c…
SEED: Domain-Specific Data Curation With Large Language Models
Zui Chen, Lei Cao, Sam Madden +7
Data curation tasks that prepare data for analytics are critical for turning data into actionable insights. However, due to the diverse requirements of applications in different do…
Structural and Disentangled Adaptation of Large Vision Language Models for Multimodal Recommendation
Zhongtao Rao, Peilin Zhou, Dading Chong +3
Multimodal recommendation enhances accuracy by leveraging visual and textual signals, and its success largely depends on learning high-quality cross-modal representations. Recent a…
Unsupervised String Transformation Learning for Entity Consolidation
Dong Deng, Wenbo Tao, Ziawasch Abedjan +7
Data integration has been a long-standing challenge in data management with many applications. A key step in data integration is entity consolidation. It takes a collection of clus…
CSTrader: A Testbed for Language-Grounded Trading in a Community-Driven Virtual Asset Market
Yao Shi, Kingfung Luo, Nan Tang +1
Niche asset markets, such as Counter-Strike 2 (CS2) weapon skins, are small, volatile, and heavily driven by community discussions and platform rules. These properties make them ha…
Realization of Ordered Magnetic Skyrmions in Thin Films at Ambient Conditions
Ryan D. Desautels, Lisa DeBeer-Schmitt, Sergio Montoya +7
Magnetic skyrmions present interesting physics due to their topological nature and hold significant promise for future information technologies. A key barrier to realizing skyrmion…
Efficient Algorithms for Approximate Single-Source Personalized PageRank Queries
Sibo Wang, Renchi Yang, Runhui Wang +5
Given a graph , a source node and a target node , the personalized PageRank (PPR) of with respect to is the probability that a random walk starting from termi…
ChartInsights: Evaluating Multimodal Large Language Models for Low-Level Chart Question Answering
Yifan Wu, Lutao Yan, Leixian Shen +3
Chart question answering (ChartQA) tasks play a critical role in interpreting and extracting insights from visualization charts. While recent advancements in multimodal large langu…
Concise Reasoning, Big Gains: Pruning Long Reasoning Trace with Difficulty-Aware Prompting
Yifan Wu, Jingze Shi, Bingheng Wu +4
Existing chain-of-thought (CoT) distillation methods can effectively transfer reasoning abilities to base models but suffer from two major limitations: excessive verbosity of reaso…
AskChart: Universal Chart Understanding through Textual Enhancement
Xudong Yang, Yifan Wu, Yizhang Zhu +2
Chart understanding tasks such as ChartQA and Chart-to-Text involve automatically extracting and interpreting key information from charts, enabling users to query or convert visual…
LakeHopper: Cross Data Lakes Column Type Annotation through Model Adaptation
Yushi Sun, Xujia Li, Nan Tang +3
Column type annotation is vital for tasks like data cleaning, integration, and visualization. Recent solutions rely on resource-intensive language models fine-tuned on well-annotat…
Rise of the Community Champions: From Reviewer Crunch to Community Power
Changlun Li, Yao Shi, Yuyu Luo +1
Academic publishing is facing a crisis driven by exponential growth in submissions and an overwhelmed peer review system, leading to inconsistent decisions and a severe reviewer sh…
The Dawn of Natural Language to SQL: Are We Fully Ready?
Boyan Li, Yuyu Luo, Chengliang Chai +2
Translating users' natural language questions into SQL queries (i.e., NL2SQL) significantly lowers the barriers to accessing relational databases. The emergence of Large Language M…
Are Large Language Models a Good Replacement of Taxonomies?
Yushi Sun, Hao Xin, Kai Sun +5
Large language models (LLMs) demonstrate an impressive ability to internalize knowledge and answer natural language questions. Although previous studies validate that LLMs perform…
A Survey of Text-to-SQL in the Era of LLMs: Where are we, and where are we going?
Xinyu Liu, Shuyu Shen, Boyan Li +7
Translating users' natural language queries (NL) into SQL queries (i.e., Text-to-SQL, a.k.a. NL2SQL) can significantly reduce barriers to accessing relational databases and support…
Technical Report: Optimizing Human Involvement for Entity Matching and Consolidation
Ji Sun, Dong Deng, Ihab Ilyas +5
An end-to-end data integration system requires human feedback in several phases, including collecting training data for entity matching, debugging the resulting clusters, confirmin…
Knowledgeable Agents by Offline Reinforcement Learning from Large Language Model Rollouts
Jing-Cheng Pang, Si-Hang Yang, Kaiyuan Li +4
Reinforcement learning (RL) trains agents to accomplish complex tasks through environmental interaction data, but its capacity is also limited by the scope of the available data. T…
KERAG: Knowledge-Enhanced Retrieval-Augmented Generation for Advanced Question Answering
Yushi Sun, Kai Sun, Yifan Ethan Xu +4
Retrieval-Augmented Generation (RAG) mitigates hallucination in Large Language Models (LLMs) by incorporating external data, with Knowledge Graphs (KGs) offering crucial informatio…
Alpha-SQL: Zero-Shot Text-to-SQL using Monte Carlo Tree Search
Boyan Li, Jiayi Zhang, Ju Fan +4
Text-to-SQL, which enables natural language interaction with databases, serves as a pivotal method across diverse industries. With new, more powerful large language models (LLMs) e…
Interleaving Pre-Trained Language Models and Large Language Models for Zero-Shot NL2SQL Generation
Zihui Gu, Ju Fan, Nan Tang +7
Zero-shot NL2SQL is crucial in achieving natural language to SQL that is adaptive to new environments (e.g., new databases, new linguistic phenomena or SQL structures) with zero an…
Are Large Language Models Good Statisticians?
Yizhang Zhu, Shiyin Du, Boyan Li +2
Large Language Models (LLMs) have demonstrated impressive capabilities across a range of scientific tasks including mathematics, physics, and chemistry. Despite their successes, th…
Three-Dimensional Structure of Hybrid Magnetic Skyrmions Determined by Neutron Scattering
WLNC Liyanage, Nan Tang, Lizabeth Quigley +9
Magnetic skyrmions are topologically protected chiral spin textures which present opportunities for next-generation magnetic data storage and logic information technologies. The to…
Text2GraphQuery-Bench: A Text to Graph Query Benchmark
Songlin Lyu, Lujie Ban, Zihang Wu +14
Graph models are fundamental to data analysis in domains rich with complex relationships. Unlike SQL, which benefits from a rel- atively unified standard and widespread familiarity…
TuneAhead: Predicting Fine-tuning Performance Before Full Training Begins
Yuxiang Luo, Haonan Long, Chen Wang +6
Fine-tuning large language models (LLMs) is compute-intensive and error-prone: model performance depends sensitively on data quality and hyperparameter choices, and naïve runs can…
RAMer: Reconstruction-based Adversarial Model for Multi-party Multi-modal Multi-label Emotion Recognition
Xudong Yang, Yizhang Zhu, Hanfeng Liu +3
Conventional Multi-modal multi-label emotion recognition (MMER) assumes complete access to visual, textual, and acoustic modalities. However, real-world multi-party settings often…
TransXSSM: A Hybrid Transformer State Space Model with Unified Rotary Position Embedding
Bingheng Wu, Jingze Shi, Yifan Wu +2
Transformers exhibit proficiency in capturing long-range dependencies, whereas State Space Models (SSMs) facilitate linear-time sequence modeling. Notwithstanding their synergistic…
Can Agentic Trading Systems Pay for Their Own Intelligence?
Qiqi Duan, Changlun Li, Chen Wang +10
Large language model (LLM) agents are increasingly used in trading systems, where model reasoning, tool use, and continual decisions incur costs that are expected to produce tradin…
MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question Answering
Teng Lin, Yuyu Luo, Honglin Zhang +4
Multi-entity question answering (MEQA) represents significant challenges for large language models (LLM) and retrieval-augmented generation (RAG) systems, which frequently struggle…
EVOQUANT: Self-Evolving Verifier-Guided Strategy Optimization for Robust Quantitative Trading
Jie Mao, Changlun Li, Xiang Li +7
EVOQUANT is a framework that uses large language models together with a verifier pipeline to automatically diagnose, edit, and improve quantitative trading strategies, achieving hi…
NL2SQL-BUGs: A Benchmark for Detecting Semantic Errors in NL2SQL Translation
Xinyu Liu, Shuyu Shen, Boyan Li +2
Natural Language to SQL (i.e., NL2SQL) translation is crucial for democratizing database access, but even state-of-the-art models frequently generate semantically incorrect SQL que…
CRAG -- Comprehensive RAG Benchmark
Xiao Yang, Kai Sun, Hao Xin +24
Retrieval-Augmented Generation (RAG) has recently emerged as a promising solution to alleviate Large Language Model (LLM)'s deficiency in lack of knowledge. Existing RAG datasets,…
Electrical observation via spin Seebeck effect of fractionalized excitations in a magnetic insulator
Nan Tang, Josef Willsher, Stephan Glamsch +9
Fractionalized excitations are among the most striking signatures of emergence in quantum matter. While widely sought in frustrated magnets, their detection and characterization re…
EXCLAIM: An Explainable Cross-Modal Agentic System for Misinformation Detection with Hierarchical Retrieval
Yin Wu, Zhengxuan Zhang, Fuling Wang +3
Misinformation continues to pose a significant challenge in today's information ecosystem, profoundly shaping public perception and behavior. Among its various manifestations, Out-…
DeepER -- Deep Entity Resolution
Muhammad Ebraheem, Saravanan Thirumuruganathan, Shafiq Joty +2
Entity resolution (ER) is a key data integration problem. Despite the efforts in 70+ years in all aspects of ER, there is still a high demand for democratizing ER - humans are heav…
DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
Boyan Li, Zhuowen Liang, Yupeng Xie +11
Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multi…
nvBench 2.0: Resolving Ambiguity in Text-to-Visualization through Stepwise Reasoning
Tianqi Luo, Chuhan Huang, Leixian Shen +5
Text-to-Visualization (Text2VIS) enables users to create visualizations from natural language queries, making data insights more accessible. However, Text2VIS faces challenges in i…
A Survey of Data Agents: Emerging Paradigm or Overstated Hype?
Yizhang Zhu, Liangwei Wang, Chenyu Yang +22
The rapid advancement of large language models (LLMs) has spurred the emergence of data agents, autonomous systems designed to orchestrate Data + AI ecosystems for tackling complex…
Automatic Database Configuration Debugging using Retrieval-Augmented Language Models
Sibei Chen, Ju Fan, Bin Wu +8
Database management system (DBMS) configuration debugging, e.g., diagnosing poorly configured DBMS knobs and generating troubleshooting recommendations, is crucial in optimizing DB…
A Plug-and-Play Natural Language Rewriter for Natural Language to SQL
Peixian Ma, Boyan Li, Runzhi Jiang +3
Existing Natural Language to SQL (NL2SQL) solutions have made significant advancements, yet challenges persist in interpreting and translating NL queries, primarily due to users' l…
Tensile and compressive strain tuning of a Kondo lattice
Soumendra Nath Panja, Anton Jesche, Nan Tang +1
We present electrical resistivity measurements on the prototypical heavy-fermion metal YbRhSi (YRS) under -axis tensile and compressive strain and focus on the evolu…
DTBench: A Synthetic Benchmark for Document-to-Table Extraction
Yuxiang Guo, Zhuoran Du, Nan Tang +3
Document-to-table (Doc2Table) extraction derives structured tables from unstructured documents under a target schema, enabling reliable and verifiable SQL-based data analytics. Alt…
ChartEditor: A Human-AI Paired Tool for Authoring Pictorial Charts
Siyu Yan, Tiancheng Liu, Weikai Yang +2
Pictorial charts are favored for their memorability and visual appeal, offering a more engaging alternative to basic charts. However, their creation can be complex and time-consumi…
TABQAWORLD: Optimizing Multimodal Reasoning for Multi-Turn Table Question Answering
Tung Sum Thomas Kwok, Xinyu Wang, Xiaofeng Lin +7
Multimodal reasoning has emerged as a powerful framework for enhancing reasoning capabilities of reasoning models. While multi-turn table reasoning methods have improved reasoning…
Long-Document QA with Chain-of-Structured-Thought and Fine-Tuned SLMs
Zhuowen Liang, Xiaotian Lin, Zhengxuan Zhang +3
Large language models (LLMs) are widely applied to data analytics over documents, yet direct reasoning over long, noisy documents remains brittle and error-prone. Hence, we study d…
Data Curation with Deep Learning [Vision]
Saravanan Thirumuruganathan, Nan Tang, Mourad Ouzzani +1
Data curation - the process of discovering, integrating, and cleaning data - is one of the oldest, hardest, yet inevitable data management problems. Despite decades of efforts from…
VLGOR: Visual-Language Knowledge Guided Offline Reinforcement Learning for Generalizable Agents
Pengsen Liu, Maosen Zeng, Nan Tang +4
Combining Large Language Models (LLMs) with Reinforcement Learning (RL) enables agents to interpret language instructions more effectively for task execution. However, LLMs typical…
Monte Carlo Tree Search for Table-to-Multimodal Report Generation
Teng Lin, Zhiyang Zhang, Yuyu Luo +1
Automatically generating professional multimodal reports comprising both textual analysis and visual charts from structured tabular data is a critical challenge in data intelligenc…
Crystal-field magnetostriction of the spin ice under ultrahigh magnetic fields
Nan Tang, Masaki Gen, Martin Rotter +9
We present a comprehensive study of the magnetoelastic properties of the Ising pyrochlore oxide HoTiO, known as spin ice, by means of high-field magnetostriction…
AnnoRetrieve: Efficient Structured Retrieval for Unstructured Document Analysis
Teng Lin, Yuyu Luo, Nan Tang
Unstructured documents dominate enterprise and web data, but their lack of explicit organization hinders precise information retrieval. Current mainstream retrieval methods, especi…
Dataset-On-Demand: Automatic View Search and Presentation for Data Discovery
Raul Castro Fernandez, Nan Tang, Mourad Ouzzani +2
Many data problems are solved when the right view of a combination of datasets is identified. Finding such a view is challenging because of the many tables spread across many datab…
Cost-Effective In-Context Learning for Entity Resolution: A Design Space Exploration
Meihao Fan, Xiaoyue Han, Ju Fan +4
Entity resolution (ER) is an important data integration task with a wide spectrum of applications. The state-of-the-art solutions on ER rely on pre-trained language models (PLMs),…
Decoding Ancient Oracle Bone Script via Generative Dictionary Retrieval
Yin Wu, Gangjian Zhang, Jiayu Chen +4
Understanding humanity's earliest writing systems is crucial for reconstructing civilization's origins, yet many ancient scripts remain undeciphered. Oracle Bone Script (OBS) from…
Deductive Optimization of Relational Data Storage
John K. Feser, Samuel Madden, Nan Tang +1
Optimizing the physical data storage and retrieval of data are two key database management problems. In this paper, we propose a language that can express a wide range of physical…
Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction
Zhuowen Liang, Zhengxuan Zhang, Jiayang Wang +2
Practical AI systems increasingly need to turn long, heterogeneous documents into queryable relational databases, not isolated spreadsheets. In domains such as finance, healthcare,…
DataPuzzle: Breaking Free from the Hallucinated Promise of LLMs in Data Analysis
Zhengxuan Zhang, Zhuowen Liang, Yin Wu +3
Large language models (LLMs) are increasingly applied to multi-modal data analysis -- not necessarily because they offer the most precise answers, but because they provide fluent,…
VerifAI: Verified Generative AI
Nan Tang, Chenyu Yang, Ju Fan +3
Generative AI has made significant strides, yet concerns about the accuracy and reliability of its outputs continue to grow. Such inaccuracies can have serious consequences such as…
Fine-Grained Knowledge Structuring and Retrieval for Visual Question Answering
Zhengxuan Zhang, Yin Wu, Yuyu Luo +1
Visual Question Answering (VQA) focuses on providing answers to natural language questions by utilizing information from images. Although cutting-edge multimodal large language mod…
Boosting Text-to-Chart Retrieval through Training with Synthesized Semantic Insights
Yifan Wu, Lutao Yan, Yizhang Zhu +6
Text-to-chart retrieval, enabling users to find relevant charts via natural language queries, has gained significant attention. However, evaluating models in real-world business in…
Quantum Griffiths phase in the kagome Kondo lattice CeRhPdSn
Nan Tang, Rajesh Tripathi, Yasuyuki Shimura +3
CeRhSn is a valence fluctuating heavy-fermion metal with a twisted Ce-kagome lattice, displaying zero-field quantum criticality, previously associated with geometrical frustration.…
CyberThreat-Eval: Can Large Language Models Automate Real-World Threat Research?
Xiangsen Chen, Xuan Feng, Shuo Chen +5
Analyzing Open Source Intelligence (OSINT) from large volumes of data is critical for drafting and publishing comprehensive CTI reports. This process usually follows a three-stage…
SketchFill: Sketch-Guided Code Generation for Imputing Derived Missing Values
Yunfan Zhang, Changlun Li, Yuyu Luo +1
Missing value is a critical issue in data science, significantly impacting the reliability of analyses and predictions. Missing value imputation (MVI) is a longstanding problem bec…
Anti-microbial properties of a multi-component alloy
Anne F. Murray, Daniel Bryan, David A. Garfinkel +8
High traffic touch surfaces such as doorknobs, countertops, and handrails can be transmission points for the spread of pathogens, emphasizing the need to develop materials that act…