Session track: Data Intelligence
Session time:
Session description:
We present a high-quality, multi-domain dataset for the Text2Cypher task which is enabling the translation of natural language (NL) questions into executable Cypher queries over graph databases. The dataset was curated using an automated pipeline. The dataset comprises 27,529 NL queries and corresponding Cyphers spanning across 11 real-world graph datasets, each accompanied by its corresponding graph database for grounded query execution. To ensure correctness, the queries are validated through a rigorous pipeline combining automated schema, runtime and value checks, along with manual review for logical correctness. Queries are further categorized by complexity to support fine-grained evaluation. We further uncover different hallucination patterns quite often seen in the LLM Cypher Generations. We have released our benchmark dataset and code to replicate our data synthesis pipeline on new graph datasets, supporting extensibility and future research for the task of Text2Cypher. The work has been published at EMNLP, 2025 (Industry Track) [https://aclanthology.org/2025.emnlp-industry.133.pdf].
Speaker

Research Fellow, Microsoft
Vashu Chauhan is a Research Fellow at Microsoft Research. He has worked with Adobe for nearly one year on multiple interesting Knowledge Graph-related projects like Brand Genome [https://brand-genome.github.io/]. He worked with IBM Research on building automation pipeline for curating enterprise level high-fidelity Knowledge Graph datasets.