Independent research: GraphRAG makes AI agents 80% more truthful | Read the report

NODES 26 — November 12, 2026

Scalable Automated Enterprise Level Data Generation

Session track: Data Intelligence

Session time:

Session description:

We present a high-quality, multi-domain dataset for the Text2Cypher task which is enabling the translation of natural language (NL) questions into executable Cypher queries over graph databases. The dataset was curated using an automated pipeline. The dataset comprises 27,529 NL queries and corresponding Cyphers spanning across 11 real-world graph datasets, each accompanied by its corresponding graph database for grounded query execution. To ensure correctness, the queries are validated through a rigorous pipeline combining automated schema, runtime and value checks, along with manual review for logical correctness. Queries are further categorized by complexity to support fine-grained evaluation. We further uncover different hallucination patterns quite often seen in the LLM Cypher Generations. We have released our benchmark dataset and code to replicate our data synthesis pipeline on new graph datasets, supporting extensibility and future research for the task of Text2Cypher. The work has been published at EMNLP, 2025 (Industry Track) [https://aclanthology.org/2025.emnlp-industry.133.pdf].

Speaker

photo of Vashu Chauhan

Vashu Chauhan

Research Fellow, Microsoft

Vashu Chauhan is a Research Fellow at Microsoft Research. He has worked with Adobe for nearly one year on multiple interesting Knowledge Graph-related projects like Brand Genome [https://brand-genome.github.io/]. He worked with IBM Research on building automation pipeline for curating enterprise level high-fidelity Knowledge Graph datasets.