Index configuration

This page describes how to configure Neo4j indexes to enhance search performance and enable full-text search. The supported index types are:

All types of indexes can be created and dropped using Cypher and they can also all be used to index both nodes and relationships. The token lookup index is the only index present by default in the database.

Range, point, text, and full-text indexes provide a mapping from a property value to an entity (node or relationship). Token lookup indexes are different and provide a mapping from labels to nodes, or relationship types to relationships, instead of between properties and entities.

When you write a Cypher query, you do not need to specify which indexes to use. Cypher’s query planner decides which of the available indexes to use.

The rest of this page provides information on the available indexes and their configuration aspects. For further details on creating, querying, and dropping indexes, see Cypher Manual → Indexes for search performance, Full-text indexes, and Vector indexes.

The type of an index can be identified according to the table below:

Index type Cypher command Core API

Range index

SHOW RANGE INDEXES

org.neo4j.graphdb.schema.IndexType#RANGE

Point index

SHOW POINT INDEXES

org.neo4j.graphdb.schema.IndexType#POINT

Text index

SHOW TEXT INDEXES

org.neo4j.graphdb.schema.IndexType#TEXT

Full-text index

SHOW FULLTEXT INDEXES

org.neo4j.graphdb.schema.IndexType#FULLTEXT

Token lookup index

SHOW LOOKUP INDEXES

org.neo4j.graphdb.schema.IndexType#LOOKUP

Vector index

SHOW VECTOR INDEXES

org.neo4j.graphdb.schema.IndexType#VECTOR

You cannot have indexes of the same type over the same properties.

Range indexes

Range indexes can be used for exact lookups on all types of values, range scans, full scans, and prefix searches.

Range indexes are the most general-purpose of the property indexes, as they support all value types and a wide range of operations.

Limitations on key size

The range index has a key size limit of around 8kB.

If a transaction reaches the key size limit for one or more of its changes, that transaction fails before committing any changes. If the limit is reached during index population, the resulting index is in a failed state, and as such is not usable for any queries.

Workarounds to address limitations

Since the text index has a key size limit of around 32kB, the key size limit of the range index can be worked around by using a text index instead. However, the text index is not a general-purpose index like the range index, so this workaround cannot be applied to all cases. For more information, see Text indexes.

Point indexes

Point indexes are a type of highly-specialized, single-property index and they only index properties with Point values, unlike range indexes.

Point indexes are designed to speed up spatial queries, specifically the distance and bounding box queries. Exact lookups are the only non-spatial query that this index type supports.

For more information on the queries a point index can be used for, refer to Cypher Manual → Query Tuning → The use of indexes.

Point indexes optionally accept configuration properties for tuning the behavior of spatial search. For more information on configuring point index, refer to Cypher Manual → Indexes for search performance.

Text indexes

Text indexes are a type of single-property index. Unlike range indexes, text indexes index only properties with string values.

Text indexes are specifically designed to deal with ENDS WITH or CONTAINS queries efficiently. They are used through Cypher and they support a smaller set of string queries. Even though text indexes do support other text queries, ENDS WITH or CONTAINS queries are the only ones for which this index type provides an advantage over a range index.

The default provider is text-2.0.

For more information on the queries a text index can be used for, refer to Cypher Manual → Query Tuning → The use of indexes.

For more information on the different index types, refer to Cypher Manual → Indexes for search performance.

Limitations

Text indexes only index single property strings.

The index has a key size limit for single property strings of around 32kB. If a transaction reaches the key size limit for one or more of its changes, that transaction fails before committing any changes. If the limit is reached during index population, the resulting index is in a failed state, and as such is not usable for any queries.

Full-text indexes

Full-text indexes are optimized for indexing and searching texts.

Even though text and full-text indexes might seem to solve very similar problems, there are essential differences. Unlike text indexes, which index only single property strings, full-text indexes can index any kind of string data. Text indexes solve substring matches and exact string matches according to the semantics defined by the Cypher language. While, full-text indexes use pluggable analyzers, many of which provide language-specific processing of the text that allows for more sophisticated queries than a simple substring match. Depending on which analyzer is used, the full-text index can be used for different text search types, such as exact matches, relevance matches, phrase queries, autocompletion, and many others. Additionally, the results are ordered by relevance.

An example of a use case for full-text indexes is parsing a book for a certain term and taking advantage of the knowledge that the book is written in a certain language. The use of an analyzer for that language enables the exclusion of stop words, such as "if" and "and", and the inclusion of word forms.

Another use case example is indexing the various address fields and text data in a corpus of emails. Using the email analyzer, you can find all emails that are sent from/to or mention a specific email account.

In contrast to range and text indexes, full-text indexes are queried using built-in procedures. They are however created and dropped using Cypher. The use of full-text indexes does require familiarity with how those indexes operate.

Full-text indexes are powered by the Apache Lucene indexing and search library. A full description of how to create and use full-text indexes is provided in the Cypher Manual → Indexes to support full-text search.

Configuring a full-text index

The following options are available for configuring full-text indexes. For a complete list of Neo4j procedures, see Built-in procedures.

db.index.fulltext.default_analyzer

The name of the default analyzer when creating a new Full-text index. Once created, the index’s analyzer is not affected by this setting.

db.index.fulltext.eventually_consistent

The default consistency model when creating a new full-text index. Once created, the index’s consistency model is not affected by this setting.

Indexes are normally fully consistent, and the committing of a transaction does not return until both the store and indexes are updated. Eventually consistent full-text indexes, on the other hand, are not updated as part of a commit but instead have their updates queued up and applied in a background thread. This means that there can be a short delay between committing a change and that change becoming visible via any eventually consistent full-text indexes. This delay is just an artifact of the queueing and is usually relatively small since eventually consistent indexes are updated "as soon as possible".

By default, this is turned off, and full-text indexes are fully consistent.

db.index.fulltext.eventually_consistent_index_update_queue_max_length

Eventually, consistent full-text indexes have their updates queued up and applied in a background thread, and this setting determines the maximum size of that update queue. If the maximum queue size is reached, then committing transactions block and wait until there is more room in the queue before adding more updates to it.

This setting applies to all eventually consistent full-text indexes, and they all use the same queue. The maximum queue length must be at least 1 index update and no more than 50 million due to heap space usage considerations.

The default maximum queue length is 10.000 index updates.

Selecting an analyzer

By default, the full-text index uses the standard-no-stop-words analyzer, specified in db.index.fulltext.default_analyzer configuration setting. This analyzer is the same as Lucene’s StandardAnalyzer, except no stop-words are filtered out.

To specify another analyzer, use the OPTIONS clause of the full-text index creation command. The list of all possible analyzers is available via the db.index.fulltext.listAvailableAnalyzers() Cypher procedure.

By default, the analyzer analyzes both the indexed values and query string. In some cases, however, using different analyzers for the indexed values and query string is more appropriate. You can do that by specifying an analyzer for the query string when using the full-text search procedures.

For detailed information on how to create and use full-text indexes, see the Cypher Manual → Indexes to support full-text search.

Per-property analyzer

A full-text index can be created over multiple properties. If different analyzers for different properties are required, the standard approach in Lucene is to create a custom Composite analyzer. The Lucene project provides PerFieldAnalyzerWrapper that can associate analyzers with specific fields. For more information, see the Lucene official documentation.

Token lookup indexes

Token lookup indexes are used to look up nodes with a specific label or relationships of a specific type. They are always created over all labels or relationship types. Therefore, databases can have a maximum of two token lookup indexes - one for nodes and one for relationships.

Use and significance

Token lookup indexes are the most important indexes as they significantly speed up the population of other indexes. They are also essential for the Cypher queries execution and Core API operations. Therefore, dropping them should be carefully considered.

The node label lookup index is important for queries that match a node by one or more labels. It can also be used for matching labels and properties of a node when there are no suitable indexes available. Likewise, the relationship type lookup index is important for queries that match relationships by their types.

Most queries are executed by matching nodes and expanding their relationships. Hence, the node label lookup index is slightly more significant than the relationship type lookup index.

Both node and relationship type lookup indexes are present by default in all databases created in 4.3 and onwards.

Databases created before 4.3

Databases created before 4.3 do not get relationship lookup index automatically, in order to preserve the backward compatibility and performance characteristics of such databases.

If needed, such databases can get a relationship type lookup index by creating it explicitly through Cypher.

Creating a relationship type lookup index on a large database can take a significant amount of time, as all relationships need to be scanned when populating such an index.

Glossary

allocator

A component in the cluster that allocates databases to servers according to the topology constraints specified and an allocation strategy.

asynchronous replication

Asynchronous replication is used by secondary copies to poll for new transactions, which means they cannot be guaranteed to have received the most recent transactions. This enables efficient scale-out of read-performance.

Aura instance

A fully-managed DBMS represented by a single instance ID, that is running in the Neo4j Aura cloud.

auto-commit transaction

An automatically committed transaction that contains a single query.

Bolt protocol

Bolt is a protocol used for interaction between Neo4j instances and drivers.

bookmark

A marker the client can request from the cluster to ensure that it is able to read its own writes so that the application’s state is consistent and only databases that have a copy of the bookmark are permitted to respond.

category (Bloom)

A category is based on a node label and is defined in a Perspective as a way of visually distinguishing nodes with the same label(s).

causal consistency

All servers in a cluster agree on the order in which transactions take place. The position of a server on the causal chain can be guaranteed using a bookmark.

cluster

A Neo4j DBMS that spans multiple servers working together to increase fault tolerance and/or read scalability. Databases on a cluster may be configured to replicate across servers in the cluster thus achieving read scalability or high availability.

client application

Software that interacts with a Neo4j server.

commit

A commit is the successful completion of a transaction, which ensures durability of any changes made. For more details, visit Operations Manual → Transaction management.

composite database

Composite databases are the means to access partitioned graph data with a single Cypher query.

constraint

Constraints are sets of data modeling rules that ensure the data is consistent and reliable.

Cypher®

Neo4j’s graph query language.

data model

A data model defines how information is organized in a database. A good data model will make querying and understanding your data easier. In Neo4j, the data models have a graph structure.

database

A database is a container used by the DBMS to manage and store graph data. The physical structure of data is controlled by the database.

database vs graph

Databases are the physical containers of graph data. Graphs are the logical structure of data in Neo4j.

Database Management System

Database Management System, or DBMS, capable of managing multiple databases. A DBMS may run on a single server, or span several servers configured as a cluster.

database schema

The prescribed property existence and datatypes for nodes and relationships.

deallocate

An act of removing a database from a server or a server from a cluster without loss of data or reduced fault tolerance.

degree (of a node)

The number of relationships of a specific node; loops are counted twice.

disaster recovery

A manual intervention to restore availability of a cluster, or databases within a cluster.

driver

A software library that provides access to Neo4j from a particular programming language.

election

In the event that the Raft leader becomes unresponsive, followers automatically trigger an election and vote for a new leader.

entity

A node or a relationship.

expression (Cypher)

A component of a Cypher query which produces values. It may be used in projections, as a predicate, or when setting properties on graph elements.

fabric

Fabric is the architectural design of a unified system that provides a single access point to local or distributed graph data.

fault tolerance

A guarantee that a cluster can maintain a database’s persistence and availability in the event of one or more servers failing.

follower

A primary copy of a database acting as a follower, receives and acknowledges synchronous writes from the leader.

Generative AI (GenAI)

A type of artificial intelligence (AI) system that generates text, images, or other media in response to prompts.

graph

A logical representation of a set of nodes where some pairs are connected by relationships.

index

Data structure that improves read performance of a database.

knowledge graph

A specific type of graph that has an organizing principle so that a user (or a computer system) can reason about the underlying data. The organizing principle provides an additional layer of structure that adds context to support knowledge discovery.

label

Marks a node as a member of a named and indexed subset. A node may be assigned zero or more labels.

leader

A single primary copy of a database is designated as the leader. It receives all write transactions from clients and replicates writes synchronously to followers and asynchronously to secondary copies of the database.

main database

In terms of Neo4j Enterprise Studio, the database(s) containing the user’s data. Can exist in the same Neo4j deployment as the tool asset database.

motif

A description of a specific pattern within a graph.

node

A node represents an entity or discrete object in your graph data model. Nodes can be connected by relationships, hold data in properties, and are classified by labels.

operator

A symbol representing a mathematical or logical operation.

parameter

Named value provided when running a Cypher statement.

path

A sequence of nodes and the relationships connecting them, that does not contain duplicate relationships. Several paths can match a pattern.

pattern

A specific arrangement of nodes and relationships that can be matched in a graph. A pattern follows a motif.

perspective (Bloom)

A Perspective defines a certain business view or domain that can be found in the target Neo4j graph. A single Neo4j graph can be viewed through different Perspectives, each tailored for a different business purpose.

primary

A copy of the database that is able to process write transactions and is eligible to be elected as a leader. It participates in fault tolerant writes as it is part of the majority required to acknowledge and commit write transactions.

primary vs secondary

In a cluster, databases can operate in either primary or secondary mode. Primary databases are able to process write and read transactions, ensuring fault tolerance. Secondary databases are replicated asynchronously from primaries, and their main purpose is to provide read scaling within the cluster.

project (Aura)

An isolated environment in the unified Aura console that contains its own database instances, configurations, and resources. Preceded by tenant in the classic Aura console.

property

Properties are key-value pairs that are used for storing data on nodes and relationships.

query (Cypher)

A statement that retrieves or writes information to a database.

Raft group

A group of servers that are participating in hosting a particular database in primary mode.

Raft group member

A server that is participating in a Raft group. A server can be a member of one or more groups.

Raft log

A shared log between all Raft group members that is guaranteed to be consistently updated and viewed by those members. The log contains both database data and operational state of the Raft group.

Raft protocol

The networking mechanism that enables a database to replicate its data across multiple servers to give high availability for accessing the data and high durability to the data stored.

read scaling

Distributing query load by creating additional database copies hosted in secondary mode (read-only).

relationship

A relationship represents a connection between nodes in your graph data model. Relationships connect a source node to a target node, hold data in properties, and are classified by type.

secondary

An asynchronously replicated copy of the database that provides read scaling within the cluster.

seed

A seed is a database dump or a full backup used to create a database on a cluster. This is sometimes called seeding.

server

A physical machine, a virtual machine, or a container running an instance of Neo4j. Servers can be standalone or part of a cluster.

session

A causally linked sequence of transactions.

session consistency

An alternative name for Neo4j’s causal consistency.

standalone

A single server running Neo4j and not part of a cluster.

synchronous replication

Synchronous replication requires the leader primary to replicate a transaction and block the commit until a quorum of the follower primaries acknowledges that the transaction is successfully replicated. Once the transaction is replicated, the commit is allowed to proceed. This ensures data durability and consistency within the cluster.

system database

A database used by Neo4j to store system information.

tenant (Aura)

An isolated environment in the classic Aura console that contains its own database instances, configurations, and resources. Replaced by project in the unified Aura console.

tool asset database

In terms of Neo4j Enterprise Studio, the database where tools' assets are stored. This can be in the same Neo4j deployment as the main database(s) or in a separate deployment.

topology

A configuration that describes how the copies of a database should be spread across the servers in a cluster, see primary mode and secondary mode.

transaction

A transaction comprises a unit of work performed against a database. It is treated in a coherent and reliable way, independent of other transactions. Transactions comply with the ACID consistency model (atomic, consistent, isolated, and durable).