Incremental import

Incremental import allows you to incorporate large amounts of data in batches into the graph. You can run this operation as part of the initial data load when it cannot be completed in a single full import. Besides, you can update your graph by importing data incrementally, which is more performant than transactional insertion of such data.

The incremental import command can be used to add:

  • New nodes with labels and properties.

    Note that you must have node property uniqueness constraints in place for the property key and label combinations that form the primary key, or the uniquely identifiable nodes. Otherwise, the command will throw an error and exit. For more information, see CSV header format.

  • New relationships between existing or new nodes.

Starting from 2025.01, the incremental import command can also be used for:

  • Adding new properties to existing nodes or relationships.

  • Updating or deleting properties in nodes or relationships.

  • Updating or deleting labels in nodes.

  • Deleting existing nodes and relationships.

These are supported only by block format. See Applying changes to data via CSV files for more information.

Incremental import requires the use of --force and can be run on an existing database only. You must stop your database, if you want to perform the incremental import within one command. If you cannot afford a full downtime of your database, split the operation into several stages:

  • prepare stage (offline)

  • build stage (offline or read-only)

  • merge stage (offline)

The database must be stopped for the prepare and merge stages. During the build stage, the database can be left online but put into read-only mode. For a detailed example, see Incremental import in stages.

It is highly recommended to back up your database before running the incremental import, as if the merge stage fails, is aborted, or crashes, it may corrupt the database.

Syntax

The syntax for importing a set of CSV files incrementally is:

neo4j-admin database import incremental [-h] [--expand-commands] [--force] [--update-all-matching-relationships]
                                        [--verbose] [--auto-skip-subsequent-headers[=true|false]] [--compress
                                        [=true|false]] [--dry-run[=true|false]] [--ignore-empty-strings[=true|false]]
                                        [--ignore-extra-columns[=true|false]] [--legacy-style-quoting[=true|false]]
                                        [--normalize-types[=true|false]] [--profile[=true|false]]
                                        [--skip-bad-entries-logging[=true|false]] [--skip-bad-relationships
                                        [=true|false]] [--skip-duplicate-nodes[=true|false]] [--strict[=true|false]]
                                        [--trim-strings[=true|false]] [--additional-config=<file>]
                                        [--array-delimiter=<char>] [--bad-tolerance=<num>] [--delimiter=<char>]
                                        [--high-parallel-io=on|off|auto] [--id-type=string|integer|actual]
                                        [--input-encoding=<character-set>] [--input-type=csv|parquet]
                                        [--max-off-heap-memory=<size>] [--path-pattern-style=regex|glob|none]
                                        [--profile-results-path=<path>] [--property-shard-count=<propertyShardCount>]
                                        [--quote=<char>] [--read-buffer-size=<size>] [--report-file=<path>]
                                        [--schema=<path>] [--stage=all|prepare|build|merge] [--target-format=<format>]
                                        [--target-location=<path>] [--temp-path=<path>] [--threads=<num>]
                                        [--vector-delimiter=<char>] [--nodes=[<label>[:<label>]...=]<files>...]...
                                        [--relationships=[<type>=]<files>...]... [--multiline-fields=true|false|<path>[,
                                        <path>] [--multiline-fields-format=v1|v2]] <database>

Description

Incremental import into an existing database.

Parameters

Table 1. neo4j-admin database import incremental parameters
Parameter Description Default

<database>

Name of the database to import. If the database into which you import does not exist prior to importing, you must create it subsequently using CREATE DATABASE.

neo4j

Options

Table 2. neo4j-admin database import incremental options
Option Description Default CSV Parquet

--additional-config=<file>[1]

Configuration file with additional configuration.

--array-delimiter=<char>

Delimiter character between array elements within a value in CSV data. Also accepts TAB and e.g. U+20AC for specifying a character using Unicode. For CSV data, this value must be different from the one specified in the --delimiter option. For Parquet data, this is only needed if the array is encoded as a string.

  • ASCII character — e.g. --array-delimiter=";".

  • \ID — Unicode character with ID, e.g. --array-delimiter="\59".

  • U+XXXX — Unicode character specified with 4 HEX characters, e.g. --array-delimiter="U+20AC".

  • \t — horizontal tabulation (HT), e.g. --array-delimiter="\t".

For horizontal tabulation (HT), use \t or the Unicode character ID \9.

Unicode character ID can be used if prepended by \.

;

--auto-skip-subsequent-headers[=true|false]

Automatically skip accidental header lines in subsequent files in file groups with more than one file.

false

--bad-tolerance=<num>

Number of bad entries before the import is aborted. The import process is optimized for error-free data. Therefore, cleaning the data before importing it is highly recommended. If you encounter any bad entries during the import process, you can set the number of bad entries to a specific value that suits your needs. However, setting a high value may affect the performance of the tool.

-1 Changed in 2025.12

--compress[=true|false][2]

Introduced in 2026.04 Request backup artifact to be compressed. Compression can yield a backup artefact many times smaller, but the exact reduction depends upon many factors, including the database format and the kind of data stored. If disabled, the size of the produced artifact will be approximately equal to the size of the backed-up database. The speed of the import operation is affected by compression, but which is faster depends upon the relative performance of CPU and storage. If import speed is important, consider evaluating both options - with compression enabled and disabled.

false

--delimiter=<char>

Delimiter character between values in CSV data. Also accepts TAB and e.g. U+002A for specifying a character using Unicode. Note that the delimiter character must be a single byte character in UTF-8.

  • ASCII character — e.g. --delimiter=",".

  • \ID — Unicode character with ID, e.g. --delimiter="\44".

  • U+XXXX — Unicode character specified with 4 HEX characters, e.g. --delimiter="U+20AC".

  • \t — horizontal tabulation (HT), e.g. --delimiter="\t".

For horizontal tabulation (HT), use \t or the Unicode character ID \9.

Unicode character ID can be used if prepended by \.

,

--dry-run[=true|false][3]

Introduced in 2026.02Flag used to indicate that a dry run of the import should be performed, i.e. no data will actually be imported, only the validation of the various arguments and estimation of size of the import will be performed and reported.

false

--expand-commands

Allow command expansion in config value evaluation.

--force

Confirm incremental import by setting this flag.

-h, --help

Show this help message and exit.

--high-parallel-io=on|off|auto

Ignore environment-based heuristics and indicate if the target storage subsystem can support parallel IO with high throughput or auto detect. Typically this is on for SSDs, large raid arrays, and network-attached storage.

auto

--id-type=string|integer|actual

Each node must provide a unique ID. This is used to find the correct nodes when creating relationships.

Possible values are:

  • string — arbitrary strings for identifying nodes.

  • integer — arbitrary integer values for identifying nodes.

  • actual — (advanced) actual node IDs.

string

--ignore-empty-strings[=true|false]

Whether or not empty string fields, i.e. "" from input source are ignored, i.e. treated as null.

false

--ignore-extra-columns[=true|false]

If unspecified columns should be ignored during the import.

false

--input-encoding=<character-set>

Character set that input data is encoded in.

UTF-8

--input-type=csv|parquet

File type to import from. Can be csv or parquet. Defaults to csv.

--legacy-style-quoting[=true|false]

Whether or not a backslash-escaped quote e.g. \" is interpreted as an inner quote.

false

--max-off-heap-memory=<size>

Maximum off-heap memory that the command can use for page cache and various caching data structures to improve performance. Use this option to tune the command memory usage; the command does not use the server.memory.pagecache.size configuration setting for this purpose. Values can be plain numbers, such as 10000000, or, for example, 20G for 20 gigabytes, or 70%, which will amount to 70% of currently free memory on the machine.

90%

--multiline-fields=true|false|<path>[,<path>] [4]

In v1, whether or not fields from an input source can span multiple lines, i.e. contain newline characters. Setting --multiline-fields=true can severely degrade the performance of the importer. Therefore, use it with care, especially with large imports. In v2, this option will specify the list of files that contain multiline fields. Files can also be specified using regular expressions.

--multiline-fields-format=v1|v2

Controls the parsing of input source that can span multiple lines, i.e. contain newline characters. When set to v1, the value for --multiline-fields can only be true or false. When set to v2, the value for --multiline-fields should be the list of files that contain multiline fields.

v1

--nodes=[<label>[:<label>]…​=]<files>…​

Node CSV header and data.

  • Multiple files will be logically seen as one big file from the perspective of the importer.

  • The first line must contain the header.

  • Multiple data sources like these can be specified in one import, where each data source has its own header.

  • Files can also be specified using regular expressions.

It is possible to import files from AWS S3 buckets, Google Cloud storage buckets, and Azure buckets using the appropriate URI as the path.

--normalize-types[=true|false]

When true, non-array property values are converted to their equivalent Cypher types. For example, all integer values will be converted to 64-bit long integers.

true

--path-pattern-style=regex|glob|none[5]

Introduced in 2026.01 Pattern style to use for matching --nodes and --relationships files.

Possible values are:

  • glob — allows you to write patterns like /some/**/deep/**/nested/structure.*.

  • regex — allows you to write regular expressions, i.e. /some/nested/structure.*.

  • none — you have to enumerate all the file paths exactly, i.e. /some/content/structure1.csv,/some/content/structure2.csv.

regex

--profile[=true|false]

Introduced in 2026.02Capture a java flight recording for the entire duration of the import.

false

--profile-results-path=<path>

Introduced in 2026.02Provide a path where to store java flight recordings captured with the --profile option. Requires --profile or --profile=true to be set to have an effect.

--property-shard-count=<propertyShardCount>[2]

Introduced in 2025.12 Infinigraph (advanced) Number of shards of property data that will be created, each shard will be its own database. Typically this option doesn’t need to be set at all, since the number of shards has already been decided during initial import.

0

--quote=<char> [6]

Character to treat as a quotation mark for values in CSV data.

For example, quotes can be escaped as per RFC 4180 by doubling them. Thus "" would be interpreted as a literal ".

You cannot escape using \. For CSV data, this value must be different from the one specified in the --delimiter option.

"

--read-buffer-size=<size>

Size of each buffer for reading input data.

It has to be at least large enough to hold the biggest single value in the input data. The value can be a plain number or a byte units string, e.g. 128k, 1m.

4194304

--relationships=[<type>=]<files>…​

Relationship CSV header and data.

  • Multiple files will be logically seen as one big file from the perspective of the importer.

  • The first line must contain the header.

  • Multiple data sources like these can be specified in one import, where each data source has its own header.

  • Files can also be specified using regular expressions.

It is possible to import files from AWS S3 buckets, Google Cloud storage buckets, and Azure buckets using the appropriate URI as the path.

--report-file=<path>

File in which to store the report of the csv-import.

The location of the import log file can be controlled using the --report-file option. If you run large imports of CSV files that have low data quality, the import log file can grow very large. For example, CSV files that contain duplicate node IDs, or that attempt to create relationships between non-existent nodes, could be classed as having low data quality. In these cases, you may wish to direct the output to a location that can handle the large log file.

If you are running on a UNIX-like system and you are not interested in the output, you can get rid of it altogether by directing the report file to /dev/null.

If you need to debug the import, it might be useful to collect the stack trace. This is done by using the --verbose option.

--schema=<path>

Introduced in 2025.02 Path to the file containing the Cypher commands for creating indexes and constraints during data import. It is possible to load commands from AWS S3 buckets, Google Cloud storage buckets, and Azure buckets using the appropriate URI as the path.

--skip-bad-entries-logging[=true|false]

When set to true, the details of bad entries are not written in the log. Disabling logging can improve performance when the data contains lots of faults. Cleaning the data before importing it is highly recommended because faults dramatically affect the tool’s performance even without logging.

false

--skip-bad-relationships[=true|false]

Whether or not to skip importing relationships that refer to missing node IDs, i.e. either start or end node ID/group referring to a node that was not specified by the node input data.

Skipped relationships will be logged if they are within the limit of entities specified by --bad-tolerance and the --skip-bad-entries-logging option is disabled.

false

--skip-duplicate-nodes[=true|false]

Whether or not to skip importing nodes that have the same ID/group.

In the event of multiple nodes within the same group having the same ID, the first encountered will be imported, whereas consecutive such nodes will be skipped.

Skipped nodes will be logged if they are within the limit of entities specified by --bad-tolerance and the --skip-bad-entries-logging option is disabled.

false

--stage=all|prepare|build|merge

Stage of incremental import.

For incremental import into an existing database use all (which requires the database to be stopped).

For semi-online incremental import run prepare (on a stopped database) followed by build (on a potentially running database) and finally merge (on a stopped database).

all

--strict[=true|false]

Whether or not the lookup of nodes referred to from relationships needs to be checked strict. If disabled, most but not all relationships referring to non-existent nodes will be detected. If enabled all those relationships will be found but at the cost of lower performance.

false

--target-format=<format>[2]

Introduced in 2025.12Enterprise edition Target format can be either a plain database directory/files structure (database) or a backup artifact (backup). Uses the --temp-path location to keep any intermediate state.

database

--target-location=<path>[2]

Introduced in 2025.12Enterprise edition Location for target backup artifact data. Used together with --target-format=backup.

--temp-path=<path>

Introduced in 2025.04 Provide a path where to store temporary files that are created and deleted during import. If not specifically provided, the default temp path will be created inside the database directory of the imported database.

--threads=<num>

(advanced) Max number of worker threads used by the importer. Defaults to the number of available processors reported by the JVM. There is a certain amount of minimum threads needed so for that reason there is no lower bound for this value. For optimal performance, this value should not be greater than the number of available processors.

--trim-strings[=true|false]

Whether or not strings should be trimmed for whitespaces.

false

--vector-delimiter=<char>

Introduced in 2025.10 Delimiter character between vector coordinates within a value in CSV data. Also accepts TAB and e.g. U+20AC for specifying a character using Unicode. For CSV data, this value must be different from the one specified in the --delimiter option. For Parquet data, this is only needed if the vector is encoded as a string.

;

--update-all-matching-relationships

Introduced in 2025.01 Whether or not to update all existing relationships that match a relationship data entry. If disabled, the relationship data entry will be logged if it is within the limit of entities specified by --bad-tolerance and the --skip-bad-entries-logging option is disabled.

false

--verbose

Enable verbose output.

2. For using this option with sharded property databases, see Property Sharding → Data import.
4. The option’s value depends on --multiline-fields-format. For details, see Importing data that spans multiple lines.
6. To escape quotation marks in the CSV data, you should double the configured character.

Usage and limitations

The following limitations apply to the incremental import command:

Importing data incrementally in a clustered environment

The importer works well on standalone servers.

To safely perform an incremental import in a clustered environment, follow these steps:

  1. Run the incremental import command on a single server in the cluster. This server can then be used as the designated seeder from which other cluster members can copy the database.

  2. Reconfigure the database topology to a single primary by running the dbms.recreateDatabase() procedure.

  3. Then stop the database using STOP DATABASE.

  4. Perform the incremental import on the server that hosts the database.

  5. Then start the database with START DATABASE.

  6. Lastly, restore the desired database topology using ALTER DATABASE.

Using both a multi-value option and a positional parameter

When using both a multi-value option, such as --nodes and --relationships, and a positional parameter (for example, in --additional-config neo4j.properties --nodes 0-nodes.csv mydatabase), the --nodes option acts "greedy" and the next option, in this case, mydatabase, is pulled in via the nodes convertor.

This is a limitation of the underlying library, Picocli, and is not specific to Neo4j Admin. For more information, see Picocli → Variable Arity Options and Positional Parameters official documentation.

To resolve the problem, use one of the following solutions:

  • Put the positional parameters first. For example, mydatabase --nodes 0-nodes.csv.

  • Put the positional parameters last, after -- and the final value of the last multi-value option. For example, nodes 0-nodes.csv — mydatabase.

Examples

Perform a dry run before importing data

Before performing the actual import, you can run a dry run to validate the input files and estimate the size of the import. The command does not write any data to the database.

neo4j@system> STOP DATABASE neo4j WAIT;
...
bin/neo4j-admin database import incremental \
--dry-run=true \
--nodes=N1=../../raw-data/incremental-import/nodes.csv \
--relationships=R1=../../raw-data/incremental-import/relationships.csv \
neo4j
Example output
Neo4j version: 2026.08.1
Checking the contents of the following files:
Nodes:
  /path/to/neo4j-enterprise-2026.08.1/raw-data/incremental-import/nodes.csv
Relationships:
  /path/to/neo4j-enterprise-2026.08.1/raw-data/incremental-import/relationships.csv

Available resources:
  Total machine memory: 64.00GiB
  Free machine memory: 9.997GiB
  Max heap memory : 14.22GiB
  Max worker threads: 10
  Configured max memory: 40.00GiB
  High parallel IO: true

Schema commands:
  Indexes to be created: 3
  Indexes to be dropped: 1
  Constraints to be created: 2
  Constraints to be dropped: 3

Estimated entity counts / sizes:
  Nodes: 2702496410
    Includes updates: true
    Labels: 2702496410
    Property count: 16348903762
    Property size: 253.6GiB
  Relationships: 18094004256
    Includes updates: false
    Property count: 6076138797
    Property size: 96.14GiB

There are two ways of importing data incrementally.

Incremental import in a single command

If downtime is not a concern, you can run a single command with the option --stage=all. This option requires the database to be stopped.

neo4j@system> STOP DATABASE neo4j WAIT;
...
bin/neo4j-admin database import incremental \
--stage=all \
--nodes=N1=../../raw-data/incremental-import/nodes.csv \
--relationships=R1=../../raw-data/incremental-import/relationships.csv \
neo4j

Incremental import in stages

If you cannot afford a full downtime of your database, you can run the import in three stages.

  1. prepare stage:

    During this stage, the import tool analyzes the CSV headers and copies the relevant data over to the new increment database path. The import command is run with the option --stage=prepare and the database must be stopped.

    1. Using the system database, stop the database neo4j with the WAIT option to ensure a checkpoint happens before you run the incremental import command. The database must be stopped to run --stage=prepare.

      STOP DATABASE neo4j WAIT
    2. Run the incremental import command with the --stage=prepare option:

      bin/neo4j-admin database import incremental \
      --stage=prepare \
      --nodes=N1=../../raw-data/incremental-import/nodes.csv \
      --relationships=R1=../../raw-data/incremental-import/relationships.csv \
      neo4j
  2. build stage:

    During this stage, the import tool imports the data, deduplicates it, and validates it in the new increment database path. This is the longest stage and you can put the database in read-only mode to allow read access. The import command is run with the option --stage=build.

    1. Put the database in read-only mode:

      ALTER DATABASE neo4j SET ACCESS READ ONLY
    2. Run the incremental import command with the --stage=build option:

      bin/neo4j-admin database import incremental \
      --stage=build \
      --nodes=N1=../../raw-data/incremental-import/nodes.csv \
      --relationships=R1=../../raw-data/incremental-import/relationships.csv \
      neo4j
  3. merge stage:

    During this stage, the import tool merges the new with the existing data in the database. It also updates the affected indexes and upholds the affected property uniqueness constraints and property existence constraints. The import command is run with the option --stage=merge and the database must be stopped. It is not necessary to include the --nodes or --relationships options when using --stage=merge.

    1. Using the system database, stop the database neo4j with the WAIT option to ensure a checkpoint happens before you run the incremental import command.

      STOP DATABASE neo4j WAIT
    2. Run the incremental import command with the --stage=merge option:

      bin/neo4j-admin database import incremental \
      --stage=merge \
      neo4j

Importing multiple input files using regular expression

During incremental import, you can use regular expressions to match multiple input files.

For example, the following command imports all CSV files in the import directory that match the pattern *.csv:

bin/neo4j-admin database import incremental \
--stage=all \
--nodes=import/*.csv \
--relationships=import/*.csv
neo4j

Importing compressed files

During incremental import, you can import files compressed with zip or gzip. Each compressed file must contain a single file.

For example, the following command imports compressed files, where actors.csv.zip is a zip file and movies.csv.gz and roles.csv.gz are gzip files:

neo4j_home$ ls import
actors-header.csv  actors.csv.zip  movies-header.csv  movies.csv.gz  roles-header.csv  roles.csv.gz
STOP DATABASE neo4j WAIT;
...
bin/neo4j-admin database import incremental \
--stage=all \
--nodes=import/movies-header.csv,import/movies.csv.gz \
--nodes=import/actors-header.csv,import/actors.csv.zip \
--relationships=import/roles-header.csv,import/roles.csv.gz
neo4j

Applying changes to data via CSV files

You can use CSV files to update existing nodes, relationships, labels, or properties during incremental import. This feature is supported only by block format.

Set an explicit action for each row

You can set an explicit action for each row in the CSV file by using the :ACTION keyword in the header file. If no action is specified, the import tool works as in full import mode, creating new data.

The following actions are supported:

  • empty = CREATE (default)

  • C, CREATE - Creates new nodes and relationships, with or without properties, as well as labels.

  • U, UPDATE - Updates existing nodes, relationships, labels, and properties.

  • D, DELETE - Deletes existing nodes or relationships. Deleting a node also deletes its relationships (DETACH DELETE).

Using actions in CSV files to update nodes
:ACTION,uid:ID(label:Person),name,:LABEL
CREATE,person1,"Keanu Reeves",Actor
UPDATE,person2,"Laurence Fishburne",Actor
DELETE,person4,,

Nodes are identified by their unique property value for the key/label combination that the header specifies.

Using actions in CSV files to update relationships
:ACTION,:START_ID,:END_ID,:TYPE,role
CREATE,person1,movie1,ACTED_IN,"Neo"
UPDATE,person2,movie1,ACTED_IN,"Morpheus"
DELETE,person3,movie1,ACTED_IN

Relationships are identified non-uniquely by their start and end node IDs, and their type.

To further narrow down selection you can tag a property column as an identifier to help out in selecting relationships uniquely (or at least more uniquely).

Using actions in CSV files to update relationships with identifier properties
:ACTION,:START_ID,:TYPE,:END_ID,p1{identifier:true},name,p4
U,person1,KNOWS,person2,abc,"Keanu Reeves","Hello Morpheus"
U,person2,KNOWS,person1,def,"Laurence Fishburne","Hello Neo"

The data in the p1 column for these relationships helps select relationships "more uniquely" if a multiple of 1,KNOWS,2 exists. There can also be multiple identifier properties defined in the header. Identifier properties match the selected relationships and will not be set on the relationships that already have them.

Update existing labels

You can add or remove one or more labels from an existing node by prepending the clause LABEL in the header with a + (default) or -:

  • :+LABEL - Add one or more labels to an existing node.

  • :-LABEL - Remove one or more labels (if they exist) from an existing node.

For example, a file could have the following format:

uid:ID(label:Person),:+LABEL,:-LABEL,name,age
person1,Actor,Producer,"Keanu Reeves",55
person2,Actor;Director,,"Laurence Fishburne",60

In this case, all labels in the second column are added and all the labels in the third column are removed (if they exist).

Remove existing properties

You can remove properties from existing nodes or relationships by a :-PROPERTY column in the header. In the contents of this field you can add zero or more property names to remove from the entity. For example:

Remove nodes' properties
:ACTION,uid:ID(label:Person),:-PROPERTY
U,person1,age;hometown

Properties age and hometown are removed from the node with the uid:ID person1.

Remove relationships' properties
:ACTION,:START_ID,:END_ID,:TYPE,:-PROPERTY
U,person1,movie1,ACTED_IN,role;description

Properties role and description are removed from the relationship with the :START_ID person1, :END_ID movie1, and :TYPE ACTED_IN.

Using actions in CSV files to update labels and properties
:ACTION,uid:ID(label:Person),:LABEL,:-LABEL,:-PROPERTY,name,height:int
U,person1,Actor,Producer,age;hometown,Henry",185

One CSV entry can specify all types of updates to one entity at the same time. In this example, the node person1 is updated with:

  • added Actor label

  • removed Producer label

  • removed age and hometown properties

  • set name="Henry" property

  • set height=185 property

Glossary

allocator

A component in the cluster that allocates databases to servers according to the topology constraints specified and an allocation strategy.

asynchronous replication

Asynchronous replication is used by secondary copies to poll for new transactions, which means they cannot be guaranteed to have received the most recent transactions. This enables efficient scale-out of read-performance.

Aura instance

A fully-managed DBMS represented by a single instance ID, that is running in the Neo4j Aura cloud.

auto-commit transaction

An automatically committed transaction that contains a single query.

Bolt protocol

Bolt is a protocol used for interaction between Neo4j instances and drivers.

bookmark

A marker the client can request from the cluster to ensure that it is able to read its own writes so that the application’s state is consistent and only databases that have a copy of the bookmark are permitted to respond.

category (Bloom)

A category is based on a node label and is defined in a Perspective as a way of visually distinguishing nodes with the same label(s).

causal consistency

All servers in a cluster agree on the order in which transactions take place. The position of a server on the causal chain can be guaranteed using a bookmark.

cluster

A Neo4j DBMS that spans multiple servers working together to increase fault tolerance and/or read scalability. Databases on a cluster may be configured to replicate across servers in the cluster thus achieving read scalability or high availability.

client application

Software that interacts with a Neo4j server.

commit

A commit is the successful completion of a transaction, which ensures durability of any changes made. For more details, visit Operations Manual → Transaction management.

composite database

Composite databases are the means to access partitioned graph data with a single Cypher query.

constraint

Constraints are sets of data modeling rules that ensure the data is consistent and reliable.

Cypher®

Neo4j’s graph query language.

data model

A data model defines how information is organized in a database. A good data model will make querying and understanding your data easier. In Neo4j, the data models have a graph structure.

database

A database is a container used by the DBMS to manage and store graph data. The physical structure of data is controlled by the database.

database vs graph

Databases are the physical containers of graph data. Graphs are the logical structure of data in Neo4j.

Database Management System

Database Management System, or DBMS, capable of managing multiple databases. A DBMS may run on a single server, or span several servers configured as a cluster.

database schema

The prescribed property existence and datatypes for nodes and relationships.

deallocate

An act of removing a database from a server or a server from a cluster without loss of data or reduced fault tolerance.

degree (of a node)

The number of relationships of a specific node; loops are counted twice.

disaster recovery

A manual intervention to restore availability of a cluster, or databases within a cluster.

driver

A software library that provides access to Neo4j from a particular programming language.

election

In the event that the Raft leader becomes unresponsive, followers automatically trigger an election and vote for a new leader.

entity

A node or a relationship.

expression (Cypher)

A component of a Cypher query which produces values. It may be used in projections, as a predicate, or when setting properties on graph elements.

fabric

Fabric is the architectural design of a unified system that provides a single access point to local or distributed graph data.

fault tolerance

A guarantee that a cluster can maintain a database’s persistence and availability in the event of one or more servers failing.

follower

A primary copy of a database acting as a follower, receives and acknowledges synchronous writes from the leader.

Generative AI (GenAI)

A type of artificial intelligence (AI) system that generates text, images, or other media in response to prompts.

graph

A logical representation of a set of nodes where some pairs are connected by relationships.

index

Data structure that improves read performance of a database.

knowledge graph

A specific type of graph that has an organizing principle so that a user (or a computer system) can reason about the underlying data. The organizing principle provides an additional layer of structure that adds context to support knowledge discovery.

label

Marks a node as a member of a named and indexed subset. A node may be assigned zero or more labels.

leader

A single primary copy of a database is designated as the leader. It receives all write transactions from clients and replicates writes synchronously to followers and asynchronously to secondary copies of the database.

main database

In terms of Neo4j Enterprise Studio, the database(s) containing the user’s data. Can exist in the same Neo4j deployment as the tool asset database.

motif

A description of a specific pattern within a graph.

node

A node represents an entity or discrete object in your graph data model. Nodes can be connected by relationships, hold data in properties, and are classified by labels.

operator

A symbol representing a mathematical or logical operation.

parameter

Named value provided when running a Cypher statement.

path

A sequence of nodes and the relationships connecting them, that does not contain duplicate relationships. Several paths can match a pattern.

pattern

A specific arrangement of nodes and relationships that can be matched in a graph. A pattern follows a motif.

perspective (Bloom)

A Perspective defines a certain business view or domain that can be found in the target Neo4j graph. A single Neo4j graph can be viewed through different Perspectives, each tailored for a different business purpose.

primary

A copy of the database that is able to process write transactions and is eligible to be elected as a leader. It participates in fault tolerant writes as it is part of the majority required to acknowledge and commit write transactions.

primary vs secondary

In a cluster, databases can operate in either primary or secondary mode. Primary databases are able to process write and read transactions, ensuring fault tolerance. Secondary databases are replicated asynchronously from primaries, and their main purpose is to provide read scaling within the cluster.

project (Aura)

An isolated environment in the unified Aura console that contains its own database instances, configurations, and resources. Preceded by tenant in the classic Aura console.

property

Properties are key-value pairs that are used for storing data on nodes and relationships.

query (Cypher)

A statement that retrieves or writes information to a database.

Raft group

A group of servers that are participating in hosting a particular database in primary mode.

Raft group member

A server that is participating in a Raft group. A server can be a member of one or more groups.

Raft log

A shared log between all Raft group members that is guaranteed to be consistently updated and viewed by those members. The log contains both database data and operational state of the Raft group.

Raft protocol

The networking mechanism that enables a database to replicate its data across multiple servers to give high availability for accessing the data and high durability to the data stored.

read scaling

Distributing query load by creating additional database copies hosted in secondary mode (read-only).

relationship

A relationship represents a connection between nodes in your graph data model. Relationships connect a source node to a target node, hold data in properties, and are classified by type.

secondary

An asynchronously replicated copy of the database that provides read scaling within the cluster.

seed

A seed is a database dump or a full backup used to create a database on a cluster. This is sometimes called seeding.

server

A physical machine, a virtual machine, or a container running an instance of Neo4j. Servers can be standalone or part of a cluster.

session

A causally linked sequence of transactions.

session consistency

An alternative name for Neo4j’s causal consistency.

standalone

A single server running Neo4j and not part of a cluster.

synchronous replication

Synchronous replication requires the leader primary to replicate a transaction and block the commit until a quorum of the follower primaries acknowledges that the transaction is successfully replicated. Once the transaction is replicated, the commit is allowed to proceed. This ensures data durability and consistency within the cluster.

system database

A database used by Neo4j to store system information.

tenant (Aura)

An isolated environment in the classic Aura console that contains its own database instances, configurations, and resources. Replaced by project in the unified Aura console.

tool asset database

In terms of Neo4j Enterprise Studio, the database where tools' assets are stored. This can be in the same Neo4j deployment as the main database(s) or in a separate deployment.

topology

A configuration that describes how the copies of a database should be spread across the servers in a cluster, see primary mode and secondary mode.

transaction

A transaction comprises a unit of work performed against a database. It is treated in a coherent and reliable way, independent of other transactions. Transactions comply with the ACID consistency model (atomic, consistent, isolated, and durable).