Back up an online database

Remember to plan your backup carefully and back up each of your databases, including the system database.

Note that it is not allowed to take a backup of a database alias, only physical databases can be backed up.

Command

A Neo4j database can be backed up in online mode using the neo4j-admin database backup command. The command must be invoked as the neo4j user to ensure the appropriate file permissions.

It is best practice, but not mandatory, to perform the backup from a server on the same network as the database, but that is not part of the cluster. You should install Neo4j on that machine to make the neo4j-admin command available. This machine is known as a backup client.

Backup artifact

The neo4j-admin database backup command produces one backup artifact file per database each time it is run. A backup artifact file is an immutable file containing the backup data of a given database along with some metadata like the database name and ID, the backup time, the lowest/highest transaction ID, etc.

Backup artifacts can be of two types:

  1. a full backup containing the whole database store or

  2. a differential backup containing a log of transactions to apply to a database store contained in a full backup artifact.

Backup chain

The first time the backup command is run, a full backup artifact is produced for a given database. On the other hand, differential backup artifacts are produced by the subsequent runs.

A backup chain consists of a full backup optionally followed by a sequence of n contiguous differential backups.

Starting with Neo4j 2026.02, the first differential backup in a chain may overlap with the previous full backup, meaning its lowest transaction ID may be less than or equal to the full backup’s highest transaction ID. Subsequent differential backups must still be contiguous with their parent.

This allows full and differential backups to be scheduled on independent cadences. See Scheduling for RPO and RTO.

backup chain
Figure 1. Backup chain

Usage

The neo4j-admin database backup command can be used for performing an online full or differential backup from a running Neo4j Enterprise server. The produced differential backup artifact contains transaction logs that can be replayed and applied to stores contained in full backup artifacts when restoring a backup chain.

Neo4j’s backup service must have been configured on the server beforehand. The command can be run both locally and remotely. However, it uses a significant amount of resources, such as memory and CPU. Therefore, it is recommended to perform the backup on a separate dedicated machine. The neo4j-admin database backup command also supports SSL/TLS. For more information, see Online backup configurations.

neo4j-admin database backup is not supported in Neo4j Aura.

Syntax

neo4j-admin database backup [-h] [--expand-commands] [--prefer-diff-as-parent] [--verbose]
                            [--compress[=true|false]] [--keep-failed[=true|false]]
                            [--parallel-download[=true|false]] [--parallel-recovery[=true|false]]
                            [--remote-address-resolution[=true|false]] [--skip-empty-diffs
                            [=true|false]] [--skip-recovery[=true|false]]
                            [--additional-config=<file>] [--include-metadata=none|all|users
                            [=user1,user2]|roles] [--inspect-path=<path>] [--pagecache=<size>]
                            [--split-archive-part-size=<splitsize>] [--temp-path=<path>] [--to-path=<path>] [--type=<type>]
                            [--from=<host:port>[,<host:port>...]]... [<database>...]

Description

Perform an online backup from a running Neo4j enterprise server. Neo4j’s backup service must have been configured on the server beforehand.

Parameters

Table 1. neo4j-admin database backup parameters
Parameter Description Default

[<database>…​]

Name(s) of the remote database(s) to backup. Supports globbing inside of double quotes, for example, "data*". (<database> is required unless --inspect-path is used.)

neo4j

If <database> is "*", neo4j-admin will attempt to back up all databases of the DBMS.

Options

Table 2. neo4j-admin database backup options
Option Description Default

--additional-config=<file>[1]

Configuration file with additional configuration.

--compress[=true|false]

Request backup artifact to be compressed. Compression can yield a backup artefact many times smaller, but the exact reduction depends upon many factors, including the database format and the kind of data stored. If disabled, the size of the produced artifact will be approximately equal to the size of the backed-up database. The speed of the backup operation is affected by compression, but which is faster depends upon the relative performance of CPU and storage. If backup speed is important, consider evaluating both options - with compression enabled and disabled.

true

--expand-commands

Allow command expansion in config value evaluation.

--from=<host:port>[,<host:port>…​]

Comma-separated list of host and port of Neo4j instances, each of which are tried in order.

-h, --help

Show this help message and exit.

--include-metadata=none|all|users[=user1,user2]|roles

Changed in 2025.10 Include metadata in the file. This cannot be used for backing up the system database. Possible values are:

  • roles - include commands to create the roles, auth rules, and privileges (for both database and graph) that affect the use of the database.

  • users - include commands to create the users that can use the database and their role assignments. If a list of users is specified (e.g. users=alice,bob,charlie), only those users are included in the backup.

  • all - include both roles and users.

  • none - does not include any metadata.

    Privileges specific to the DBMS and not to the backed-up database are not included in the backup. For instance, GRANT ROLE MANAGEMENT ON DBMS TO $role will not be backed up.

Accordingly, roles and users that do not have database-related privileges are not included in the backup (e.g. those with only DBMS or no privileges).

It is recommended to use SHOW USERS, SHOW ROLES, and SHOW ROLE $role PRIVILEGES AS COMMANDS to get the complete list of users, roles and privileges in these situations.

all

--inspect-path=<path>

List and show the metadata of the backup artifact(s). Accepts a folder or a file.

--keep-failed[=true|false]

Request failed backup to be preserved for further post-failure analysis. If enabled, a directory with the failed backup database is preserved.

false

--pagecache=<size>

The size of the page cache to use for the backup process.

--parallel-download[=true|false]

Introduced in 2025.11 Download backup data from multiple Neo4j instances in parallel.

false

--parallel-recovery[=true|false]

Allow multiple threads to apply pulled transactions to a backup in parallel. For some databases and workloads, this may reduce backup times significantly. Note: this is an EXPERIMENTAL option. Consult Neo4j support before use.

false

--prefer-diff-as-parent

Introduced in 2025.04 When performing a differential backup, prefer the latest non-empty differential backup as the parent instead of the latest backup.

false

--remote-address-resolution[=true|false]

Introduced in 2025.09 Allow the DBMS to automatically determine which servers are eligible to serve as backup sources, instead of requiring manual selection.

false

--skip-empty-diffs[=true|false]

Introduced in 2026.08 When performing a differential backup and there are no new transactions to back-up, do not produce any backup file. This option can only be used if metadata is explicitly omitted (--include-metadata=none).

false

--skip-recovery[=true|false]

Introduced in 2025.11 Skip recovery as part of the full backup. Skipping recovery might result in faster backups, but recovery will have to be done during restore time.

false

--split-archive-part-size=<splitsize>

Introduced in 2026.09 Splits the resulting backup artifact into multiple files of the specified size. The size can be specified in bytes or with a unit suffix (e.g. 5G, 100g, 1TiB). The minimum split size is 1GiB. If not specified the default value of 0 means the backup is not split and is written as a single file.

0

--temp-path=<path>

Provide a path to a temporary empty directory for storing backup files until the command is completed. The files will be deleted once the command is finished.

--to-path=<path>

Directory to place backup in (required unless --inspect-path is used). It is possible to back up databases into AWS S3 buckets, Google Cloud storage buckets, and Azure using the appropriate URI as the path.

--type=<type>

Type of backup to perform. Possible values are: FULL, DIFF, AUTO. If none is specified, the type is automatically determined based on the existing backups. If you want to force a full backup, use FULL.

AUTO

--verbose

Enable verbose output.

The --to-path=<path> option can also back up databases into AWS S3 buckets, Google Cloud storage buckets, and Azure buckets. For more information, see Back up a database to a cloud storage.

Even when --to-path points to a cloud storage bucket, the backup process requires a temporary local directory to store the database store files and transaction logs before creating and streaming the backup artifact to the cloud destination. The temporary directory must therefore have enough free space to hold the entire backup.

The temporary path can be specified by the --temp-path option. If --temp-path is unspecified, Neo4j creates the temporary directory in its current working directory. If that location does not have sufficient free disk space, the backup can fail.

To avoid this issue, it is strongly recommended to specify --temp-path and point it to a local directory with sufficient free disk space, especially when backing up to cloud storage.

Exit codes

Depending on whether the backup was successful or not, neo4j-admin database backup exits with different codes. The error codes include details of what error was encountered.

Table 3. Neo4j Admin backup exit codes when backing up one database
Code Description

0

Success.

1

Backup failed, or succeeded but encountered problems such as some servers being uncontactable. See logs for more details.

Table 4. Neo4j Admin backup exit codes when backing multiple databases
Code Description

0

All databases are backed up successfully.

1

One or several backups failed, or succeeded with problems.

Online backup configurations

Checkpointing

When a full backup is requested, it always triggers a checkpoint. The backup cannot proceed until the checkpoint finishes.

While the server is checkpointing, the backup job receives no data, which may lead to the backup timeout. To extend the backup timeout, modify the dbms.cluster.network.client_inactivity_timeout setting, which restricts the network inactivity. It controls the timeout duration of the catchup protocol, which is the underlying protocol of multiple catchup processes, including backups.

You can also tune up the Checkpoint settings or check that your disks are performant enough to handle the load. For more information, see Checkpoint IOPS limit.

To read more about checkpointing, see Database internals → Checkpointing and log pruning.

Server configuration

The table below lists the basic server parameters relevant to backups. Note that by default, the backup service is enabled but only listens on localhost (127.0.0.1). This needs to be changed if backups are to be taken from another machine.

Table 5. Server parameters for backups
Parameter name Default value Description

server.backup.enabled

true

Enable support for running online backups.

server.backup.listen_address

127.0.0.1:6362

Listening server for online backups.

Memory configuration

You can configure the memory allocated to the backup client in several ways:

  • Configure heap size:

    HEAP_SIZE configures the maximum heap size allocated for the backup process. Define the variable HEAP_SIZE before starting the operation. If not specified, the Java Virtual Machine chooses a value based on the server resources.

  • Configure page cache:

    Set the page cache size with the --pagecache option of the neo4j-admin database backup command.

  • Override the memory settings:

    Use the --additional-config option of the neo4j-admin database backup command to override the memory configurations in the neo4j.conf file.

You should give the Neo4J page cache as much memory as possible, as long as it satisfies the following constraint:

Neo4J page cache + OS page cache < available RAM, where 2 to 4GB should be dedicated to the operating system’s page cache.

For example, if your current database has a Total mapped size of 128GB as per the debug.log, and you have enough free space (meaning you have left aside 2 to 4 GB for the OS), then you can set --pagecache to 128GB.

Computational resource configurations

Transaction log files

The transaction log files, which keep track of recent changes, are rotated and pruned based on a provided configuration. For example, setting db.tx_log.rotation.retention_policy=3 files keeps 3 transaction log files in the backup. Because recovered servers do not need all of the transaction log files that have already been applied, it is possible to further reduce storage size by reducing the size of the files to the bare minimum. This can be done by setting db.tx_log.rotation.size=1M and db.tx_log.rotation.retention_policy=3 files. You can use the --additional-config parameter to override the configurations in the neo4j.conf file.

Removing transaction logs manually can result in a broken backup.

Security configurations

Securing your backup network communication with an SSL policy and a firewall protects your data from unwanted intrusion and leakage. When using the neo4j-admin database backup command, you can configure the backup server to require SSL/TLS, and the backup client to use a compatible policy. For more information on how to configure SSL in Neo4j, see SSL framework.

Configuration for the backup server should be added to the neo4j.conf file and configuration for backup client to the neo4j-admin.conf file. The easiest way to ensure compatibility is to use the same SSL policy configuration for both the server and the client. For production environments, it is recommended to use Certificate Authorities (CAs) to sign certificates rather than self-signed certificates.

The default backup port is 6362, configured with key server.backup.listen_address. The SSL configuration policy has the key of dbms.ssl.policy.backup.

As an example, add the following content to your neo4j.conf and neo4j-admin.conf files:

Server configuration in neo4j.conf
dbms.ssl.policy.backup.enabled=true
dbms.ssl.policy.backup.client_auth=REQUIRE
dbms.ssl.policy.backup.tls_versions=TLSv1.2,TLSv1.3
dbms.ssl.policy.backup.ciphers=TLS_ECDHE_ECDSA_WITH_AES_256_GCM_SHA384,TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384
Client configuration in neo4j-admin.conf
dbms.ssl.policy.backup.enabled=true
dbms.ssl.policy.backup.client_auth=REQUIRE
dbms.ssl.policy.backup.tls_versions=TLSv1.2,TLSv1.3
dbms.ssl.policy.backup.ciphers=TLS_ECDHE_ECDSA_WITH_AES_256_GCM_SHA384,TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384

Neo4j also supports TLSv1.3. To use both TLSv1.2 and TLSv1.3 versions, you must specify which ciphers to be enforced for each version. Otherwise, Neo4j could use every possible cipher in the JVM for those versions, leading to a less secure configuration.

For a detailed list of recommendations regarding security in Neo4j, see Security checklist.

It is very important to ensure that there is no external access to the port specified by the setting server.backup.listen_address. Failing to protect this port may leave a security hole open by which an unauthorized user can make a copy of the database onto a different machine. In production environments, external access to the backup port should be blocked by a firewall.

The cluster transaction port 6000 inherits its bind address from server.default_listen_address. If this is set to 0.0.0.0, the cluster port is network-reachable and provides unauthenticated backup/replication capability. Protect this port with SSL (dbms.ssl.policy.cluster) and/or a firewall.

Cluster configurations

In a cluster topology, it is possible to take a backup from any server hosting the database to backup, and each server has two configurable ports capable of serving a backup. These ports are configured by server.backup.listen_address and server.cluster.listen_address respectively. Functionally, they are equivalent for backups, but separating them can allow some operational flexibility, while using just a single port can simplify the configuration.

It is generally recommended to select secondary database copies to act as backup resources since they are more numerous than primary copies in typical cluster deployments. Furthermore, the possibility of performance issues on a secondary database allocation, caused by a large backup, does not affect the performance or redundancy of the primaries. If a secondary is not available, then a primary can be selected based on factors, such as its physical proximity, bandwidth, performance, and liveness.

Use the SHOW DATABASES command to learn which database is hosted on which server.

To avoid taking a backup from a cluster member that is lagging behind, you can look at the transaction IDs by exposing Neo4j metrics or via Neo4j Browser. To view the latest processed transaction IDs (and other metrics) in Neo4j Browser, type :sysinfo at the prompt.

Targeting multiple servers

It is recommended to provide a list of multiple target servers when taking a backup from a cluster, since that may allow a backup to succeed even if some server is down, or not all databases are hosted on the same servers. If the command finds one or more servers that do not respond, it continues trying to backup from other servers and continues backing up other requested databases, but the exit code of the command is non-zero, to alert the user to the fact there is a problem. If a name pattern is used for the database together with multiple target servers, all servers contribute to the list of matching databases.

Using --remote-address-resolution

Starting from 2025.09, the --remote-address-resolution option is available. When enabled, the DBMS automatically selects the most appropriate servers to act as backup sources for a given database.

By default, the online backup command requires the user to determine which servers host the target database and direct the command to one of them. With remote address resolution enabled, the DBMS performs this mapping automatically, removing the need for manual server selection.

The server selection occurs in the following order:

  • The DBMS first selects all servers hosting the database in secondary mode.

  • If it is not possible to back up from one of the secondaries, the DBMS attempts to take a backup from the primary followers before finally trying the database primary writer.

Using --parallel-download

Starting from 2025.11, the --parallel-download option is available. When enabled, the backup process pulls data from multiple servers which have been either defined in the --from option or determined from --remote-address-resolution.

Note that this will speed up the backup process if sufficient throughput is provisioned for both the backup process and the source servers from which data is being pulled.

Scheduling for RPO and RTO

Starting with Neo4j 2026.02, the first differential backup in a backup chain may overlap with its parent full backup. This allows you to schedule full and differential backups independently, tuning each cadence to a different recovery objective:

Recovery Point Objective (RPO)

The maximum amount of data loss you can tolerate, measured in time. Drive RPO with the differential backup schedule — a differential backup every 15 minutes yields an RPO of roughly 15 minutes.

Recovery Time Objective (RTO)

The maximum time a restore is allowed to take. Drive RTO with the full backup schedule — a more recent full backup shortens a differential chain that must be replayed at restore time, enabling faster recovery.

Before 2026.02, a full backup taken between two scheduled differential backups would break the differential chain, because the next differential backup had to start exactly where the new full backup ended. In practice this forced the two schedules to be coupled, and a full backup would interrupt the differential cadence, degrading RPO.

With overlap allowed, two independent processes can run safely:

  • A differential schedule set to the desired RPO (for example, every 15 minutes).

  • A full schedule set to the desired RTO (for example, daily or weekly).

The first differential backup after each full backup naturally overlaps with it; the chain remains restorable.

A differential backup can only succeed if there is already a backup chain to build upon — that is, at least one full backup must be present in the target location.

You have two options when setting up the schedules:

  • Seed the target location with an initial full backup (or ensure the full schedule has produced one) before the differential schedule starts running.

  • Run the differential schedule with --type=AUTO (the default) instead of --type=DIFF. With AUTO, the backup client falls back to a full backup when no chain is present. Therefore, the first run produces a full backup and subsequent runs produce differential backups.

Consider also running neo4j-admin backup aggregate to collapse a chain into a single recovered full artifact when restore time matters most.

Examples

The following are examples of how to perform a backup of a single database and multiple databases. The target directory /mnt/backups/neo4j must exist before calling the command and the database(s) must be online.

Back up a single database

You do not need to use the --type option to specify the type of backup. By default, the type is automatically determined based on the existing backups.

bin/neo4j-admin database backup --to-path=/path/to/backups/neo4j neo4j

Perform a forced full backup of a single database

If you want to force a full backup after several differential backups, you can use the --type=full option.

bin/neo4j-admin database backup --type=full --to-path=/path/to/backups/neo4j neo4j

Back up multiple databases

To back up multiple databases that match a name pattern, you can use name globbing. For example, to backup all databases that start with n in your three-server cluster, run:

bin/neo4j-admin database backup --from=192.168.1.34:6362,192.168.1.35:6362,192.168.1.36:6362 --to-path=/mnt/backups/neo4j --pagecache=4G "n*"

Back up a list of databases

To back up several databases by name, you can provide a list of database names.

neo4j-admin database backup --from=192.168.1.34:6362,192.168.1.35:6362,192.168.1.36:6362 --to-path=/mnt/backups/neo4j --pagecache=4G "test*" "neo4j"

Back up a database to a cloud storage

In Neo4j 2025.03, new cloud integration settings are introduced to provide better support for deployment and management in cloud ecosystems. For details, refer to Configuration settings → Cloud storage integration settings.

The following examples show how to back up a database to a cloud storage bucket using the --to-path option.

Backing up directly to cloud storage still requires local disk space on the machine running the backup command. The store files and transaction logs are first copied to a local temporary directory (configured with --temp-path) before being streamed to the cloud destination, so the temporary location must have free space at least equal to the size of the database being backed up. See Options for details on the --temp-path option.

Neo4j uses the AWS SDK v2 to call the APIs on AWS using AWS URLs. Alternatively, you can override the endpoints so that the AWS SDK can communicate with alternative storage systems, such as Ceph, Minio, or LocalStack, using the system variables aws.endpointUrls3, aws.endpointUrlS3, or aws.endpointUrl, or the environments variables AWS_ENDPOINT_URL_S3 or AWS_ENDPOINT_URL.

  1. Install the AWS CLI by following the instructions in the AWS official documentation — Install the AWS CLI version 2.

  2. Create an S3 bucket and a directory to store the backup files using the AWS CLI:

    aws s3 mb --region=us-east-1 s3://myBucket
    aws s3api put-object --bucket myBucket --key myDirectory/

    For more information on how to create a bucket and use the AWS CLI, see the AWS official documentation — Use Amazon S3 with the AWS CLI and Use high-level (s3) commands with the AWS CLI.

  3. Verify that the ~/.aws/config file is correct by running the following command:

    cat ~/.aws/config

    The output should look like this:

    [default]
    region=us-east-1
  4. Configure the access to your AWS S3 bucket by setting the aws_access_key_id and aws_secret_access_key in the ~/.aws/credentials file and, if needed, using a bucket policy. For example:

    1. Use aws configure set aws_access_key_id aws_secret_access_key command to set your IAM credentials from AWS and verify that the ~/.aws/credentials is correct:

      cat ~/.aws/credentials

      The output should look like this:

      [default]
      aws_access_key_id=this.is.secret
      aws_secret_access_key=this.is.super.secret
    2. Additionally, you can use a resource-based policy to grant access permissions to your S3 bucket and the objects in it. Create a policy document with the following content and attach it to the bucket. Note that both resource entries are important to be able to download and upload files.

      {
          "Version": "2012-10-17",
          "Id": "Neo4jBackupAggregatePolicy",
          "Statement": [
              {
                  "Sid": "Neo4jBackupAggregateStatement",
                  "Effect": "Allow",
                  "Action": [
                      "s3:ListBucket",
                      "s3:GetObject",
                      "s3:PutObject",
                      "s3:DeleteObject"
                  ],
                  "Resource": [
                      "arn:aws:s3:::myBucket/*",
                      "arn:aws:s3:::myBucket"
                  ]
              }
          ]
      }
  5. Run the neo4j-admin database backup command to back up your database to your AWS S3 bucket:

    bin/neo4j-admin database backup --to-path=s3://myBucket/myDirectory/ mydatabase
  1. Ensure you have a Google account and a project created in the Google Cloud Platform (GCP).

    1. Install the gcloud CLI by following the instructions in the Google official documentation — Install the gcloud CLI.

    2. Create a service account and a service account key using Google official documentation — Create service accounts and Creating and managing service account keys.

    3. Download the JSON key file for the service account.

    4. Set the GOOGLE_APPLICATION_CREDENTIALS and GOOGLE_CLOUD_PROJECT environment variables to the path of the JSON key file and the project ID, respectively:

      export GOOGLE_APPLICATION_CREDENTIALS="/path/to/keyfile.json"
      export GOOGLE_CLOUD_PROJECT=YOUR_PROJECT_ID
    5. Authenticate the gcloud CLI with the e-mail address of the service account you have created, the path to the JSON key file, and the project ID:

      gcloud auth activate-service-account [email protected] --key-file=$GOOGLE_APPLICATION_CREDENTIALS --project=$GOOGLE_CLOUD_PROJECT

      For more information, see the Google official documentation — gcloud auth activate-service-account.

    6. Create a bucket in the Google Cloud Storage using Google official documentation — Create buckets.

    7. Verify that the bucket is created by running the following command:

      gcloud storage ls

      The output should list the created bucket.

  2. Run neo4j-admin database backup command to back up your database to your Google bucket:

    bin/neo4j-admin database backup --to-path=gs://myBucket/myDirectory/ mydatabase
  1. Ensure you have an Azure account, an Azure storage account, and a blob container.

    1. You can create a storage account using the Azure portal.
      For more information, see the Azure official documentation on Create a storage account.

    2. Create a blob container in the Azure portal.
      For more information, see the Azure official documentation on Quickstart: Upload, download, and list blobs with the Azure portal.

  2. Install the Azure CLI by following the instructions in the Azure official documentation — Azure official documentation.

  3. Authenticate the neo4j or neo4j-admin process against Azure using the default Azure credentials.
    See the Azure official documentation on default Azure credentials for more information.

    az login

    Then you should be ready to use Azure URLs in either neo4j or neo4j-admin.

  4. To validate that you have access to the container with your login credentials, run the following commands:

    # Upload a file:
    az storage blob upload --file someLocalFile  --account-name accountName - --container someContainer --name remoteFileName  --auth-mode login
    
    # Download the file
    az storage blob download  --account-name accountName --container someContainer --name remoteFileName --file downloadedFile --auth-mode login
    
    # List container files
    az storage blob list  --account-name someContainer --container someContainer  --auth-mode login
  5. Run neo4j-admin database backup command to back up your database to your Azure container:

    bin/neo4j-admin database backup --to-path=azb://myStorageAccount/myContainer/myDirectory/ mydatabase

Perform a differential backup using the --prefer-diff-as-parent option

By default, a differential backup (--type=DIFF) uses the most recent non-empty backup — whether full or differential — in the directory as its parent.

The --prefer-diff-as-parent option changes this behavior and forces the backup job to use the latest differential backup as the parent, even if a newer full backup exists.

This approach allows you to maintain a chain of differential backups for all transactions and restore to any point in time. Without this option, the transactions between the last full backup and a previous differential backup cannot be backed up as individual transactions.

To use the --prefer-diff-as-parent option, set it to true.

The following examples cover different scenarios for using the --prefer-diff-as-parent option.

Let’s assume that you write 10 transactions to the neo4j database every hour, except from 12:30 to 13:30, when you do not write any transactions.

There is a backup job that takes a backup every hour and a full backup every four hours. An empty backup has no transactions, meaning that both the lower transaction ID and the upper transaction ID are zero.

Imagine you have the following backup chain:

Timestamp Backup name Backup type Lower Transaction ID Upper Transaction ID

10:30

backup1

FULL

1

10

11:30

backup2

DIFF

11

20

12:30

backup3

DIFF

21

30

13:30

backup4

DIFF

0

0

14:30

backup5

FULL

1

40

At 15:30, you execute the following backup command:

neo4j-admin database backup --from=<address:port> --to-path=<targetPath> --type=DIFF neo4j

The result would be:

15:30

backup6

DIFF

41

50

The result means you have chosen backup5 as the parent for your differential backup6 since the backup5 is the latest non-empty backup.

However, if you execute the following command with the --prefer-diff-as-parent option:

neo4j-admin database backup --from=<address:port> --to-path=<targetPath> --type=DIFF --prefer-diff-as-parent neo4j

The result would be:

15:30

backup6

DIFF

31

50

In this case, the backup3 is selected as the parent since it is the latest non-empty differential backup.

Let’s assume that you write 10 transactions to the neo4j database every hour and trigger an hourly full backup.

Timestamp Backup name Backup type Lower Transaction ID Upper Transaction ID

10:30

backup1

FULL

1

10

11:30

backup2

FULL

11

20

In this case, there is no differential backup. Therefore, the --prefer-diff-as-parent option has no effect and the behaviour is the same as the default one.

neo4j-admin database backup \
--from=<address:port> --to-path=<targetPath> \
--type=DIFF --prefer-diff-as-parent \
neo4j

The result would be (with or without the --prefer-diff-as-parent option):

12:30

backup3

DIFF

21

30

Split a backup archive into multiple files

Starting with Neo4j 2026.09, you can split the backup archive created by neo4j-admin database backup into multiple files using the --split-archive-part-size option.

Run the following command to split the backup into parts of up to 500GB:

neo4j-admin database backup \
    --type=full \
    --to-path=/path/to/backup \
    --temp-path=/path/to/preferred/staging/area \
    --split-archive-part-size=500G mydatabase

The --split-archive-part-size option defines the maximum size of each split part. The size can be specified in bytes or with a unit suffix, such as 5G, 100g, or 1TiB.

The minimum split size is 1GiB. If the option is not specified, or is set to 0, the backup is written as a single file.

When splitting is enabled, the backup is written as a set of files in the same directory:

  • file.backup

  • file.backup.1

  • file.backup.2

  • …​

The first file (.backup) is a metadata file that contains information about the split archive, such as its ID and the number of files. The remaining files contain the backup data.

You must keep all files belonging to a split archive in the same location. If you move some of the files, all files for that backup artifact must be moved together.

When loading a split archive, specify the first file (.backup) as the backup artifact, in the same way as for an unsplit backup.

You can also configure a default split size using the server.split_archive.part_size setting. This is useful when you regularly run neo4j-admin commands and want to use the same split size without specifying --split-archive-part-size for every command.

For example:

server.split_archive.part_size=500GiB

When server.split_archive.part_size is configured, neo4j-admin database backup uses that value by default.

If both server.split_archive.part_size and --split-archive-part-size are specified, the command-line option overrides the setting’s value.

Both the server.split_archive.part_size setting and the --split-archive-part-size option default to 0, which disables splitting.

For both the configuration setting and the command-line option, a non-zero value must be at least 1GiB. Values smaller than 1GiB cause the command to fail.

Glossary

allocator

A component in the cluster that allocates databases to servers according to the topology constraints specified and an allocation strategy.

asynchronous replication

Asynchronous replication is used by secondary copies to poll for new transactions, which means they cannot be guaranteed to have received the most recent transactions. This enables efficient scale-out of read-performance.

Aura instance

A fully-managed DBMS represented by a single instance ID, that is running in the Neo4j Aura cloud.

auto-commit transaction

An automatically committed transaction that contains a single query.

Bolt protocol

Bolt is a protocol used for interaction between Neo4j instances and drivers.

bookmark

A marker the client can request from the cluster to ensure that it is able to read its own writes so that the application’s state is consistent and only databases that have a copy of the bookmark are permitted to respond.

category (Bloom)

A category is based on a node label and is defined in a Perspective as a way of visually distinguishing nodes with the same label(s).

causal consistency

All servers in a cluster agree on the order in which transactions take place. The position of a server on the causal chain can be guaranteed using a bookmark.

cluster

A Neo4j DBMS that spans multiple servers working together to increase fault tolerance and/or read scalability. Databases on a cluster may be configured to replicate across servers in the cluster thus achieving read scalability or high availability.

client application

Software that interacts with a Neo4j server.

commit

A commit is the successful completion of a transaction, which ensures durability of any changes made. For more details, visit Operations Manual → Transaction management.

composite database

Composite databases are the means to access partitioned graph data with a single Cypher query.

constraint

Constraints are sets of data modeling rules that ensure the data is consistent and reliable.

Cypher®

Neo4j’s graph query language.

data model

A data model defines how information is organized in a database. A good data model will make querying and understanding your data easier. In Neo4j, the data models have a graph structure.

database

A database is a container used by the DBMS to manage and store graph data. The physical structure of data is controlled by the database.

database vs graph

Databases are the physical containers of graph data. Graphs are the logical structure of data in Neo4j.

Database Management System

Database Management System, or DBMS, capable of managing multiple databases. A DBMS may run on a single server, or span several servers configured as a cluster.

database schema

The prescribed property existence and datatypes for nodes and relationships.

deallocate

An act of removing a database from a server or a server from a cluster without loss of data or reduced fault tolerance.

degree (of a node)

The number of relationships of a specific node; loops are counted twice.

disaster recovery

A manual intervention to restore availability of a cluster, or databases within a cluster.

driver

A software library that provides access to Neo4j from a particular programming language.

election

In the event that the Raft leader becomes unresponsive, followers automatically trigger an election and vote for a new leader.

entity

A node or a relationship.

expression (Cypher)

A component of a Cypher query which produces values. It may be used in projections, as a predicate, or when setting properties on graph elements.

fabric

Fabric is the architectural design of a unified system that provides a single access point to local or distributed graph data.

fault tolerance

A guarantee that a cluster can maintain a database’s persistence and availability in the event of one or more servers failing.

follower

A primary copy of a database acting as a follower, receives and acknowledges synchronous writes from the leader.

Generative AI (GenAI)

A type of artificial intelligence (AI) system that generates text, images, or other media in response to prompts.

graph

A logical representation of a set of nodes where some pairs are connected by relationships.

index

Data structure that improves read performance of a database.

knowledge graph

A specific type of graph that has an organizing principle so that a user (or a computer system) can reason about the underlying data. The organizing principle provides an additional layer of structure that adds context to support knowledge discovery.

label

Marks a node as a member of a named and indexed subset. A node may be assigned zero or more labels.

leader

A single primary copy of a database is designated as the leader. It receives all write transactions from clients and replicates writes synchronously to followers and asynchronously to secondary copies of the database.

main database

In terms of Neo4j Enterprise Studio, the database(s) containing the user’s data. Can exist in the same Neo4j deployment as the tool asset database.

motif

A description of a specific pattern within a graph.

node

A node represents an entity or discrete object in your graph data model. Nodes can be connected by relationships, hold data in properties, and are classified by labels.

operator

A symbol representing a mathematical or logical operation.

parameter

Named value provided when running a Cypher statement.

path

A sequence of nodes and the relationships connecting them, that does not contain duplicate relationships. Several paths can match a pattern.

pattern

A specific arrangement of nodes and relationships that can be matched in a graph. A pattern follows a motif.

perspective (Bloom)

A Perspective defines a certain business view or domain that can be found in the target Neo4j graph. A single Neo4j graph can be viewed through different Perspectives, each tailored for a different business purpose.

primary

A copy of the database that is able to process write transactions and is eligible to be elected as a leader. It participates in fault tolerant writes as it is part of the majority required to acknowledge and commit write transactions.

primary vs secondary

In a cluster, databases can operate in either primary or secondary mode. Primary databases are able to process write and read transactions, ensuring fault tolerance. Secondary databases are replicated asynchronously from primaries, and their main purpose is to provide read scaling within the cluster.

project (Aura)

An isolated environment in the unified Aura console that contains its own database instances, configurations, and resources. Preceded by tenant in the classic Aura console.

property

Properties are key-value pairs that are used for storing data on nodes and relationships.

query (Cypher)

A statement that retrieves or writes information to a database.

Raft group

A group of servers that are participating in hosting a particular database in primary mode.

Raft group member

A server that is participating in a Raft group. A server can be a member of one or more groups.

Raft log

A shared log between all Raft group members that is guaranteed to be consistently updated and viewed by those members. The log contains both database data and operational state of the Raft group.

Raft protocol

The networking mechanism that enables a database to replicate its data across multiple servers to give high availability for accessing the data and high durability to the data stored.

read scaling

Distributing query load by creating additional database copies hosted in secondary mode (read-only).

relationship

A relationship represents a connection between nodes in your graph data model. Relationships connect a source node to a target node, hold data in properties, and are classified by type.

secondary

An asynchronously replicated copy of the database that provides read scaling within the cluster.

seed

A seed is a database dump or a full backup used to create a database on a cluster. This is sometimes called seeding.

server

A physical machine, a virtual machine, or a container running an instance of Neo4j. Servers can be standalone or part of a cluster.

session

A causally linked sequence of transactions.

session consistency

An alternative name for Neo4j’s causal consistency.

standalone

A single server running Neo4j and not part of a cluster.

synchronous replication

Synchronous replication requires the leader primary to replicate a transaction and block the commit until a quorum of the follower primaries acknowledges that the transaction is successfully replicated. Once the transaction is replicated, the commit is allowed to proceed. This ensures data durability and consistency within the cluster.

system database

A database used by Neo4j to store system information.

tenant (Aura)

An isolated environment in the classic Aura console that contains its own database instances, configurations, and resources. Replaced by project in the unified Aura console.

tool asset database

In terms of Neo4j Enterprise Studio, the database where tools' assets are stored. This can be in the same Neo4j deployment as the main database(s) or in a separate deployment.

topology

A configuration that describes how the copies of a database should be spread across the servers in a cluster, see primary mode and secondary mode.

transaction

A transaction comprises a unit of work performed against a database. It is treated in a coherent and reliable way, independent of other transactions. Transactions comply with the ACID consistency model (atomic, consistent, isolated, and durable).