Create a graph
The main entry point to every graph operation is the Graph object.
There are several ways to create a graph with the Python client:
-
From a Neo4j database with a graph projection
-
From a distribution or previously created graph (see Other ways to generate graphs)
-
From Pandas DataFrames with the
constructmethod -
From a NetworkX graph
Once created, the Graph object can be used to inspect the projected graph, run algorithms, or train machine learning models.
Neo4j graph projection
As for the Cypher API, the Python client offers two ways of projecting a graph object:
-
Native projection for most use cases
-
Cypher projection for more complex use cases that require custom Cypher queries
|
The examples below assume that a |
Preparation
The following query creates an example graph in the Neo4j database.
gds.run_cypher(
"""
CREATE
(m: City {name: "Malmö", population: 360000}),
(l: City {name: "London", population: 8800000}),
(s: City {name: "San Mateo", population: 105000}),
(m)-[:FLY_TO {cost: 200}]->(l),
(l)-[:FLY_TO {cost: 200}]->(m),
(l)-[:FLY_TO {cost: 1000}]->(s),
(s)-[:FLY_TO {cost: 1000}]->(l)
"""
)
Native projection
The native projection in the client is similar to its Cypher API counterpart.
G, result = gds.graph.project.native(
graph_name="offices", # Graph name
node_label_filter=["City"], # Node projection
relationship_type_filter=["FLY_TO"], # Relationship projection
node_properties=["population"], # Node properties
relationship_properties=["cost"] # Relationship properties
)
G is a Graph object, and result is a ProjectionResult object containing metadata from the underlying procedure call.
You can use the wildcard * in the node_label_filter or relationship_type_filter arguments to project all available node labels or relationship types, respectively (for instance node_label_filter=["*"]).
G, result = gds.graph.project.native(
graph_name="offices", # Graph name
node_projection=["City"], # Node projection
relationship_projection=["FLY_TO"], # Relationship projection
node_properties=["population"], # Node properties
relationship_properties=["cost"], # Relationship properties
read_concurrency=4 # Configuration
)
G is a Graph object, and result is a GraphProjectResult object containing metadata from the underlying procedure call.
Note that all projection syntax variants are supported by way of specifying a Python dict or list for the node and relationship projection arguments.
To specify configuration parameters corresponding to the keys of the procedure’s configuration map, we give named keyword arguments, like for read_concurrency=4 above.
Read more about the syntax in the GDS manual.
Cypher projection
The Cypher projection in the client is similar to its Cypher API counterpart.
G, result = gds.graph.project.cypher(
# The query has to use the .remote endpoint in AGA,
# which does not include the graph name
query="""
MATCH (source:City)-[rel:FLY_TO]->(target:City)
RETURN gds.graph.project.remote(
source,
target,
{
sourceNodeLabels: labels(source),
targetNodeLabels: labels(target),
sourceNodeProperties: { population: source.population },
targetNodeProperties: { population: target.population },
relationshipType: type(rel),
relationshipProperties: { cost: rel.cost }
}
)
""",
graph_name="offices-cypher" # Method parameter
)
G, result = gds.graph.project.cypher(
query="""
MATCH (source:City)-[rel:FLY_TO]->(target:City)
RETURN gds.graph.project(
'offices-cypher',
source,
target,
{
sourceNodeLabels: labels(source),
targetNodeLabels: labels(target),
sourceNodeProperties: { population: source.population },
targetNodeProperties: { population: target.population },
relationshipType: type(rel),
relationshipProperties: { cost: rel.cost }
}
)
"""
)
Unlike gds.run_cypher but very much like gds.graph.project.native, it returns a tuple of a Graph object and a GraphCypherProjectResult object containing metadata from the Cypher execution.
The method does not modify the Cypher query in any way, all projection configuration must be done in the query itself.
It expects the query to end with a single RETURN gds.graph.project(…) aggregation, and raises a ValueError if the query returns anything else.
If your query needs additional aggregations or a renamed result row, use gds.run_cypher together with gds.graph.get to achieve the same result.
Other ways to generate graphs
To generate a graph from a distribution, use gds.graph.generate.
To get a graph object that represents a graph that has already been projected into the graph catalog, one can call the client-side only get method and passing it a name:
G = gds.graph.get("offices")
For users who are GDS admins, gds.graph.get will resolve graph names into Graph objects also when the provided name refers to another user’s graph projection.
In addition to those aforementioned there are more methods that create graph objects:
-
gds.graph.filter -
gds.graph.sample.rwr -
gds.graph.sample.cnarw
Their Cypher signatures map to Python in much the same way as gds.graph.project above.
Context management
The graph object also implement the context managment protocol, i.e., is usable inside with clauses.
On exiting the with block, the graph projection will be automatically dropped on the server side.
# We use the example graph from the `Projecting a graph object` section
with gds.graph.project.native(
"tmp_offices", # Graph name
["City"], # Node projection
["FLY_TO"], # Relationship projection
)[0] as G_tmp:
assert G_tmp.exists()
# Outside of the with block the Graph does not exist
assert not gds.graph.exists("tmp_offices")
Using a graph object
The primary use case for a graph object is to pass it to algorithms, but it’s also the input to most methods of the GDS Graph Catalog:
The graph catalog
All procedures of the GDS Graph Catalog have corresponding Python methods in the client.
Of those catalog procedures that take a graph name string as input, their Python client equivalents instead take a Graph object, with the exception of gds.graph.exists which still takes a graph name string.
Below are some examples of how the GDS Graph Catalog can be used via the client:
result = gds.degree_centrality.mutate(G, mutate_property="degree")
assert result.centrality_distribution is not None
# List graphs in the catalog
list_result = gds.graph.list()
# Check for existence of a graph in the catalog
assert gds.graph.exists("offices")
# Stream the node property 'degree'
result = gds.graph.node_properties.stream(G, node_properties=["degree"])
# Drop a graph; same as G.drop()
gds.graph.drop(G)
Inspecting a graph object
The Graph object includes convenience methods to retrieve information on the projected graph.
-
name: The name of the projected graph -
database: Name of the database in which the graph has been projected -
node_count: The node count of the projected graph -
relationship_count: The relationship count of the projected graph -
node_labels: A list of the node labels present in the graph -
relationship_types: A list of the relationship types present in the graph -
node_properties: Returns a dictionary mapping every node label to a list of the properties present on nodes with that label -
relationship_properties: Returns a dictionary mapping every relationship type to a list of the properties present on relationships with that type -
degree_distribution: The average out-degree of generated nodes -
density: Density of the graph -
size_in_bytes: Number of bytes used in the Java heap to store the graph -
memory_usage: Human-readable description ofsize_in_bytes -
exists: ReturnsTrueif the graph exists in the GDS Graph Catalog, otherwiseFalse -
drop: Removes the graph from the GDS Graph Catalog -
configuration: The configuration used to project the graph in memory -
creation_time: Time when the graph was projected -
modification_time: Time when the graph was last modified
For example, to get the node count and node properties of a graph G, run the following:
n = G.node_count()
props = G.node_properties()["City"]