A taxi dataset is already a graph: locations are nodes, trips are directed weighted edges. The interesting part was not the algorithms. It was running the identical algorithms through two completely different deployment models and getting the same answers out.
Phase one: everything at build time
The Dockerfile builds a working Neo4j end to end — installs the GDS and APOC plugins, configures networking, and runs a loader that filters 3.6 million rows of March 2022 TLC data down to the 1,530 valid Bronx trips and writes them in as 42 location nodes and 1,530 relationships. No manual steps after the build.
Loading went in through parameterised Cypher UNWIND statements rather than
row-by-row inserts. PageRank and BFS both run through the GDS library against
an in-memory projection.
Phase two: the same graph, streaming
The batch loader was replaced by a Python Kafka producer streaming trips in real time through Zookeeper and Kafka into a Neo4j Kafka Sink Connector, with Neo4j itself deployed by Helm onto Minikube. The graph algorithms sit behind the same Python interface in both phases, which is what makes the two results comparable at all.
Three failures worth keeping
The rewrite broke in ways that only distributed systems break.
A Kubernetes Service label selector did not match its pods, so port-forwarding silently succeeded and forwarded to nothing. The connector was OOM-killed because its container memory limit was smaller than the JVM's default heap — the JVM was behaving correctly and the container killed it for it. And a third-party image shipped with baked-in credentials, overridden through a ConfigMap volume mount rather than a rebuild.
