Deploy Kafka on Kubernetes with Strimzi (KRaft Mode)

This guide shows you how to deploy Kafka on Kubernetes using the Strimzi operator in KRaft mode. The deployment provides separate broker and controller node pools, JMX metrics, integrated Cruise Control for automated rebalancing, and namespaced operator RBAC. All configuration is driven from a single environment file in version control.

Why KRaft mode

This deployment uses Kafka KRaft mode, in which a dedicated controller quorum manages cluster metadata instead of ZooKeeper. The result is fewer moving parts and independent scaling of brokers and controllers. ZooKeeper is not deployed and not required.

Prerequisites

Before you begin, confirm that:

  • You have a Kubernetes cluster with kubectl configured against it.
  • You have permissions to create namespaces, custom resource definitions, and namespaced operator resources.
  • The cluster does not already have a Strimzi operator running. The bundle in this repository will conflict with an existing operator.
  • Your storage class supports the disk sizes set in BROKER_STORAGE_SIZE (default 100 Gi per broker) and CONTROLLER_STORAGE_SIZE (default 20 Gi per controller).

Reviewing config.env Before Deploying

Open kafka/env/config.env and review the values that drive the deployment. The most important groups are:

GroupVariablesDefault
VersionsSTRIMZI_VERSION, KAFKA_VERSION0.47.0 / 4.0.0
Node countsBROKER_REPLICAS, CONTROLLER_REPLICAS2 / 2 (development)
StorageBROKER_STORAGE_SIZE, BROKER_STORAGE_CLASS, CONTROLLER_STORAGE_SIZE, CONTROLLER_STORAGE_CLASS100Gi / standard / 20Gi / standard
Broker resourcesBROKER_CPU_REQUEST, BROKER_CPU_LIMIT, BROKER_MEM_REQUEST, BROKER_MEM_LIMIT, BROKER_JVM_XMS, BROKER_JVM_XMX1 / 1 / 4Gi / 4Gi / 2g / 2g
Controller resourcesCONTROLLER_CPU_REQUEST, CONTROLLER_CPU_LIMIT, CONTROLLER_MEM_REQUEST, CONTROLLER_MEM_LIMIT1 / 1 / 2Gi / 2Gi
Connect worker (optional)CONNECT_REPLICAS, CONNECT_MEM_LIMIT, CONNECT_JVM_XMX, CONNECT_MIN_REPLICAS, CONNECT_MAX_REPLICAS2 / 8Gi / 5g / 1 / 3

The defaults target development

The default node counts are BROKER_REPLICAS=2 and CONTROLLER_REPLICAS=2. These are appropriate for development and smoke testing only. For production, set both to 3 or higher. Controllers must be deployed with an odd number of replicas (1, 3, 5) to maintain quorum during rolling restarts and partial failures.

Deploying the Stack

From the repository root, apply the Kafka overlay. This installs the Strimzi operator with namespaced RBAC, the broker and controller node pools, and the Kafka custom resource with Cruise Control:

kubectl apply -k kafka/

Wait for the operator deployment to become Available before proceeding:

kubectl wait --for=condition=Available deployment/strimzi-cluster-operator \
  -n kafka --timeout=300s

Once the operator is Ready, it begins reconciling the broker and controller node pools and the Kafka cluster itself.

Optional: Enabling Kafka Connect

The repository ships a 04-connect/ overlay with Kafka Connect and a ClickHouse sink connector. By default this overlay is commented out. Enable it only when you need Connect for the migration pipeline or for streaming Kafka topics into ClickHouse outside of the dedicated migration overlay.

Verifying the Initial Deployment

Confirm the operator and the Kafka custom resources are reconciling cleanly:

kubectl -n kafka get pods
kubectl -n kafka get kafka,kafkanodepool

Expected pods:

  • One pod per broker (named cly-kafka-brokers-N).
  • One pod per controller (named cly-kafka-controllers-N).
  • One Cruise Control pod (cly-kafka-cruise-control-…).
  • One Strimzi operator pod (strimzi-cluster-operator-…).

The Kafka custom resource should report True KRaft True in its status conditions, which confirms the cluster is running in KRaft mode and is ready.

Connecting Producers and Consumers

The bootstrap service for in-cluster clients is:

cly-kafka-kafka-bootstrap:9092

Test by exec-ing into a broker pod and using the bundled CLI:

kubectl -n kafka exec cly-kafka-brokers-0 -- \
  /opt/kafka/bin/kafka-topics.sh --bootstrap-server localhost:9092 --list

Cluster is up

With pods Ready and the topics command returning successfully, the cluster is ready. Run the full validation checklist documented separately to confirm metrics, replication, and Cruise Control are healthy before depending on the cluster in production.

Common Issues and Gotchas

Operator pod is in CrashLoopBackOff

Most common causes are missing RBAC subresource permissions or a missing image map in the operator ConfigMap. Confirm RBAC:

kubectl auth can-i --as=system:serviceaccount:kafka:strimzi-cluster-operator \
  -n kafka list kafkas

Then read the operator logs to identify the specific permission error:

kubectl -n kafka logs deployment/strimzi-cluster-operator --tail=100
Pods stuck in Pending

The PVC has not yet bound, or the cluster does not have nodes that can satisfy the resource request. Inspect the PVCs and the unschedulable pod:

kubectl -n kafka get pvc
kubectl -n kafka describe pod <pod-name>

Check storage class and node taints. Adjust BROKER_STORAGE_CLASS in config.env if the default standard class is not present in your cluster.

Configuration not applied after editing config.env

Confirm the ConfigMap was regenerated with the new values:

kubectl get configmap kafka-config -n kafka -o yaml

Then confirm the Kafka custom resource picked up the values via Kustomize replacements:

kubectl get kafka cly-kafka -n kafka -o yaml

If the values are present in the ConfigMap but not in the Kafka resource, re-apply the overlay:

kubectl apply -k kafka/
Brokers cannot reach controllers

Read the broker logs and the controller logs side by side:

kubectl -n kafka logs cly-kafka-brokers-0
kubectl -n kafka logs cly-kafka-controllers-0

Confirm that Kubernetes NetworkPolicies, if any, allow inter-pod communication on the KRaft controller port. The most common cause of this symptom is a NetworkPolicy applied to the kafka namespace that does not include the Strimzi-managed selectors.

Was this page helpful?
Reach out to us for any other questions.
Helpful?

Looking For More Help?