This guide shows you how to deploy Kafka on Kubernetes using the Strimzi operator in KRaft mode. The deployment provides separate broker and controller node pools, JMX metrics, integrated Cruise Control for automated rebalancing, and namespaced operator RBAC. All configuration is driven from a single environment file in version control.
Why KRaft mode
This deployment uses Kafka KRaft mode, in which a dedicated controller quorum manages cluster metadata instead of ZooKeeper. The result is fewer moving parts and independent scaling of brokers and controllers. ZooKeeper is not deployed and not required.
Prerequisites
Before you begin, confirm that:
- You have a Kubernetes cluster with
kubectlconfigured against it. - You have permissions to create namespaces, custom resource definitions, and namespaced operator resources.
- The cluster does not already have a Strimzi operator running. The bundle in this repository will conflict with an existing operator.
- Your storage class supports the disk sizes set in
BROKER_STORAGE_SIZE(default 100 Gi per broker) andCONTROLLER_STORAGE_SIZE(default 20 Gi per controller).
Reviewing config.env Before Deploying
Open kafka/env/config.env and review the values that drive the deployment. The most important groups are:
| Group | Variables | Default |
|---|---|---|
| Versions | STRIMZI_VERSION, KAFKA_VERSION | 0.47.0 / 4.0.0 |
| Node counts | BROKER_REPLICAS, CONTROLLER_REPLICAS | 2 / 2 (development) |
| Storage | BROKER_STORAGE_SIZE, BROKER_STORAGE_CLASS, CONTROLLER_STORAGE_SIZE, CONTROLLER_STORAGE_CLASS | 100Gi / standard / 20Gi / standard |
| Broker resources | BROKER_CPU_REQUEST, BROKER_CPU_LIMIT, BROKER_MEM_REQUEST, BROKER_MEM_LIMIT, BROKER_JVM_XMS, BROKER_JVM_XMX | 1 / 1 / 4Gi / 4Gi / 2g / 2g |
| Controller resources | CONTROLLER_CPU_REQUEST, CONTROLLER_CPU_LIMIT, CONTROLLER_MEM_REQUEST, CONTROLLER_MEM_LIMIT | 1 / 1 / 2Gi / 2Gi |
| Connect worker (optional) | CONNECT_REPLICAS, CONNECT_MEM_LIMIT, CONNECT_JVM_XMX, CONNECT_MIN_REPLICAS, CONNECT_MAX_REPLICAS | 2 / 8Gi / 5g / 1 / 3 |
The defaults target development
The default node counts are BROKER_REPLICAS=2 and CONTROLLER_REPLICAS=2. These are appropriate for development and smoke testing only. For production, set both to 3 or higher. Controllers must be deployed with an odd number of replicas (1, 3, 5) to maintain quorum during rolling restarts and partial failures.
Deploying the Stack
From the repository root, apply the Kafka overlay. This installs the Strimzi operator with namespaced RBAC, the broker and controller node pools, and the Kafka custom resource with Cruise Control:
kubectl apply -k kafka/
Wait for the operator deployment to become Available before proceeding:
kubectl wait --for=condition=Available deployment/strimzi-cluster-operator \ -n kafka --timeout=300s
Once the operator is Ready, it begins reconciling the broker and controller node pools and the Kafka cluster itself.
Optional: Enabling Kafka Connect
The repository ships a 04-connect/ overlay with Kafka Connect and a ClickHouse sink connector.
By default this overlay is commented out. Enable it only when you need Connect for the migration pipeline or
for streaming Kafka topics into ClickHouse outside of the dedicated migration overlay.
Verifying the Initial Deployment
Confirm the operator and the Kafka custom resources are reconciling cleanly:
kubectl -n kafka get pods kubectl -n kafka get kafka,kafkanodepool
Expected pods:
- One pod per broker (named
cly-kafka-brokers-N). - One pod per controller (named
cly-kafka-controllers-N). - One Cruise Control pod (
cly-kafka-cruise-control-…). - One Strimzi operator pod (
strimzi-cluster-operator-…).
The Kafka custom resource should report True KRaft True in its status conditions, which confirms the cluster is running in KRaft mode and is ready.
Connecting Producers and Consumers
The bootstrap service for in-cluster clients is:
cly-kafka-kafka-bootstrap:9092
Test by exec-ing into a broker pod and using the bundled CLI:
kubectl -n kafka exec cly-kafka-brokers-0 -- \ /opt/kafka/bin/kafka-topics.sh --bootstrap-server localhost:9092 --list
Cluster is up
With pods Ready and the topics command returning successfully, the cluster is ready. Run the full validation checklist documented separately to confirm metrics, replication, and Cruise Control are healthy before depending on the cluster in production.
Common Issues and Gotchas
CrashLoopBackOffMost common causes are missing RBAC subresource permissions or a missing image map in the operator ConfigMap. Confirm RBAC:
kubectl auth can-i --as=system:serviceaccount:kafka:strimzi-cluster-operator \ -n kafka list kafkas
Then read the operator logs to identify the specific permission error:
kubectl -n kafka logs deployment/strimzi-cluster-operator --tail=100
PendingThe PVC has not yet bound, or the cluster does not have nodes that can satisfy the resource request. Inspect the PVCs and the unschedulable pod:
kubectl -n kafka get pvc kubectl -n kafka describe pod <pod-name>
Check storage class and node taints. Adjust BROKER_STORAGE_CLASS in config.env if the default standard class is not present in your cluster.
config.envConfirm the ConfigMap was regenerated with the new values:
kubectl get configmap kafka-config -n kafka -o yaml
Then confirm the Kafka custom resource picked up the values via Kustomize replacements:
kubectl get kafka cly-kafka -n kafka -o yaml
If the values are present in the ConfigMap but not in the Kafka resource, re-apply the overlay:
kubectl apply -k kafka/
Read the broker logs and the controller logs side by side:
kubectl -n kafka logs cly-kafka-brokers-0 kubectl -n kafka logs cly-kafka-controllers-0
Confirm that Kubernetes NetworkPolicies, if any, allow inter-pod communication on the KRaft controller port. The most common cause of this symptom is a NetworkPolicy applied to the kafka namespace that does not include the Strimzi-managed selectors.