Failure modes you are likely to meet while backfilling drill events on a Kubernetes deployment, and what to do about each one.
The migration dashboard's Help & Recovery tab covers the same ground with the fix one click away, and is usually faster than working from here. The guiding property of this flow is worth keeping in mind while you read: after cutover, no failure anywhere in it can touch live data. Every response below is restart, resume, or redo: never clean up or restore.
The Run Starts but Nothing Happens: pauseReason: boundary-unset
This is the startup guard, not a fault. A fresh run found live data already arriving in its target ClickHouse with no cd bound set, and is holding before mapping anything, because that situation is either a duplicate factory or a deliberate choice.
If a mirror is running (scenario 2), set the bound with POST /control/set-boundary and the run releases itself. If nothing mirrors traffic (scenarios 1 and 3), click Proceed unbounded or call POST /control/allow-unbounded, which releases every held pod at once. A plain Resume is deliberately ignored while the question is open, which is why pressing it appears to do nothing.
Pods Are Up but Idle: pauseReason: not-started
Expected, not a fault: LEDGER_START_PAUSED: "true" is set, which is what the prerequisites article tells you to do so you can run preflight and the dry run without starting a migration. The pods are holding before mapping: nothing has been read, written, or indexed.
Press Start on the dashboard, or call POST /control/resume, when you actually want the backfill to begin. The gate lives in the ledger, so one Start covers the whole fleet and pods that join later start immediately.
Do not confuse this with boundary-unset above. not-started means you have not started the run; boundary-unset means the run started and stopped to ask a question.
CrashLoopBackOff With a Config Error
Only MONGO_URI and CLICKHOUSE_URL are required; everything else has a default. The service also rejects a few semantically impossible combinations at startup rather than misbehaving later:
-
CLICKHOUSE_RETRY_BASE_DELAY_MSgreater thanCLICKHOUSE_RETRY_MAX_DELAY_MS -
BACKPRESSURE_PARTITION_PCT_LOWnot belowBACKPRESSURE_PARTITION_PCT_HIGH, or the same for theTOTAL_PCTpair -
LEDGER_CD_UPPER_BOUNDthat is neither an ISO date nor epoch milliseconds
The previous container's log names the rejected setting: kubectl logs deploy/drill-migrator --previous | head -40.
ImagePullBackOff
The image or tag does not exist, or registry credentials are missing. Verify the tag, and add an imagePullSecrets entry to the pod spec for a private registry.
Liveness or Readiness Probe Failing
Both probes hit /healthz. If it is failing, one of the two backing services is unreachable. Check what the pod is actually pointed at:
kubectl exec deploy/drill-migrator -- env | grep -E 'MONGO|CLICKHOUSE' kubectl logs deploy/drill-migrator | tail -40
A NetworkPolicy that does not allow egress to the mongodb and clickhouse namespaces is a common cause in a locked-down cluster. Note that /healthz is the only health endpoint: a manifest probing /readyz will fail every probe.
Pod Exits With "No collections found matching prefix"
The source database has no drill_events* collections. Confirm the restore landed in the right database and that MONGO_DB matches (it defaults to countly_drill). Check the value with kubectl exec deploy/drill-migrator -- env | grep MONGO.
Log Warns "No apps found in countly.apps — hash map will be empty"
The connection you gave the migration has no reachable countly database, so it cannot resolve the SHA1 app and event hashes in old drill documents. Point MONGO_COUNTLY_DB at the right database on that same connection, and make sure it really is the one belonging to this dataset: a countly database from a different deployment produces hashes that never match.
Preflight Fails: "Source frozen & clocks sane"
The newest source cd is within 60 seconds of ClickHouse server time. Either old ingestion is still running, or the two machines' clocks are skewed. Old ingestion usually survives because the cutover's "stop old ingestion" step has not actually taken effect: the ingress rule covers /i but not /i/bulk or the feedback paths.
Both causes invalidate the boundary between migrated and live data, so fix them before migrating. In bounded (mirror) mode this check is replaced by a bound-sanity check, because the source is expected to keep growing.
Preflight Fails: "Old ingestion stopped (source frozen)"
A four-second probe saw a collection's newest cd or estimated count advance. The detail line names which collections grew. The cause is the same as above.
Preflight Fails: Disk Headroom Below 10%
The most preventable mid-migration incident there is. Both sides need room: MongoDB keeps its data throughout, and ClickHouse needs space for the migrated rows plus the per-chunk staging tables. Expand the PVCs before starting rather than hitting the wall 80% through a multi-day run.
Backfill Is Very Slow
Almost always MongoDB, and almost always memory, not CPU. The backfill is one long paged scan over collections far larger than RAM, so throughput is governed by how much of the index and working set stays resident in the WiredTiger cache. Starve the cache and every page turn becomes a disk read; adding cores to that changes nothing.
Check in this order:
- The WiredTiger cache.
- The pod's memory limit. Keep it at least twice the cache: WiredTiger needs headroom beyond its cache, and a limit set too close to it gets the pod OOMKilled rather than making it faster.
- The
{cd, _id}indexes. - Only then CPU, which is rarely the constraint.
Confirm the diagnosis rather than guess at it. Compare what WiredTiger holds against what it is allowed to hold:
kubectl exec -n mongodb <mongodb-pod> -- mongosh --quiet --eval '
const c = db.serverStatus().wiredTiger.cache;
print("configured : " + (c["maximum bytes configured"]/1e9).toFixed(1) + " GB");
print("in cache : " + (c["bytes currently in the cache"]/1e9).toFixed(1) + " GB");
print("read in : " + (c["bytes read into cache"]/1e9).toFixed(1) + " GB cumulative");'A cache sitting at its configured ceiling while bytes read into cache climbs steadily is cache starvation: the working set does not fit, so pages are evicted and re-read continuously. Only more cache fixes that.
Scaling migration replicas will not help if MongoDB reads are already saturated, and neither will giving MongoDB more cores.
Also check where the pods landed with kubectl get pods -l app=drill-migrator -o wide. A single pod saturates roughly four cores on BSON decode, so several pods scheduled onto the same node mostly contend with each other. If the scheduler has packed them, spread them across nodes with anti-affinity.
No Rows Are Moving at All on a First Run
The service is almost certainly still building {cd, _id} indexes. It does that itself for any collection that lacks one and will not process a collection until its index is ready. On a 10 TB dataset that can take one to three days, which looks exactly like a stuck migration. Watch /api/index-progress; the pod logs report each build as it completes. Pre-building during the preparation phase avoids this entirely.
Migration Paused a Long Time on Backpressure
ClickHouse has too many active parts, usually a merge backlog. Confirm with curl -s localhost:8080/stats | jq .clickhouse.
Wait for merges, lower MONGO_PAGE_SIZE or LEDGER_INSERT_INFLIGHT, or raise BACKPRESSURE_PARTS_TO_THROW_INSERT. The service force-resumes on its own after BACKPRESSURE_MAX_PAUSE_EPISODE_MS (three minutes by default). Persistent backpressure means ClickHouse is genuinely under-resourced: raise its CPU and memory in the ClickHouseInstallation, and consider fewer migration replicas until it catches up.
The Engine Paused Itself: Circuit Breaker
Three separate breakers can pause the run, and each is a stop-and-decide moment rather than an error to retry blindly:
| Breaker | Trips when | Response |
|---|---|---|
| Per-chunk fail rate | More than 5% of a chunk's documents fail (LEDGER_BREAKER_PCT) |
The DLQ already names the error. Investigate, fix the transform rule, then POST /control/retry-failed
|
| Consecutive failures | Three chunks fail in a row (LEDGER_BREAKER_CONSECUTIVE) |
Same response: this one catches a systematic bug that a per-chunk threshold would miss |
| Mass DLQ | The DLQ passes one million documents (LEDGER_DLQ_PAUSE_THRESHOLD) |
Usually orphan documents without uid. Inspect a few samples, waive or replay, then resume |
Structured skips do not count toward the fail-rate breaker, so a chunk that was entirely skippable documents completes as done rather than tripping it.
A Chunk Keeps Failing, and the Pod Crashes Each Time
This is the poison-pill path, and it handles itself. After three crash-retries the chunk is split instead of retried; repeated splitting converges on a window of roughly a minute, which is quarantined as a tiny failed chunk while everything else migrates. Inspect the few source documents in that chunk's cd window, fix or remove them, then POST /control/retry-failed.
This is also why backoffLimit: 50 in k8s/job.yaml is deliberate: crash-redo is normal operation here, not failure.
The Invariant Monitor Flagged a Done Chunk
A background monitor spot-checks completed chunks against the live table every 15 minutes by default. If rows for a done chunk are lost or corrupted, it detects the count mismatch, pauses, and flags the chunk. POST /control/retry-failed purges that chunk's cd window and redoes it.
OOMKilled Pods
Lower MONGO_PAGE_SIZE or LEDGER_INSERT_INFLIGHT for smaller per-cycle memory, or raise resources.limits.memory. The image caps the Node heap at 4 GB, which is why the shipped manifest limits memory at 6 Gi. Setting the limit at or below 4 Gi guarantees kills under load.
An OOMKill is not a data problem: the chunk is redone from its staging table when the pod comes back.
A Pod Died, Was Evicted, or Was Drained
Nothing else notices, by design. In-flight chunks are redone from their staging tables, and the dead pod's lease expires so other pods reclaim them. That is why the shipped Deployment has no preStop hook and only a 30-second termination grace period: there is nothing to drain.
There is no lock to force-release, no stale-pod entry to remove, and no endpoint for either. Recovery is the pod coming back, or another pod taking the lease; nothing else is required of you.
The Ledger Is Lost or Looks Wrong
The ledger (mig_ranges) is rebuildable from the two databases: Help & Recovery > Rebuild ledger from data, or POST /control/rebuild-ledger, with progress at /api/rebuild. Recovery never trusts the ledger blindly in any case: every claim it makes is verified against actual row counts before anything irreversible happens.
The DLQ (mig_dlq_docs) is not rebuildable. Back it up before dropping MANIFEST_DB.
Two Pods Appear to Be Working the Same Data
They are not. Chunk claims are atomic and leased in MongoDB, and the exclusive cd bounds of a chunk mean two pods cannot overlap even transiently. What usually looks like this is one pod having reclaimed another's expired lease, with the old pod's last log lines still in view. Check /api/chunks for the current claim.
Final Check Says FAIL
Every red line names its own action. The common ones:
-
Failed chunks: run
POST /control/retry-failed, wait, and re-run the check. - Unresolved DLQ entries: replay or waive them; sign-off requires zero pending.
- Count mismatches against the source: investigate before doing anything else; do not decommission.
- Active chunk claims: the check refuses while any pod still holds one. Let the run finish, pause it, or scale to zero.
A quick check is capped at PASS WITH NOTES, and the note names what it did not re-prove. Run the deep tier once, as the gate before deleting the source, and run it while the old cluster is still up, because the source is the reference.
Duplicate Rows After a Mirrored Cutover
If a tee was running and the migration ran without LEDGER_CD_UPPER_BOUND, every event in the overlap window exists twice. This is recoverable, but only while the old cluster still exists, because its MongoDB is what separates the two copies. See Tee-Overlap Dedupe in Running the Migration. Run the dry run first; an empty dry run means there are no duplicates, and the correct response is to skip the step rather than widen the window.
The frequent root cause is a pod that came up without the bound in its ConfigMap, which is why the bounded · cd < … badge should be checked on every pod, not just the one you happened to port-forward to.
Dashboards Are Empty After Reclaiming Disk Space
The report path is falling back to MongoDB for data that now only lives in ClickHouse. Confirm the ClickHouse adapter registered (see Confirming Queries Are Routing to ClickHouse in Running the Migration), and confirm you dropped only drill_events* and not the whole countly_drill database.
A Specific Date Range Shows Nothing in Countly
Filter by ts, the event time, which is what the drill UI uses. cd is the server-side creation timestamp: migrated rows keep their historical value and live-ingested rows get post-cutover ones, which is how the tool tells them apart. It is not an event-time column.
If the range is genuinely empty, check whether those chunks are done yet. Chunks are claimed newest-first, so older ranges land last.
Unique User Counts Differ Slightly From the Old Deployment
Small discrepancies in unique-user figures (user profiles, drill unique counts, funnels, and revenue) are usually not a migration fault. v26.01 counts unique users in ClickHouse with an approximate function by default: uniqCombined64(20) rather than uniqExact. It is fast and memory-cheap over billions of rows and accurate to within a fraction of a percent, but it is an estimate, so it will not always match a v25.x figure to the last digit.
The switch is the clickhouse_use_approximate_uniq setting on the drill plugin, which defaults to true and can be overridden per app:
| Value | Function used | Trade-off |
|---|---|---|
true (default) |
uniqCombined64(20) |
Fast and bounded in memory at any data volume; counts are estimates |
false |
uniqExact |
Exact counts; memory grows with the number of distinct users in the range, and large queries become substantially slower or fail outright |
Before switching it off, confirm the difference really is estimation error rather than missing data: check that day-by-day event coverage is complete, and compare a small date range where the exact and approximate answers should be close. A gap in coverage shows up as counts that are consistently and substantially low, not as a fraction of a percent.
If you do set it to false, do it per app rather than globally, and only where exact figures genuinely matter. On a large dataset, uniqExact is the difference between a query that returns and one that exhausts the ClickHouse memory limit.
Retention TTL Keeps Deleting on the Old Side During a Validation Window
Expected, not a defect. In scenario 2 the old deployment keeps running its own retention policy throughout the parallel period, so the source shrinks under you. The deep source audit classifies this as deletion drift and spot-checks it rather than assuming it is benign; it does not report it as data loss.
Live ClickHouse Itself Has to Be Rebuilt
Not a disaster. Live events still sit in the Kafka log, and history still sits in the frozen MongoDB. Recreate the table, reset only the ClickHouse-sink connector's offsets to earliest (leave the aggregator consumer groups alone), and re-run the migrator. This is the reason you raised Kafka drill-events retention during preparation.