The end-to-end procedure for backfilling drill events into ClickHouse on a Kubernetes deployment: cutting over, running the backfill, verifying it, rebuilding the aggregates that depend on it, and reclaiming the MongoDB space afterwards.
This article assumes you have completed Migration Prerequisites and chosen a scenario from the introduction.
Cutting Over
The order here is what keeps the ingestion pause down to minutes and keeps the data consistent afterwards.
- Stop old ingestion at whatever terminates SDK traffic (see below).
-
Sync the stateful-set delta since the pre-copy: changed users by last-seen, and the aggregated data.
app_usersmust be complete before the next step, or new ingestion mints colliding uids. Aggregated data must land before new ingestion writes current-period documents. - Enable ingestion on the new stack and point SDKs at it. They deliver everything they queued.
The old MongoDB is now frozen. That is what makes everything after this point safe to redo.
Pausing Ingestion: Return an Error, Never 200
On Kubernetes this belongs at whatever terminates SDK traffic (your ingress controller, load balancer, or CDN), applied to /i and everything under /i/. Leave the read endpoints (/o, /o/) untouched so dashboards keep working.
Make sure the rule really covers every write path, including /i/bulk, /i/feedback/input, and /i/feedback/inputs. Countly's own nginx routes those as exact-match locations, so a rule written only for /i leaves them ingesting throughout the "pause", which removes the property the whole cutover rests on.
Do not return 200. Returning 200 while discarding the request body causes every SDK to treat the event as delivered and drop it from its queue permanently. The introduction explains why.
Scenario 2: Configuring the Mirror
In this topology the new deployment is primary and the old one is kept alive as the rollback net, so the mirror is configured at whatever fronts the new deployment and every write it receives is also sent to the old one. Flip it inside a short ingestion pause (about 60 seconds): pause the old API, enable the mirror, point SDKs at the new deployment, and resume. That pause creates a sharp boundary: every old-cluster document whose cd predates the pause was written by the old deployment itself, and everything after it arrived through the mirror. Record any timestamp inside the pause window as the bound.
If a pause is not possible, the bound is the flip time plus the old architecture's worst drill-write latency, and you accept that the few seconds of mirrored traffic inside that margin will be double-counted once. Pick a quiet hour.
Preferred: a proxy you control. If SDK traffic reaches the deployment through an nginx, load balancer, or CDN that you configure directly, put the mirror there. It is ordinary nginx configuration, it is scoped to exactly the paths you choose, and it survives chart upgrades. Treat ingress-level mirroring as the fallback, not the first choice: the annotation route below has several ways to fail silently.
If you must do it at the F5 NGINX Ingress Controller. It has no dedicated mirror annotation, but it does support snippets, and the shipped values already use them. That is the trap: nginx.org/server-snippets and nginx.org/location-snippets in environments/reference/countly.yaml already carry required configuration (OpenTelemetry tracing, the X-Forwarded-* and X-Request-* headers, request buffering, upstream retry policy, and timeouts). Helm values do not merge multi-line strings, so setting either key to just your mirror directives silently discards everything already in it.
Append to the existing content instead of replacing it. Copy the current values verbatim and add the mirror lines at the end:
ingress:
annotations:
# Keep every line the reference values already set, then append the mirror.
nginx.org/server-snippets: |
otel_trace on;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_set_header X-Forwarded-Port $server_port;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Scheme $scheme;
proxy_set_header Connection "";
proxy_set_header X-Request-ID $request_id;
proxy_set_header X-Request-Start $msec;
client_header_timeout 30s;
# --- appended for the migration mirror ---
location = /_mirror_i {
internal;
proxy_pass https://old-deployment.example.com/i$is_args$args;
}
nginx.org/location-snippets: |
proxy_request_buffering on;
proxy_next_upstream error timeout http_502 http_503 http_504;
proxy_next_upstream_timeout 30s;
proxy_next_upstream_tries 3;
proxy_temp_file_write_size 1m;
client_body_timeout 120s;
# --- appended for the migration mirror ---
mirror /_mirror_i;
mirror_request_body on;Three further points:
- Snippets must be enabled on the controller, or the annotations are ignored silently.
-
location-snippetsapply to every path served by that Ingress, so scope the mirror to the write paths with a separate mergeable minion Ingress (nginx.org/mergeable-ingress-type: minion, the pattern the canary and Dex charts already use) rather than mirroring dashboard reads. - Re-check these annotations after any chart upgrade, since the reference values are the upstream source of the lines you copied.
The old deployment re-ingests each mirrored request independently, so the same event exists on both sides under a different _id and cd. Nothing downstream can deduplicate across that seam. The only protection is the bound: set it before the backfill runs.
Deploying the Migration
If you followed the prerequisites, the pods are already up and holding at their start gate with LEDGER_START_PAUSED: "true", having read nothing. Opening the gate is the whole step: press Start on the dashboard, or:
kubectl port-forward svc/drill-migrator 8080:8080
curl -s -X POST localhost:8080/control/resume -H 'content-type: application/json' -d '{}'One Start covers the whole fleet: the gate lives in the ledger, so every pod begins, pods that join later begin immediately, and a pod that restarts afterwards stays started. Mapping runs now (the chunk grid is cut and any missing {cd, _id} index is built), then chunks start moving, newest first.
Starting from nothing instead. If the pods are not up yet, remove LEDGER_START_PAUSED from the ConfigMap (or set it to "false") and apply. The run begins immediately, with no gate to open:
kubectl apply -f k8s/migration.yaml kubectl get pods -l app=drill-migrator kubectl logs -f deploy/drill-migrator
Either way, the dashboard takes over from here, and chunk state lives in MongoDB, so any pod's dashboard shows the whole run: the Migration Guide tab walks the procedure for your scenario, and Help & Recovery covers every failure case with the fix one click away.
Deployment or Job? The Deployment is recommended: when the run completes, pods stay up serving the dashboard for verification and sign-off, and you scale to zero when you are done. k8s/job.yaml uses EXIT_ON_COMPLETE for fire-and-forget automation, but the dashboard disappears with the pods when the Job finishes. It sets backoffLimit: 50 deliberately: crash-redo is normal operation here, not failure.
Answering the Startup Guard
A fresh run that finds live data already arriving in its target ClickHouse with no bound set will hold before mapping anything, showing pauseReason: boundary-unset. This is deliberate: that situation is either a duplicate factory or a legitimate choice, and the service cannot tell which from the data.
-
Scenario 2 (a mirror is running): set the bound, and the run releases itself.
# detect, and apply automatically when the seam is an exact ingestion-pause gap curl -s -X POST localhost:8080/control/set-boundary -H 'content-type: application/json' -d '{}' curl -s localhost:8080/api/boundary | jq # read the outcome; the receipt lands in .applied # no exact gap? review the report, then accept the anchor explicitly curl -s -X POST localhost:8080/control/set-boundary -H 'content-type: application/json' -d '{"acceptAnchor": true}' # or set a known timestamp directly (applies immediately, across all pods) curl -s -X POST localhost:8080/control/set-boundary -H 'content-type: application/json' -d '{"boundMs": 1789966140000}' -
Scenarios 1 and 3 (nothing mirrors traffic): click Proceed unbounded in the banner, run
curl -X POST localhost:8080/control/allow-unbounded(which releases every held pod), or deploy withLEDGER_UNBOUNDED_OK: "1"so the question never arises.
An exact gap applies unattended; an anchor (a quantified ambiguity rather than a clean seam) is never auto-applied without acceptAnchor. A plain Resume is deliberately ignored while the question is open.
Once a bound is set, check every pod's dashboard header for the bounded · cd < … badge. If it is missing on any pod, stop that pod: it is migrating past the boundary.
Monitoring Progress
The dashboard is the primary view: cluster-wide throughput, per-collection progress, the chunk map, the pod roster, DLQ state, and the sign-off gates, with the control actions as buttons.
For terminal-only operation, three things give you the same picture without a browser:
# 1. the dashboard as text, refreshed watch -n 5 'curl -s localhost:8080/status.txt' # 2. a structured progress line every minute, straight to stdout kubectl logs -f deploy/drill-migrator | grep 'progress heartbeat' # 3. the raw JSON curl -s localhost:8080/stats | jq # cluster rate, run times, backpressure, progress curl -s localhost:8080/report | jq # skips, coercions per key, DLQ summary curl -s localhost:8080/api/pods | jq curl -s localhost:8080/api/chunks | jq
Because the heartbeat goes to stdout, it flows into whatever log pipeline collects container output (such as Loki, ELK, or CloudWatch) with no extra plumbing.
A ClickHouse-side sanity count:
kubectl exec -n clickhouse <clickhouse-pod> -- \ clickhouse-client --password "$CLICKHOUSE_PASSWORD" \ --query "SELECT formatReadableQuantity(count()) FROM countly_drill.drill_events"
Because chunks are claimed newest-first, the last 30 days typically appear within hours even on a multi-day backfill. Progress is reported against a frozen denominator, so "to go" counts down rather than drifting.
Scaling Out
Scaling is raising the replica count. There is nothing to configure: pods coordinate through chunk leases in MongoDB, and POD_ID defaults to the hostname, which in Kubernetes is the unique pod name. Run kubectl scale deployment/drill-migrator --replicas=5, or change the replica count through your GitOps values, which is what you want in a managed cluster.
Scale across nodes, not on one machine. A single pod saturates roughly four cores on BSON decode, so more pods on the same node just contend. Beyond a handful of pods, ClickHouse backpressure or MongoDB read capacity becomes the limit rather than the pod count.
Chunks for all collections are mapped upfront and claimed globally, in collection order and newest-data-first within each, so pods spill into the next collection the moment the current one has nothing claimable. A dataset made of many small collections parallelizes just as well as one large one: there is no separate "range-parallel" mode to enable and no document threshold to tune.
Abrupt kills and evictions are safe by design. A killed pod's chunks are redone from their staging tables, and its lease expires so other pods reclaim them. That is why the shipped Deployment has no preStop hook and no long termination grace period: there is nothing to drain and no lock to release.
Controlling the Run
curl -s -X POST localhost:8080/control/pause -H 'content-type: application/json' -d '{}'
curl -s -X POST localhost:8080/control/resume -H 'content-type: application/json' -d '{}'
curl -s -X POST localhost:8080/control/retry-failed -H 'content-type: application/json' -d '{}'
curl -s -X POST localhost:8080/control/replay-dlq -H 'content-type: application/json' -d '{}'
curl -s -X POST localhost:8080/control/waive-dlq -H 'content-type: application/json' -d '{}'Retry, replay, and waive only queue work while the engine is paused: finish with /control/resume. Any pod answers; the state is shared, so there is no per-pod vs. global distinction to reason about.
retry-failed purges the failed chunks' cd windows and redoes them, then resumes. It is the answer to almost every "this chunk is wrong" situation, including a chunk the invariant monitor has flagged for a count mismatch.
Pause is the right lever when the cluster is under load from something else: it stops after the current chunk and keeps all state.
HTTP Endpoint Reference
| Method | Path | Description |
|---|---|---|
| GET |
/ · /viz
|
The dashboard |
| GET | /healthz |
Liveness and readiness (both probes use it) |
| GET | /status.txt |
Human-readable snapshot for the terminal |
| GET | /stats |
Throughput, cluster rate, run times, backpressure, and progress |
| GET | /report |
Skips, coercions per key, and DLQ summary |
| GET |
/api/pods · /api/chunks · /api/config
|
Pod roster, chunk table, and effective configuration |
| GET |
/api/preflight · /api/index-progress · /api/dryrun
|
Readiness checks, index builds, and rehearsal |
| POST |
/control/pause · /control/resume
|
Pause and resume the run |
| POST | /control/retry-failed |
Purge and redo failed chunks, then resume |
| POST |
/control/replay-dlq · /control/waive-dlq
|
Resolve DLQ entries; progress at /api/replay, state at /api/dlq
|
| POST |
/control/build-indexes · /control/dry-run
|
Start index builds or a sampled rehearsal |
| POST | /control/rebuild-ledger |
Reconstruct the ledger from the data; progress at /api/rebuild
|
| POST | /control/set-boundary |
Detect and apply the cd bound in one call; read at /api/boundary
|
| POST |
/control/detect-boundary · /control/apply-bound
|
The two halves separately |
| POST | /control/allow-unbounded |
Release the startup guard across every held pod |
| POST |
/control/verify · /control/audit-source · /control/audit-content
|
Individual sign-off checks; poll /api/verify, /api/audit-source, and /api/audit-content
|
| POST | /control/final-check |
The composite sign-off check; read /final-check.txt or /api/final-check
|
| POST | /control/dedupe-overlap |
Tee-overlap dedupe; results at /api/dedupe-overlap
|
Every endpoint is cluster-wide: state is shared through MongoDB, so a control call to any pod applies to the whole run. There are no per-pod variants and nothing to address a single pod with.
Handling the Dead-Letter Queue
Documents that cannot be converted or inserted are isolated by bisection and stored in mig_dlq_docs with their full raw source, so nothing is silently dropped. The run continues around them. List them with curl -s localhost:8080/api/dlq | jq.
Each entry is resolved one of two ways, and sign-off requires zero pending:
-
Fix and replay: correct the transform rule (platform-first, syncing the goldens) or fix the stored raw document, then
POST /control/replay-dlqand watch/api/replay. Documents that keep failing stay pending with an updated error. -
Waive:
POST /control/waive-dlqexplicitly accepts that they will not be migrated. The raw documents are retained as the record, the waive is counted, and the source audit attributes each window's shortfall to its waived documents, so sign-off stays exact. Pass{"ids":[...]}to waive a subset instead of all pending.
skip:missing_uid: Orphan Documents of Deleted Users
Old deployments accumulate drill documents with no uid (often no ts or did either, with cd near the epoch). These are orphans of app users that were deleted in the old system: the user record and its uid link are gone, the event document remained. They cannot be attributed to any user in either system.
The standard call is to waive them. Consider a sentinel-uid replay instead only if the affected volume is large enough to distort historical event totals for an app and the documents carry usable ts/did. Check a few samples in the DLQ panel first.
Chunks that were entirely such documents complete as done, because structured skips do not trip the fail-rate breaker. The mass-DLQ pause that fires when millions accumulate is the built-in stop-and-decide moment; after waiving, resume.
Verifying the Migration
Running the Final Check
Do not interpret audit buckets by hand. The Final check runs everything (chunk states, DLQ, verification of the target against the ledger, cd-checksum fingerprints, and sampled content comparison), applies the tee and cutover rules itself, and answers the only question that matters: is it safe to decommission the old cluster? The verdict is PASS, PASS WITH NOTES, or FAIL, in plain sentences, with the action named on every red line.
There are two tiers:
| Tier | Runtime | What it does | When |
|---|---|---|---|
| Quick (default) | Minutes | Chunk states, DLQ, every migrated window's live count against the recorded count with duplicate attribution, and random content samples against the source. Catches everything that can happen after reading. Capped at PASS WITH NOTES, and the note names what it did not re-prove | Routinely, while the source still exists |
| Deep (opt-in) | Hours on large runs | Additionally recounts every window against the source with cd-checksum fingerprints and sampled identity coverage. The only check that derives everything from the two databases alone, with no reliance on the tool's own records |
Once, as the gate before the source is deleted |
# quick
curl -s -X POST localhost:8080/control/final-check -H 'content-type: application/json' -d '{}'
# deep — before deleting the source
curl -s -X POST localhost:8080/control/final-check -H 'content-type: application/json' -d '{"deep": true}'
# read the verdict; re-run until it says PASS or FAIL (it shows progress while running)
curl -s localhost:8080/final-check.txtOn the dashboard this is the Final check card; tick deep source recount for the pre-teardown gate.
A stored or environment cd bound is picked up automatically as the cutover time. On a mirrored run with no stored bound, pass the cutover explicitly: {"cutoverMs": <epoch ms of the tee flip>}. Post-cutover source windows are then excluded and explained in a note: divergence there is the mirror still feeding the old side, not data loss.
Run it while the old cluster is still up. The source is the reference.
Deep is not distrust of the ledger: chunk reads are already recounted against the source at migration time. It is the only layer that can see two specific things:
| Failure class | Caught by |
|---|---|
| Under-read at read time (source count ≠ read tally) | The migration itself, per chunk |
| Rows lost or duplicated in ClickHouse after attach | Quick: ledger vs. target verification |
| Skipped documents | DLQ accounting, both tiers; unresolved means FAIL |
| Wrong content in migrated rows | Quick: random content samples against the source |
Documents written into already-done windows later (imports, restores, backdated cd) |
Deep only: the ledger is blind to them by design |
Count-preserving identity swaps (same count, same cd-sum, different documents) |
Deep only: checksum plus sampled id coverage |
| Routine retention deleting source documents | Deep classifies it exactly, spot-checked rather than assumed benign |
Spot-Checking the Totals Yourself
SELECT count() AS total, uniqExact(_id) AS distinct_ids FROM countly_drill.drill_events;
For date coverage, filter on ts: the event time, which is what the drill UI filters on and part of the table's sort key. Always use full ISO timestamps with a trailing Z; a bare string literal is parsed in the ClickHouse server's local timezone and can silently shift the window:
SELECT toDate(ts) AS day, count()
FROM countly_drill.drill_events
WHERE ts >= parseDateTime64BestEffort('2026-01-01T00:00:00.000Z')
AND ts < parseDateTime64BestEffort('2026-02-01T00:00:00.000Z')
GROUP BY day ORDER BY day;Every day that had traffic in the source should have a non-zero row.
What cd means now. Migrated rows keep their historical cd: the service emits it explicitly rather than letting the column's DEFAULT now64(3) apply, falling back to ts only for documents that predate cd entirely. That is deliberate, and it is what makes provenance decidable: a migrated row carries a pre-cutover cd, a live-ingested row carries a post-cutover one. The migration chunks, verifies, and audits by cd for exactly this reason. So cd is meaningful for telling migrated from live data, and ts is what you filter on for event-time coverage.
Tee-Overlap Dedupe: Fixing a Missing Bound After the Fact
Only relevant if a mirrored cutover ran without LEDGER_CD_UPPER_BOUND. In that case the migration copied the mirror's re-ingested documents on top of natively ingested rows, and every event in the overlap window (tee flip to migration completion) exists twice in ClickHouse.
The copies are separable: the migrated copy's _id exists in the old cluster's MongoDB, the native one's does not. This must run before the old cluster is decommissioned, because the old MongoDB is the separator.
An id match alone is not proof of duplication. If the tee, or the new side's ingestion, dropped a request, the migrated row is the only copy of that event. So every hour bucket needs count-evidence that native counterparts exist before anything in it is deleted. Buckets that fall short are skipped and reported as unsafe: review those hours for a tee outage or a wrong start time rather than forcing them. Collections without their own scope have no usable evidence and are always reported unsafe, never deleted.
# 1. DRY RUN (counts only). fromMs = the tee flip; toMs = migration completion
curl -s -X POST localhost:8080/control/dedupe-overlap -H 'content-type: application/json' \
-d '{"fromMs": 1789700000000, "toMs": 1789794970435}'
curl -s localhost:8080/api/dedupe-overlap | jq '.totals.chMatched' # the duplicates
# 2. EXECUTE — refused unless a dry run over the SAME window completed first
curl -s -X POST localhost:8080/control/dedupe-overlap -H 'content-type: application/json' \
-d '{"fromMs": 1789700000000, "toMs": 1789794970435, "execute": true}'
# 3. re-run the Final check with the same cutover to confirmThere is a dashboard card for this under Overview; Delete duplicates unlocks only after a dry run. The check is strict by default; slackPct (up to 5) may be passed consciously to absorb ingest-timing straddle at bucket edges.
An empty dry run means there are no duplicates. Skip the step: never widen the window to make it match something. Both the dedupe and the Final check refuse to run while any pod still holds an active chunk claim.
Rebuilding Aggregates and Flushing Caches
Copying events is not the whole job. Some dashboard data is precomputed, and those precomputed values need to be rebuilt from ClickHouse.
Confirming Queries Are Routing to ClickHouse
v26.01 routes drill-derived queries through a query adapter, in the order given by COUNTLY_CONFIG__DATABASE_ADAPTERPREFERENCE (["clickhouse","mongodb"] in the shipped chart values). Confirm the ClickHouse adapter registered:
kubectl logs -n countly deploy/countly-api | grep -E "Database 'clickhouse'|Disabling ClickHouse adapter"
A healthy startup logs Database 'clickhouse' (clickhouse) registered with common object. If instead you see Disabling ClickHouse adapter in database config, the plugin could not reach ClickHouse and turned itself off. Queries then fall back to MongoDB, and reports will look empty once you reclaim MongoDB space at the end of this article.
Flushing the Active-Users Cache
The active_users collection caches daily, weekly, and monthly active users. These do refresh on their own (today's values every ten minutes, older days at least once a day), but a dashboard opened right after the backfill can show stale or empty numbers for up to a day. Clear the cache to force an immediate rebuild:
curl -G "https://<your-host>/i/active_users/clear_active_users_cache" \ --data-urlencode "api_key=<ADMIN_API_KEY>" \ --data-urlencode "app_id=<APP_ID>"
Repeat per app. To clear every app at once, empty the collection directly:
kubectl exec -n mongodb <mongodb-pod> -- mongosh countly --quiet \
--eval 'db.active_users.deleteMany({})'Regenerating View, Event, and Session Aggregates
Aggregate collections whose schema changed between versions (notably app_viewdata) cannot be carried over from the old deployment: they have to be rebuilt from the migrated ClickHouse data. Use the drill regeneration endpoint:
curl -X POST "https://<your-host>/i/drill/regeneration" \ --data-urlencode "api_key=<ADMIN_API_KEY>" \ --data-urlencode "app_id=<APP_ID>" \ --data-urlencode "method=views" \ --data-urlencode 'period=["1704067200000","1735689600000"]'
| Parameter | Notes |
|---|---|
method |
Required. One of views, events, or sessions
|
app_id |
Required |
event |
Required when method=events
|
period |
A named period (30days, month, day, yesterday, hour, or prevMonth), a [startMs, endMs] array, or {"since": ms}. Defaults to 30days
|
view_id |
Optional with method=views; omit to rebuild all views |
wait_to_finish |
Optional. Responds only after all writes commit |
Regeneration replaces the aggregates for the requested period, so it is safe to re-run: it does not double-count.
Omit wait_to_finish for long periods
wait_to_finish=true holds the connection open, and ingress controllers typically time out after about 60 seconds. For more than a couple of days of data, omit it. The request then returns immediately and the work runs as a background task, visible under Manage > Tasks (type regeneration).
Repeat per app and per method, then spot-check dashboard date ranges that predate the migration.
After Cutover: Confirming Coverage
Verify day-by-day coverage in ClickHouse by ts across the whole period, not just the range you expected the backfill to cover. Because ingestion is paused across the cutover, coverage should be continuous: every day that had traffic in the source should be present, with no step down around the switch. If a day is short or missing, the old deployment still holds those events for the length of your fallback window, so raise it with Countly support before you decommission it.
Check the MongoDB aggregates for the same window. A gap around cutover is not only a drill-events gap: the precomputed dashboard collections in the countly database were being written on the old deployment during that window, so they may be missing from what you copied. Regeneration rebuilds views, events, and sessions from ClickHouse, but not the other breakdowns (users, device_details, browser, sources, cities, and similar). Open the dashboard on a date inside that window and check those specific reports; if they are empty, those monthly summary documents have to be copied across from the old deployment.
GDPR erasures and app-user merges executed on the old system during a validation window apply only there. Re-apply them through the new architecture before sign-off.
Reclaiming Disk Space
Only after the Final check (ideally the deep tier) and the aggregate rebuild look correct.
Drop only the drill_events* collections. Do not drop the countly_drill database. That database still holds drill_meta* and drill_bookmarks, which v26.01 reads from MongoDB, and (unless you overrode MANIFEST_DB) the ledger and DLQ as well.
kubectl exec -i -n mongodb <mongodb-pod> -- mongosh countly_drill --quiet <<'EOF'
db.getCollectionNames().filter(n => n.startsWith("drill_events")).forEach(n => {
print("dropping " + n);
db[n].drop();
});
print("remaining: " + db.getCollectionNames().join(", "));
EOFThe remaining collection list should still contain your drill_meta* and drill_bookmarks collections, plus mig_ranges and mig_dlq_docs if the ledger lives here.
Then remove the migration workload with kubectl delete -f k8s/migration.yaml and revert the Kafka drill-events retention you raised during preparation.
Keep the original snapshot or dump somewhere safe until you have run at least one full reporting cycle on the new deployment. Back up mig_dlq_docs before dropping MANIFEST_DB: unlike the ledger, it cannot be rebuilt from the data.
Starting Over
Nothing up to the reclaim step is destructive: the MongoDB source is untouched, and the migration only appends to ClickHouse.
To discard a backfill and start fresh:
# 1. Stop the workers kubectl scale deployment/drill-migrator --replicas=0 # 2. Empty the target table (add ON CLUSTER '<cluster>' if ClickHouse is replicated) kubectl exec -n clickhouse <clickhouse-pod> -- \ clickhouse-client --password "$CLICKHOUSE_PASSWORD" \ --query "TRUNCATE TABLE countly_drill.drill_events" # 3. Drop the migration state (back up mig_dlq_docs first if it holds anything) kubectl exec -n mongodb <mongodb-pod> -- mongosh countly_drill --quiet --eval ' db.mig_ranges.drop(); db.mig_dlq_docs.drop(); db.mig_run_config.drop(); db.mig_collection_est.drop();' # 4. Start again kubectl scale deployment/drill-migrator --replicas=1
A lighter alternative to step 3 is to change LEDGER_RUN_ID to a new value, which starts a separate run and leaves the previous one's records in place.
TRUNCATE also removes live events
TRUNCATE also removes events the live pipeline has ingested since the stack came up. Only use it while the new deployment is not yet receiving production traffic. Otherwise the recovery path is to reset the ClickHouse-sink connector's offsets and replay from Kafka, which is why you raised the retention.
If the ledger is lost or damaged but the data is fine, you do not need to start over: Help & Recovery > Rebuild ledger from data, or POST /control/rebuild-ledger, reconstructs it from the two databases.
In-Place Upgrades (Scenario 3)
When MongoDB stays where it is and only the drill store changes, the prepare and cutover phases collapse to a configuration flip: no stateful copy, and an easy rollback while the old drill collections still exist. What to watch instead:
- Resource contention. The migrator competes with production. Throttle it, read from a secondary, and build indexes off-peak.
- Peak disk. MongoDB keeps its data while ClickHouse and the staging tables grow beside it. Drop old per-event collections only after their chunks are done and signed off.
- Hard memory limits on the new components. An OOM here is a production incident, not a migration inconvenience. Backpressure protects production ClickHouse, but give it room.
Troubleshooting
Common failure modes and their fixes are collected in the troubleshooting article. The dashboard's Help & Recovery tab covers the same ground with the fix one click away.