The end-to-end procedure for backfilling drill events into ClickHouse alongside a Docker Compose deployment: cutting over, running the backfill, verifying it, rebuilding the aggregates that depend on it, and reclaiming the MongoDB space afterwards.
This article assumes you have completed Migration Prerequisites and chosen a scenario from the introduction.
Cutting Over
The order here is what keeps the ingestion pause down to minutes and keeps the data consistent afterwards.
- Stop old ingestion. Configure the write endpoints on whichever deployment your SDKs currently target to return an error (see below). SDKs queue and retry.
-
Sync the stateful-set delta since the pre-copy: changed users by last-seen, and the aggregated data.
app_usersmust be complete before the next step, or new ingestion mints colliding uids. Aggregated data must land before new ingestion writes current-period documents. - Enable ingestion on the new stack and point SDKs at it. They deliver everything they queued.
The old MongoDB is now frozen. That is what makes everything after this point safe to redo.
Pausing Ingestion: Return an Error, Never 200
Edit the vhost file, not nginx.conf. The top-level nginx.conf is not mounted into the container. The nginx entrypoint assembles the running config from vhosts/, picking the file that matches your setup: countly.conf with your own certificate, countly-le.conf with Let's Encrypt, or the -blocked variant of either when API_TRAFFIC_ACCESS / LE_API_TRAFFIC_ACCESS is restricted.
There are five write locations, not one. nginx exact-match (=) locations inherit nothing from each other or from the ^~ /i/ prefix location, so each has to be short-circuited individually:
location = /i { return 503; }
location = /i/bulk { return 503; }
location = /i/feedback/input { return 503; }
location = /i/feedback/inputs { return 503; }
location ^~ /i/ { return 503; }Leaving /i/bulk or the feedback endpoints open means the deployment is still ingesting during the "pause", which removes the property the whole cutover rests on. The -blocked vhosts have no general ^~ /i/ location; they expose ^~ /i/campaign/click/ instead, so block that one in its place.
Leave the read locations (/o, /o/) alone so dashboards keep working, then apply with docker compose restart nginx.
Returning 200 while discarding the request body causes every SDK to treat the event as delivered and drop it from its queue permanently. The introduction explains why.
Scenario 2: Configuring the Mirror
In this topology the new deployment is primary and the old one is kept alive as the rollback net, so the mirror is configured on the new deployment (the one your SDKs now talk to), and every write it receives is also sent to the old one. Flip it inside a short ingestion pause (about 60 seconds): pause the old API, enable the mirror, point SDKs at the new deployment, and resume. That pause creates a sharp boundary: every old-cluster document whose cd predates the pause was written by the old deployment itself, and everything after it arrived through the mirror. Record any timestamp inside the pause window as the bound.
If a pause is not possible, the bound is the flip time plus the old architecture's worst drill-write latency, and you accept that the few seconds of mirrored traffic inside that margin will be double-counted once. Pick a quiet hour.
The same five locations need the mirror, for the same reason:
location = /i { mirror /_mirror_i; mirror_request_body on; proxy_pass http://countly_ingestor; }
location = /i/bulk { mirror /_mirror_i; mirror_request_body on; proxy_pass http://countly_ingestor; }
location = /i/feedback/input { mirror /_mirror_i; mirror_request_body on; proxy_pass http://countly_ingestor; }
location = /i/feedback/inputs { mirror /_mirror_i; mirror_request_body on; proxy_pass http://countly_ingestor; }
location ^~ /i/ { mirror /_mirror_i; mirror_request_body on; proxy_pass http://countly_api; }
location = /_mirror_i {
internal;
proxy_pass https://old-deployment.example.com$request_uri;
}Using $request_uri in the mirror target forwards each request to the same path on the old deployment, so one mirror location covers all five. Keep the rest of each location block as it already is (the proxy_set_header lines and the text/ping guard), then reload with docker compose restart nginx.
Read endpoints (/o, /o/) do not need mirroring.
The old deployment re-ingests each mirrored request independently, so the same event exists on both sides under a different _id and cd. Nothing downstream can deduplicate across that seam. The only protection is the bound: set it before the backfill runs.
Starting the Backfill
If you followed the prerequisites, the service is already up and holding at its start gate with LEDGER_START_PAUSED=true, having read nothing. Opening the gate is the whole step: press Start on the dashboard at http://localhost:8080, or:
curl -s -X POST localhost:8080/control/resume -H 'content-type: application/json' -d '{}'One Start covers the whole run: the gate lives in the ledger, so every instance begins, instances added later begin immediately, and a container that restarts afterwards stays started. Mapping runs now (the chunk grid is cut and any missing {cd, _id} index is built), then chunks start moving, newest first.
Starting from nothing instead. If the service is not up yet, remove LEDGER_START_PAUSED from .env (or set it to false) and bring it up. It begins immediately, with no gate to open:
docker compose up --build -d docker compose logs -f
Either way, the dashboard takes over from here: the Migration Guide tab walks the procedure for your scenario, and Help & Recovery covers every failure case with the fix one click away.
Large backfills run for days. Because the service runs as a detached container with restart: unless-stopped, an SSH disconnect does not interrupt it: no tmux or nohup needed.
Answering the Startup Guard
A fresh run that finds live data already arriving in its target ClickHouse with no bound set will hold before mapping anything, showing pauseReason: boundary-unset. This is deliberate: that situation is either a duplicate factory or a legitimate choice, and the service cannot tell which from the data.
-
Scenario 2 (a mirror is running): set the bound, and the run releases itself.
# detect, and apply automatically when the seam is an exact ingestion-pause gap curl -s -X POST localhost:8080/control/set-boundary -H 'content-type: application/json' -d '{}' curl -s localhost:8080/api/boundary | jq # read the outcome; the receipt lands in .applied # no exact gap? review the report, then accept the anchor explicitly curl -s -X POST localhost:8080/control/set-boundary -H 'content-type: application/json' -d '{"acceptAnchor": true}' # or set a known timestamp directly (applies immediately, cluster-wide) curl -s -X POST localhost:8080/control/set-boundary -H 'content-type: application/json' -d '{"boundMs": 1789966140000}' -
Scenarios 1 and 3 (nothing mirrors traffic): click Proceed unbounded in the banner, run
curl -X POST localhost:8080/control/allow-unbounded, or deploy withLEDGER_UNBOUNDED_OK=1so the question never arises.
An exact gap applies unattended; an anchor (a quantified ambiguity rather than a clean seam) is never auto-applied without acceptAnchor. A plain Resume is deliberately ignored while the question is open.
Once a bound is set, check every pod's dashboard header for the bounded · cd < … badge. If it is missing on any pod, stop that pod: it is migrating past the boundary.
Monitoring Progress
The dashboard at http://localhost:8080 is the primary view: cluster-wide throughput, per-collection progress, the chunk map, the pod roster, DLQ state, and the sign-off gates, with the control actions as buttons.
For terminal-only operation, three things give you the same picture:
# 1. the dashboard as text, refreshed watch -n 5 'curl -s localhost:8080/status.txt' # 2. a structured progress line every minute, straight to stdout docker compose logs -f | grep 'progress heartbeat' # 3. the raw JSON curl -s localhost:8080/stats | jq # cluster rate, run times, backpressure, progress curl -s localhost:8080/report | jq # skips, coercions per key, DLQ summary curl -s localhost:8080/api/pods | jq # pod table curl -s localhost:8080/api/chunks | jq
Because the heartbeat goes to stdout, it flows into whatever log pipeline collects container output (such as Loki, ELK, or CloudWatch) with no extra plumbing.
A ClickHouse-side sanity count, reading credentials from your Countly .env:
set -a; . ./.env; set +a
docker exec countly-clickhouse clickhouse-client \
--user "${CLICKHOUSE_USER:-countly}" --password "$CLICKHOUSE_PASSWORD" \
--query "SELECT formatReadableQuantity(count()) FROM countly_drill.drill_events"Because chunks are claimed newest-first, the last 30 days typically appear within hours even on a multi-day backfill. Progress is reported against a frozen denominator, so "to go" counts down rather than drifting.
Scaling Out
Scaling is adding instances with the same .env. There is no coordinator to configure and no Redis: pods claim chunks through leases in MongoDB, and POD_ID defaults to the container hostname, which is already unique.
docker run -d --env-file .env --name drill-migrator-2 \ -p 8081:8080 countly/countly-migration:<tag>
Scale across machines, not on one host. A single container saturates roughly four cores on BSON decode, so extra containers on the same box mostly contend with each other. Running the same command on another machine is the whole procedure: each container's dashboard shows the entire run.
Chunks for all collections are mapped upfront and claimed globally, in collection order and newest-data-first within each, so pods spill into the next collection the moment the current one has nothing claimable. A dataset made of many small collections parallelizes just as well as one large one.
If a pod dies, its lease expires and another pod reclaims the chunk, drops the staging table, and redoes it. No manual cleanup exists in this flow.
Controlling the Run
curl -s -X POST localhost:8080/control/pause -H 'content-type: application/json' -d '{}'
curl -s -X POST localhost:8080/control/resume -H 'content-type: application/json' -d '{}'
curl -s -X POST localhost:8080/control/retry-failed -H 'content-type: application/json' -d '{}'
curl -s -X POST localhost:8080/control/replay-dlq -H 'content-type: application/json' -d '{}'
curl -s -X POST localhost:8080/control/waive-dlq -H 'content-type: application/json' -d '{}'Retry, replay, and waive only queue work while the engine is paused: finish with /control/resume. Any pod answers; the state is shared.
retry-failed purges the failed chunks' cd windows and redoes them, then resumes. It is the answer to almost every "this chunk is wrong" situation, including a chunk the invariant monitor has flagged for a count mismatch.
Handling the Dead-Letter Queue
Documents that cannot be converted or inserted are isolated by bisection and stored in mig_dlq_docs with their full raw source, so nothing is silently dropped. The run continues around them. List them with curl -s localhost:8080/api/dlq | jq.
Each entry is resolved one of two ways, and sign-off requires zero pending:
-
Fix and replay: correct the transform rule (platform-first, syncing the goldens) or fix the stored raw document, then
POST /control/replay-dlqand watch/api/replay. Documents that keep failing stay pending with an updated error. -
Waive:
POST /control/waive-dlqexplicitly accepts that they will not be migrated. The raw documents are retained as the record, the waive is counted, and the source audit attributes each window's shortfall to its waived documents, so sign-off stays exact. Pass{"ids":[...]}to waive a subset instead of all pending.
skip:missing_uid: Orphan Documents of Deleted Users
Old deployments accumulate drill documents with no uid (often no ts or did either, with cd near the epoch). These are orphans of app users that were deleted in the old system: the user record and its uid link are gone, the event document remained. They cannot be attributed to any user in either system.
The standard call is to waive them. Consider a sentinel-uid replay instead only if the affected volume is large enough to distort historical event totals for an app and the documents carry usable ts/did. Check a few samples in the DLQ panel first.
Chunks that were entirely such documents complete as done, because structured skips do not trip the fail-rate breaker. The mass-DLQ pause that fires when millions accumulate is the built-in stop-and-decide moment; after waiving, resume.
Verifying the Migration
Running the Final Check
Do not interpret audit buckets by hand. The Final check runs everything (chunk states, DLQ, verification of the target against the ledger, cd-checksum fingerprints, and sampled content comparison), applies the tee and cutover rules itself, and answers the only question that matters: is it safe to decommission the old cluster? The verdict is PASS, PASS WITH NOTES, or FAIL, in plain sentences, with the action named on every red line.
There are two tiers:
| Tier | Runtime | What it does | When |
|---|---|---|---|
| Quick (default) | Minutes | Chunk states, DLQ, every migrated window's live count against the recorded count with duplicate attribution, and random content samples against the source. Catches everything that can happen after reading. Capped at PASS WITH NOTES, and the note names what it did not re-prove | Routinely, while the source still exists |
| Deep (opt-in) | Hours on large runs | Additionally recounts every window against the source with cd-checksum fingerprints and sampled identity coverage. The only check that derives everything from the two databases alone, with no reliance on the tool's own records |
Once, as the gate before the source is deleted |
# quick
curl -s -X POST localhost:8080/control/final-check -H 'content-type: application/json' -d '{}'
# deep — before deleting the source
curl -s -X POST localhost:8080/control/final-check -H 'content-type: application/json' -d '{"deep": true}'
# read the verdict; re-run until it says PASS or FAIL (it shows progress while running)
curl -s localhost:8080/final-check.txtOn the dashboard this is the Final check card; tick deep source recount for the pre-teardown gate.
A stored or environment cd bound is picked up automatically as the cutover time. On a mirrored run with no stored bound, pass the cutover explicitly: {"cutoverMs": <epoch ms of the tee flip>}, or type it into the dashboard field. Post-cutover source windows are then excluded and explained in a note: divergence there is the mirror still feeding the old side, not data loss.
Run it while the old cluster is still up. The source is the reference.
Deep is not distrust of the ledger: chunk reads are already recounted against the source at migration time. It is the only layer that can see two specific things:
| Failure class | Caught by |
|---|---|
| Under-read at read time (source count ≠ read tally) | The migration itself, per chunk |
| Rows lost or duplicated in ClickHouse after attach | Quick: ledger vs. target verification |
| Skipped documents | DLQ accounting, both tiers; unresolved means FAIL |
| Wrong content in migrated rows | Quick: random content samples against the source |
Documents written into already-done windows later (imports, restores, backdated cd) |
Deep only: the ledger is blind to them by design |
Count-preserving identity swaps (same count, same cd-sum, different documents) |
Deep only: checksum plus sampled id coverage |
| Routine retention deleting source documents | Deep classifies it exactly, spot-checked rather than assumed benign |
The individual checks are also available on their own: /control/verify, /control/audit-source, and /control/audit-content, each a background task you POST to start and GET to poll at /api/verify, /api/audit-source, and /api/audit-content. The Final check is the composite, and is what sign-off should use.
Spot-Checking the Totals Yourself
SELECT count() AS total, uniqExact(_id) AS distinct_ids FROM countly_drill.drill_events;
For date coverage, filter on ts: the event time, which is what the drill UI filters on and part of the table's sort key. Always use full ISO timestamps with a trailing Z; a bare string literal is interpreted in the ClickHouse server's local timezone and can silently shift your window:
SELECT toDate(ts) AS day, count()
FROM countly_drill.drill_events
WHERE ts >= parseDateTime64BestEffort('2026-01-01T00:00:00.000Z')
AND ts < parseDateTime64BestEffort('2026-02-01T00:00:00.000Z')
GROUP BY day ORDER BY day;Every day that had traffic in the source should have a non-zero row.
What cd means now. Migrated rows keep their historical cd: the service emits it explicitly rather than letting the column's DEFAULT now64(3) apply, falling back to ts only for documents that predate cd entirely. That is deliberate, and it is what makes provenance decidable: a migrated row carries a pre-cutover cd, a live-ingested row carries a post-cutover one. The migration chunks, verifies, and audits by cd for exactly this reason. So cd is meaningful for telling migrated from live data, and ts is what you filter on for event-time coverage.
Tee-Overlap Dedupe: Fixing a Missing Bound After the Fact
Only relevant if a mirrored cutover ran without LEDGER_CD_UPPER_BOUND. In that case the migration copied the mirror's re-ingested documents on top of natively ingested rows, and every event in the overlap window (tee flip to migration completion) exists twice in ClickHouse.
The copies are separable: the migrated copy's _id exists in the old cluster's MongoDB, the native one's does not. This must run before the old cluster is decommissioned, because the old MongoDB is the separator.
An id match alone is not proof of duplication. If the tee, or the new side's ingestion, dropped a request, the migrated row is the only copy of that event. So every hour bucket needs count-evidence that native counterparts exist before anything in it is deleted. Buckets that fall short are skipped and reported as unsafe: review those hours for a tee outage or a wrong start time rather than forcing them. Collections without their own scope have no usable evidence and are always reported unsafe, never deleted.
# 1. DRY RUN (counts only). fromMs = the tee flip; toMs = migration completion
curl -s -X POST localhost:8080/control/dedupe-overlap -H 'content-type: application/json' \
-d '{"fromMs": 1789700000000, "toMs": 1789794970435}'
curl -s localhost:8080/api/dedupe-overlap | jq '.totals.chMatched' # the duplicates
# 2. EXECUTE — refused unless a dry run over the SAME window completed first
curl -s -X POST localhost:8080/control/dedupe-overlap -H 'content-type: application/json' \
-d '{"fromMs": 1789700000000, "toMs": 1789794970435, "execute": true}'
# 3. re-run the Final check with the same cutover to confirmThere is a dashboard card for this under Overview; Delete duplicates unlocks only after a dry run. The check is strict by default; slackPct (up to 5) may be passed consciously to absorb ingest-timing straddle at bucket edges.
An empty dry run means there are no duplicates. Skip the step: never widen the window to make it match something. Both the dedupe and the Final check refuse to run while any pod still holds an active chunk claim.
Rebuilding Aggregates and Flushing Caches
Copying events is not the whole job. Some dashboard data is precomputed, and those precomputed values need to be rebuilt from ClickHouse.
Confirming Queries Are Routing to ClickHouse
v26.01 routes drill-derived queries through a query adapter, in the order given by COUNTLY_CONFIG__DATABASE_ADAPTERPREFERENCE (["clickhouse","mongodb"] in the shipped Compose files). Confirm the ClickHouse adapter actually registered:
docker logs countly-api | grep -E "Database 'clickhouse'|Disabling ClickHouse adapter"
A healthy startup logs Database 'clickhouse' (clickhouse) registered with common object. If instead you see Disabling ClickHouse adapter in database config, the plugin could not reach ClickHouse and turned itself off. Queries then fall back to MongoDB, and reports will look empty once you reclaim MongoDB space at the end of this article.
Flushing the Active-Users Cache
The active_users collection caches daily, weekly, and monthly active users. These do refresh on their own (today's values every ten minutes, older days at least once a day), but a dashboard opened right after the backfill can show stale or empty numbers for up to a day. Clear the cache to force an immediate rebuild:
curl -G "https://<your-host>/i/active_users/clear_active_users_cache" \ --data-urlencode "api_key=<ADMIN_API_KEY>" \ --data-urlencode "app_id=<APP_ID>"
Repeat per app. To clear every app at once, empty the collection directly:
docker exec countly-mongodb mongosh countly --quiet --eval 'db.active_users.deleteMany({})'Regenerating View, Event, and Session Aggregates
Aggregate collections whose schema changed between versions (notably app_viewdata) cannot be carried over from the old deployment: they have to be rebuilt from the migrated ClickHouse data. Use the drill regeneration endpoint:
curl -X POST "https://<your-host>/i/drill/regeneration" \ --data-urlencode "api_key=<ADMIN_API_KEY>" \ --data-urlencode "app_id=<APP_ID>" \ --data-urlencode "method=views" \ --data-urlencode 'period=["1704067200000","1735689600000"]'
| Parameter | Notes |
|---|---|
method |
Required. One of views, events, or sessions
|
app_id |
Required |
event |
Required when method=events
|
period |
A named period (30days, month, day, yesterday, hour, or prevMonth), a [startMs, endMs] array, or {"since": ms}. Defaults to 30days
|
view_id |
Optional with method=views; omit to rebuild all views |
wait_to_finish |
Optional. Responds only after all writes commit |
Regeneration replaces the aggregates for the requested period, so it is safe to re-run: it does not double-count.
Omit wait_to_finish for long periods
wait_to_finish=true holds the HTTP connection open, and nginx or a load balancer will typically time out after about 60 seconds. For anything beyond a couple of days of data, omit it. The request then returns immediately and the work runs as a background task, visible under Manage > Tasks (type regeneration).
Repeat per app, and per method for any aggregate that looks wrong. Then open the dashboard and spot-check a few date ranges that predate the migration.
After Cutover: Confirming Coverage
Verify day-by-day coverage in ClickHouse by ts across the whole period, not just the range you expected the backfill to cover. Because ingestion is paused across the cutover, coverage should be continuous: every day that had traffic in the source should be present, with no step down around the switch. If a day is short or missing, the old deployment still holds those events for the length of your fallback window, so raise it with Countly support before you decommission it.
Check the MongoDB aggregates for the same window. A gap around cutover is not only a drill-events gap: the precomputed dashboard collections in the countly database were being written on the old deployment during that window, so they may be missing from what you copied. Regeneration rebuilds views, events, and sessions from ClickHouse, but not the other breakdowns (users, device_details, browser, sources, cities, and similar). Open the dashboard on a date inside that window and check those specific reports; if they are empty, those monthly summary documents have to be copied across from the old deployment.
GDPR erasures and app-user merges executed on the old system during a validation window apply only there. Re-apply them through the new architecture before sign-off.
Reclaiming Disk Space
Only after the Final check (ideally the deep tier) and the aggregate rebuild look correct.
Drop only the drill_events* collections. Do not drop the countly_drill database. That database still holds drill_meta* and drill_bookmarks, which v26.01 reads from MongoDB, and (unless you overrode MANIFEST_DB) the ledger and DLQ as well.
docker exec -i countly-mongodb mongosh countly_drill --quiet <<'EOF'
db.getCollectionNames().filter(n => n.startsWith("drill_events")).forEach(n => {
print("dropping " + n);
db[n].drop();
});
print("remaining: " + db.getCollectionNames().join(", "));
EOFThe remaining collection list should still contain your drill_meta* and drill_bookmarks collections, plus mig_ranges and mig_dlq_docs if the ledger lives here.
Then stop the migration service with docker compose down and revert the Kafka drill-events retention you raised during preparation.
Keep the original dump or snapshot somewhere safe until you have run at least one full reporting cycle on the new deployment. Back up mig_dlq_docs before dropping MANIFEST_DB: unlike the ledger, it cannot be rebuilt from the data.
Starting Over
Nothing up to the reclaim step is destructive: the MongoDB source data is untouched, and the migration only appends to ClickHouse.
To discard a backfill and start fresh:
# 1. Stop the migration
docker compose down
# 2. Empty the target table
docker exec countly-clickhouse clickhouse-client \
--user "${CLICKHOUSE_USER:-countly}" --password "$CLICKHOUSE_PASSWORD" \
--query "TRUNCATE TABLE countly_drill.drill_events"
# 3. Drop the migration state (back up mig_dlq_docs first if it holds anything)
docker exec countly-mongodb mongosh countly_drill --quiet --eval '
db.mig_ranges.drop(); db.mig_dlq_docs.drop();
db.mig_run_config.drop(); db.mig_collection_est.drop();'
# 4. Start again
docker compose up --build -dA lighter alternative to step 3 is to change LEDGER_RUN_ID to a new value, which starts a separate run and leaves the previous one's records in place.
TRUNCATE also removes live events
TRUNCATE also removes any events the live pipeline has ingested since the stack came up. Only use it while the new deployment is not yet receiving production traffic. Otherwise the recovery path is to reset the ClickHouse-sink connector's offsets and replay from Kafka, which is why you raised the retention.
If the ledger is lost or damaged but the data is fine, you do not need to start over: Help & Recovery > Rebuild ledger from data, or POST /control/rebuild-ledger, reconstructs it from the two databases.
Troubleshooting
Common failure modes and their fixes are collected in the troubleshooting article. The dashboard's Help & Recovery tab covers the same ground with the fix one click away.