Countly v26.01 stores raw events in ClickHouse instead of MongoDB. If you are moving an existing Countly deployment (v25.x or earlier) onto v26.01, your historical events have to be copied out of MongoDB and into ClickHouse before your dashboards will show them.
This article explains what the migration involves, how the migration service works, and how to choose a cutover scenario. Read it before you start: the scenario decision sets one configuration value that cannot be corrected silently afterwards, and it is the one choice that can leave permanently duplicated data if you get it wrong. When you have made that decision, continue with the guide for your deployment type.
Who These Guides Are For
Use this guide set when all of the following are true:
- You have an existing Countly deployment on v25.x or earlier, where events live in MongoDB (
countly_drill.drill_events*). - You are moving that data onto a Countly v26.01 stack, on either Docker Compose or Kubernetes.
- The migration can reach your existing
drill_events*collections, either in place on the old deployment or on a copy you bring into the new one. The prerequisites guide covers the options.
Read Architecture Changes for v26.01 first so the component names used throughout are familiar.
What Actually Gets Migrated
Only raw events move to a new database engine. Everything else stays in MongoDB and comes across with your operational databases.
| Data | Where it lived in v25.x | Where it lives in v26.01 | How it gets there |
|---|---|---|---|
| Raw events (drill) | MongoDB countly_drill.drill_events*
|
ClickHouse countly_drill.drill_events
|
The migration service |
| Drill metadata and bookmarks | MongoDB countly_drill.drill_meta*, drill_bookmarks
|
Unchanged: still MongoDB countly_drill
|
Moves with the operational databases (see the prerequisites guide) |
| Apps, app keys, app users, event definitions, dashboard users, and plugin config | MongoDB countly
|
Unchanged: still MongoDB countly
|
Bulk pre-copy before cutover, then a delta sync at cutover |
| Aggregated dashboard data (sessions, users, events, views, devices, …) | MongoDB countly
|
Unchanged: still MongoDB countly
|
Must land before new ingestion writes current-period documents |
| Live event ingestion | SDK → API → MongoDB | SDK → API → Kafka → ClickHouse | Nothing to migrate; starts working when the stack is up |
Never drop the whole countly_drill database
drill_meta and drill_bookmarks stay in MongoDB permanently in v26.01, and several code paths still read them directly from MongoDB. When you reclaim disk space, drop only the drill_events* collections: see Reclaiming Disk Space in the Running the Migration guide.
Order matters at cutover. app_users must be complete on the new stack before new ingestion starts, or ingestion mints colliding uids. Aggregated data has to land before new ingestion writes current-period documents, or the two overwrite each other.
How the Migration Service Works
The migration is a separate service you run alongside your deployment for the duration of the backfill. Its only dependencies are the source MongoDB and the target ClickHouse. Everything it needs to survive a crash lives in MongoDB.
The unit of work is a chunk: one cd range of one collection. Every collection is mapped into chunks upfront, and pods claim them globally, newest data first.
Nothing reaches the live table until its chunk has been counted and verified in staging, so a partially-copied chunk is never visible to the dashboard. What that buys you:
- Newest data first. The last 30 days are usually visible within hours; the full backfill then runs for days with no impact on live ingestion.
- Crash recovery without cleanup. A killed pod's chunks are redone from their staging tables, and its lease expires so other pods reclaim them. Abrupt kills, evictions, and OOMs are safe by design. Every incident response in this flow is restart or resume: never clean up or restore.
- Automatic isolation of bad documents. A document that cannot be inserted or converted is bisected out, stored in the dead-letter queue (DLQ) with its full raw source, and the run continues. Nothing is silently dropped.
- Poison-pill containment. A document that crashes the process is not retried forever: after three crash-retries the chunk is split, and repeated splitting converges on a window of roughly a minute that is quarantined as a tiny failed chunk while everything else migrates.
- Circuit breakers. The engine pauses itself when more than 5% of a chunk's documents fail, after three consecutive failed chunks, or when the DLQ passes one million documents. These are the built-in "stop and decide" moments.
- Backpressure. Inserts pause when ClickHouse is behind on merges, then resume.
- Multi-pod scaling with no coordinator. Pods claim chunks through leases in MongoDB. Any pod's dashboard shows the whole run.
The service also carries its own operations interface. Once it is running, http://localhost:8080 serves a dashboard whose Migration Guide tab walks the whole procedure and whose Help & Recovery tab covers every failure scenario with the fix one click away. Everything the dashboard does is plain HTTP, so an SSH-only operator can drive the same actions with curl.
Where the State Lives
| Store | Holds | If it is lost |
|---|---|---|
MongoDB mig_ranges (the ledger) |
Chunk state, cd bounds, counts, claims, and leases |
Rebuildable from the data: Help & Recovery > Rebuild ledger from data |
MongoDB mig_dlq_docs (the DLQ) |
Documents that could not be migrated, with their full raw source | Not rebuildable: back it up before dropping MANIFEST_DB
|
MongoDB mig_run_config, mig_collection_est
|
The stored cd bound, the start gate, and collection size estimates |
Re-derived or re-entered |
| ClickHouse staging tables | One per in-flight chunk, dropped on promotion | Redone automatically |
ClickHouse drill_events
|
The migrated data | Re-runnable while the source exists |
All of the MongoDB state lives in MANIFEST_DB, which defaults to countly_drill. Recovery never trusts the ledger blindly: every claim it makes is verified against actual row counts before anything irreversible happens.
Choosing Your Scenario
This is the decision that can leave permanently duplicated data, so make it before you deploy anything.
The ClickHouse drill_events table is a MergeTree-family table. It does not collapse duplicate rows. The only question that changes the configuration is whether a tee (a reverse proxy duplicating the same SDK requests into both stacks) is keeping the old deployment fed as a rollback net. A tee re-ingests each request independently, so the same event ends up in both systems under a different _id and cd. Nothing downstream can detect that. The only protection is a time bound on what the migration is allowed to copy.
That bound is LEDGER_CD_UPPER_BOUND: documents at or after it are never migrated.
| # | Topology | LEDGER_CD_UPPER_BOUND |
Ingestion switch | New data arriving in the old MongoDB | Sign-off |
|---|---|---|---|---|---|
| 1 | Two clusters, no mirroring (plain switch) | Unset | Before the migration, or after the bulk pass with a final drain | Migrated: top-up passes chase it until the drain finds nothing | Final check, DLQ resolved |
| 2 | Two clusters, mirror new → old (new primary, old is the rollback net) | Set to the moment new became primary | Already happened at the flip | Never migrated past the bound: it is the mirror's copy | Final check for the pre-bound region; dashboard comparison for post-bound |
| 3 | Single cluster, in-place upgrade | Unset | The upgrade itself is the switch; the old drill collections freeze | Transition tail drained by top-up; no tee, so nothing to duplicate | Final check, DLQ resolved |
The scenario is also selectable on the dashboard's Migration Guide tab, which renders the per-scenario checklist and states the bound requirement.
Rolling back differs by scenario rather than by a separate plan. In scenario 2 the old stack is still receiving the same requests through the mirror, so rollback is repointing SDK traffic at it. In scenarios 1 and 3 the old data is frozen and intact until you reclaim disk space, so rollback is repointing traffic and accepting that events collected on the new stack since the switch stay there.
The Bound Is Opt-In, and Getting It Wrong Is Caught
The boundary detector only suggests a value; it is never applied automatically in unbounded mode, where applying one would orphan new arrivals. To stop the obvious mistake, a startup guard holds any fresh run that finds live data already arriving in its target ClickHouse with no bound set. That is exactly the situation where an unset bound is either a duplicate factory (a mirror is running) or a deliberate choice (scenarios 1 and 3). The service cannot tell those apart from the data, so it asks once, before mapping anything, and holds with pauseReason: boundary-unset until you answer:
-
A mirror is active: set the bound (Detect boundary on the dashboard,
POST /control/set-boundary, orLEDGER_CD_UPPER_BOUND). The run releases itself. -
Nothing mirrors traffic: click Proceed unbounded,
POST /control/allow-unbounded, or deploy withLEDGER_UNBOUNDED_OK=1.
A plain Resume is deliberately ignored while that question is open. Resumed runs, and runs whose target holds no recent data, never trip the guard.
When the bound is set, every pod's dashboard header shows a bounded · cd < … badge. If that badge is missing on any pod, stop that pod: it is migrating past the boundary.
How Ingestion Is Paused, and for How Long
Ingestion pauses exactly once, for minutes, at cutover: never for the length of the migration. The backfill itself runs against a frozen (or bounded) source with live traffic flowing into the new stack the whole time.
Countly SDKs queue unsent requests locally rather than discarding them. Each request is written to a queue persisted on the device, so it survives application restarts, and the SDK retries it later. A request is removed from the queue only when the server returns a 2xx response; any other outcome, including a connection failure, leaves it queued.
So the response your endpoints return during the pause decides whether data is retained. Configure the write endpoints (/i and everything under /i/) to return an error. Returning 200 while discarding the request body causes every SDK to treat the event as delivered and drop it from its queue permanently.
The queue is bounded. Each SDK stores at most a fixed number of requests per device; past that limit, further requests are dropped and cannot be recovered. A pause is therefore safe only while every device stays below its own limit, which is roughly the queue size divided by that device's request rate. Low-traffic devices tolerate a much longer pause than high-volume ones. The specific limit differs between SDKs; see A Deeper Look at SDK Concepts.
The Shape of a Migration
Both deployment guides follow the same six phases. The prerequisites article covers phases one to three; the running article covers four to six.
-
Prepare: deploy the new stack alongside the old, set Kafka
drill-eventsretention to cover the migration window, and bulk pre-copy the stateful set (apps and app keys,app_users, event definitions, dashboard users, plugin config, and aggregated data). No user-facing impact. -
Index: start
{cd:1, _id:1}builds on alldrill_events*collections now. Roughly 1–3 days for 10 TB. This must not sit inside the post-cutover window. No collection consolidation is ever needed. - Rehearse: a dry run against a throwaway target, then review the report with whoever owns sign-off.
- Cutover: stop old ingestion, sync the stateful-set delta since the pre-copy, and enable ingestion on the new stack. SDK offline queues absorb the window. The old MongoDB is now frozen, which is what makes everything after this safe to redo.
- Migrate: start the service and scale it with pods. Watch the dashboard; the invariant monitor spot-checks continuously.
- Finish: all chunks done, Final check green, sign-off, revert Kafka retention, and decommission the old cluster.
Continuing With Your Deployment Type
The remaining articles come in a set of three for each deployment type: prerequisites, running the migration, and troubleshooting.