Docker Compose: Migration Prerequisites

Everything you need in place before starting the drill-events backfill alongside a Docker Compose deployment. Work through this article in order, then move on to running the migration.

If you have not read Migrating to Countly v26.01: Introduction yet, start there: it covers what the migration moves, how the service works, and how to choose a cutover scenario.

Before You Begin

  • The v26.01 stack is installed and healthy. Follow Single-Host Docker Compose first. make ps should show every service (healthy).
  • You have chosen a scenario. The three topologies in the introduction decide whether LEDGER_CD_UPPER_BOUND is set. Decide before you start the service, not after.
  • Your host has headroom. The backfill is read-heavy on MongoDB and write-heavy on ClickHouse at the same time. See Choose a Docker Compose Deployment Tier. If your host sits at the edge of a tier, pick the next one up for the duration of the migration.
  • You have disk space for both copies. During the migration the same events exist in MongoDB and ClickHouse. Budget the migrated data at roughly 10–20% of the MongoDB size after compression, plus transient headroom for the per-chunk staging tables, until you reclaim the MongoDB space at the end. Your own ratio will vary with event size and how much segmentation you send; the service's preflight measures actual free space on both sides and fails below 10%.
  • You can pull the migration image: countly/countly-migration, or a tag you build yourself from the migration repository.

Commands that act on the Countly stack run from the deploy/compose/ directory of your deployment. Commands that act on the migration service run from your checkout of the migration repository.

Preparing the New Stack

Bringing Up v26.01 and Confirming the ClickHouse Schema

The migration service writes into an existing ClickHouse table. It never creates that table. The clickhouse plugin creates the schema on API startup, so the stack has to come up at least once first.

make up
make verify

Confirm the target table exists:

docker exec countly-clickhouse clickhouse-client \
  --user "${CLICKHOUSE_USER:-countly}" --password "$CLICKHOUSE_PASSWORD" \
  --query "SHOW TABLES FROM countly_drill"

You should see drill_events in the output. If you do not, check docker logs countly-api for schema bootstrap errors before going any further. The service's own preflight checks this too, and fails if the table is missing.

Setting Kafka Retention to Cover the Migration Window

Live events flow through Kafka into ClickHouse. If the ClickHouse table ever has to be rebuilt during the migration, the recovery path is to reset the ClickHouse-sink connector's offsets and replay from the Kafka log, which only works while the log still holds the window. Raise drill-events retention to cover the whole migration (the default is 14 days) and revert it at sign-off.

Replication factor is a separate, deliberate choice. RF ≥ 2 is recommended for large instances; if you run RF=1, record the accepted risk, because one broker disk loss forfeits the replay guarantee.

Pre-Copying the Stateful Set

The migration service moves drill events and nothing else. Everything the dashboard needs to interpret them has to be brought across separately, and most of it can be copied well before cutover with no user-facing impact:

  • Apps and app keys
  • app_users
  • Event definitions
  • Dashboard users
  • Plugin configuration
  • Aggregated data

At cutover you sync only the delta since this pre-copy, which is what keeps the ingestion pause down to minutes. Two ordering constraints apply at that point, both covered in the running guide: app_users must be complete before new ingestion starts, and aggregated data must land before new ingestion writes current-period documents.

Making the Source Data Reachable

Before the backfill can run, the migration has to be able to read your old drill_events* collections. How you arrange that is the single biggest planning decision in this migration, because it depends on how much data you have.

Start by separating two different needs, because they have different answers:

What needs it Which data Access required Typical share of total size
The migration service countly_drill.drill_events* Read only, from anywhere it can reach over the network The bulk of your data
Countly v26.01 itself countly (apps, members, aggregates) and countly_drill's drill_meta* / drill_bookmarks Read/write, in this deployment's own countly-mongodb The smaller share

That split is what makes the large-data approaches possible: the events never have to be copied into this host at all, only read out of wherever they already live.

One more thing shapes every option below: the migration uses one MongoDB connection string for three different things.

Setting Default What the migration does with it
MONGO_DB countly_drill Reads the drill_events* collections: the actual source data
MONGO_COUNTLY_DB countly Reads apps and events to resolve per-event collection hashes during transform
MANIFEST_DB countly_drill Writes the ledger and DLQ (mig_ranges, mig_dlq_docs, mig_run_config, and mig_collection_est)

The second of these is easily overlooked, and it is not optional. Old drill documents identify their app and event with a 40-character SHA1 hash, and the migration reconstructs the real values by reading every app from countly.apps and every custom event name from countly.events, hashing eventName + appId, and building a lookup table at startup. If that database is missing or empty on the connection you give it, the lookup table is empty and hashed events cannot be resolved.

So whichever option you pick, the connection string you hand the migration must expose both countly_drill and a matching countly, and it must accept writes for MANIFEST_DB.

Option A is the right starting point for most deployments. Options B and C exist for cases where copying the event collections is impractical or unnecessary.

Option A: Cloning the Disk and Mounting It (Recommended)

The default choice. Only the drill events move to ClickHouse: v26.01 still runs on MongoDB for apps, members, aggregates, drill metadata, and bookmarks, so that data has to exist on this host regardless. A disk clone brings countly and countly_drill across together in one operation, which is exactly what both Countly and the migration need. This is a block-level copy, so it moves terabytes in the time your storage layer takes to clone a volume rather than the time mongorestore takes to re-insert every document and rebuild every index.

The Compose stack bind-mounts the MongoDB data directory from the host, so a cloned disk is simply a matter of mounting it and pointing .env at it:

  1. Snapshot the source MongoDB data volume with your cloud provider or storage layer. For a replica set, snapshot one secondary, after stopping writes to it or using a filesystem-consistent snapshot.
  2. Create a disk from the snapshot, attach it to the new host, and mount it. 

     See Data disk setup in the Docker Compose deployment guide for the fstab conventions and the mount-check guard.

  3. Point .env at the mount and make sure MongoDB can write to it:
MONGODB_DATA_DIR=/mongodb-data
MONGODB_REPLICA_SET=rs0   # must match the replica set name on the cloned volume
chown -R 999:999 /mongodb-data   # MongoDB runs as UID 999
docker compose up -d mongodb

Two issues commonly arise here:

  • Replica set identity travels with the volume. The local database on the cloned disk still contains the old replica set name and member hostnames. If MONGODB_REPLICA_SET does not match, the node will not come up as a usable member. Either match the old name, or start it standalone, drop the local database, and let the stack's replica-set initiator re-initiate.
  • MongoDB version compatibility. Data files can be opened by the same major version or upgraded one major at a time. Bring the clone up on the version that wrote it, confirm it is healthy, and only then upgrade.

If you clone inside the ingestion pause (pause old ingestion, clone, then resume ingestion on the new stack), the clone can never hold a mirror copy of a natively ingested event. That is the cleanest possible run: no bound, no duplicates, and nothing to deduplicate later. A clone taken after ingestion resumed has a duplicated tail and needs the dedupe pass described in the running guide.

Note that a frozen clone changes what the audits compare against: the tool sees the clone, not the live old cluster, so zero counts after the clone moment mean "clone taken here", not lost data. Scope any comparison against the live old deployment to cd below the clone moment.

Option B: Pointing the Migration at Your Existing MongoDB (No Copy)

For very large event volumes, in combination with Option A or C. Nothing is copied for the events themselves: the migration streams them straight out of the old cluster into ClickHouse. Read from a secondary so you do not add load to the primary that is still serving your live deployment:

MONGO_URI=mongodb://migrator:PASS@10.0.0.11:27017,10.0.0.12:27017/admin?replicaSet=rs0
MONGO_DB=countly_drill
MONGO_COUNTLY_DB=countly
MANIFEST_DB=countly_migration_manifest

Requirements and caveats:

  • Network reachability. The migration container must reach every member the replica set advertises, not just the seed hosts, because the driver connects using the hostnames in the replica set config. If those are internal names your new host cannot resolve, add extra_hosts entries to the migration service.
  • The migration needs write access to the old cluster. There is only one connection string, so the ledger and DLQ are written through the same URI, into MANIFEST_DB on the old cluster. Grant the migration user read on countly_drill and countly, and write on that one database. Override the default here. MANIFEST_DB defaults to countly_drill, which would write migration state into the old production drill database; a distinct name like countly_migration_manifest keeps it identifiable and safe to drop later.
  • Read preference is automatic. On a replica set the engine selects secondaryPreferred by itself. Because the source is frozen after cutover, secondary reads are exact. Only set MONGO_READ_PREFERENCE to override that deliberately: preflight warns if you have forced primary.
  • Read load on a live cluster. Reading a few thousand documents per second from a secondary is real load. Prefer a dedicated hidden or analytics secondary if your replica set has one.
  • This host still needs its own countly. Option B only solves reading the events. Countly v26.01 serves the dashboard from countly-mongodb, so countly (apps, members, aggregates) and countly_drill's drill_meta* / drill_bookmarks still have to exist there. Those are the smaller share of your data, so bring them across with Option A or Option C. That combination (operational databases cloned or restored, events read in place) is the usual reason to choose this option.

Option C: Dumping and Restoring

For small deployments, and for the operational databases alongside Option A or B. A dev or staging environment, a single small app, or a total dataset in the tens of gigabytes restores fine this way. Beyond that, mongorestore becomes impractical: it re-inserts every document and rebuilds every index, which takes far longer than the equivalent disk clone and needs scratch space for the archive on both ends.

# On the OLD server — dump both databases
mongodump --db countly       --out /backup/countly-dump
mongodump --db countly_drill --out /backup/countly-dump

# Copy /backup/countly-dump to the new host, then restore into the running container
docker cp /backup/countly-dump countly-mongodb:/tmp/dump
docker exec countly-mongodb mongorestore --drop /tmp/dump

MongoDB in the Compose stack runs without authentication on the internal network, so no credentials are needed for mongosh or mongorestore inside the container. See the mongorestore documentation for the details.

If you are pairing Option C with Option B, you only need the countly dump plus the drill_meta* and drill_bookmarks collections, not the drill_events* collections, which the migration reads in place.

Building the Read Index

The migration pages through each collection on a {cd: 1, _id: 1} compound index. It builds this index itself for any collection that lacks one, and it will not start processing a collection until that collection's index is ready. On a large deployment that means a fresh run can spend its first hours (or, at 10 TB, its first one to three days) building indexes with no rows moving, which looks exactly like a stuck migration.

Start these builds during the preparation phase, not after cutover. Build them in the background, throttled, on secondaries where possible:

docker exec -i countly-mongodb mongosh countly_drill --quiet <<'EOF'
db.getCollectionNames().filter(n => n.startsWith("drill_events")).forEach(n => {
  print("indexing " + n);
  db[n].createIndex({ cd: 1, _id: 1 });
});
EOF

The service can also start the builds for you once it is running (Build indexes on the dashboard, or POST /control/build-indexes, with progress at /api/index-progress), but starting early overlaps the wait with work you were doing anyway.

No collection consolidation is ever needed. The service discovers every collection matching the drill_events prefix and maps chunks across all of them.

Giving MongoDB and ClickHouse Enough CPU and Memory

MongoDB reads become the bottleneck long before ClickHouse writes do, and an under-provisioned MongoDB is the single most common cause of a slow backfill. The Compose fallback for MONGODB_CPUS is 1.0, but the tier files raise it (T32-128.env sets 7.0), so check what your deployment is actually running before assuming.

For the duration of the migration, raise the MongoDB allocation in .env above your tier. Values used on a 32 vCPU / 126 GB production host:

MONGODB_CPUS=8
MONGODB_MEM_LIMIT=70g
MONGODB_CACHE_GB=20

Scale those to your host, then apply them with docker compose up -d mongodb.

Use up -d, not restart

docker compose restart does not apply CPU or memory changes: only up -d recreates the container with the new limits. Verify what took effect:

docker inspect countly-mongodb --format '{{.HostConfig.Memory}} {{.HostConfig.NanoCpus}}'

ClickHouse needs headroom as well. It absorbs a sustained insert stream throughout the backfill, holds a staging table per in-flight chunk, and merges the parts it creates; if it is under-resourced it spends the migration in backpressure, pausing the inserts while it catches up. Raise its allocation in .env too:

CLICKHOUSE_CPUS=8
CLICKHOUSE_MEM_LIMIT=34g
CLICKHOUSE_MAX_SERVER_MEMORY_GB=26

Keep CLICKHOUSE_MAX_SERVER_MEMORY_GB below CLICKHOUSE_MEM_LIMIT: the first is the ceiling ClickHouse enforces on itself, the second is the hard Docker limit, and the gap absorbs allocations ClickHouse does not account for. Apply with docker compose up -d clickhouse, for the same reason as MongoDB.

Lower both back to your tier defaults after the migration completes.

Configuring the Migration Service

The migration service ships its own Compose file and connects to your existing MongoDB and ClickHouse over the network. It is not part of the Countly stack's Compose project.

git clone https://github.com/Countly/migration.git
cd migration
cp .env.example .env

Only two variables are required:

MONGO_URI=mongodb://mongodb:27017/?replicaSet=rs0
CLICKHOUSE_URL=http://clickhouse:8123

From a container, localhost on the host machine is host.docker.internal. If you run the migration on the Countly stack's own network, attach it to countly-network so the service names resolve.

The settings you are most likely to touch:

Variable Default What it controls
LEDGER_RUN_ID ledger-v1 The resume key. Keep it identical across every restart and every pod of the same migration; changing it starts a separate run
MANIFEST_DB countly_drill Where the ledger and DLQ live. Override when reading from a live old cluster (Option B)
LEDGER_CD_UPPER_BOUND unset Set only in scenario 2 (mirrored cutover). Epoch ms or ISO. Documents at or after it are never migrated
LEDGER_UNBOUNDED_OK false Declares up front that nothing mirrors traffic, so the startup guard does not hold the run
LEDGER_START_PAUSED false Deploy now, start later: the service comes up and serves the dashboard without reading, mapping, or indexing until Start is pressed once for the whole run. Set this: it is how you rehearse without starting a migration
POD_ID container hostname Unique per instance. The default is already unique per container
EXIT_ON_COMPLETE false true makes the container exit 0 when every chunk is terminal, for fire-and-forget runs
SERVICE_PORT 8080 Where the dashboard binds
LEDGER_CHUNK_DOCS_TARGET 2000000 Target documents per chunk
LEDGER_MAX_CHUNK_DAYS 7 Upper bound on a chunk's time span: guards against a bad estimatedDocumentCount producing one enormous chunk
MONGO_PAGE_SIZE 10000 Documents per read page
LEDGER_INSERT_INFLIGHT 3 Concurrent inserts into the staging table
DRY_RUN / DRY_RUN_SAMPLE_PCT false / 2 Sampled rehearsal against a throwaway target; nothing is stored

.env.example in the repository is the full commented reference. There is no RERUN_MODE in this service: resuming is what LEDGER_RUN_ID does, and starting over is covered in the running guide.

Starting the Service Without Starting the Run

Bring the container up with its start gate closed, so it serves the dashboard and answers its API without touching your data. Set LEDGER_START_PAUSED=true in .env, then start it:

docker compose up -d --build
docker compose logs -f

The log line confirming the gate is holding:

LEDGER_START_PAUSED: holding before mapping — press Start on the dashboard (POST /control/resume) to begin

The hold is before mapping, deliberately. Mapping is what builds the {cd, _id} index on the source and cuts the chunk grid, and a run nobody has started yet should be doing neither. In this state the service has read nothing, written nothing, and indexed nothing; it reports pauseReason: not-started and waits.

What it will do while held is serve the dashboard on http://localhost:8080 and answer three things: preflight, index builds, and the dry run. Those are exactly what you need before cutover, and all three are safe to run against a live old deployment.

The gate lives in the ledger rather than in the container, so a container that restarts while held stays held, and one Start later covers every instance of the run. Nothing begins until you press Start: that step belongs to the running guide.

Running the Built-In Preflight

With the service up and holding, run preflight from the dashboard's Migration Guide tab, or from the shell:

curl -s localhost:8080/api/preflight | jq '.checks[] | {label, status, detail}'

It is read-only, so run it as often as you like. It reports pass, warn, or fail for each of:

Check What a failure means
MongoDB source reachable Wrong URI, credentials, or network path; also reports how many drill_events* collections it found
{cd,_id} index on all collections Warn only: the service will build the missing ones, but you will wait
Estimated documents to migrate Informational; the estimate can be off after an unclean mongod shutdown, which is why chunk sizing is span-guarded
Source frozen & clocks sane Fail. The newest source cd is within 60 s of ClickHouse server time, so the migrated/live boundary is not trustworthy: old ingestion is still running, or the clocks are skewed
Old ingestion stopped Fail. A four-second probe saw a collection still growing. In bounded (mirror) mode this check is replaced by a bound-sanity check instead, because the source is expected to keep growing
New ingestion flowing into ClickHouse Warn only: either the SDK flip has not happened yet or traffic is genuinely zero
Documents without cd Warn: a dedicated sweep chunk migrates them, strictly after all regular chunks
Replica set detected Warn if you forced MONGO_READ_PREFERENCE=primary; remove it and let the engine pick secondaryPreferred
MongoDB / ClickHouse disk headroom Fail below 10% free, warn below 20%. The most preventable mid-migration incident there is
ClickHouse target table Fail. drill_events does not exist: start the new stack first
Insert-dedup canary Warn if dedup tokens are inert on this target. Safe either way, because chunk redo covers it
Dry run Warn until you have run one

Rehearsing With a Dry Run

A dry run maps a sampled subset of the work and takes it through the full transform and ClickHouse validation path against a Null-engine clone of the target, storing nothing. It is the last prerequisite, and it is safe against a live old deployment.

Trigger it on the service you already have up with the dashboard's Dry run button, or:

curl -s -X POST localhost:8080/control/dry-run -H 'content-type: application/json' -d '{}'
curl -s localhost:8080/api/dryrun | jq          # poll until status is "completed"

You do not need to set DRY_RUN in .env, and you do not need a second container. The endpoint opens its own source and target connections so it can never disturb the main run's state, and the rehearsal is recorded under its own run id (<LEDGER_RUN_ID>-dry) with its own start gate, so rehearsing can never silently authorize the real run.

Two conditions apply. The main run must not be actively copying (holding at the gate is fine, and so is paused), and only one dry run may be in flight. If either is violated the call returns {"started": false, "reason": "…"} rather than queueing.

To change the sample size, set DRY_RUN_SAMPLE_PCT in .env before starting the container. It accepts 0.1 to 5 and defaults to 2.

Then review the report (curl -s localhost:8080/report | jq) with whoever owns sign-off. It lists skip reasons, per-key coercions, and a DLQ summary. What you are looking for is anything systematic, such as a whole event type coercing or a skip reason with a large count, because that is far cheaper to fix now than mid-backfill.

Re-run preflight afterwards: its Dry run check turns green once the sampled chunks are done, which is the signal that the prerequisites are complete.

Next Steps

With the prerequisites met, continue to Running the Migration.

Was this page helpful?
Reach out to us for any other questions.
Helpful?

Looking For More Help?