Failure modes you are likely to meet while backfilling drill events alongside a Docker Compose deployment, and what to do about each one.
The migration dashboard's Help & Recovery tab covers the same ground with the fix one click away, and is usually faster than working from here. The guiding property of this flow is worth keeping in mind while you read: after cutover, no failure anywhere in it can touch live data. Every response below is restart, resume, or redo: never clean up or restore.
The Service Is Up but Idle: pauseReason: not-started
Expected, not a fault: LEDGER_START_PAUSED=true is set, which is what the prerequisites article tells you to do so you can run preflight and the dry run without starting a migration. The service is holding before mapping: nothing has been read, written, or indexed.
Press Start on the dashboard, or POST /control/resume, when you actually want the backfill to begin. The gate lives in the ledger, so one Start covers every instance, and a container that restarts afterwards stays started.
Do not confuse this with boundary-unset below. not-started means you have not started the run; boundary-unset means the run started and stopped to ask a question.
The Run Starts but Nothing Happens: pauseReason: boundary-unset
This is the startup guard, not a fault. A fresh run found live data already arriving in its target ClickHouse with no cd bound set, and is holding before mapping anything because that situation is either a duplicate factory or a deliberate choice.
If a mirror is running (scenario 2), set the bound with POST /control/set-boundary and the run releases itself. If nothing mirrors traffic (scenarios 1 and 3; see Choosing Your Scenario in the Introduction), click Proceed unbounded or POST /control/allow-unbounded. A plain Resume is deliberately ignored while the question is open: that is why pressing Resume appears to do nothing.
Container Exits Immediately With a Config Error
Only MONGO_URI and CLICKHOUSE_URL are required; everything else has a default. The service also rejects a few semantically impossible combinations at startup rather than misbehaving later:
-
CLICKHOUSE_RETRY_BASE_DELAY_MSgreater thanCLICKHOUSE_RETRY_MAX_DELAY_MS -
BACKPRESSURE_PARTITION_PCT_LOWnot belowBACKPRESSURE_PARTITION_PCT_HIGH, or the same for theTOTAL_PCTpair -
LEDGER_CD_UPPER_BOUNDthat is neither an ISO date nor epoch milliseconds
Migration Exits With ClickHouse Connection Errors
Check CLICKHOUSE_URL, CLICKHOUSE_USERNAME, and CLICKHOUSE_PASSWORD. Note the variable name: the Countly stack uses CLICKHOUSE_USER in its own .env, while the migration service expects CLICKHOUSE_USERNAME. Copying the Countly value under the Countly name silently leaves the migration on the default user.
Migration Exits With "No collections found matching prefix"
The source database has no drill_events* collections. Confirm the source data really is where you think it is, and that MONGO_DB matches (it defaults to countly_drill). Check what the container is actually pointed at with docker compose exec migration env | grep MONGO.
Log Warns "No apps found in countly.apps — hash map will be empty"
The connection you gave the migration has no reachable countly database, so it cannot resolve the SHA1 app/event hashes in old drill documents. Point MONGO_COUNTLY_DB at the right database on that same connection, and make sure it really is the one belonging to this dataset: a countly from a different deployment produces hashes that never match.
Image Pull Fails ("manifest unknown" / "unauthorized")
Pin a known-good tag, or build the image locally from the migration repository with docker compose up --build.
Preflight Fails: "Source frozen & clocks sane"
The newest source cd is within 60 seconds of ClickHouse server time. Either old ingestion is still running (the cutover's "stop old ingestion" step has not actually taken effect, often because only /i was blocked and not the other four write locations), or the two machines' clocks are skewed. Both invalidate the boundary between migrated and live data, so fix it before migrating. In bounded (mirror) mode this check is replaced by a bound-sanity check, because the source is expected to keep growing.
Preflight Fails: "Old ingestion stopped (source frozen)"
A four-second probe saw a collection's newest cd or estimated count advance. The detail line names which collections grew. Same cause as above.
Preflight Fails: Disk Headroom Below 10%
The most preventable mid-migration incident there is. Both sides need room: MongoDB keeps its data throughout, and ClickHouse needs space for the migrated rows plus the per-chunk staging tables. Add capacity before starting rather than hitting it at 80% through a multi-day run.
Backfill Is Very Slow
Almost always MongoDB, and almost always memory, not CPU. The backfill is one long paged scan over collections far larger than RAM, so what governs throughput is how much of the index and working set stays resident in the WiredTiger cache. Starve the cache and every page turn becomes a disk read; adding cores to that changes nothing.
Check in this order:
-
MONGODB_CACHE_GB: the WiredTiger cache. This is the lever that matters, and the first one to raise. -
MONGODB_MEM_LIMIT. Keep it at least twice the cache. WiredTiger needs real headroom beyond its cache for connections, sorts, and page reconstruction, and a limit set too close to the cache gets the container OOM-killed instead of running faster. -
The
{cd, _id}indexes from the prerequisites guide. Without them the scan cannot page efficiently at all, and no amount of memory compensates. -
MONGODB_CPUS. Worth having, but rarely the constraint. MongoDB pinned at roughly 100% of a single core indocker statsis the exception, not the rule: if you see it, raise this; if you do not, raising it will not help.
Apply any of these with docker compose up -d mongodb. docker compose restart does not apply resource changes, so verify what actually took effect:
docker inspect countly-mongodb --format '{{.HostConfig.Memory}} {{.HostConfig.NanoCpus}}'To confirm the diagnosis rather than guess at it, compare what WiredTiger is holding against what it is allowed to hold:
docker exec countly-mongodb mongosh --quiet --eval '
const c = db.serverStatus().wiredTiger.cache;
print("configured : " + (c["maximum bytes configured"]/1e9).toFixed(1) + " GB");
print("in cache : " + (c["bytes currently in the cache"]/1e9).toFixed(1) + " GB");
print("read in : " + (c["bytes read into cache"]/1e9).toFixed(1) + " GB cumulative");'A cache sitting at its configured ceiling while bytes read into cache climbs steadily is cache starvation: the working set does not fit, so pages are being evicted and re-read continuously. That is a memory problem, and only more cache fixes it.
Two things that will not help when MongoDB is the constraint: adding migration containers, and adding cores. A single container also saturates about four cores on BSON decode, so extra containers on the same host mostly contend with each other: scale across machines, and only once MongoDB has room.
No Rows Are Moving at All on a First Run
The service is almost certainly still building {cd, _id} indexes. It does that itself for any collection that lacks one and will not process a collection until its index is ready. On a 10 TB dataset that can take one to three days, which looks exactly like a stuck migration. Watch /api/index-progress; the logs report each build as it completes. Pre-building during the preparation phase avoids this entirely.
Migration Paused a Long Time on Backpressure
ClickHouse has too many active parts, usually a merge backlog. Check with curl -s localhost:8080/stats | jq .clickhouse.
Wait for merges, lower MONGO_PAGE_SIZE or LEDGER_INSERT_INFLIGHT, or raise BACKPRESSURE_PARTS_TO_THROW_INSERT. The service force-resumes on its own after BACKPRESSURE_MAX_PAUSE_EPISODE_MS (three minutes by default). Persistent backpressure means ClickHouse is genuinely under-resourced: raise its CPU and memory as the prerequisites guide describes.
The Engine Paused Itself: Circuit Breaker
Three separate breakers can pause the run, and each is a stop-and-decide moment rather than an error to retry blindly:
| Breaker | Trips when | Response |
|---|---|---|
| Per-chunk fail rate | More than 5% of a chunk's documents fail (LEDGER_BREAKER_PCT) |
The dead-letter queue (DLQ) already names the error. Investigate, fix the transform rule, then POST /control/retry-failed
|
| Consecutive failures | Three chunks fail in a row (LEDGER_BREAKER_CONSECUTIVE) |
Same: this one catches a systematic bug that a per-chunk threshold would miss |
| Mass DLQ | The DLQ passes one million documents (LEDGER_DLQ_PAUSE_THRESHOLD) |
Usually orphan documents without uid. Inspect a few samples, waive or replay, then resume |
Structured skips do not count toward the fail-rate breaker, so a chunk that was entirely skippable documents completes as done rather than tripping it.
A Chunk Keeps Failing, and the Process Crashes Each Time
This is the poison-pill path, and it handles itself. After three crash-retries the chunk is split instead of retried; repeated splitting converges on a window of roughly a minute, which is quarantined as a tiny failed chunk while everything else migrates. Inspect the few source documents in that chunk's cd window, fix or remove them, then POST /control/retry-failed.
The Invariant Monitor Flagged a Done Chunk
A background monitor spot-checks completed chunks against the live table every 15 minutes by default. If rows for a done chunk are lost or corrupted, it detects the count mismatch, pauses, and flags the chunk. POST /control/retry-failed purges that chunk's cd window and redoes it.
Migration Container OOM-Killed
Lower MONGO_PAGE_SIZE or LEDGER_INSERT_INFLIGHT for smaller per-cycle memory, or raise the container's memory limit. The image caps the Node heap at 4 GB, so give the container meaningfully more than that: around 6 GB is a reasonable limit.
MongoDB OOM-Killed During the Backfill
Raise MONGODB_MEM_LIMIT and keep it at least twice MONGODB_CACHE_GB. Apply with docker compose up -d mongodb: restart will not pick up the change.
A Pod Died and Its Chunks Look Stuck
Nothing else notices, by design. In-flight chunks are redone from their staging tables, and the dead pod's lease expires so other pods reclaim them. Restart the container. There is no lock to release, no dead-pod entry to remove, and no cleanup step: those belonged to the previous Redis-based engine.
The Ledger Is Lost or Looks Wrong
The ledger (mig_ranges) is rebuildable from the two databases: Help & Recovery > Rebuild ledger from data, or POST /control/rebuild-ledger, with progress at /api/rebuild. Recovery never trusts the ledger blindly in any case: every claim it makes is verified against actual row counts before anything irreversible happens.
The DLQ (mig_dlq_docs) is not rebuildable. Back it up before dropping MANIFEST_DB.
Final Check Says FAIL
Every red line names its own action. The common ones:
-
Failed chunks: run
POST /control/retry-failed, wait, and re-run the check. - Unresolved DLQ entries: replay or waive them; sign-off requires zero pending.
- Count mismatches against the source: investigate before doing anything else; do not decommission.
- Active chunk claims: the check refuses while any pod still holds one. Let the run finish or pause it.
A quick check is capped at PASS WITH NOTES and the note names what it did not re-prove. Run the deep tier once, as the gate before deleting the source, and run it while the old cluster is still up, because the source is the reference.
Duplicate Rows After a Mirrored Cutover
If a tee was running and the migration ran without LEDGER_CD_UPPER_BOUND, every event in the overlap window exists twice. This is recoverable, but only while the old cluster still exists, because its MongoDB is what separates the two copies. See Tee-Overlap Dedupe in Running the Migration. Run the dry run first; an empty dry run means there are no duplicates, and the correct response is to skip the step rather than widen the window.
Dashboards Are Empty After Reclaiming Disk Space
The report path is falling back to MongoDB for data that now only exists in ClickHouse. Confirm the ClickHouse adapter registered (see Confirming Queries Are Routing to ClickHouse in Running the Migration), and confirm you dropped only drill_events* and not the whole countly_drill database.
A Specific Date Range Shows Nothing in Countly
Filter by ts, the event time, which is what the drill UI uses. cd is the server-side creation timestamp; migrated rows keep their historical value and live-ingested rows get post-cutover ones, which is how the tool tells them apart: it is not an event-time column.
If the range is genuinely empty, check whether those chunks are done yet. Chunks are claimed newest-first, so older ranges land last.
Unique User Counts Differ Slightly From the Old Deployment
Small discrepancies in unique-user figures (user profiles, drill unique counts, funnels, and revenue) are usually not a migration fault. v26.01 counts unique users in ClickHouse with an approximate function by default: uniqCombined64(20) rather than uniqExact. It is fast and memory-cheap over billions of rows and accurate to within a fraction of a percent, but it is an estimate, so it will not always match a v25.x figure to the last digit.
The switch is the clickhouse_use_approximate_uniq setting on the drill plugin, which defaults to true and can be overridden per app:
| Value | Function used | Trade-off |
|---|---|---|
true (default) |
uniqCombined64(20) |
Fast and bounded in memory at any data volume; counts are estimates |
false |
uniqExact |
Exact counts; memory grows with the number of distinct users in the range, and large queries become substantially slower or fail outright |
Before switching it off, confirm the difference really is estimation error rather than missing data: check that day-by-day event coverage is complete, and compare a small date range where the exact and approximate answers should be close. A gap in coverage shows up as counts that are consistently and substantially low, not as a fraction of a percent.
If you do set it to false, do it per app rather than globally, and only where exact figures genuinely matter: on a large dataset uniqExact is the difference between a query that returns and one that exhausts the ClickHouse memory limit.
Live ClickHouse Itself Has to Be Rebuilt
Not a disaster. Live events still sit in the Kafka log, and history still sits in the frozen MongoDB. Recreate the table, reset only the ClickHouse-sink connector's offsets to earliest (leave the aggregator consumer groups alone), and re-run the migrator. This is the reason you raised Kafka drill-events retention during preparation.