Skill
mongodb
mongodb · current version v2
Troubleshoot and support a MongoDB server — connect, diagnose, and resolve problems across 5.0/6.0/7.0/8.0, on Docker, systemd, Atlas, or a Kubernetes operator. Covers mongosh/mongod arguments, replica sets, sharding, slow queries, memory, and recovery. Use when someone reports MongoDB is down, slow, unreachable, out of memory, out of disk, failing auth, lagging, or refusing to elect a primary.
13 downloads · published 2026-09-02
What this grants
- skill mongodb
Skill Card
- License or terms: check the skill's own repository for license details (mongodb on GitHub).
Security Audits
- NanoInfra Scanner PASS no issues found
- VirusTotal PASS no engines flagged this file (full report)
Version history
| Version | Published | Status |
|---|---|---|
| v2 | 2026-09-02 | published |
| v1 | 2026-09-02 | published |
Files
SKILL.md(14431 bytes)
SKILL.md
raw | preview
MongoDB — Troubleshooting & Support
A support runbook, not a tutorial. Work top to bottom: identify what you are actually on, read before you write, and treat every mutating step as one that needs a reason. Nothing here changes data or config until a section says so and says why.
Version dates and end-of-life move; confirm any lifecycle claim against https://www.mongodb.com/support-policy/lifecycles. Knowledge here is current to early 2026: the newest major is 8.0; there is no MongoDB 9.
0. Identify what you are actually on
Half of "MongoDB is broken" is "MongoDB is not what you think it is." Establish these five before diagnosing anything:
# Version + build (from a shell already connected)
db.version()
db.serverBuildInfo().version
# Or from the binary
mongod --version
# Topology: standalone, replica set, or sharded?
db.hello() # isWritablePrimary, setName present => replica set; msg:"isdbgrid" => mongos
rs.status() # errors on a standalone; describes members on a replica set
sh.status() # only meaningful through a mongos (sharded cluster)
# Storage engine + edition
db.serverStatus().storageEngine.name # expect "wiredTiger"
db.serverBuildInfo().modules # ["enterprise"] on Enterprise, [] on Community
Establish the host first with the detect-platform skill — OS/package family, init system,
bare metal vs VM vs container vs pod, the cgroup memory limit, and the security module. This
skill assumes you have that profile; it decides which of the environments in §2 you are in and
whether the bare-metal host checks apply. (A mongodb+srv:// connection string with no host
access is a managed cluster — §2.)
The four environments covered in §2: Docker container, bare-metal or VM under systemd, managed cluster, Kubernetes operator. Bare metal has an extra surface nothing else does — the host OS tuning MongoDB depends on — so it gets its own checklist there.
1. Versions at a glance (why the version matters for support)
| Major | GA | Support-relevant fact |
|---|---|---|
| 5.0 | 2021 | LTS, near/at end of life — treat as "upgrade path" territory |
| 6.0 | 2022 | The legacy mongo shell is removed. mongosh is the only shell |
| 7.0 | 2023 | LTS, supported |
| 8.0 | 2024 | LTS, current |
Two traps that produce confusing tickets:
mongovsmongosh. A runbook or script that callsmongofails on 6.0+ with "command not found". Usemongosheverywhere; the query language inside is the same.- featureCompatibilityVersion (FCV) lags the binary. A 7.0 server can still run with
FCV: "6.0", which disables 7.0 features and blocks a jump to 8.0. Check it early:db.adminCommand({ getParameter: 1, featureCompatibilityVersion: 1 })
2. Getting a shell / connecting, per deployment
Self-managed — Docker
docker ps --filter ancestor=mongo --format '{{.Names}}\t{{.Image}}\t{{.Status}}'
docker exec -it <container> mongosh -u <user> -p --authenticationDatabase admin
docker logs --tail 200 -f <container> # server log
docker exec <container> cat /etc/mongod.conf.orig 2>/dev/null || true
Self-managed — systemd (bare metal or VM)
systemctl status mongod
journalctl -u mongod -n 200 --no-pager
mongosh "mongodb://<user>:<pass>@127.0.0.1:27017/?authSource=admin"
# config + data + log paths come from the unit or the config file:
grep -E 'dbPath|systemLog|bindIp|port|replication|security' /etc/mongod.conf
systemctl cat mongod | grep -E 'LimitNOFILE|ExecStart|User' # limits/args the unit sets
Bare metal — the host checks nothing else needs
On bare metal (and, mostly, on a VM) MongoDB depends on OS settings that a container inherits
from its image and a managed service handles for you. mongod prints most of these as
STARTUP WARNINGS in its first log lines — read those first, they are the answer more often
than anything inside the database:
mongosh --quiet --eval 'db.adminCommand({getLog:"startupWarnings"}).log.forEach(l=>print(l))'
The usual culprits, each a known cause of latency spikes, stalls, or refused connections:
# Transparent Huge Pages — MUST be disabled for MongoDB; enabled THP causes latency stalls.
cat /sys/kernel/mm/transparent_hugepage/enabled # want [never] (or madvise); [always] is the bug
cat /sys/kernel/mm/transparent_hugepage/defrag
# Open-file / process limits — a busy server hits the default 1024 and refuses connections.
cat /proc/$(pgrep -x mongod)/limits | grep -E 'open files|processes' # want ~64000 files
ulimit -n
# NUMA — mongod should run interleaved; a single-node bind starves half of RAM.
numactl --hardware 2>/dev/null && echo "check the unit runs: numactl --interleave=all mongod …"
# Readahead on the data disk — WiredTiger wants it low (8–32 sectors), not the distro default.
blockdev --getra /dev/<data-disk>
# Swappiness — high swappiness swaps out the working set under cache pressure.
cat /proc/sys/vm/swappiness # 1 is the usual recommendation for a DB host
# Filesystem — XFS is recommended for WiredTiger; ext4 works, others invite trouble.
findmnt -no FSTYPE,OPTIONS <dbPath> # want xfs, mounted noatime
# Clock — replica-set elections and oplog ordering assume synced clocks.
timedatectl | grep -E 'synchronized|NTP'
Fixes are host config, not MongoDB config: THP via a tuned profile or a systemd drop-in, files
via LimitNOFILE in the unit, NUMA via the ExecStart line, readahead via udev/blockdev,
swappiness via sysctl. All of these are deliberate host changes — see §6.
Managed — MongoDB Atlas
No host shell and no mongod.conf — you cannot touch the process. You can still connect,
and the panel is where server-side diagnosis happens.
mongosh "mongodb+srv://<user>:<pass>@<cluster>.mongodb.net/?retryWrites=true"
For load, slow queries and connections use Atlas Metrics, Profiler, and Real-Time;
for restarts/EOL/upgrades it is the Atlas UI or atlas CLI, not this shell.
Kubernetes — operators (Percona PSMDB, MongoDB Community/Enterprise Operator)
kubectl get pods -l app.kubernetes.io/name=percona-server-mongodb # or the operator's label
kubectl exec -it <pod> -c mongod -- mongosh -u <user> -p --authenticationDatabase admin
kubectl logs <pod> -c mongod --tail 200 -f
# Credentials live in a Secret the operator manages, not in a file:
kubectl get secret <cluster>-secrets -o jsonpath='{.data.MONGODB_DATABASE_ADMIN_USER}' | base64 -d
Config is the CR (kubectl get psmdb <name> -o yaml), not a file on the pod — edit the CR,
let the operator reconcile; hand-editing inside the pod is reverted.
3. The tools and the arguments you will reach for
mongosh (the shell)
--host --port | --username/-u --password/-p | --authenticationDatabase <db>
--eval "<js>" run one statement and exit (scripting)
--quiet suppress the banner (clean output for parsing)
--tls --tlsCAFile --tlsCertificateKeyFile TLS/mTLS
--retryWrites=false connect string flag when a single node is not a replica set
mongod (the server — you rarely run this by hand; know the flags to read a unit)
--dbpath <dir> --port --bind_ip <ips> --replSet <name> --config /etc/mongod.conf
--logpath --logappend --wiredTigerCacheSizeGB <n>
--repair LAST RESORT, offline, can discard unrecoverable data — see §6
Backup / move data
mongodump --uri "mongodb://…" --archive=dump.gz --gzip [--oplog] [--nsInclude db.coll]
mongorestore --uri "mongodb://…" --archive=dump.gz --gzip [--drop] [--nsInclude db.coll]
mongoexport --uri "…" -d db -c coll --out coll.json # JSON/CSV, per collection
mongoimport --uri "…" -d db -c coll --file coll.json
Note: --drop on restore deletes the target collection first. --oplog gives a
point-in-time-consistent dump of a replica set.
Live counters (no shell state, safe to run)
mongostat --uri "…" 2 # ops/s, dirty %, used cache, conns — one line every 2s
mongotop --uri "…" 2 # read/write time per collection
4. Diagnostics — read-only, run these first
db.serverStatus() // connections, mem, wiredTiger cache, opcounters, asserts
db.serverStatus().connections // current / available / totalCreated
db.serverStatus().wiredTiger.cache // "bytes currently in the cache" vs "maximum bytes configured"
rs.status() // member states, health, lastHeartbeat
rs.printSecondaryReplicationInfo() // replication lag per secondary (seconds behind primary)
db.currentOp({ "secs_running": { $gt: 5 } }) // operations running longer than 5s
db.getSiblingDB("admin").aggregate([{ $currentOp: {} }]) // full picture incl. idle sessions
Slow queries:
db.setProfilingLevel(1, { slowms: 100 }) // log ops slower than 100ms
db.system.profile.find().sort({ ts: -1 }).limit(5)
db.<coll>.find({ … }).explain("executionStats") // look for stage:"COLLSCAN" and docsExamined ≫ nReturned
Turn the profiler back off when done: db.setProfilingLevel(0).
Logs by deployment: docker logs, journalctl -u mongod, kubectl logs … -c mongod, or
Atlas → Logs. MongoDB log lines are JSON since 4.4 — filter with jq 'select(.s=="E" or .s=="W")'.
5. Common problems → resolution
Cannot connect / authentication failed
Authentication failed: the user is almost always in a different DB than you named. Add--authenticationDatabase admin(or?authSource=adminin the URI) — the user is defined where it was created, not in the DB you are opening.connection refused: checkbindIpinmongod.conf(a server bound to127.0.0.1is unreachable from another host — bind the real interface and secure it, do not open0.0.0.0without auth), the port, and any firewall/security group.MongoServerSelectionErroragainst one node: it may be a replica set member that is not primary. Add?replicaSet=<name>ordirectConnection=truefor a deliberate single-node connection.- TLS errors: mismatched
--tls*files or an expired cert;openssl s_client -connect host:27017.
Replica set: no primary / stuck SECONDARY / lag
rs.status() // who is PRIMARY? any member in (RECOVERING, DOWN, STARTUP2)?
rs.printSecondaryReplicationInfo() // how many seconds behind is each secondary?
- No primary usually means no majority is reachable (network partition, an even member count,
or arbiter down). Restore quorum; do not force-reconfigure unless you have lost members
permanently, and then
rs.reconfig(cfg, {force:true})is the deliberate, data-risking step. - Growing lag: a secondary is I/O- or CPU-starved, or a long-running write on the primary is
serialising through the oplog. Check
mongostat, disk, anddb.currentOp()on the primary.
High memory / OOMKilled
- WiredTiger defaults its cache to ~50% of (RAM − 1 GB), which is invisible to a container
memory limit. A
mongodin a 4 GB container can try to size a cache for the host's RAM and get OOMKilled. Pin it:--wiredTigerCacheSizeGB(orstorage.wiredTiger.engineConfig.cacheSizeGB) to roughly half the container limit, and set the container limit above cache + connections. - Confirm the cap the process actually sees before blaming Mongo:
cat /sys/fs/cgroup/memory.max.
Slow queries
explain("executionStats")showingCOLLSCANanddocsExamined≫nReturned= a missing index. Add it, but on a replica set/production build it with care (background is default in 4.2+).- A query that was fast and turned slow: check for a plan-cache flip or data growth crossing a
threshold;
$indexStatsshows which indexes are actually used.
Disk full
- MongoDB does not release space to the OS on delete; it reuses it internally. If the OS disk is
full, freeing documents will not immediately free filesystem space.
db.<coll>.stats()showsstorageSizevssize.compact(per collection, blocks that collection, one node at a time on a replica set) reclaims it; on a replica set, resync-from-scratch of a secondary is the cleaner reclaim. Both are deliberate, disruptive operations.
Corruption / won't start
mongod --repairis offline, single-node, and can discard data it cannot recover. Take a filesystem copy ofdbPathfirst. On a replica set, never repair — remove the member and resync it from a healthy one, which is safe and usually faster.db.<coll>.validate({ full: true })inspects a collection without changing it.
Upgrades (the FCV trap)
- You cannot skip a major. 6.0 → 8.0 goes through 7.0. At each step: upgrade binaries, verify the
set is healthy, then raise FCV:
db.adminCommand({ setFeatureCompatibilityVersion: "7.0" }). Raising FCV before every node runs the new binary, or skipping it, is how an upgrade half-lands.
6. Before you run anything that writes
These change data, availability, or config, and on a nanoinfra deployment they resolve to a
mutate.remote capability — they ask for approval interactively and need a standing grant to run
unattended. Say what you are about to do and why, on the specific target, before you do it:
rs.reconfig(..., {force:true}),rs.stepDown()— change who serves writes.db.killOp(<opid>)— kill an in-flight operation (safe for a runaway read; a killed write rolls back).mongod --repair,compact,--dropon restore,db.dropDatabase(), index drops — destructive.setFeatureCompatibilityVersion— one-way in practice; downgrade needs a documented dance.
When in doubt, capture state (mongodump, a dbPath copy, rs.status() output) first. A read
that tells you what is wrong is always cheaper than a write that turns out to be the wrong fix.