Skill

mongodb

mongodb · current version v2

Download v2

Troubleshoot and support a MongoDB server — connect, diagnose, and resolve problems across 5.0/6.0/7.0/8.0, on Docker, systemd, Atlas, or a Kubernetes operator. Covers mongosh/mongod arguments, replica sets, sharding, slow queries, memory, and recovery. Use when someone reports MongoDB is down, slow, unreachable, out of memory, out of disk, failing auth, lagging, or refusing to elect a primary.

13 downloads · published 2026-09-02

What this grants

Skill Card

Security Audits

Version history

VersionPublishedStatus
v2 2026-09-02 published
v1 2026-09-02 published

Files

SKILL.md

raw | preview

MongoDB — Troubleshooting & Support

A support runbook, not a tutorial. Work top to bottom: identify what you are actually on, read before you write, and treat every mutating step as one that needs a reason. Nothing here changes data or config until a section says so and says why.

Version dates and end-of-life move; confirm any lifecycle claim against https://www.mongodb.com/support-policy/lifecycles. Knowledge here is current to early 2026: the newest major is 8.0; there is no MongoDB 9.

0. Identify what you are actually on

Half of "MongoDB is broken" is "MongoDB is not what you think it is." Establish these five before diagnosing anything:

# Version + build (from a shell already connected)
db.version()
db.serverBuildInfo().version

# Or from the binary
mongod --version

# Topology: standalone, replica set, or sharded?
db.hello()               # isWritablePrimary, setName present => replica set; msg:"isdbgrid" => mongos
rs.status()              # errors on a standalone; describes members on a replica set
sh.status()              # only meaningful through a mongos (sharded cluster)

# Storage engine + edition
db.serverStatus().storageEngine.name   # expect "wiredTiger"
db.serverBuildInfo().modules           # ["enterprise"] on Enterprise, [] on Community

Establish the host first with the detect-platform skill — OS/package family, init system, bare metal vs VM vs container vs pod, the cgroup memory limit, and the security module. This skill assumes you have that profile; it decides which of the environments in §2 you are in and whether the bare-metal host checks apply. (A mongodb+srv:// connection string with no host access is a managed cluster — §2.)

The four environments covered in §2: Docker container, bare-metal or VM under systemd, managed cluster, Kubernetes operator. Bare metal has an extra surface nothing else does — the host OS tuning MongoDB depends on — so it gets its own checklist there.

1. Versions at a glance (why the version matters for support)

| Major | GA | Support-relevant fact | |---|---|---| | 5.0 | 2021 | LTS, near/at end of life — treat as "upgrade path" territory | | 6.0 | 2022 | The legacy mongo shell is removed. mongosh is the only shell | | 7.0 | 2023 | LTS, supported | | 8.0 | 2024 | LTS, current |

Two traps that produce confusing tickets:

2. Getting a shell / connecting, per deployment

Self-managed — Docker

docker ps --filter ancestor=mongo --format '{{.Names}}\t{{.Image}}\t{{.Status}}'
docker exec -it <container> mongosh -u <user> -p --authenticationDatabase admin
docker logs --tail 200 -f <container>            # server log
docker exec <container> cat /etc/mongod.conf.orig 2>/dev/null || true

Self-managed — systemd (bare metal or VM)

systemctl status mongod
journalctl -u mongod -n 200 --no-pager
mongosh "mongodb://<user>:<pass>@127.0.0.1:27017/?authSource=admin"
# config + data + log paths come from the unit or the config file:
grep -E 'dbPath|systemLog|bindIp|port|replication|security' /etc/mongod.conf
systemctl cat mongod | grep -E 'LimitNOFILE|ExecStart|User'   # limits/args the unit sets

Bare metal — the host checks nothing else needs

On bare metal (and, mostly, on a VM) MongoDB depends on OS settings that a container inherits from its image and a managed service handles for you. mongod prints most of these as STARTUP WARNINGS in its first log lines — read those first, they are the answer more often than anything inside the database:

mongosh --quiet --eval 'db.adminCommand({getLog:"startupWarnings"}).log.forEach(l=>print(l))'

The usual culprits, each a known cause of latency spikes, stalls, or refused connections:

# Transparent Huge Pages — MUST be disabled for MongoDB; enabled THP causes latency stalls.
cat /sys/kernel/mm/transparent_hugepage/enabled   # want [never] (or madvise); [always] is the bug
cat /sys/kernel/mm/transparent_hugepage/defrag

# Open-file / process limits — a busy server hits the default 1024 and refuses connections.
cat /proc/$(pgrep -x mongod)/limits | grep -E 'open files|processes'   # want ~64000 files
ulimit -n

# NUMA — mongod should run interleaved; a single-node bind starves half of RAM.
numactl --hardware 2>/dev/null && echo "check the unit runs: numactl --interleave=all mongod …"

# Readahead on the data disk — WiredTiger wants it low (8–32 sectors), not the distro default.
blockdev --getra /dev/<data-disk>

# Swappiness — high swappiness swaps out the working set under cache pressure.
cat /proc/sys/vm/swappiness            # 1 is the usual recommendation for a DB host

# Filesystem — XFS is recommended for WiredTiger; ext4 works, others invite trouble.
findmnt -no FSTYPE,OPTIONS <dbPath>    # want xfs, mounted noatime

# Clock — replica-set elections and oplog ordering assume synced clocks.
timedatectl | grep -E 'synchronized|NTP'

Fixes are host config, not MongoDB config: THP via a tuned profile or a systemd drop-in, files via LimitNOFILE in the unit, NUMA via the ExecStart line, readahead via udev/blockdev, swappiness via sysctl. All of these are deliberate host changes — see §6.

Managed — MongoDB Atlas No host shell and no mongod.conf — you cannot touch the process. You can still connect, and the panel is where server-side diagnosis happens.

mongosh "mongodb+srv://<user>:<pass>@<cluster>.mongodb.net/?retryWrites=true"

For load, slow queries and connections use Atlas Metrics, Profiler, and Real-Time; for restarts/EOL/upgrades it is the Atlas UI or atlas CLI, not this shell.

Kubernetes — operators (Percona PSMDB, MongoDB Community/Enterprise Operator)

kubectl get pods -l app.kubernetes.io/name=percona-server-mongodb   # or the operator's label
kubectl exec -it <pod> -c mongod -- mongosh -u <user> -p --authenticationDatabase admin
kubectl logs <pod> -c mongod --tail 200 -f
# Credentials live in a Secret the operator manages, not in a file:
kubectl get secret <cluster>-secrets -o jsonpath='{.data.MONGODB_DATABASE_ADMIN_USER}' | base64 -d

Config is the CR (kubectl get psmdb <name> -o yaml), not a file on the pod — edit the CR, let the operator reconcile; hand-editing inside the pod is reverted.

3. The tools and the arguments you will reach for

mongosh (the shell)

--host --port  |  --username/-u --password/-p  |  --authenticationDatabase <db>
--eval "<js>"       run one statement and exit (scripting)
--quiet             suppress the banner (clean output for parsing)
--tls --tlsCAFile --tlsCertificateKeyFile        TLS/mTLS
--retryWrites=false connect string flag when a single node is not a replica set

mongod (the server — you rarely run this by hand; know the flags to read a unit)

--dbpath <dir>  --port  --bind_ip <ips>  --replSet <name>  --config /etc/mongod.conf
--logpath --logappend  --wiredTigerCacheSizeGB <n>
--repair            LAST RESORT, offline, can discard unrecoverable data — see §6

Backup / move data

mongodump  --uri "mongodb://…" --archive=dump.gz --gzip [--oplog] [--nsInclude db.coll]
mongorestore --uri "mongodb://…" --archive=dump.gz --gzip [--drop] [--nsInclude db.coll]
mongoexport --uri "…" -d db -c coll --out coll.json      # JSON/CSV, per collection
mongoimport --uri "…" -d db -c coll --file coll.json

Note: --drop on restore deletes the target collection first. --oplog gives a point-in-time-consistent dump of a replica set.

Live counters (no shell state, safe to run)

mongostat --uri "…" 2       # ops/s, dirty %, used cache, conns — one line every 2s
mongotop  --uri "…" 2       # read/write time per collection

4. Diagnostics — read-only, run these first

db.serverStatus()                    // connections, mem, wiredTiger cache, opcounters, asserts
db.serverStatus().connections        // current / available / totalCreated
db.serverStatus().wiredTiger.cache   // "bytes currently in the cache" vs "maximum bytes configured"
rs.status()                          // member states, health, lastHeartbeat
rs.printSecondaryReplicationInfo()   // replication lag per secondary (seconds behind primary)
db.currentOp({ "secs_running": { $gt: 5 } })   // operations running longer than 5s
db.getSiblingDB("admin").aggregate([{ $currentOp: {} }])  // full picture incl. idle sessions

Slow queries:

db.setProfilingLevel(1, { slowms: 100 })          // log ops slower than 100ms
db.system.profile.find().sort({ ts: -1 }).limit(5)
db.<coll>.find({ … }).explain("executionStats")   // look for stage:"COLLSCAN" and docsExamined ≫ nReturned

Turn the profiler back off when done: db.setProfilingLevel(0).

Logs by deployment: docker logs, journalctl -u mongod, kubectl logs … -c mongod, or Atlas → Logs. MongoDB log lines are JSON since 4.4 — filter with jq 'select(.s=="E" or .s=="W")'.

5. Common problems → resolution

Cannot connect / authentication failed

Replica set: no primary / stuck SECONDARY / lag

rs.status()                          // who is PRIMARY? any member in (RECOVERING, DOWN, STARTUP2)?
rs.printSecondaryReplicationInfo()   // how many seconds behind is each secondary?

High memory / OOMKilled

Slow queries

Disk full

Corruption / won't start

Upgrades (the FCV trap)

6. Before you run anything that writes

These change data, availability, or config, and on a nanoinfra deployment they resolve to a mutate.remote capability — they ask for approval interactively and need a standing grant to run unattended. Say what you are about to do and why, on the specific target, before you do it:

When in doubt, capture state (mongodump, a dbPath copy, rs.status() output) first. A read that tells you what is wrong is always cheaper than a write that turns out to be the wrong fix.