Skill
mongodb
mongodb · current version v2
Troubleshoot and support a MongoDB server — connect, diagnose, and resolve problems across 5.0/6.0/7.0/8.0, on Docker, systemd, Atlas, or a Kubernetes operator. Covers mongosh/mongod arguments, replica sets, sharding, slow queries, memory, and recovery. Use when someone reports MongoDB is down, slow, unreachable, out of memory, out of disk, failing auth, lagging, or refusing to elect a primary.
13 downloads · published 2026-09-02
What this grants
- skill mongodb
Skill Card
- License or terms: check the skill's own repository for license details (mongodb on GitHub).
Security Audits
- NanoInfra Scanner PASS no issues found
- VirusTotal PASS no engines flagged this file (full report)
Version history
| Version | Published | Status |
|---|---|---|
| v2 | 2026-09-02 | published |
| v1 | 2026-09-02 | published |
Files
SKILL.md(14431 bytes)
SKILL.md
raw | preview
---
name: mongodb
description: Troubleshoot and support a MongoDB server — connect, diagnose, and resolve problems across 5.0/6.0/7.0/8.0, on Docker, systemd, Atlas, or a Kubernetes operator. Covers mongosh/mongod arguments, replica sets, sharding, slow queries, memory, and recovery. Use when someone reports MongoDB is down, slow, unreachable, out of memory, out of disk, failing auth, lagging, or refusing to elect a primary.
---
# MongoDB — Troubleshooting & Support
A support runbook, not a tutorial. Work top to bottom: identify what you are actually
on, read before you write, and treat every mutating step as one that needs a reason.
Nothing here changes data or config until a section says so and says why.
Version dates and end-of-life move; confirm any lifecycle claim against
<https://www.mongodb.com/support-policy/lifecycles>. Knowledge here is current to early 2026:
the newest major is **8.0**; there is no MongoDB 9.
## 0. Identify what you are actually on
Half of "MongoDB is broken" is "MongoDB is not what you think it is." Establish these
five before diagnosing anything:
```bash
# Version + build (from a shell already connected)
db.version()
db.serverBuildInfo().version
# Or from the binary
mongod --version
# Topology: standalone, replica set, or sharded?
db.hello() # isWritablePrimary, setName present => replica set; msg:"isdbgrid" => mongos
rs.status() # errors on a standalone; describes members on a replica set
sh.status() # only meaningful through a mongos (sharded cluster)
# Storage engine + edition
db.serverStatus().storageEngine.name # expect "wiredTiger"
db.serverBuildInfo().modules # ["enterprise"] on Enterprise, [] on Community
```
Establish the host first with the **detect-platform** skill — OS/package family, init system,
bare metal vs VM vs container vs pod, the cgroup memory limit, and the security module. This
skill assumes you have that profile; it decides which of the environments in §2 you are in and
whether the bare-metal host checks apply. (A `mongodb+srv://` connection string with no host
access is a managed cluster — §2.)
The four environments covered in §2: **Docker container**, **bare-metal or VM under systemd**,
**managed cluster**, **Kubernetes operator**. Bare metal has an extra surface nothing else does
— the host OS tuning MongoDB depends on — so it gets its own checklist there.
## 1. Versions at a glance (why the version matters for support)
| Major | GA | Support-relevant fact |
|---|---|---|
| 5.0 | 2021 | LTS, near/at end of life — treat as "upgrade path" territory |
| 6.0 | 2022 | **The legacy `mongo` shell is removed.** `mongosh` is the only shell |
| 7.0 | 2023 | LTS, supported |
| 8.0 | 2024 | LTS, current |
Two traps that produce confusing tickets:
- **`mongo` vs `mongosh`.** A runbook or script that calls `mongo` fails on 6.0+ with
"command not found". Use `mongosh` everywhere; the query language inside is the same.
- **featureCompatibilityVersion (FCV)** lags the binary. A 7.0 server can still run with
`FCV: "6.0"`, which disables 7.0 features and blocks a jump to 8.0. Check it early:
```bash
db.adminCommand({ getParameter: 1, featureCompatibilityVersion: 1 })
```
## 2. Getting a shell / connecting, per deployment
**Self-managed — Docker**
```bash
docker ps --filter ancestor=mongo --format '{{.Names}}\t{{.Image}}\t{{.Status}}'
docker exec -it <container> mongosh -u <user> -p --authenticationDatabase admin
docker logs --tail 200 -f <container> # server log
docker exec <container> cat /etc/mongod.conf.orig 2>/dev/null || true
```
**Self-managed — systemd (bare metal or VM)**
```bash
systemctl status mongod
journalctl -u mongod -n 200 --no-pager
mongosh "mongodb://<user>:<pass>@127.0.0.1:27017/?authSource=admin"
# config + data + log paths come from the unit or the config file:
grep -E 'dbPath|systemLog|bindIp|port|replication|security' /etc/mongod.conf
systemctl cat mongod | grep -E 'LimitNOFILE|ExecStart|User' # limits/args the unit sets
```
**Bare metal — the host checks nothing else needs**
On bare metal (and, mostly, on a VM) MongoDB depends on OS settings that a container inherits
from its image and a managed service handles for you. `mongod` prints most of these as
**STARTUP WARNINGS** in its first log lines — read those first, they are the answer more often
than anything inside the database:
```bash
mongosh --quiet --eval 'db.adminCommand({getLog:"startupWarnings"}).log.forEach(l=>print(l))'
```
The usual culprits, each a known cause of latency spikes, stalls, or refused connections:
```bash
# Transparent Huge Pages — MUST be disabled for MongoDB; enabled THP causes latency stalls.
cat /sys/kernel/mm/transparent_hugepage/enabled # want [never] (or madvise); [always] is the bug
cat /sys/kernel/mm/transparent_hugepage/defrag
# Open-file / process limits — a busy server hits the default 1024 and refuses connections.
cat /proc/$(pgrep -x mongod)/limits | grep -E 'open files|processes' # want ~64000 files
ulimit -n
# NUMA — mongod should run interleaved; a single-node bind starves half of RAM.
numactl --hardware 2>/dev/null && echo "check the unit runs: numactl --interleave=all mongod …"
# Readahead on the data disk — WiredTiger wants it low (8–32 sectors), not the distro default.
blockdev --getra /dev/<data-disk>
# Swappiness — high swappiness swaps out the working set under cache pressure.
cat /proc/sys/vm/swappiness # 1 is the usual recommendation for a DB host
# Filesystem — XFS is recommended for WiredTiger; ext4 works, others invite trouble.
findmnt -no FSTYPE,OPTIONS <dbPath> # want xfs, mounted noatime
# Clock — replica-set elections and oplog ordering assume synced clocks.
timedatectl | grep -E 'synchronized|NTP'
```
Fixes are host config, not MongoDB config: THP via a `tuned` profile or a systemd drop-in, files
via `LimitNOFILE` in the unit, NUMA via the `ExecStart` line, readahead via `udev`/`blockdev`,
swappiness via `sysctl`. All of these are deliberate host changes — see §6.
**Managed — MongoDB Atlas**
No host shell and no `mongod.conf` — you cannot touch the process. You can still connect,
and the panel is where server-side diagnosis happens.
```bash
mongosh "mongodb+srv://<user>:<pass>@<cluster>.mongodb.net/?retryWrites=true"
```
For load, slow queries and connections use Atlas **Metrics**, **Profiler**, and **Real-Time**;
for restarts/EOL/upgrades it is the Atlas UI or `atlas` CLI, not this shell.
**Kubernetes — operators (Percona PSMDB, MongoDB Community/Enterprise Operator)**
```bash
kubectl get pods -l app.kubernetes.io/name=percona-server-mongodb # or the operator's label
kubectl exec -it <pod> -c mongod -- mongosh -u <user> -p --authenticationDatabase admin
kubectl logs <pod> -c mongod --tail 200 -f
# Credentials live in a Secret the operator manages, not in a file:
kubectl get secret <cluster>-secrets -o jsonpath='{.data.MONGODB_DATABASE_ADMIN_USER}' | base64 -d
```
Config is the CR (`kubectl get psmdb <name> -o yaml`), not a file on the pod — edit the CR,
let the operator reconcile; hand-editing inside the pod is reverted.
## 3. The tools and the arguments you will reach for
**`mongosh`** (the shell)
```
--host --port | --username/-u --password/-p | --authenticationDatabase <db>
--eval "<js>" run one statement and exit (scripting)
--quiet suppress the banner (clean output for parsing)
--tls --tlsCAFile --tlsCertificateKeyFile TLS/mTLS
--retryWrites=false connect string flag when a single node is not a replica set
```
**`mongod`** (the server — you rarely run this by hand; know the flags to read a unit)
```
--dbpath <dir> --port --bind_ip <ips> --replSet <name> --config /etc/mongod.conf
--logpath --logappend --wiredTigerCacheSizeGB <n>
--repair LAST RESORT, offline, can discard unrecoverable data — see §6
```
**Backup / move data**
```bash
mongodump --uri "mongodb://…" --archive=dump.gz --gzip [--oplog] [--nsInclude db.coll]
mongorestore --uri "mongodb://…" --archive=dump.gz --gzip [--drop] [--nsInclude db.coll]
mongoexport --uri "…" -d db -c coll --out coll.json # JSON/CSV, per collection
mongoimport --uri "…" -d db -c coll --file coll.json
```
Note: `--drop` on restore deletes the target collection first. `--oplog` gives a
point-in-time-consistent dump of a replica set.
**Live counters (no shell state, safe to run)**
```bash
mongostat --uri "…" 2 # ops/s, dirty %, used cache, conns — one line every 2s
mongotop --uri "…" 2 # read/write time per collection
```
## 4. Diagnostics — read-only, run these first
```javascript
db.serverStatus() // connections, mem, wiredTiger cache, opcounters, asserts
db.serverStatus().connections // current / available / totalCreated
db.serverStatus().wiredTiger.cache // "bytes currently in the cache" vs "maximum bytes configured"
rs.status() // member states, health, lastHeartbeat
rs.printSecondaryReplicationInfo() // replication lag per secondary (seconds behind primary)
db.currentOp({ "secs_running": { $gt: 5 } }) // operations running longer than 5s
db.getSiblingDB("admin").aggregate([{ $currentOp: {} }]) // full picture incl. idle sessions
```
Slow queries:
```javascript
db.setProfilingLevel(1, { slowms: 100 }) // log ops slower than 100ms
db.system.profile.find().sort({ ts: -1 }).limit(5)
db.<coll>.find({ … }).explain("executionStats") // look for stage:"COLLSCAN" and docsExamined ≫ nReturned
```
Turn the profiler back off when done: `db.setProfilingLevel(0)`.
Logs by deployment: `docker logs`, `journalctl -u mongod`, `kubectl logs … -c mongod`, or
Atlas → Logs. MongoDB log lines are JSON since 4.4 — filter with `jq 'select(.s=="E" or .s=="W")'`.
## 5. Common problems → resolution
**Cannot connect / authentication failed**
- `Authentication failed`: the user is almost always in a different DB than you named.
Add `--authenticationDatabase admin` (or `?authSource=admin` in the URI) — the user is
defined where it was created, not in the DB you are opening.
- `connection refused`: check `bindIp` in `mongod.conf` (a server bound to `127.0.0.1` is
unreachable from another host — bind the real interface and secure it, do not open `0.0.0.0`
without auth), the port, and any firewall/security group.
- `MongoServerSelectionError` against one node: it may be a replica set member that is not
primary. Add `?replicaSet=<name>` or `directConnection=true` for a deliberate single-node
connection.
- TLS errors: mismatched `--tls*` files or an expired cert; `openssl s_client -connect host:27017`.
**Replica set: no primary / stuck SECONDARY / lag**
```javascript
rs.status() // who is PRIMARY? any member in (RECOVERING, DOWN, STARTUP2)?
rs.printSecondaryReplicationInfo() // how many seconds behind is each secondary?
```
- No primary usually means no majority is reachable (network partition, an even member count,
or arbiter down). Restore quorum; do **not** force-reconfigure unless you have lost members
permanently, and then `rs.reconfig(cfg, {force:true})` is the deliberate, data-risking step.
- Growing lag: a secondary is I/O- or CPU-starved, or a long-running write on the primary is
serialising through the oplog. Check `mongostat`, disk, and `db.currentOp()` on the primary.
**High memory / OOMKilled**
- WiredTiger defaults its cache to ~50% of (RAM − 1 GB), which is invisible to a **container
memory limit**. A `mongod` in a 4 GB container can try to size a cache for the host's RAM and
get OOMKilled. Pin it: `--wiredTigerCacheSizeGB` (or `storage.wiredTiger.engineConfig.cacheSizeGB`)
to roughly half the container limit, and set the container limit above cache + connections.
- Confirm the cap the process actually sees before blaming Mongo: `cat /sys/fs/cgroup/memory.max`.
**Slow queries**
- `explain("executionStats")` showing `COLLSCAN` and `docsExamined` ≫ `nReturned` = a missing
index. Add it, but on a replica set/production build it with care (background is default in 4.2+).
- A query that was fast and turned slow: check for a plan-cache flip or data growth crossing a
threshold; `$indexStats` shows which indexes are actually used.
**Disk full**
- MongoDB does not release space to the OS on delete; it reuses it internally. If the OS disk is
full, freeing documents will not immediately free filesystem space. `db.<coll>.stats()` shows
`storageSize` vs `size`. `compact` (per collection, blocks that collection, one node at a time
on a replica set) reclaims it; on a replica set, resync-from-scratch of a secondary is the
cleaner reclaim. Both are deliberate, disruptive operations.
**Corruption / won't start**
- `mongod --repair` is offline, single-node, and **can discard data it cannot recover**. Take a
filesystem copy of `dbPath` first. On a replica set, never repair — remove the member and
resync it from a healthy one, which is safe and usually faster.
- `db.<coll>.validate({ full: true })` inspects a collection without changing it.
**Upgrades (the FCV trap)**
- You cannot skip a major. 6.0 → 8.0 goes through 7.0. At each step: upgrade binaries, verify the
set is healthy, then raise FCV: `db.adminCommand({ setFeatureCompatibilityVersion: "7.0" })`.
Raising FCV before every node runs the new binary, or skipping it, is how an upgrade half-lands.
## 6. Before you run anything that writes
These change data, availability, or config, and on a nanoinfra deployment they resolve to a
`mutate.remote` capability — they ask for approval interactively and need a standing grant to run
unattended. Say what you are about to do and why, on the specific target, before you do it:
- `rs.reconfig(..., {force:true})`, `rs.stepDown()` — change who serves writes.
- `db.killOp(<opid>)` — kill an in-flight operation (safe for a runaway read; a killed write rolls back).
- `mongod --repair`, `compact`, `--drop` on restore, `db.dropDatabase()`, index drops — destructive.
- `setFeatureCompatibilityVersion` — one-way in practice; downgrade needs a documented dance.
When in doubt, capture state (`mongodump`, a `dbPath` copy, `rs.status()` output) first. A read
that tells you what is wrong is always cheaper than a write that turns out to be the wrong fix.