Capacity & resource limits¶
On memory-constrained nodes, several failures look like application bugs but are really capacity and scheduling effects. The usual signals mislead — know the real ones.
A memory limit is not a reservation¶
Only a container's requests count against a node's allocatable memory and drive scheduling.
The limit is just the burst ceiling. So a limit can be generous at no steady-state cost —
but for a cache-sizing application the limit is memory it will actually consume. Some
components (a metrics TSDB, for instance) size their caches to a percentage of their memory
limit, so raising the limit raises real occupancy. Treat such a limit as occupancy, not
headroom.
Some services derive concurrency from their limits — and fail quietly if you don't pin them¶
A metrics database can derive its query concurrency (and cache size) from its CPU/memory limit. Leave the limit unset or wrong and it mis-sizes silently — no error, just degraded or failing queries. Pin CPU and memory limits explicitly on components that self-tune from them.
Before raising any request/limit, check the target node's real free memory¶
Raising a pod's requests/limits is both a placement change and a consumption change:
- A bigger request can relocate the pod — the scheduler may move it to a different node (least-allocated scoring). "It fits somewhere in the cluster" is not enough; check which node it will actually land on.
- A bigger limit, for a cache-sizing app, is memory that node will actually give up.
Before merging a bump, check the winning node's actual free memory (available bytes, not its request sum), and treat a cache-sizing app's whole limit as occupancy. Getting this wrong can push a node into kernel OOM.
MemoryPressure lags the kernel OOM killer — read exit codes instead¶
The kubelet's MemoryPressure condition trails the kernel OOM killer, so it can read False
while the kernel is already killing containers. Never use MemoryPressure to rule out
memory exhaustion. The real signal is co-tenant containers dying with exit code 137 — if
unrelated pods on one node are OOM-killed together, a memory bump or a fat cache is the
suspect, not those apps.
Over-packed nodes cause probe restart storms, not OOM¶
When several unrelated pods on one node restart-loop together, and the restarts show exit 0 / SIGTERM (a liveness kill) rather than 137, the cause is usually CPU contention, not the apps or memory: the node is over-committed on CPU requests, probes miss their deadlines, the kubelet kills and restarts pods, and the restarts burn more CPU — a self-reinforcing storm that will not settle on its own. Fix it by cordoning the node and rescheduling the churning pods onto lighter nodes (delete tightly-coupled pods one at a time to avoid a service gap); don't wait it out. Prevent recurrence with soft pod-anti-affinity so co-dependent replicas don't re-stack on one node. (Pod-anti-affinity is namespace-scoped by default — it won't spread identically-labelled pods that live in different namespaces.)
Small control-plane nodes and the API server's memory ceiling
If you run small control-plane nodes and cap the API server's Go memory
(GOMEMLIMIT) to keep it from OOMing them, size that cap above the API server's
steady-state working set with headroom. Set too close to the working set, the Go GC clamps
its target under the live heap and thrashes — chronic high CPU that "resets on reboot"
(really the multi-day cache-refill ramp before it re-hits the ceiling) but is neither load
nor a leak. The only lever that relaxes both the GC thrash and the memory pressure is more
control-plane RAM.
See also: Storage · Node upgrade & maintenance · Operations.