A Kubernetes worker can have gigabytes free and still be one resource away from failure. The missing resource is not memory, CPU, or even disk blocks. It is the finite set of inode records the filesystem uses to name and describe files.

That is the trap in Onkar Raghunath Shelke's CNCF incident analysis. One node showed 103 GB used and 22 GB available, about 83% byte utilization. Its inode table looked less alarming: 11 million used and 5.3 million free, about 67%. Yet the alert was about inodes, and it was right. Their consumption rate pointed toward exhaustion while Kubernetes' routine image cleanup was still watching a different gauge.

The node was not misconfigured. Kubelet knew how many inodes remained. The problem was timing: its background image garbage collection reacts to bytes, while its default inode protection belongs to node-pressure eviction. By the time that second mechanism acts, the quiet maintenance window has already become an emergency.

A filesystem has two kinds of room: space for file contents and space for the files themselves. Kubernetes does not clean both on the same schedule.

One Disk, Two Ledgers

An inode stores filesystem metadata for an object: ownership, permissions, timestamps, size, and pointers toward its data. A directory entry connects a name to an inode. Every regular file consumes an inode whether it holds a multi-gigabyte model or one byte.

On ext4, the inode count is normally chosen when the filesystem is created. It does not grow merely because free blocks remain. A common formatting ratio allocates one inode for every 16,384 bytes of filesystem capacity. Meanwhile, a non-empty tiny file typically occupies at least one 4 KiB block.

That arithmetic creates a counterintuitive ceiling. A 1 GiB ext4 filesystem with one inode per 16 KiB gets 65,536 inodes and 262,144 4 KiB blocks. Fill it with one-block files and all inodes are consumed after 65,536 blocks, only one quarter of the nominal data blocks. Journal and filesystem metadata change the exact observed percentage, but not the shape of the failure.

1 GiB filesystem
262,144 data blocks x 4 KiB
 65,536 inodes      x 1 file

tiny-file ceiling
 65,536 files x 4 KiB = 256 MiB of file data
 inodes: 100% used
 blocks: roughly 25% used before metadata
Free blocks cannot help once the fixed inode table has no entry left for another file.

At that point, writes that create files fail even though ordinary disk-capacity tools still report room. The machine is empty and full at the same time. Watching only byte utilization is like counting open shelf space after the card catalog has run out of cards.

Kubelet Has An Early Byte Alarm

Kubelet's image garbage collector has a deliberately calm operating range. The default imageGCHighThresholdPercent is 85. Once image storage crosses that byte-usage threshold, kubelet removes unused images, oldest first, until usage approaches the default low threshold of 80. Both settings are explicitly percentages of disk usage.

Inodes are covered elsewhere. On Linux, Kubernetes' default hard-eviction thresholds include nodefs.inodesFree<5% and imagefs.inodesFree<5%. The same defaults watch free bytes at 10% on nodefs and 15% on imagefs. Those signals can refer to separate filesystems or to the same underlying mount, depending on how the node and container runtime are laid out.

The asymmetry is the important part. Bytes get routine image garbage collection at 85% used and a later node-pressure threshold. Inodes get the latter. At 5% free, kubelet enters its reclaim and eviction path. It first tries node-level cleanup, including removing unused images where appropriate. If that does not recover enough, it can evict workload pods. A hard threshold has no grace period and does not honor a PodDisruptionBudget.

So it is accurate to say kubelet watches inodes. It is misleading to assume that means inode growth receives the same early maintenance response as byte growth. The default safety net is close to the floor.

Container Images Can Spend Inodes Quietly

The incident's inode bill lived under containerd's overlayfs snapshot store. One dependency directory, @mui/icons-material, contained 21,553 files in each affected snapshot. The full application image carried more than 40,000 files. Retain several versions on a node and the count moves from inconvenient to millions without requiring a proportionate number of bytes.

This is not a Node.js-only problem. Python virtual environments, vendored PHP trees, Ruby gems, generated headers, localization bundles, and other forests of small files create the same resource shape. Container registries and content stores may deduplicate identical compressed layers by digest, but the runtime still needs unpacked filesystem trees for distinct snapshots. A new layer digest can therefore become another large set of directory entries and inodes on every node that pulls it.

A single-stage build makes the problem easy to ship. Build tools, source, development dependencies, caches, and the final artifact all survive into the runtime image. A missing .dockerignore can copy a host-side dependency tree into another layer. Rebuilding dependency layers without reproducible metadata or a useful cache can produce new digests that leave more snapshots behind.

None of those mistakes has to make the image look enormous in bytes. Forty thousand tiny files can hide inside a size that seems ordinary, which is precisely why the byte-triggered collector may wait.

The First Fix Belongs In The Artifact

The strongest repair is to stop deploying build-time file trees. A multi-stage image installs dependencies and produces the application in one stage, then copies only the runtime result into a smaller final stage. A compiled web frontend may need a web server and a modest set of static assets, not its source tree and every development package used to create it.

FROM node:20 AS build
WORKDIR /app
COPY package*.json ./
RUN npm ci
COPY . .
RUN npm run build

FROM nginx:alpine
COPY --from=build /app/build /usr/share/nginx/html
The exact runtime varies, but the boundary is the point: ship the product of the build, not the workshop.

A good image review should count files as well as bytes. Compare candidates with a command such as find / -xdev -type f | wc -l inside a disposable container. Inspect layers for repeated dependency trees. Keep node_modules, local virtual environments, build caches, and repository metadata out of the build context unless the runtime truly needs them.

Existing nodes still need deliberate cleanup. crictl rmi --prune removes images no current container uses, after which the runtime can pull them again if needed. More aggressive cleanup can erase stopped-container evidence such as previous logs, so it should remain an incident action rather than a blind cron job.

Monitor The Gap, Not Just Each Gauge

The upstream NodeFilesystemFilesFillingUp alert did valuable work in the reported incident. It uses a projected trend, so it can page while absolute utilization still looks comfortable. That is why 67% inode use was meaningful: the slope said the remaining supply would not last.

Trend alerts can quiet down when growth pauses, so operators also need a direct threshold with enough margin to respond before 95% usage. The exact warning point depends on node size, image churn, and response time. An 80% inode threshold held for a short period is a reasonable starting discussion, not a universal constant.

A dashboard becomes more useful when it compares the two ledgers. Subtract byte-use percentage from inode-use percentage for each relevant filesystem. A strongly positive result highlights nodes where file count is racing ahead of capacity consumed. Those are the machines for which the familiar disk bar is least informative.

When one appears, the investigation is pleasantly concrete:

  • Use df -h and df -i on the same mount to compare bytes and inodes.
  • Run du --inodes -xS /var/lib/containerd | sort -rh | head to find directories with the most entries without crossing filesystems.
  • Map large snapshot trees back to images, then inspect how those images are layered and built.
  • Verify the node's effective kubelet configuration. Kubernetes warns that overriding one hard-eviction setting can zero omitted defaults unless default merging is enabled, so partial threshold edits deserve care.

The Emergency Brake Is Not Maintenance

Kubernetes has mechanisms for both resources. That fact alone can give a false sense of symmetry. Image garbage collection prevents ordinary byte pressure from reaching an incident. Inode protection, by default, joins the path designed to save a node that is already under pressure.

The distinction matters because container packaging determines which gauge moves first. A lean artifact lets bytes and file count tell roughly the same story. A sprawling dependency tree can make inode use accelerate while the disk still looks healthy. Kubernetes cannot infer that the image is wasteful from its size alone.

Capacity planning is often treated as one number. Filesystems have at least two, and container platforms layer their own cleanup rules on top. The practical lesson is small: count the files you ship. The operational lesson is larger: know whether the mechanism watching a resource is routine housekeeping or the last thing that happens before workloads are removed.


Sources