# When Your Kubernetes Nodes Run Out of Inodes but Have Plenty of Disk Space
## The Silent Crisis Beneath the Surface
A Kubernetes worker node triggers a `NodeFilesystemFilesFillingUp` alert. An operator checks disk usage immediately — it reads 83% full, 103 GiB consumed out of roughly 125 GiB. That’s concerning but not yet catastrophic. Then they look at inodes: 67% consumed, 11 million used, 5.3 million remaining. The inode table is what actually paged.
Here’s the unsettling part: the disk metric looks worse than the inode metric, yet the inode count is the one that triggered the alert. This isn’t a contradiction — it’s how the alerting system actually works, and understanding why it works this way is the key to preventing a node crash.
## How kubelet Protects a Node
Kubernetes’ kubelet agent has built-in mechanisms to keep nodes from running out of storage, but those mechanisms operate differently depending on what resource is under pressure.
### Byte-Based Defenses
For disk capacity measured in bytes, kubelet has two layers of protection. The first is image garbage collection, which runs quietly in the background. When the image filesystem crosses 85% capacity by default, kubelet starts deleting unused container images, oldest first, until usage drops back to 80%. This is a background process that requires no pod disruption.
The second layer is hard eviction. If node disk free space falls below 10% or the image filesystem free space drops below 15%, kubelet begins forcefully evicting pods to reclaim resources. This is disruptive — pods get terminated.
So bytes get two watches: an early background cleanup at 85%, and a hard eviction floor at either 10% or 15% free.
### The Inode Gap
Inodes receive a different level of attention. Kubelet does track inode consumption on Linux, with a hard eviction threshold set at 5% free inodes. The default configuration for Linux nodes includes five eviction signals, covering both memory, node filesystem, and image filesystem for both bytes and inodes.
However — and this is the critical detail — there is no early trigger for inodes. The only mechanism that responds to inode exhaustion is the eviction floor at 5% free. Image garbage collection, which runs smoothly in the background for byte pressure, never once examines file counts. It operates entirely on byte percentages.
This means that workloads generating enormous numbers of small files get zero early intervention from kubelet. The first time anything responds to an inode shortage is when the node is already at the emergency eviction point.
## Why Small Files Create a Different Kind of Problem
To understand the math, consider a simple thought experiment. Take a filesystem and fill it exclusively with one-byte files. What percentage of the disk would be consumed by the time inodes run out?
The answer is smaller than most people expect. The ext4 filesystem assigns a fixed number of inodes at creation time, determined by the inode ratio — one inode per N bytes of capacity. The stock Linux configuration (`/etc/mke2fs.conf`) sets this ratio to 16,384 bytes per inode. This number never changes after the filesystem is created.
Meanwhile, every non-empty file on ext4 occupies at least one 4 KiB block, regardless of how small the file’s actual content is. So a one-byte file consumes 4,096 bytes of block space but only one inode.
A quick demonstration confirms the math. Creating a 1 GiB filesystem image with the stock ratio produces roughly 65,536 inodes and 262,144 blocks. If every file consumes one block, all 65,536 inodes are exhausted after 256 MiB — just 25% of the total disk capacity. The remaining 75% of blocks are still free, but the filesystem can no longer store new files.
### Watching It Fail in Practice
The failure is not hypothetical and can be reproduced without root privileges. Create a directory containing 4,000 one-byte files, then build an ext4 filesystem from that directory into a 64 MiB image file. The resulting filesystem shows 97.9% of inodes consumed but only 32.6% of blocks used. The files themselves account for roughly 24.4 percentage points of block usage; the remaining 8.2 points come from journal metadata and filesystem overhead.
Now add 200 more files. The build fails: `Could not allocate inode in ext2 filesystem while populating file system`. At approximately file 4,085, the inodes are gone. From that point forward, free blocks become meaningless — the filesystem cannot accept new files regardless of available space.
On modern e2fsprogs versions, the entire sequence completes in under two seconds, making this an excellent experiment to run locally.
## Reverse-Engineering an Actual Node’s Configuration
Looking at a real incident, the node had roughly 16.3 million total inodes (11 million used plus 5.3 million free) on a filesystem of approximately 124 to 133 GiB. At the stock 16,384 bytes-per-inode ratio, a filesystem of that size would only allocate about 8.1 to 8.7 million inodes — far fewer than the 16.3 million already present.
This tells us the filesystem was not formatted with the default ratio. Dividing the filesystem size by the inode count yields roughly 7,500 to 9,000 bytes per inode, aligning closely with a ratio of 8,192. At that ratio, an all-small-files scenario exhausts inodes at roughly 50% of disk usage rather than 25%.
A production node will never contain exclusively tiny files. Layer blobs, logs, and other large objects mean that disk usage typically outpaces inode consumption. But the asymmetry between monitoring tools means the inode trajectory can be invisible right up until the moment it becomes critical.
## Locating the Files Responsible
Discovering where millions of files accumulated requires the right `du` command. The two flags that matter are `-x`, which restricts the scan to a single filesystem boundary (avoiding tmpfs mounts and kubelet-managed volume directories under `/var/lib/kubelet/pods`), and `-S`, which reports each directory’s own file count without recursively adding subdirectory totals. Without `-S`, parent directories like `/var` and `/var/lib` dominate the output, obscuring the actual directories of interest.
On the affected node, the top consumer was containerd’s overlayfs snapshot store at `/var/lib/containerd/io.containerd.snapshotter.v1.overlayfs/snapshots/`. Each numbered subdirectory represented one unpacked image layer. Several of those snapshot directories contained the exact same path: `app/node_modules/@mui/icons-material`, each carrying 21,553 files.
That is one package directory within a single container image. The full image contained more than 40,000 files total — the icon package represents over half of that count, with the remainder distributed across the rest of the dependency tree, application source code, and build output, all bundled into the same image. Multiply 40,000 files by every image version still held on the node, and millions of inodes consumed begins to make complete sense.
### The Deduplication Illusion
containerd’s content store deduplicates compressed layers by content digest, but the overlayfs snapshotter unpacks each distinct layer into its own independent directory and shares nothing between them. Two builds that place an identical `node_modules` tree into layers bearing different digests result in two complete, separate copies on disk.
The `node_modules` directory is the well-known culprit in this pattern, but it is not the only offender. Any container image that ships a dependency tree composed of many small files exhibits identical behavior — whether that is a Python virtual environment, vendored PHP packages, or a collection of Ruby gems.
## Root Cause: The Image Build Process
The cluster configuration itself was not at fault. Kubelet, containerd, and the node operating system were all running with their default settings. The problem originated in how the container images were constructed.
The image in question used a single-stage build, meaning everything required during the build process — source code, build tools, intermediate artifacts, and all dependencies — ended up in the final image that gets deployed. A typical single-stage Dockerfile of this type looks like:
“`dockerfile
FROM node:20
WORKDIR /app
COPY package*.json ./
RUN npm install
COPY . .
RUN npm run build
“`
There are two compounding issues here. First, the `COPY . .` instruction includes everything in the build context, including a local `node_modules` directory, unless a `.dockerignore` file explicitly excludes it. If `node_modules` is not excluded, every commit that touches the build context produces a new layer containing another copy of that directory.
Second, even with a properly configured `.dockerignore`, a CI pipeline that runs without a warm layer cache re-executes `npm install` on every build. The resulting layer carries new file timestamps even if `package.json` has not changed, producing a unique digest every time. Every new digest becomes a distinct snapshot on every node that pulls the image.
containerd retains these snapshots until the image itself is removed, and kubelet only removes unused images when image garbage collection fires at the 85% byte threshold. The disk usage metric remained comfortably below that threshold throughout, while the inode count climbed silently.
## Remediation Strategies
### Build Process Correction
The primary fix belongs in the Dockerfile itself. A multi-stage build separates the dependency installation and compilation phase from the runtime image. The first stage installs dependencies and compiles the application. The final stage pulls only a lightweight base image and copies the compiled output from the build stage.
“`dockerfile
FROM node:20 AS build
WORKDIR /app
COPY package*.json ./
RUN npm install
COPY . .
RUN npm run build
FROM nginx:alpine
COPY –from=build /app/build /usr/share/nginx/html
“`
A React build output deployed this way typically results in a few dozen files rather than tens of thousands. The deployed image’s file count drops by orders of magnitude. To verify, run `docker run –rm
### Monitoring Adjustments
The existing `NodeFilesystemFilesFillingUp` alert performed its intended function by catching the trajectory at 67% inode usage. However, it is a predictive alert based on growth rate, meaning it goes quiet when file growth temporarily stalls.
A flat threshold alert provides a complementary safety net that fires well before kubelet’s 95% eviction point:
“`
1 – node_filesystem_files_free{mountpoint=”/”} > 0.80
“`
A 30-minute evaluation window prevents transient spikes during image pulls or rollouts from triggering false alarms. The 80% threshold is a judgment call — choose a value that leaves sufficient time to act before reaching the 95% eviction floor.
For teams running Prometheus, a useful dashboard metric compares inode consumption against byte consumption per node, highlighting systems where inodes are depleting faster than disk space:
“`
sort_desc(
(1 – node_filesystem_files_free{mountpoint=”/”})
–
(1 – node_filesystem_avail_bytes{mountpoint=”/”} / node_filesystem_size_bytes{mountpoint=”/”})
)
“`
This subtracts free bytes ratio from free files ratio, placing nodes that burn through inodes faster than bytes near the top of the list. These are the systems where byte-based thresholds will ultimately fail.
### Cleanup Operations
Neither fix removes files that already exist on nodes. Old snapshots persist until their parent images are removed. Running `crictl rmi –prune` removes every image that has no running containers referencing it. Those images are pulled again the next time a container needs them.
The command `crictl rm –all` clears every stopped container. It works effectively in emergencies, but removing it from a recurring cron job is wise, as it also clears logs accessed via `kubectl logs –previous` for recently restarted pods.
## Frequently Asked Questions
### Why doesn’t kubelet’s image garbage collection look at file counts?
Image garbage collection was designed around byte-based capacity planning. Container images are primarily measured by their disk footprint, and byte thresholds map naturally to storage provisioning decisions. File counts are a secondary concern that only becomes critical when images contain enormous numbers of small files — a pattern that is common in JavaScript-heavy dependency trees but was not the primary driver behind the original kubelet design.
### Can I change kubelet’s inode eviction threshold?
Yes. The `nodefs.inodesFree` and `imagefs.inodesFree` values in kubelet’s eviction configuration can be adjusted. Setting a higher threshold (for example, 10% free instead of 5% free) provides more runway to react before pod eviction begins. However, this does not solve the underlying problem of having no early trigger — it only moves the emergency braking point further out.
### What happens when a node runs out of inodes?
Once all inodes are consumed, the filesystem can no longer create new files, even if significant disk space remains free. New container image layers cannot be unpacked, new logs cannot be written, and new sockets cannot be created. kubelet responds by evicting pods to try to free resources, but this is a last-resort measure rather than a graceful degradation path.
### Are Windows nodes affected by this issue?
Windows nodes use a different kubelet configuration (`defaults_windows.go` and `defaults_others.go`) that does not include inode-based eviction thresholds. The byte-based thresholds still apply. The specific inode exhaustion pattern described here is a Linux ext4/xfs phenomenon.
### How can I check my own node’s inode ratio?
Run `tune2fs -l /dev/your-root-device | grep -E ‘^(Inode count|Block count|Block size):’` on the target node. Divide the inode count by the block count to understand how many bytes each inode represents. Compare this against the stock ratio of 16,384 bytes to determine whether your filesystem was formatted with a custom setting.
### Is `node_modules` the only common cause of inode exhaustion?
No. Any pattern that bundles thousands of small files into container images can produce the same result. Python virtual environments, vendored PHP package directories, Ruby gem trees, and statically compiled language build outputs with many object files all exhibit the same behavior.
## Conclusion
The gap between how kubelet monitors bytes and how it monitors inodes creates a blind spot for file-heavy workloads. Disk usage percentages can look entirely healthy while inode consumption climbs toward a cliff edge with no early warning mechanism to intervene. The problem is not misconfiguration — it is a structural feature of how container images are built and how the node agent watches resources.
Multi-stage builds, proper `.dockerignore` files, and inode-aware monitoring together close this gap. The fix belongs in the CI pipeline before it belongs in the cluster, because once millions of small files are sitting on a node, cleanup becomes a disruption rather than a prevention.
The lesson is straightforward: if your workloads carry dependency trees full of small files, the inode count deserves the same monitoring attention you give to disk bytes.
Thank you for reading



