Systems and Operations

The Complete Linux Low-Disk-Space and Inode Exhaustion Troubleshooting Guide: Safe Cleanup, Log Rotation, and Releasing Ghost File Handles

Compiled by the VPSMap Editorial Team · Updated 2026-09-26 · 24-minute read · Plain-Text Version

Almost every engineer administering cloud servers has encountered No space left on device(No space left on device) emergency alert. In lightweight production environments, for example, BandwagonHost( BandwagonHost ) Common entry-level KVM plans (typically with a quota of 20GB40GB SSD) or DMIT entry-level cloud servers (such as 10GB models in the PVM series40GB of high-speed NVMe storage), disk space is often tightly budgeted. The high cost of premium BGP/CN2 GIA network bandwidth leaves less room for spare storage capacity. If logs grow unchecked, caches accumulate, or tiny files proliferate, available disk space can easily be exhausted within hours.

When storage actually reaches 100% full, the problem goes beyond applications being unable to write: the system may not even be able to create PAM authentication session files,/tmp 临时文件或 /run/user sockets cannot be allocated, causing ordinary remote SSH connections to be rejected outright.

This article explores the underlying mechanisms of Linux storage and the VFS (Virtual File System), focusing on data-block exhaustion and inode depletion. It provides a production-tested systematic troubleshooting process, a method for truncating ghost files held by process handles, and automated log-protection strategies.


One. Emergency Access: Recovery When a Full Disk Prevents SSH Login#

当磁盘写入量彻底锁死在 100% 时,OpenSSH 守护进程在处理公钥或密码登录时可能因无法写入系统审计日志、无法为 sshd allocating a PTY terminal for the session and returns Connection closed by remote host. Do not blindly force a power cycle at this point: memory buffers that have not been flushed to disk and damaged metadata may cause the filesystem check (fsck) to get stuck in read-only mode.

Use the Provider's Out-of-Band Management Console:

1. BandwagonHost: Bypass the Network Layer with KiwiVM Interactive Console#

BandwagonHost's cloud servers are managed by its proprietary KiwiVM Control Panel under management. In extreme cases where standard SSH access is unavailable:

  • 登录 KiwiVM 面板,在左侧导航栏找到 Interactive Console or Root shell - interactive。
  • Note: the KiwiVM control panel administration password is independent of both your client area account password and the operating system's root password.
  • Through a browser-emulated hardware-level terminal, you can operate without relying on the system's /run a temporary session network, directly obtain a shell and perform emergency space-recovery operations.

2. DMIT: Open the Web Console Through the Console's Access Panel#

During initial deployment, DMIT system templates generally disable remote root password login by default and use SSH key pairs. If a full disk prevents public-key authentication from writing temporary files:

  • Log into the DMIT client area, open the relevant instance's management interface, and switch to Access label.
  • Click Console open the interactive web terminal.
  • 控制台提供基础的 TTY 登录环境,可在 Networking 层服务异常或系统 Storage 锁死时直接以本地终端形态切入排障。

Two. Core Mechanisms: Separate Accounting for Disk Blocks and Inodes#

Linux filesystems, such as the widely used Ext4 and XFS, divide physical storage into two independent structures:

Commands / Configuration
                           Linux 存储空间计量体系
                 ┌───────────────────────────────────────┐
                 │          文件系统 (Ext4 / XFS)        │
                 └──────────────────┬────────────────────┘
                                    │
           ┌────────────────────────┴────────────────────────┐
           ▼                                                 ▼
┌──────────────────────┐                         ┌──────────────────────┐
│  数据块 (Data Block) │                         │   索引节点 (Inode)   │
├──────────────────────┤                         ├──────────────────────┤
│ · 存储文件真实内容   │                         │ · 存储文件元数据     │
│ · 尺寸通常为 4KB/块  │                         │ · 大小、权限、时间戳 │
│ · 指标查看:df -h    │                         │ · 指向数据块的指针   │
│ · 症状:大文件撑爆   │                         │ · 指标查看:df -i    │
│                      │                         │ · 症状:海量小文件   │
└──────────────────────┘                         └──────────────────────┘
  1. 数据块空间耗尽(Block Exhaustion): This appears as df -h shows mount-point usage reaching 100%. This is usually caused by uncontrolled growth of one or a few huge files, such as unrotated database logs, Docker image layers, or core dump crash files.
  2. 索引节点耗尽(Inode Exhaustion): This appears as df -h shows several GB or even tens of GB of free space, but any touch or mkdir operations all return No space left on device. At this point, run df -i, you will find IUse% 达到了 100%。
    • 成因:Ext4 等文件系统在格式化之初就固定了 Inode 表的绝对总数(默认通常按每 16KB 数据块分配 1 Inode)。如果程序意外生成了数百万个几十字节的微小文件(例如失效的 PHP/Node.js Session 文件、异常堆积的 Postfix 邮件队列、前端打包产生的深度依赖层),虽然实际只消耗了数百兆磁盘容量,但 Inode resources 池已枯竭,文件系统随之瘫痪。

三、 第一阶段排障:全盘体检与病灶定位#

Whatever the alert, the first step after entering the terminal is to determine the nature of the problem and identify the target directory.

1. Identify the Type of Alert#

Run the following two commands to compare the data:

bash
# 检查物理块容量占用
df -hT -x tmpfs -x devtmpfs

# 检查 Inode 节点消耗量
df -ihT -x tmpfs -x devtmpfs

Parameter Descriptions:-x tmpfs -x devtmpfs filters out memory-resident virtual mount points to keep the investigation focused.

2. Identify Space Hogs: Safely Scan Large Directories#

Avoid Running Indiscriminate Commands Directly from the Root Directory: du -sh /*, because unrestricted recursion on systems with many small files or remote mounts can cause severe I/O blocking or even exhaust memory. Use commands limited to a single filesystem:

bash
# 仅在当前物理文件系统内查找(-x 避免遍历 /proc, /sys 等伪文件系统),输出最大的前 15 个目录/文件
du -ahx / 2>/dev/null | sort -rh | head -n 15

If temporary package installation is allowed, use the modern tool ncdu 能够极大提升分析效率:

bash
# Debian / Ubuntu 环境
apt-get install -y ncdu

# 扫描根目录并进入交互式图形界面(-x 限制在当前文件系统内)
ncdu -x /

3. Locate Inode Consumers: Efficiently Measure File Density#

Finding which directory contains huge numbers of tiny files is a difficult part of traditional troubleshooting. Conventional ls 遇到数百万级文件的目录可能会直接无响应。利用 find 的流式输出结合 Sort 可以在更低开销下定位:

bash
# 统计各目录下的文件总数并按降序排名前 15
find / -xdev -type f -printf '%h\n' 2>/dev/null | sort | uniq -c | sort -rn | head -n 15

四、 第二阶段实操:常见系统级“空间杀手”安全清理#

On small VPS instances, most disk-full problems occur in a few predictable areas: system logs, package-manager leftovers, old kernels, and container data. The following procedures prioritize production safety and avoid blindly using rm -rf /var/log/* 引发服务崩溃。

1. Reduce Systemd Journal Disk Usage#

Modern Linux distributions (Debian 11/12, Ubuntu 20.04/22.04/24.04) use the following by default: systemd-journald 收集各类系统及服务 stdout/stderr。在未加配额限制的情况下,日志很容易吞掉 2GB~4GB 空间。

bash
# 检查当前 journal 日志实际磁盘消耗
journalctl --disk-usage

# 生产级清理:强制截断日志,仅保留最近 3 天或上限 150MB
journalctl --vacuum-time=3d
journalctl --vacuum-size=150M

Long-Term Hardening Configuration: Modify /etc/systemd/journald.conf, set hard limits to prevent the problem from recurring:

ini
[Journal]
SystemMaxUse=200M
SystemMaxFileSize=50M

After saving, run systemctl restart systemd-journald for it to take effect.

2. Clean Up Package Manager Caches and Orphaned Dependencies#

On systems maintained with APT, the files downloaded during each upgrade .deb packages are all stored in full in /var/cache/apt/archives/ in:

bash
# 查看 apt 缓存占用的容量
du -sh /var/cache/apt/archives/

# 安全清空全部 deb 包缓存(不影响任何已安装软件)
apt-get clean

# 自动卸载已被淘汰、不再被任何现行软件依赖的库文件与孤立包
apt-get autoremove --purge -y

3. Remove Old Kernel Versions and Clean Up /boot#

On a basic VPS configuration, if the system is /boot 建立了较小的独立分区(常见于某些特定系统模板,仅分配 500MB1GB), the system repeatedly apt upgrade 3 left behind afterward5 old kernel images (vmlinuz, initrd.img) will directly cause /boot fill it completely, interrupting any subsequent updates or patch installations.

bash
# 1. 确认当前正在运行的内核版本(务必确认,绝对不可删除运行中的内核!)
uname -r

# 2. 列出系统内所有已注册的内核相关包
dpkg -l | grep -E 'linux-image|linux-headers|linux-modules'

# 3. 使用 APT 安全移除未使用的旧内核(现代发行版首选方式)
apt-get --purge autoremove -y

If the system, because of /boot has completely run out of free space, causing apt itself fails and becomes stuck, you can use lower-level tools to manually remove the oldest kernels and break the deadlock:

bash
# 假设 5.15.0-80 为陈旧未使用内核,5.15.0-105 为当前内核
dpkg --purge linux-image-5.15.0-80-generic

4. Thorough Cleanup of the Docker Container Environment#

Docker dangling image layers (Dangling images), obsolete build caches, and container logs continuously written to standard output are the biggest space consumers in containerized deployments.

bash
# 全局查看 Docker 各子组件的占用细则
docker system df

# 一键清除所有已停止的容器、未使用的网络和虚悬镜像
docker system prune -f

# 进阶清除:如果需要清理所有未被现存容器使用的镜像(包括长期不用的基础镜像)
docker system prune -a -f

# 清理无名卷(Dangling Volumes,注意确认其中没有包含重要未挂载数据)
docker volume prune -f

Truncation Method for Uncontrolled Docker Container Logs: Many long-running containers lack log-driver size limits, so a single *-json.log may grow beyond 10GB:

bash
# 定位最大的容器日志文件
find /var/lib/docker/containers/ -name "*-json.log" -exec du -h {} + | sort -rh | head -n 10

# 在容器运行中进行清空(必须使用截断方式,不能直接 rm,原因见第五章)
truncate -s 0 /var/lib/docker/containers/<container_id>/<container_id>-json.log

Five. Tackling the Hard Part: Releasing “Ghost Deleted Files” Held by Open Handles#

Here is a classic scenario that puzzles many administrators: You discover /var/log/nginx/ under access.log occupied a full 8GB, so you confidently ran rm -f /var/log/nginx/access.log. After execution ls no longer shows the file, but then running df -h,find that used space has not changed at all and the disk still shows a critical 100% usage level。

1. The Operating System Mechanism Behind the Behavior#

In Linux VFS, a file's existence is maintained by two components:

  1. Directory Entries (dentry) and Hard Link Count (nlink):rm The System Call Essentially unlink, it merely removes the file's directory entry and changes the Inode reference counter nlink decreases by 1.
  2. Process Open File Descriptor Count (Open File Descriptors): As long as an active process, such as Nginx Master/Worker, Java, or Python, retains a file descriptor (fd) for the file, the Inode's reference count remains greater than 0.

Conclusion: only when nlink == 0 and 所有持有该 fd 的进程均关闭了句柄或进程退出 only then does the filesystem actually release the corresponding data Blocks back to the free storage pool.

Commands / Configuration
执行 rm 操作后的物理状态:
[用户执行 rm] ──> 目录项断开 (ls 无法查见)
                        │
                        ▼
      [持有该文件句柄的进程 (如 Nginx PID: 12480)]
                        │
                        ▼
    Inode 引用仍存在 ──> 磁盘扇区锁定 ──> 容量绝不归还系统!

2. Identify Processes Holding Ghost Files Open#

Use lsof(List Open Files)命令可以准确抓取所有“已被 flag 为 deleted 但句柄尚未释放”的文件:

bash
# 安装 lsof (若系统缺失)
apt-get install -y lsof 2>/dev/null || yum install -y lsof 2>/dev/null

# 查找处于 (deleted) 状态但依然被进程持有的文件
lsof +L1
# 或使用通用过滤:
lsof | grep -iE 'delete|deleted'

Typical output is as follows:

text
COMMAND   PID USER   FD   TYPE DEVICE    SIZE/OFF NODE NAME
nginx   12480 root    7w   REG  252,1  8589934592 131075 /var/log/nginx/access.log (deleted)

The output clearly shows: the PID is 12480 of nginx process is still using the file descriptor 7w holds a deleted log file that still occupies 8.5GB of physical space.

3. Two Ways to Reclaim Space Safely#

Method A: Gracefully Reload or Restart the Associated Service (Recommended Standard Approach)#

If the service can tolerate a brief restart, reset the process lifecycle directly:

bash
# 优雅重载服务(如果服务支持重新打开日志文件,如 Nginx 的 reopen)
nginx -s reopen || systemctl reload nginx

# 或直接重启服务以彻底关闭历史文件描述符
systemctl restart nginx

Method B: Truncate Through the Process's Open File Descriptor While Running (Seamless Zero-Downtime Approach)#

在核心业务(如高吞吐 Gateway 、支付核心或不能中断的长连接服务)不可重启的环境下,可直接 Passed Linux 内核的 /proc pseudo-filesystem to locate the process's file descriptor and truncate it in place:

bash
# 格式规范:: > /proc/<PID>/fd/<FD_NUM>
# 或使用 truncate 命令:truncate -s 0 /proc/<PID>/fd/<FD_NUM>

# 针对前文示例中的 PID 12480 和 FD 7:
: > /proc/12480/fd/7

原理解析: The redirection operation : > sends the active handle a O_TRUNC a truncation signal, the system immediately returns the physical data blocks to the free pool and reduces the file size to 0 bytes, whileThe Process Itself Does Not Throw Any Invalid-Handle Errors. After completion, run again df -h, and disk space will immediately return to normal.


Six. Prevent Problems Before They Start: Build Automated Logrotate Rotation and Self-Recovery#

The fundamental way to avoid being awakened at night by a full disk is to make log rotation and quota-based circuit breakers part of the system's routine background management.

1. Configure Standard Logrotate Rules#

直接删除活动日志是运维大忌。规范的方式是配置 logrotate Periodically rotate, compress, and move archived logs.

For custom applications, such as those running in /data/app/logs workloads), write standardized rules /etc/logrotate.d/my-app:

ini
/data/app/logs/*.log {
    daily                   # 每日检查切分
    rotate 7                # 仅保留最近 7 份归档
    missingok               # 若日志文件丢失不抛出报警
    notifempty              # 文件为空时不执行切分
    compress                # 使用 gzip 压缩归档文件,削减 90% 存储
    delaycompress           # 延迟到下一次轮转周期再压缩,保障写入安全
    sharedscripts
    copytruncate            # 关键指令:将源日志截断为 0,而非直接重命名移走,避免产生幽灵句柄
}

Important Note copytruncate: For those that do not support HUP signal to reopen log files without restarting, such services must declare copytruncate. It copies the current log contents to a backup, then immediately truncates the original file in place, ensuring that the process retains the same valid file descriptor and eliminating ghost handles.

2. Prevent Inode Exhaustion: Regularly Clean Up Sessions and Temporary Caches#

对于产生海量短生命周期小文件的系统,应配置系统定时任务自动 Clear 陈旧文件,而不是等 Inode 报警。

Add lifecycle management logic to crontab (for example, cleaning up /data/cache files under it that have not been modified for more than 7 days):

bash
# 编辑 root 定时任务
crontab -e

# 每天凌晨 3:30 执行一次清理,并直接在文件系统层删除,避免一次性载入内存
30 3 * * * find /data/cache/ -type f -mtime +7 -delete

Performance Advantages:-delete the operation is performed by find execute directly at the kernel level unlink, far more efficient than find ... | xargs rm, and will not produce the following error from an excessively long argument list: Argument list too long error.


Seven. Completing the Operations Cycle: Quick Reference for Common Inspection Commands#

Add the following checks to your routine inspection process:

排查目标Key Troubleshooting Commands核心价值 / 适用场景
整体容量df -hT -x tmpfs -x devtmpfsQuickly Identify Which Physical Partition Is Full
Inodesdf -ihT -x tmpfs -x devtmpfsDiagnose an Inode Shortage When “Cannot Create File” Appears Despite Free Disk Space
大文件定位du -ahx / 2>/dev/null | sort -rh | head -n 15Find Large Directories and Individual Files Within a Single Filesystem, Working from Deeper Levels Upward
Small-File Accumulationfind / -xdev -type f -printf '%h\n' 2>/dev/null | sort | uniq -c | sort -rn | head -n 15Locate Directories Where Huge Numbers of Tiny Files Cause Inode Exhaustion
Ghost File Handleslsof +L1 or lsof | grep -i deletedFind those that have been rm 移除但仍被服务进程锁死物理空间的僵尸文件
Truncate While Online: > /proc/<PID>/fd/<FD>Instantly Reclaim Physical Storage in Production Without Restarting Services or Causing Downtime
Log Truncationjournalctl --vacuum-size=100MQuickly Reclaim Several GB of System Logs Consumed by systemd

Mastering this approach, from out-of-band panel access and dual-track filesystem diagnosis to truncating ghost file handles and robust Logrotate management, ensures long-term system stability even on small VPS instances with limited hardware quotas, eliminating the risk of service interruptions caused by uncontrolled storage growth.