在跨境运维与 Overseas Cloud servers 管理过程中, Networking 不可达是发生频率最高、定位成本最重的问题之一。许多运维人员在遭遇连接中断时,往往仅凭本地终端的 ping a timeout and blindly conclude that “the server is blocked by the firewall” or “the provider's data center is down,” then hastily request an IP replacement or reinstall the system. This unsystematic troubleshooting wastes time and can easily misdiagnose deeper protocol negotiation failures.
实际上,从最底层的虚拟化 instance 状态,到跨境 BGP 路由、中间防火墙策略,再到 TCP 握手与链路 MTU 分片协商,任何一个层级的异常都会表现为客户端侧的“失联”。本手册以主流境外优化 Networks Providers BandwagonHost( BandwagonHost ) and DMIT using its actual infrastructure and console mechanisms as references to build a bottom-up, full-layer troubleshooting model and practical command set.
Quick Self-Check: A Decision Diagram for Troubleshooting Every Layer#
When an overseas VPS suddenly becomes unreachable, avoid haphazard testing. Follow the standard layered diagnostic process below, working through each layer to narrow down the source of the problem:
flowchart TD
Start([客户端连接 VPS 失败]) --> L0{第 0 层: 实例生命周期<br/>机器在运行吗?账单过期吗?}
L0 -- 关机 / 欠费暂停 / 死机 --> Fix0[进入控制台启动实例 / 核对账单 / VNC 查看内核崩溃]
L0 -- 实例运行正常 (Running) --> L1{第 1 层: 公网 IP 连通性<br/>全球与大陆 ICMP 状态?}
L1 -- 全球多点均全红超时 --> Fix1[宿主机路由故障 / 触发 DDoS 黑洞 / 联系机房工单]
L1 -- 海外全绿 / 国内三网全红 --> Fix1_GFW[IP 遭遇跨境阻断<br/>按服务商规则申请更换或重构路由]
L1 -- 海外全绿 / 国内正常响应 --> L2{第 2 层: 传输层端口<br/>nc / tcping 目标端口是否可达?}
L2 -- 瞬时收到 RST / 连接超时 --> Fix2[端口被拦截 / 修改高位端口或检查安全组]
L2 -- Connection Refused --> L3{第 3 层: 系统内部服务<br/>sshd 监听 0.0.0.0 吗?防火墙放行吗?}
L2 -- TCP 握手正常但操作卡死 --> L4{第 4 层: MTU / MSS 协商<br/>小包通、大包死?TLS 握手长久挂起?}
L3 -- 仅监听 127.0.0.1 / 防火墙 DROP --> Fix3[修正 ListenAddress / 调整本地防火墙规则]
L4 -- ping 大包受阻 / PMTUD 丢失 --> Fix4[配置 TCPMSS 自动钳制 / 调整网卡 MTU]Layer 0: Instance Lifecycle and Out-of-Band Console Diagnostics (Host and Billing Status)#
Before analyzing network packets, first rule out an instance that is not actually running because of hardware problems, billing suspension, or a frozen operating system.
1. 账单周期与服务状态确认#
Overseas VPS providers generally use strict automated billing. Nonpayment and automatic protection mechanisms can directly disable network interfaces:
- BandwagonHost: renewal invoices are generally generated 7 days before service expiration. The system does not automatically charge a linked credit card or PayPal account; automatic payment requires sufficient prepaid account credit. Missing the payment window causes the instance to be automatically suspended.
- DMIT: monthly invoices are also issued 7 days before expiration. After an expired instance becomes Suspended, the provider generally allows only a 3-day grace period. If payment remains outstanding, the host destroys the instance and permanently erases its data disk.
排障第一步:登录相应客户后台,确认该服务的状态是否为绿色的 Active / Running。
2. 带外救援终端(VNC / Serial Console)实操#
When SSH is completely unavailable and the server cannot be probed over the public internet, use the provider's out-of-band management connection to access the system and check for OOM (Out of Memory) or Kernel Panic:
BandwagonHost (KiwiVM Architecture):
- Log in to the KiwiVM control panel. Note that the separate KiwiVM panel password, client area credentials, and internal system root password are independent.
- Observe the main interface's
Status是否为运行中,若处于Stopped则点击Start。 - Expand the left-hand menu and go to Interactive Console, this feature provides browser-based out-of-band Shell/VNC access.
- 若终端黑屏或无响应,直接在 KiwiVM 点击
Force Stopbefore starting again; if the screen displays an obvious kernel failure stack (Call Trace:/Kernel panic - not syncing), you can use the panel's Install new OS reimage the server (note that this formats the system disk).
DMIT (Console & Access Architecture):
- Log in to DMIT's instance details panel. A green status indicator means the KVM virtualization process is running.
- Click the following in the top action bar: Console, opening the emergency browser-based VNC terminal.
- Key Authentication Differences: To meet its security baseline, DMIT's official system templatesRemote root Password Login Disabled by Default, requiring an SSH key pair. If the local private key is lost and password entry fails in the out-of-band Console, switch to the instance panel's Access tab to reset the password or push the Public Key again.
- 重点 Details 规范: After changing a public key or password in the DMIT panel's Access tab,You must reboot (Reboot) the instance as instructed by the panel, allowing Cloud-init or the underlying configuration script to complete key injection and reload authentication.
Layer 1: Diagnose Cross-Border Public IP Connectivity (ICMP Layer)#
ICMP Echo Test (pingis the first key distinction in diagnosing routing boundaries. The central task is to distinguishPhysical Connectivity Failure WorldwideandCross-Border Routing Blocked Only in the Mainland China Direction。
1. 破除认知盲区:ICMP 的误导性#
- Dropped ICMP Packets Do Not Mean the Server Is Down: to defend against ICMP Flood attacks, some backbone nodes or host firewalls assign Echo Requests extremely low-priority Rate Limits or drop them entirely. Severe Ping packet loss does not necessarily indicate a TCP connection failure.
- Working ICMP Does Not Mean the Application Is Available:跨境防火墙具备细粒度识别能力,能够只针对 TCP 端口 Reset 或丢弃,而完全放行 ICMP 报文。
2. Global Multi-Location Topology Probe Matrix#
Use Third-Party Multi-Location Probes, Such As ping.pe、itdog.cn or boce.com)对 VPS 的 public network IP 执行双向探测,对照下表得出确定性结论:
| Overseas Node Responses (Europe/Americas/Asia-Pacific) | Responses from Mainland China's Three Major Carriers (China Telecom / China Unicom / China Mobile) | Isolate the Root Cause | Next Steps and Provider Policy Restrictions |
|---|---|---|---|
| All Requests Time Out (100% Packet Loss) | All Requests Time Out (100% Packet Loss) | Instance Powered Off, Kernel Hung, Upstream BGP Offline, or IP Blackholed by DDoS Protection | Return to Layer 0 to troubleshoot through the console; if the system is running normally, open a support ticket to verify the data center host's routing |
| Normal Connectivity (0% Packet Loss) | All Requests Time Out (100% Packet Loss) | The IP Is Blackholed by Mainland China's Great Firewall (GFW) in One or Both Directions | See the Provider's IP Replacement Rules Below |
| Normal Connectivity (0% Packet Loss) | Timeouts / Extremely High Packet Loss on Certain Carriers | Evening Peak Congestion on International Backbones, Indirect Return Routes, or Degraded Local Peering | Inspect MTR Hops to Check Whether Return Routes for the Three Carriers Have Deviated from Premium Paths |
| Normal Connectivity (0% Packet Loss) | 正常通畅 (延迟正常) | IP-Layer Routing Is Fully Healthy | The Fault Is at the Transport or Application Layer; Go Directly to Layer 2 Troubleshooting |
3. Strategies for IP Blocking and Practical Provider-Specific Pitfalls#
Once the IP is confirmed to be accessible overseas but completely blocked within mainland China, replace the IP:
- Constraints on BandwagonHost Data Center Migration: Many users know that BandwagonHost offers “Migrate to another DC” and mistakenly assume switching data centers will also give them a new usable IP.Explicit Official Rules: When a VPS's current IP is externally Blacklisted, KiwiVM automatically locks or strictly limits data center migration to avoid contaminating other IP pools. Free migration cannot resolve this situation; use the control center's paid IP replacement process.
- DMIT IP Replacement and Guarantee Terms: DMIT has strict IP service terms for different route tiers, such as the premium CN2 GIA/AS9929/CMIN2 Pro series and standard promotional routes. Some higher-guarantee plans allow periodic paid or conditional replacements. New instances qualify for a full refund within 3 days with traffic usage below 30GB. Beyond that window, or after substantial abuse, request a usable replacement IP through an Addon in the client area.
Layer 2: Transport-Layer Port Filtering and Connection Blocking (TCP Layer)#
If the IP responds normally to Ping but SSH or application connections remain unsuccessful, network-layer routing is working and the blockage is at the TCP transport layer.
1. Client-Side Transport Handshake Probing#
In your local terminal (use built-in tools on macOS/Linux/WSL; on Windows, use PowerShell or tcping) to probe the application port:
# 使用 netcat 探测目标主机指定端口的 TCP 握手(超时限制 3 秒)
nc -zv -w 3 154.xx.xx.xx 22
# 或使用 curl 验证 TLS / HTTP 端口的连接握手过程
curl -Iv https://154.xx.xx.xx:443 --connect-timeout 32. Analyze Handshake Results as a State Machine#
客户端 SYN 报文发出
│
├─► 超过 3 秒毫无响应 (Operation timed out)
│ └── 原因:中间防火墙(GFW)将该端口精确列入黑名单,实施静默丢包(DROP)。
│
├─► 瞬时收到 RST 报文 (Connection reset by peer)
│ └── 原因:链路中间设备检测到敏感特征或明文握手,伪造 TCP RST 中断连接。
│
├─► 收到 ICMP 端口不可达 / Connection refused
│ └── 原因:报文成功到达 VPS,但宿主机上没有进程监听该端口,或本机防火墙 REJECT。
│
└─► 提示 Connected to ... (TCP 握手成功)
└── 原因:传输层完全正常,故障在认证层或 MTU 协商层。3. 规避端口封锁:默认端口策略与修改规范#
- BandwagonHost's Default High-Numbered Port Mechanism: when BandwagonHost reloads a new system image through KiwiVM, it usuallyWill Not将 SSH 绑定在 Standard 22 端口,而是由系统安装脚本随机分配一个高位端口(如
28765), which is displayed after reinstallation and sent by email. If you habitually connect to port 22 when troubleshooting, you will encounterConnection refused。 - DMIT's Default Policy: DMIT's official preconfigured images use standard port 22 by default. If port 22 is DROPped or RST in mainland China but appears open from overseas probes, only that port on the IP is blocked. Use the out-of-band Console to change SSH to a nonstandard high-numbered port (such as
40000-65535within that range):
# 登录带外 Console,修改 SSH 配置文件
sed -i 's/^#\?Port 22/Port 48222/' /etc/ssh/sshd_config
# 验证配置语法有效性
sshd -t
# 重启 SSH 守护进程
systemctl restart sshd || systemctl restart ssh第 3 层: Operating system 内部监听状态与本地防火墙(Host 层)#
If the probe result is Connection refused, indicating that the VPS physically received the packet but its kernel network stack actively rejected it.
1. 检查套接字绑定范围(Socket Binding)#
Even if the service is running, if it listens only on the loopback interface (127.0.0.1), public network 接口将无法接入:
# 查看系统内所有监听中的 TCP 套接字及关联进程
ss -tulpn | grep -E 'sshd|nginx|caddy'- Abnormal Output:
LISTEN 0 128 127.0.0.1:22or127.0.0.53:53. This indicates that the service is bound only to the local loopback interface and cannot receive public internet traffic. - Normal Output:
LISTEN 0 128 0.0.0.0:22orLISTEN 0 128 [::]:22(listening on all IPv4/IPv6 addresses). - 修复方法: Check
/etc/ssh/sshd_configinListenAddressconfiguration, explicitly declaring it asListenAddress 0.0.0.0, then restart the service.
2. Audit Local Firewall Rules#
on Debian/Ubuntu ufw 以及 CentOS/AlmaLinux 上的 firewalld or the underlying iptables can easily be misconfigured during initialization or panel setup to drop incoming traffic by default:
# 1. 检查 ufw 状态与放行规则
ufw status verbose
# 若发现 ufw 处于 active 且默认 incoming 为 deny,临时放行目标端口:
ufw allow 48222/tcp
# 2. 检查 iptables INPUT 链规则与计数器
iptables -L INPUT -n -v --line-numbers
# 关键线索:若出现 DROP 策略且报文匹配数(pkts 计数器)在持续增加,
# 表明入站流量正在被本机过滤。排查阶段可临时插入放行规则至链首:
iptables -I INPUT 1 -p tcp --dport 48222 -j ACCEPTLayer 4: MTU Blackholes and PMTUD Negotiation Failure (The Packet Fragmentation Trap)#
This is the most elusive and damaging protocol-level fault encountered in troubleshooting, yet it is often misdiagnosed as an “unstable connection.”
1. Typical Failure Symptoms#
If your VPS shows the following unusually contrasting symptoms, there is a 99% chance it has encountered Path MTU (Path Maximum Transmission Unit) Blackhole:
- Interaction Initially Works, but Freezes as Soon as Enter Is Pressed:SSH 能够正常提示输入密码,完成密钥认证;但一旦输入密码进入终端,执行
ls -la /usr/lib、ip aorcata longer block of text, the cursor immediately freezes and keystrokes produce no output until the connection times out. - Small Pages Load Instantly, but Large Resources or TLS Negotiation Hang: When accessing a web service hosted on the VPS through a browser, plain text or HTTP 301 redirects return instantly, but an HTTPS site with a large certificate chain remains stuck at
Client HelloandServer Hello; alternatively, a file download may drop to zero speed and stall after transferring tens of KB.
2. Failure Mechanism: Why Does PMTUD Fail?#
- Standard Ethernettypically has an MTU of
1500bytes. - Premium cross-border dedicated routes, including BandwagonHost's China Telecom CN2 GIA/China Unicom AS9929 and DMIT's CMIN2/CN2 optimized return paths, commonly apply the following at the infrastructure level: MPLS 标签嵌套、GRE 隧道或 IPsec 加密封装. This leaves some intermediate routers on cross-border links with an effective MTU below 1500 (for example, only
1450or1420)。 - When the client sends packets exceeding this limit with the following flag set: DF(Don't Fragment,不可分片) flagged TCP packets, an intermediate router cannot forward them directly. It must drop them and, as required by the RFC, send the source a
ICMP Type 3, Code 4(Fragmentation Needed,需要分片并给出下一跳 MTU)报文。 - How the Critical Problem Arises: Many routers along the public path or the user's local firewall block all ICMP for “security,” preventing these notification packets from reaching the sender. The sender assumes large packets are being lost and repeatedly retransmits after timeouts; intermediate nodes keep dropping them, creating a complete communication deadlock. This is Path MTU Discovery(PMTUD)黑洞。
[客户端] (MTU 1500) [中间隧道路由器] (MTU 1420) [海外 VPS]
│ │ │
├─── 1. TCP 握手 (SYN 小包 64字节, 携带 MSS=1460) ──────────►│─── 顺利通过 ─────────────────────────────►│
│◄── 2. TCP 握手 (SYN-ACK 小包 64字节) ◄─────────────────────│◄── 顺利通过 ──────────────────────────────┤
├─── 3. 连接建立,SSH 交互少量字节 ────────────────────────►│─── 顺利通过 ─────────────────────────────►│
│ │ │
│ │◄── 4. 执行 ls / 传输大证书 (1500字节, DF=1) ─┤
│ │ │
│ [丢弃大包 1500 > 1420] │
│ │ │
│ X ◄── 5. 发送 ICMP Type 3 Code 4 (沿途被防火墙丢弃) ───┤ │
│ │ │
│ [客户端无感知,持续等待] │
│ [VPS 超时重传,再次被丢] │
│ [最终连接僵死、超时中断] │3. Precisely Determine the Link's Path MTU#
From your local terminal, send probe packets with the DF flag to the VPS IP, then use a binary search to determine the largest packet size that can traverse the path:
# macOS 探测命令 (-D 设置不可分片,-s 指定载荷字节)
ping -D -s 1472 154.xx.xx.xx
# Linux 探测命令 (-M do 设置不可分片,-s 指定载荷字节)
ping -c 4 -M do -s 1472 154.xx.xx.xx
# Windows 探测命令 (-f 设置不可分片,-l 指定载荷字节)
ping -f -l 1472 154.xx.xx.xx注:实际 MTU = ICMP 载荷长度 + 20 字节(IP 头)+ 8 字节(ICMP 头)。当测试 1472 提示 Packet needs to be fragmented but DF set , gradually reduce the value (for example, test 1400、1372), until replies are received normally. If the maximum working payload is 1392, the path's actual physical MTU is only 1420。
4. Production-Grade Permanent Solutions#
Method A: Deploy TCP MSS Clamping on the VPS (Automatic Clamping, Recommended First Choice)#
Without modifying the client, forcibly rewrite the MSS (Maximum Segment Size) advertised during the TCP handshake at the VPS kernel's forwarding layer, forcing persistent connections in both directions to use smaller packets:
# 利用 iptables mangle 表,在 SYN 协商阶段强制将 MSS 钳制为路径 MTU 期望值
iptables -t mangle -A POSTROUTING -p tcp --tcp-flags SYN,RST SYN -j TCPMSS --clamp-mss-to-pmtu
# 针对特别恶劣的封装链路,可强行硬编码锁定 MSS 为保守安全值 (如 1360 字节):
iptables -t mangle -A POSTROUTING -p tcp --tcp-flags SYN,RST SYN -j TCPMSS --set-mss 1360
# 安装 iptables-persistent 固化规则,确保重启不丢失
apt-get install -y iptables-persistent && netfilter-persistent saveOption B: Directly Adjust the VPS's Primary Network Interface MTU#
Reduce the MTU of the system's physical network interface to a safe range, such as 1420 or 1380):
# 1. 查找当前主网卡名称(通常为 eth0 或 ens3)
ip route show | grep default
# 输出示例: default via 154.xx.xx.1 dev eth0 onlink
# 2. 临时下调网卡 MTU
ip link set dev eth0 mtu 1420
# 3. 验证是否生效
ip link show eth0If you are using a newer version of Ubuntu/Debian, persist the configuration in Netplan (/etc/netplan/ the corresponding yaml file) or /etc/network/interfaces:
# Netplan 示例: /etc/netplan/50-cloud-init.yaml
network:
version: 2
ethernets:
eth0:
dhcp4: true
mtu: 1420After saving, run netplan apply to make it permanent.
Comprehensive Practical Troubleshooting Quick Reference#
综合上述各层级排障逻辑,遇到异常时可对照下表进行对号入座与快速处置:
| 故障现象 | Key Console Symptoms | Root Cause Isolation | Definitive Resolution Path |
|---|---|---|---|
| Local Terminal Reports Connection timed out | instance 状态显示 Running, but ping.pe Shows All Mainland China Nodes in Red and All Overseas Nodes in Green | Public IP Blackholed on Cross-Border Routes (GFW Blocking) | BandwagonHost: Do not force a data center switch; request a replacement with a usable IP through the control panel; DMIT: Check whether you are within the 3-day/30GB new-purchase refund window, or submit an IP replacement request under the policy. |
| 本地终端提示 Connection refused | The Instance Is Running and Responds Normally to ICMP Ping from Mainland China | 1. Connecting to the wrong port (such as port 22); 2. The target service is not running; 3. The service listens only on 127.0.0.1 | 1. If it is BandwagonHost, check the provisioning email for the randomly assigned high-numbered port; 2. Log into the Out-of-Band Console and Run ss -tulpn confirm that the port is listening;3. Run systemctl restart sshd。 |
| SSH Key Login Reports Permission denied (publickey) | The Instance Is Running and the TCP Handshake Succeeds | The Remote Host Lacks the Corresponding Public Key, or Password Login Is Disabled | If it is DMIT, since passwords are disabled by default, you need to use the client area Access reset the public key or password in the panel,Then Be Sure to Perform a Hard Reboot of the Instancetrigger key synchronization. |
| SSH Freezes After Connecting, or the Cursor Stops During Commands with Large Output | ICMP Works, the TCP Port Is Open, and Small-Packet Handshakes Succeed | 跨国优质 Networks (CN2 GIA/AS9929)经过隧道封装引发 Path MTU Blackhole | Run on the VPS:iptables -t mangle -A POSTROUTING -p tcp --tcp-flags SYN,RST SYN -j TCPMSS --clamp-mss-to-pmtu或将网卡 MTU 调整为 1420。 |
| 内外网所有节点均无法 Ping 通,TCP 亦全军覆没 | 控制台状态显示未知或 instance 已停止 | 1. Suspension Due to an Overdue Invoice; 2. Physical host failure; 3. Linux Kernel Crash (Kernel Panic) | 1. Check whether the invoice is overdue (both providers generate invoices 7 days before expiration; note that DMIT destroys data after 3 days overdue); 2. Open Interactive Console / VNC, inspect on-screen kernel errors, and restart. |