Systems and Operations

Building an Overseas Cloud Server Backup and Disaster Recovery (DR) System: The 3-2-1 Rule, Rclone Off-Site Cold Backups, and Snapshots in Practice

Compiled by the VPSMap Editorial Team · Updated 2026-09-26 · 29-minute read · Plain-Text Version

In production operations on overseas cloud servers (VPS), the risks of service interruption and data loss arise not only from system failures and human error but also from uncertainties including cross-border network fluctuations, upstream data center routing failures, abnormal IP connectivity, and provider lifecycle policies. Backup plans that have never been tested in practice often prove useless when a real disaster strikes.

Building a highly available, disaster-resilient Disaster Recovery (DR) system with dependable recovery requires moving beyond excessive reliance on a single cloud infrastructure point. Based on the 3-2-1 data protection principle, this article examines the practical differences between BandwagonHost and DMIT in underlying architecture, control panel permissions, and instance operations. Combining production-grade MySQL logical exports without table locks with Rclone end-to-end encrypted off-site cold backups, it provides a complete, directly deployable disaster recovery solution.


One. Disaster Recovery Models for Overseas Cloud Hosting and Implementing the 3-2-1 Rule#

The classic 3-2-1 data protection principle requires:

  1. 3 Complete Copies of the Data:一份生产数据,两份备份副本;
  2. 2 Different Media Types / Storage Systems: For example, a high-performance local NVMe array and independent cloud object storage;
  3. 1 Off-Site Cold-Backup Copy:物理机房、地理 Region 甚至上游云厂商层面的完全隔离。
Commands / Configuration
                              ┌── 副本 1 (生产源):VPS 本地磁盘 (/var/www, /var/lib/mysql)
                              │
3-2-1 容灾架构拓扑 ───────────┼── 副本 2 (机房近线):宿主机底层快照 (KiwiVM / DMIT 快照)
                              │
                              └── 副本 3 (离岸冷备):跨厂商加密对象存储 (Cloudflare R2 / AWS S3)

为什么机房快照不能替代异地冷备?#

Many operations engineers treat a provider panel's “one-click snapshot” as an all-purpose safeguard. However, in certain overseas cloud hosting scenarios, relying solely on data center snapshots creates a serious single point of failure:

  1. 同机房物理故障与 Storage 池崩溃: Provider snapshot volumes are usually stored alongside cloud servers in the same data center's large distributed storage pool (Ceph/SAN). A power outage, severed fiber link, or storage hardware failure can make snapshots unreachable together with production disks.
  2. Instance Lifecycle and Billing Policy Risks: Providers enforce very strict data-cleanup policies for overdue instances.
    • DMIT billing system generates renewal invoices 7 days before an instance expires. If payment is overdue, there is usually only a 3-day grace period after suspension, after which the data is automatically and permanently destroyed.
    • BandwagonHost Invoices are issued 7 days before expiration. By default, no unauthorized automatic charge is initiated against a credit card or PayPal account (automatic payment is possible only when the account balance is sufficient). If payment becomes overdue because an overseas credit card expires or an email is missed, the underlying snapshots are also deleted when the server is reclaimed.
    • Under both providers' Terms of Service, once a refund is approved, the associated instance and snapshot data are immediately deleted by the system.
  3. Blocked IP Connectivity and Migration Restrictions: Under BandwagonHost's operating model, if an instance IP is blocked or blacklisted, the panel's Data Center Migration feature is restricted. Without an independent external cold backup, administrators may be unable either to back up normally or to migrate data centers without data loss.

Therefore,Data Center Snapshots Provide “Failure Rollback in Seconds and Deployment Protection”; Off-Site Cold Backups Provide “Protection Against Physical Destruction and Account-Level Disasters”, the two complement each other and neither should be neglected.


Two. The Provider's First Line of Defense: Snapshot Mechanisms and Differences in Emergency Panel Operations#

Before implementing automation scripts, fully understand the management tools and security boundaries of the underlying hosting platform.

DimensionBandwagonHost ( BandwagonHost )DMIT
Infrastructure Control PanelIndependent, Proprietary KiwiVM System基于现代化 Client Area 的 instance 控制台
Secure Credential IsolationKiwiVM 管理密码、客户中心账户密码、系统 OS root 密码三者彼此独立Remote root Password Login Disabled by Default, with SSH Key Pairs and Separate Panel Access Controls
Snapshot Policy and Retention支持全盘快照,系统默认保留 30 days;支持手动勾选“Sticky”永久锁定Snapshot/backup support depends on the specific plan and is managed through the instance panel
Recovery Console for an Unreachable ServerKiwiVM 提供的 Interactive Console (VNC/Serial)Emergency Console on the Instance Management Page
Connectivity Restriction Policy目标 IP 状态异常时,跨机房迁移功能受阻Network Priority Is Strictly Defined by Plan; Key Changes Require an Instance Restart to Take Effect

1. BandwagonHost (KiwiVM) Production Operations Guidelines#

  • Sticky Snapshot Retention: Standard KiwiVM snapshots are automatically rotated out and deleted after 30 days. Before high-risk system changes, such as major Linux kernel or database upgrades, mark the baseline snapshot in the panel as Sticky(受保护状态),防止其在容灾周期内被自动化机制 Clear 。
  • Three-Layer Credential Isolation:KiwiVM 控制台采用独立的面板密码体系。即使业务系统的 root 密码由于漏洞外泄,攻击者若无 KiwiVM 独立凭据,仍无法 Passed 底层直接销毁机器;同样,在主机系统完全挂死、SSH 拒绝服务时,运维人员需 Passed 独立 KiwiVM 凭据登录并启动 Interactive Console to perform recovery.

2. DMIT Instance Access Control and Emergency Access#

  • When SSH Key Changes Take Effect:DMIT 的云主机在交付初始化时默认关闭密码远程登录,全面采用 SSH Public Key 机制。在 DMIT instance 面板的 Access section to replace or inject a new public key, the underlying metadata is not immediately applied to the currently running memory environment,You Must Cold-Reboot the Instance (Reboot) as Instructed by the Panel, allowing the system configuration script to reload the key file.
  • Recover an Unresponsive Server with Emergency Console: if you misconfigure ufw / nftables 防火墙或误修改 /etc/ssh/sshd_config causes SSH to disconnect, go directly to the DMIT console and open the Web Console, then use the privileged terminal to inspect network interface configuration and security policies.

Three. The Second Line of Defense: Production MySQL Logical Exports Without Table Locks#

Never copy underlying files directly from a running production database (such as directly compressing /var/lib/mysql). Without a maintenance window, directly copying files being written concurrently can corrupt data files, damage pages (Page Corruption), and create transaction inconsistencies.

The standard production-grade approach uses mysqldump Works with InnoDB's transaction snapshot reads to provide a Non-locking export.

1. Key Export Parameters Explained#

  • --single-transaction: Creates a Consistent Read snapshot, allowing the entire export to proceed without blocking external application reads or writes (InnoDB only).
  • --quick: Makes mysqldump retrieve records from the server row by row instead of first loading the entire result set into memory, avoiding the OOM Killer on low-memory overseas VPS instances.
  • --routines --triggers --events: Export all stored procedures, triggers, and scheduled events to preserve the complete application logic after migration.
  • --master-data=2(for primary-replica replication): record the Binlog Position at the time of export to support subsequent Point-in-Time Recovery (PITR).

2. Highly Reliable Local Hot-Backup Script:/usr/local/bin/backup_db.sh#

bash
#!/usr/bin/env bash
set -euo pipefail

# ==============================================================================
# MySQL/MariaDB 生产级无锁热备脚本
# ==============================================================================
export PATH="/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin"

# 基础环境配置
BACKUP_ROOT="/data/backups/mysql"
DATE_STR="$(date +'%Y%m%d_%H%M%S')"
TARGET_DIR="${BACKUP_ROOT}/${DATE_STR}"
RETENTION_DAYS=7

# 数据库安全连接凭证配置 (建议通过 ~/.my.cnf 或独立权限账号连接)
DB_USER="dr_backup"
DB_PASS="YourSuperSecureBackupPassword"
DB_NAME="production_db"

# 创建安全隔离的归档目录
mkdir -p "${TARGET_DIR}"
chmod 700 "${TARGET_DIR}"

DUMP_FILE="${TARGET_DIR}/${DB_NAME}_${DATE_STR}.sql.gz"
SHA_FILE="${TARGET_DIR}/${DB_NAME}_${DATE_STR}.sha256"

echo "[$(date +'%F %T')] [INFO] 开始执行数据库 [${DB_NAME}] 事务级热备..."

# 执行流式导出与动态 gzip 压缩
mysqldump \
    --user="${DB_USER}" \
    --password="${DB_PASS}" \
    --host="127.0.0.1" \
    --single-transaction \
    --quick \
    --routines \
    --triggers \
    --default-character-set=utf8mb4 \
    "${DB_NAME}" | gzip -9 > "${DUMP_FILE}"

# 生成文件校验和,确保备份介质完整性
sha256sum "${DUMP_FILE}" > "${SHA_FILE}"

echo "[$(date +'%F %T')] [INFO] 导出成功: ${DUMP_FILE}"
echo "[$(date +'%F %T')] [INFO] SHA256 校验和已写入: ${SHA_FILE}"

# 清理本地超过保留周期的历史归档
find "${BACKUP_ROOT}" -mindepth 1 -maxdepth 1 -type d -mtime +"${RETENTION_DAYS}" -exec rm -rf {} +
echo "[$(date +'%F %T')] [INFO] 本地超 ${RETENTION_DAYS} 天旧归档已清理完成。"

Four. The Third Line of Defense: Build End-to-End Encrypted Off-Site Cold Backups with Rclone#

本地生成的压缩包依然暴露在单机物理损毁的风险之中。利用 rclone Uploading encrypted backups to object storage across the ocean, such as Cloudflare R2 or AWS S3, is crucial for protecting against physical failures and billing-related shutdowns.

1. Overseas Data Center Transfer Characteristics and Traffic Policy Trade-Offs#

在选择异地传输路径时,必须考虑 Providers 的流量计费模式:

  • DMIT premium-network products, such as the Pro / EB series, provide exceptionally high-quality CN2 GIA or CMIN2 bandwidth optimized for China. However, monthly allowances are usually fixed, and overage throttling or blocking policies vary by plan.切勿 Passed 昂贵的回国优化通道将 GB 级别的冷备传回境内服务器。
  • Cloudflare R2 offers excellent public-network connection speeds to overseas VPS locations such as the US West Coast, Hong Kong, China, and Japan, with native S3 compatibility, andZero Egress Fees,是境外云主机 Storage 异地冷备的高性价比方案。

2. Rclone Crypt (Transparent Client-Side End-to-End Encryption)#

为了防止对象 Storage 凭据泄露导致明文 Databases 外泄,必须在传输前由本地客户端完成 AES-256-GCM 加密。即使云 Storage 节点被非法入侵,攻击者获取的也仅是无法破解的高熵乱码。

Configuration Steps:#

Run rclone config, create two levels of Remote:

  1. Base Storage Layer (Such As r2-raw):配置 S3 compatible 协议,填入 Cloudflare R2 的 Account ID、Access Key ID 与 Secret Access Key。
  2. a secure encryption layer (such as r2-secure):
    • Choose the Storage Type crypt;
    • The Target Points to a Path Within the Base-Layer Bucket:remote = r2-raw:vps-disaster-recovery/node-backup;
    • configure strong filename encryption (filename_encryption = standard) and directory-name encryption;
    • Enter and securely store the encryption key Passphrase and Salt.

3. Full Off-Site Cold-Backup Orchestration Script:/usr/local/bin/sync_disaster_recovery.sh#

bash
#!/usr/bin/env bash
set -euo pipefail

# ==============================================================================
# 全量异地容灾冷备编排脚本
# ==============================================================================
export PATH="/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin"

WORK_DIR="/data/backups"
APP_DIR="/var/www/production_app"
NGINX_CONF="/etc/nginx"
DATE_STR="$(date +'%Y%m%d_%H%M%S')"
TMP_ARCHIVE_DIR="${WORK_DIR}/tmp_${DATE_STR}"
LOG_FILE="/var/log/dr_sync.log"

mkdir -p "${TMP_ARCHIVE_DIR}"

log() {
    echo "[$(date +'%F %T')] $1" | tee -a "${LOG_FILE}"
}

cleanup() {
    rm -rf "${TMP_ARCHIVE_DIR}"
}
trap cleanup EXIT

log "[START] 启动全量容灾数据归档流程..."

# 1. 触发数据库热备
/usr/local/bin/backup_db.sh

# 2. 打包应用源码与核心环境配置
log "[ARCHIVE] 正在打包 Web 静态资产与配置文件..."
tar -czpf "${TMP_ARCHIVE_DIR}/app_code_${DATE_STR}.tar.gz" -C "${APP_DIR}" .
tar -czpf "${TMP_ARCHIVE_DIR}/nginx_conf_${DATE_STR}.tar.gz" -C "${NGINX_CONF}" .

# 3. 将本地数据推送到 Rclone 加密远程存储
# 关键参数解释:
# --transfers: 控制并发传输连接数,防止占用过多 VPS 内存
# --bwlimit: 施加带宽上限,避免瞬时吃满端口影响生产流量
# --checksum: 基于文件哈希验证是否已存在相同文件,跳过重复上传
log "[SYNC] 正在加密上传至异地对象存储..."
rclone copy "${WORK_DIR}/mysql" "r2-secure:mysql" \
    --checksum \
    --transfers=2 \
    --bwlimit="15M" \
    --log-file="${LOG_FILE}" \
    --log-level=INFO

rclone copy "${TMP_ARCHIVE_DIR}" "r2-secure:system_assets/${DATE_STR}" \
    --checksum \
    --transfers=2 \
    --bwlimit="15M" \
    --log-file="${LOG_FILE}" \
    --log-level=INFO

# 4. 异地冷备生命周期管理:仅保留远端 30 天以内的历史备份
log "[PURGE] 正在清理异地存储超过 30 天的失效冷备..."
rclone delete --min-age 30d "r2-secure:mysql" || true
rclone delete --min-age 30d "r2-secure:system_assets" || true
rclone rmdirs --leave-root "r2-secure:system_assets" || true

log "[SUCCESS] 全量容灾同步任务顺利收敛。"

4. Automated Task Scheduling: Use a Systemd Timer Instead of Crontab#

In modern Linux operations, the recommended tool is systemd.timer instead of the traditional cron, with the advantage of capturing complete standard output and error logs in journalctl, supports dependency configuration and precise control over retry policies after task failures.

Create the Service Unit /etc/systemd/system/dr-backup.service:#

ini
[Unit]
Description=Automated Disaster Recovery Backup and Cloud Sync
After=network.target

[Service]
Type=oneshot
User=root
ExecStart=/usr/local/bin/sync_disaster_recovery.sh
StandardOutput=journal
StandardError=journal

创建定时器单元 /etc/systemd/system/dr-backup.timer:#

ini
[Unit]
Description=Run Disaster Recovery Backup at 03:30 AM Daily

[Timer]
OnCalendar=*-*-* 03:30:00
RandomizedDelaySec=600
Persistent=true

[Install]
WantedBy=timers.target

Enable and start the timer:

bash
systemctl daemon-reload
systemctl enable --now dr-backup.timer
systemctl list-timers --all | grep dr-backup

五、 灾难恢复演练 SOP(Recovery Verification)#

There is an iron rule in disaster recovery:An Unverified Backup Is No Backup at All. During a system crash, panic, missing recovery dependencies, or lost decryption credentials can easily cause a second incident. Production teams must establish standardized quarterly emergency recovery drills (SOPs).

Commands / Configuration
           [灾难发生: 实例损毁/逾期清退]
                        │
                        ▼
       [在备用节点初始化全新的干净 Linux 实例]
                        │
                        ▼
     [拉取 rclone.conf 并挂载 r2-secure 加密存储]
                        │
                        ▼
   [同步指定时间戳的 DB 与 Web 资产归档至临时工作区]
                        │
                        ▼
     [核对 SHA256 校验和 ── 校验失败? ──> 报警人工介入]
                        │ 校验通过
                        ▼
    [导入 MySQL 实例] ──> [还原 Nginx/Web 代码目录]
                        │
                        ▼
             [启动应用与数据库引擎]
                        │
                        ▼
            [业务可用性与数据完整性校验]
                        │
                        ▼
             [切换 DNS 解析,完成恢复]

Drill Scenario: Simulate Complete Destruction of the Production VPS and Fully Rebuild the Service on a New Node#

Step 1: Initialize the New Node and Prepare Credentials#

在任意可用的备用节点(无论是临时重装后的当前主机,还是不同 Providers 的新建 instance )准备最简环境:

bash
# 安装基础工具链
apt-get update && apt-get install -y curl tar gzip mysql-client rclone

# 将离线保管的安全凭证还原至本地
mkdir -p ~/.config/rclone
# 部署经过严格密钥保管的 rclone.conf (务必包含正确的 crypt 密码)
chmod 600 ~/.config/rclone/rclone.conf

Step 2: Retrieve the Backup for a Specific Date from Encrypted Storage#

bash
RECOVERY_TARGET_DATE="20260925_033000"
RESTORE_DIR="/tmp/dr_restore"
mkdir -p "${RESTORE_DIR}"

# 拉取数据库 dump 及对应的 SHA256 校验和文件
rclone copy "r2-secure:mysql/${RECOVERY_TARGET_DATE}" "${RESTORE_DIR}/db"

# 严格执行哈希校验,防止传输中出现比特翻转或不完整写入
cd "${RESTORE_DIR}/db"
sha256sum -c *.sha256

Troubleshooting Notes: If sha256sum Report FAILED, indicating that the object was corrupted during transfer or writing. Do not continue the restore; immediately retrieve the previous day's backup archive for comparison.

Step 3: Rebuild the Database and Import Data#

bash
# 假设目标环境已完成 MySQL/MariaDB 安装并启动
TARGET_DB="production_db"
MYSQL_ROOT_PASS="NewHostSecureRootPassword"

# 创建干净的目标库
mysql -uroot -p"${MYSQL_ROOT_PASS}" -e "
    DROP DATABASE IF EXISTS ${TARGET_DB};
    CREATE DATABASE ${TARGET_DB} CHARACTER SET utf8mb4 COLLATE utf8mb4_unicode_ci;
"

# 解压并流式导入数据
zcat "${RESTORE_DIR}/db/production_db_"*.sql.gz | mysql -uroot -p"${MYSQL_ROOT_PASS}" "${TARGET_DB}"

Step 4: Mount Static Assets and Configuration Files#

bash
# 拉取系统环境资产
rclone copy "r2-secure:system_assets/${RECOVERY_TARGET_DATE}" "${RESTORE_DIR}/assets"

# 还原 Web 站点文件,并恢复文件所有权
mkdir -p /var/www/production_app
tar -xzpf "${RESTORE_DIR}/assets/app_code_"*.tar.gz -C /var/www/production_app/
chown -R www-data:www-data /var/www/production_app/

# 还原并检查 Nginx 规则
tar -xzpf "${RESTORE_DIR}/assets/nginx_conf_"*.tar.gz -C /etc/nginx/
nginx -t && systemctl reload nginx

Step 5: Data Consistency and Transaction Validation Checklist#

After completing restoration, perform logical integrity checks inside the database. Do not simply change DNS and put the service live:

  1. 行数比对: Query core application tables (such as users、orders、posts) of SELECT COUNT(*), and compare it with monitoring statistics or business metrics from before the disaster;
  2. Verify the Maximum Primary Key Value:执行 SELECT MAX(id), MAX(created_at) FROM orders;,确认最后一次入库的业务记录与灾备设定的恢复点目标(RPO)一致;
  3. Verify System Permissions and Service Startup: Start application daemons such as PHP-FPM / Python Gunicorn / Docker Containers, then use curl -I http://127.0.0.1 Confirm that the page returns 200 and that the logs contain no database connection-pool errors;
  4. Clean Up After the Recovery Drill: after the test passes, securely erase /tmp/dr_restore directory to prevent sensitive production data from remaining on temporary disks.

Six. Comprehensive Disaster Recovery Checklist (DR Checklist)#

To keep the disaster recovery system healthy and ready for immediate use, perform a quick monthly review using the checklist below:

  • Billing and Lifecycle:
    • 掌握各 instance 账单日(两家 Providers 均在到期前 7 days生成账单);
    • 账户预存足够余额或设置高优先级提醒,规避 DMIT 逾期 3 days即自动销毁数据等严苛策略。
  • Keep Credentials Offline:
    • BandwagonHost's separate KiwiVM management password has been recorded in the team's password vault;
    • The private key for DMIT SSH access has an offline cold backup, and console Emergency Access is available;
    • Store the Rclone Crypt Passphrase and Salt in a secure password manager physically isolated from the production cloud server.
  • Underlying Snapshot Status:
    • Baseline snapshots for critical nodes in KiwiVM are marked “Sticky” and have not been deleted by automatic 30-day rotation;
    • An immediate snapshot has been taken before any major system upgrade or network policy change.
  • Cold-Backup Data and Network Health:
    • Local /data/backups Disk usage is normal, expired archives are automatically reclaimed, and the disk has not filled up (No Space Left on Device);
    • systemctl status dr-backup.timer shows normal scheduling with no unexplained silent exits;
    • A complete copy has been created in the off-site storage bucket of .sql.gz and its corresponding .sha256 checksum files;
    • 上传任务配置了限速与并发控制,未对 DMIT 等有限配额主机的month度流量造成非预期透支。
  • 恢复确定性:
    • At Least One End-to-End Recovery Drill Has Been Completed in the Last 90 Days, Retrieving Data from R2 Object Storage and Rebuilding Services on an Empty Node.