I depend on Internet access. It should always be fast and it should always work. For a few years a dual WAN Cisco router took care of that, with a Ziggo cable line and a KPN fiber line. I love Debian, so when I replaced it I chose an N100 mini PC with six gigabit ports and an NVMe disk, and turned it into a router with Debian 13. This post describes the setup and the things that went wrong along the way.

An N100 mini PC with six ports replaces the Cisco

The N100 has more than enough power to route a few gigabit lines and it uses little power. Three of the six ports are in use:

Interface Connected to
enp1s0 Ziggo modem (bridge mode)
enp2s0 KPN fiber media converter
enp5s0 16-port PoE LAN switch

The ports enp3s0, enp4s0 and enp7s0 are spare. The network configuration is done with ifupdown. The main /etc/network/interfaces file only sets up the loopback interface and includes a directory:

source /etc/network/interfaces.d/*

auto lo
iface lo inet loopback

In /etc/network/interfaces.d there is one file per port, numbered after the port on the case: 0_ziggo, 1_kpn and 4_lan. The numbers keep the files in port order and make it easy to see which cable goes where.

Ziggo is a plain DHCP interface

The Ziggo modem hands out the public address with DHCP, so /etc/network/interfaces.d/0_ziggo is short:

# WAN - ETH0

allow-hotplug enp1s0
iface enp1s0 inet dhcp
iface enp1s0 inet6 auto

On my Debian 13 install the DHCP client is dhcpcd and not dhclient. That matters later on, as the hooks and the configuration file of dhcpcd are in different places.

KPN needs VLAN 6, PPPoE and a larger MTU

KPN delivers internet on VLAN 6 and wants you to log in with PPPoE. On Debian you need the ppp and vlan packages for that, and ifup fails with unclear errors until they are installed:

sudo apt install ppp vlan

/etc/network/interfaces.d/1_kpn sets up the physical port, the VLAN on top of it and the PPPoE session on top of the VLAN:

# WAN - ETH1

# Physical port to the media converter
auto enp2s0
iface enp2s0 inet manual
    mtu 1512

# KPN internet VLAN
auto enp2s0.6
iface enp2s0.6 inet manual
    vlan-raw-device enp2s0
    mtu 1508

# PPPoE session
auto kpn
iface kpn inet ppp
    provider kpn

PPPoE adds an 8 byte header to every packet. With a normal MTU of 1500 on the VLAN that leaves 1492 bytes for the PPPoE session. KPN supports “baby jumbo frames” (RFC 4638), so I raised the MTU of the VLAN to 1508 and of the port to 1512 (the VLAN tag needs 4 more bytes). This gives the PPPoE session a full 1500 bytes, just like the Ziggo line.

The session itself is configured in /etc/ppp/peers/kpn:

plugin pppoe.so
nic-enp2s0.6

user "internet"
password "internet"
noauth
hide-password

mtu 1500
mru 1500

noipdefault
#defaultroute
#usepeerdns
+ipv6

persist
maxfail 0
holdoff 5
lcp-echo-interval 10
lcp-echo-failure 6

noaccomp
nopcomp
novj
novjccomp
nodeflate
nobsdcomp

KPN does not check the username and password, “internet” for both is fine. The options persist, maxfail 0 and holdoff 5 make pppd reconnect forever, with 5 seconds between attempts. The LCP echo options make pppd send a keepalive every 10 seconds and give up on the session after 6 missed replies. I will get to the two options that are commented out below.

PPPoE must not take the default route

The first time the KPN session came up, pppd took the default route away from Ziggo, because of the defaultroute option. I commented it out, as I want to decide myself which line carries the traffic. The routing setup comes later in this post.

A router needs to forward packets, so /etc/sysctl.d/99-router.conf contains:

net.ipv4.ip_forward=1

Run sudo sysctl --system to load it without a reboot.

dnsmasq serves DHCP and DNS on the LAN

The LAN port has a static address in /etc/network/interfaces.d/4_lan:

# LAN ETH4

auto enp5s0
iface enp5s0 inet static
    address 192.168.1.1/24

dnsmasq hands out addresses on the LAN and answers DNS queries. It forwards those to Cloudflare and Google instead of to the DNS servers of one of the providers, so DNS keeps working whichever line is active. This is /etc/dnsmasq.d/lan.conf:

interface=enp5s0
bind-dynamic
dhcp-range=192.168.1.10,192.168.1.250,255.255.255.0,12h
dhcp-option=option:router,192.168.1.1
dhcp-option=option:dns-server,192.168.1.1
domain-needed
bogus-priv
no-resolv
server=1.1.1.1
server=8.8.8.8

nftables does the firewall and NAT for both lines

This is /etc/nftables.conf:

#!/usr/sbin/nft -f

flush ruleset

define ZIGGO = "enp1s0"
define KPN = "ppp0"
define WAN = { $ZIGGO, $KPN }
define LAN = "enp5s0"

table inet filter {
    chain input {
        type filter hook input priority filter; policy drop;

        ct state established,related accept
        ct state invalid drop
        iif lo accept

        # WAN: DHCP replies from Ziggo
        iifname $ZIGGO udp sport 67 udp dport 68 accept

        # LAN: DHCP and DNS (dnsmasq)
        iifname $LAN udp dport { 53, 67 } accept
        iifname $LAN tcp dport 53 accept

        # SSH on all interfaces, rate limited per source address
        tcp dport 22 ct state new meter ssh4 { ip saddr limit rate 10/minute burst 5 packets } accept
        tcp dport 22 ct state new meter ssh6 { ip6 saddr limit rate 10/minute burst 5 packets } accept

        # Ping and IPv6 control traffic
        meta l4proto { icmp, ipv6-icmp } accept
    }

    chain forward {
        type filter hook forward priority filter; policy drop;

        # Clamp TCP MSS to the route MTU (safety net for PPPoE)
        tcp flags syn tcp option maxseg size set rt mtu

        ct state established,related accept
        ct state invalid drop
        iifname $LAN oifname $WAN accept
    }

    chain output {
        type filter hook output priority filter; policy accept;
    }
}

table ip nat {
    chain postrouting {
        type nat hook postrouting priority srcnat; policy accept;
        oifname $WAN masquerade
    }
}

The WAN define is a set of both lines, so forwarding and NAT work for either of them. Only the DHCP rule uses ZIGGO, as that is the only line that uses DHCP. The MSS clamping rule normally does nothing, because the PPPoE session has a full 1500 byte MTU. If KPN ever falls back to 1492, it prevents TCP connections that hang as soon as a large packet is sent. It has to come before the accept rules, as an accept ends the evaluation of the chain.

I want to reach the router with SSH from home, so port 22 is open on all interfaces. SSH only accepts keys, and the meter limits every source address to 10 new connections per minute with a burst of 5. Note that every SYN counts, including the retries of your SSH client, so when you reconnect a few times in a row you may get a “Connection timed out” for a minute.

Apply the file with:

sudo nft -c -f /etc/nftables.conf && sudo systemctl restart nftables

The -c checks the file first, so a typo does not leave you with half a firewall.

When I started I had also installed ufw, and it ran next to nftables. A packet has to pass both rule sets, so this is confusing at best. When I purged ufw my SSH connections from outside stopped working. The allow 22 of ufw had been the only rule that let SSH in from the internet, my nftables configuration only allowed SSH from the LAN. So if you remove ufw, add the SSH rule to nftables first. I could still log in from the LAN to fix it.

You can test the KPN path before you send any real traffic over it, by routing a single address over ppp0 and running a traceroute from the LAN:

sudo ip route add 1.0.0.1 dev ppp0
traceroute -n 1.0.0.1
sudo ip route del 1.0.0.1 dev ppp0

The first hop after the router was the KPN peer address, 195.190.228.120, so forwarding and NAT over KPN worked.

Each line gets its own routing table

When the router sends a packet from its KPN address, that packet must leave through KPN, whatever the default route is. Ziggo drops packets that leave with a KPN source address and the other way around. This matters for SSH (the reply to a connection that came in on KPN must go out through KPN) and for the health checks of the failover script, which have to test each line separately.

The solution is a routing table per line, with a rule that selects the table by source address. On Debian 13 there is no /etc/iproute2 directory by default, the defaults live in /usr/share/iproute2. Your own table names go in /etc/iproute2/rt_tables.d:

sudo mkdir -p /etc/iproute2/rt_tables.d
printf '100 ziggo\n200 kpn\n' | sudo tee /etc/iproute2/rt_tables.d/wan.conf

The addresses of both lines are dynamic, so the tables and rules are filled by hooks. For Ziggo that is /etc/dhcpcd.exit-hook, which dhcpcd runs after every change in the lease:

# Keep routing table "ziggo" in sync with the DHCP lease on enp1s0
if [ "$interface" = "enp1s0" ]; then
    case "$reason" in
        BOUND|RENEW|REBIND|REBOOT)
            ip route replace "$new_network_number/$new_subnet_cidr" dev enp1s0 table ziggo
            ip route replace default via "${new_routers%% *}" dev enp1s0 table ziggo
            while ip rule del lookup ziggo 2>/dev/null; do :; done
            ip rule add from "$new_ip_address" lookup ziggo priority 100
            ;;
        EXPIRE|FAIL|RELEASE|STOP|NOCARRIER)
            ip route flush table ziggo
            while ip rule del lookup ziggo 2>/dev/null; do :; done
            ;;
    esac
fi

For KPN pppd runs the scripts in /etc/ppp/ip-up.d and /etc/ppp/ip-down.d. Here is my /etc/ppp/ip-down.d/kpn-table:

#!/bin/sh
[ "$PPP_IFACE" = "ppp0" ] || exit 0
ip route flush table kpn
while ip rule del lookup kpn 2>/dev/null; do :; done

I show the ip-up.d script in the next section, as it also sets the default route. Make both executable, and do not put a dot in the file name, because run-parts skips those files. You can test both lines from the router with:

ping -c3 -I <ziggo-address> 1.1.1.1
ping -c3 -I <kpn-address> 1.1.1.1
curl -4 --interface <kpn-address> ifconfig.me

The last command should print the KPN address.

KPN is the primary line because it is faster

At first the Ziggo line carried all traffic and KPN was the backup. Then I looked at the numbers: a ping to 1.1.1.1 takes 4 ms over KPN and 12 ms over Ziggo. After a reboot the KPN session was up within 13 seconds, while the Ziggo modem needed about 80 seconds to hand out a lease. So I made KPN the primary line.

The switching is done with route metrics. dhcpcd installs the Ziggo default route with metric 1002. The ip-up.d script adds a default route over ppp0 with metric 100, and the lowest metric wins. This is the updated /etc/ppp/ip-up.d/kpn-table:

#!/bin/sh
# Keep routing table "kpn" in sync with the PPPoE session
[ "$PPP_IFACE" = "ppp0" ] || exit 0
ip route replace default dev ppp0 table kpn
while ip rule del lookup kpn 2>/dev/null; do :; done
ip rule add from "$PPP_LOCAL" lookup kpn priority 200

# KPN is the primary line: lower metric than the Ziggo route from dhcpcd (1002)
ip route replace default dev ppp0 metric 100

When the PPPoE session drops, the ppp0 interface disappears and the kernel removes its routes, so traffic falls back to Ziggo without any script. When the session comes back, the hook adds the route again. The main table then looks like this:

default dev ppp0 scope link metric 100
default via <ziggo-gateway> dev enp1s0 proto dhcp src <ziggo-address> metric 1002 mtu 1500

Connections that the LAN had open over Ziggo were set up with the Ziggo address by NAT, and they break when the route changes. You can delete them from the connection tracking table, so the clients reconnect right away:

sudo conntrack -D -s 192.168.1.0 --mask-src 255.255.255.0

This only matches connections from the LAN. I first tried conntrack -D --src-nat, but that also deleted entries of incoming pings that were never translated.

The router’s own DNS must not depend on Ziggo

My first switch to KPN failed. The LAN worked, but curl ifconfig.me on the router itself hung. The cause was /etc/resolv.conf, which dhcpcd had filled with the DNS servers of Ziggo. Those do not answer queries from a KPN address:

dig +time=3 +tries=1 -b <kpn-address> @62.179.104.196 ifconfig.me
;; communications error to 62.179.104.196#53: timed out

The LAN did not suffer from this, as dnsmasq forwards to 1.1.1.1 and 8.8.8.8. So the fix was to let the router use dnsmasq as well. dnsmasq also listens on the loopback interface when you configure an interface, so /etc/resolv.conf becomes:

nameserver 127.0.0.1

To stop dhcpcd from overwriting this file, add this to /etc/dhcpcd.conf:

nohook resolv.conf

And for the same reason usepeerdns is commented out in the KPN peer file. Otherwise the Debian script /etc/ppp/ip-up.d/0000usepeerdns adds the KPN DNS servers to the file on every connect.

A running dhcpcd does not read its configuration again by itself. It rewrites /etc/resolv.conf on every network event, including the IPv6 router advertisements of Ziggo, so my edit was gone within minutes. dhcpcd -n enp1s0 reloads the configuration, but it also rebinds the Ziggo lease, and I was connected through Ziggo. Thanks to the routing tables I could SSH in through the KPN address instead, which does not depend on Ziggo at all.

A script handles a KPN line that is up but broken

The route metrics cover the case where the PPPoE session drops. They do not cover the case where ppp0 is up, but no traffic gets through. pppd also needs up to a minute (6 missed keepalives at 10 second intervals) before it gives up on a dead session. For these cases I wrote a small script, /usr/local/sbin/wan-failover:

#!/bin/sh
# KPN (ppp0) is the primary line, Ziggo (enp1s0) the backup.
# If the PPPoE session drops, the kernel removes the KPN route by itself.
# This script covers the other case: ppp0 is up, but KPN has no internet.

TARGETS="1.1.1.1 9.9.9.9 8.8.8.8"
INTERVAL=5      # seconds between checks
FAIL_AFTER=3    # failed KPN checks before switching to Ziggo
BACK_AFTER=6    # good KPN checks before switching back

# IPv4 address of an interface, empty if it has none
wan_ip() {
    ip -4 -o addr show "$1" 2>/dev/null | awk '{print $4}' | cut -d/ -f1
}

# A line is healthy when at least one target answers from its address.
# The source address selects the per-line routing table.
healthy() {
    src=$(wan_ip "$1")
    [ -n "$src" ] || return 1
    for t in $TARGETS; do
        ping -n -q -c 1 -W 2 -I "$src" "$t" >/dev/null 2>&1 && return 0
    done
    return 1
}

on_kpn() {
    ip route show default | grep -q 'dev ppp0'
}

after_switch() {
    # Drop LAN connections that are still NAT-ed to the old line
    conntrack -D -s 192.168.1.0 --mask-src 255.255.255.0 >/dev/null 2>&1
}

fails=0
goods=0
ziggo_ok=unknown

while true; do
    if healthy ppp0; then
        fails=0
        goods=$((goods + 1))
    else
        goods=0
        fails=$((fails + 1))
    fi

    if healthy enp1s0; then z=yes; else z=no; fi
    if [ "$z" != "$ziggo_ok" ]; then
        echo "Ziggo healthy: $z"
        ziggo_ok=$z
    fi

    if on_kpn; then
        if [ "$fails" -ge "$FAIL_AFTER" ]; then
            if [ "$ziggo_ok" = yes ]; then
                echo "KPN failed $fails checks, switching to Ziggo"
                ip route del default dev ppp0 metric 100
                after_switch
            elif [ "$fails" -eq "$FAIL_AFTER" ]; then
                echo "KPN failed $fails checks, but Ziggo is down too, staying on KPN"
            fi
        fi
    elif [ -n "$(wan_ip ppp0)" ] && [ "$goods" -ge "$BACK_AFTER" ]; then
        echo "KPN healthy for $goods checks, switching back to KPN"
        ip route replace default dev ppp0 metric 100
        after_switch
    fi

    sleep "$INTERVAL"
done

The script pings three targets from the address of each line, and a line is healthy when one of them answers. Because of the routing tables these checks work whether a line carries the default route or not. After three failed rounds it removes the KPN default route, so Ziggo takes over. A failed round takes about 11 seconds (5 seconds of sleep and three pings with a 2 second timeout), so that is about half a minute. It only switches back after KPN has been healthy for six rounds in a row, so a flapping line does not cause constant switching.

The script does not keep track of the active line itself, it reads the routing table every round. This keeps it in line with the ip-up.d hook, which adds the KPN route on every reconnect. When both lines are down, it does nothing and logs that once.

It runs as a systemd service, /etc/systemd/system/wan-failover.service:

[Unit]
Description=WAN failover (KPN primary, Ziggo backup)
After=network-online.target
Wants=network-online.target

[Service]
ExecStart=/usr/local/sbin/wan-failover
Restart=always
RestartSec=5

[Install]
WantedBy=multi-user.target

Enable it with:

sudo apt install conntrack
sudo systemctl daemon-reload
sudo systemctl enable --now wan-failover

The output of the script ends up in the journal, which you can follow with journalctl -u wan-failover -f.

After a reboot I looked at the boot log. The Ziggo link went down and up twice while the modem and the network card agreed on the link, and it took 80 seconds before Ziggo handed out a lease. In the meantime dhcpcd had given enp1s0 a link-local address (169.254.x.x) and a default route over it. That route leads nowhere. To prevent this, add this line to /etc/dhcpcd.conf:

noipv4ll

Dynamic DNS follows the active line

To reach the router with SSH I use a DNS name with a short TTL. A cron job on the router calls dynip.php, a small PHP script on my web server. The script takes the address that the request came from, and when it changed, updates the A record through the TransIP API:

<?php
$key = trim(file_get_contents('transip.key'));
$ip = $_SERVER['REMOTE_ADDR'];
$token = $_GET['token'];
if ($token != trim(file_get_contents('token.txt'))) {
    die('ko');
}
if (trim(file_get_contents('ip.txt')) == $ip) {
    die('ok');
}
function curlReq($method, $url, $body, $headers)
{
    $ch = curl_init();
    curl_setopt($ch, CURLOPT_URL, $url);
    curl_setopt($ch, CURLOPT_POSTFIELDS, json_encode($body));
    curl_setopt($ch, CURLOPT_HTTPHEADER, array_merge(['Content-Type: application/json'], $headers));
    curl_setopt($ch, CURLOPT_CUSTOMREQUEST, $method);
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
    $data = json_decode(curl_exec($ch), true);
    curl_close($ch);
    if ($data['error'] ?? '') {
        throw new \Exception($data['error']);
    }
    return $data;
}
function getToken($key)
{
    $method = 'POST';
    $url = 'https://api.transip.nl/v6/auth';
    $body = [
        'login' => 'tip_username',
        'nonce' => random_int(0, PHP_INT_MAX),
        'read_only' => false,
        'expiration_time' => '10 seconds',
        'label' => 'dynip',
        'global_key' => false,
    ];
    openssl_sign(json_encode($body), $signature, $key, OPENSSL_ALGO_SHA512);
    $headers = ['Signature: ' . base64_encode($signature)];
    $data = curlReq($method, $url, $body, $headers);
    return $data['token'];
}
function setDns($token, $ip)
{
    $method = 'PATCH';
    $url = 'https://api.transip.nl/v6/domains/domeinnaam.nl/dns';
    $body = [
        'dnsEntry' => [
            'name' => 'router',
            'expire' => 300,
            'type' => 'A',
            'content' => $ip,
        ]
    ];
    $headers = ['Authorization: Bearer ' . $token];
    curlReq($method, $url, $body, $headers);
}
setDns(getToken($key), $ip);
file_put_contents('ip.txt', $ip);
die('ok');

The request leaves the router through the line that has the default route, so the name follows the active line. Thanks to the routing tables SSH works on both addresses, so it does not matter much which one the name points to.

Testing failover without unplugging cables

To test the script without physical access to the router, I blocked the pings of the router over ppp0 in a separate nftables table. LAN traffic and the PPPoE session keep working, but the health checks fail:

sudo nft add table inet failtest
sudo nft add chain inet failtest out '{ type filter hook output priority 0; }'
sudo nft add rule inet failtest out oifname "ppp0" icmp type echo-request drop

The script switched to Ziggo after three failed rounds. After deleting the table again with sudo nft delete table inet failtest, it switched back to KPN after six good rounds. Restarting nftables also removes the table, as the configuration file starts with flush ruleset.

Future work: unplug tests, IPv6 and load balancing

There are a few things left to do. I still want to unplug each line and measure how long the LAN is offline, and to check a full reboot with the new DNS settings. The LAN is IPv4 only for now: Ziggo gives the router a single IPv6 address and KPN only a link-local one. Getting IPv6 on the LAN requires DHCPv6 prefix delegation and router advertisements on enp5s0, and a review of the firewall for IPv6. Load balancing over both lines may come after that, but for now failover does what is needed.

Enjoy!