I depend on Internet access. It should always be fast and it should always work. For a few years a dual WAN Cisco router took care of that, with a Ziggo cable line and a KPN fiber line. I love Debian, so when I replaced it I chose an N100 mini PC with six gigabit ports and an NVMe disk, and turned it into a router with Debian 13. This post describes the setup and the things that went wrong along the way.
An N100 mini PC with six ports replaces the Cisco
The N100 has more than enough power to route a few gigabit lines and it uses little power. Three of the six ports are in use:
| Interface | Connected to |
|---|---|
| enp1s0 | Ziggo modem (bridge mode) |
| enp2s0 | KPN fiber media converter |
| enp5s0 | 16-port PoE LAN switch |
The ports enp3s0, enp4s0 and enp7s0 are spare. The network configuration is done
with ifupdown. The main /etc/network/interfaces file only sets up the loopback
interface and includes a directory:
source /etc/network/interfaces.d/*
auto lo
iface lo inet loopback
In /etc/network/interfaces.d there is one file per port, numbered after the
port on the case: 0_ziggo, 1_kpn and 4_lan. The numbers keep the files in
port order and make it easy to see which cable goes where.
Ziggo is a plain DHCP interface
The Ziggo modem hands out the public address with DHCP, so
/etc/network/interfaces.d/0_ziggo is short:
# WAN - ETH0
allow-hotplug enp1s0
iface enp1s0 inet dhcp
iface enp1s0 inet6 auto
On my Debian 13 install the DHCP client is dhcpcd and not dhclient. That matters later on, as the hooks and the configuration file of dhcpcd are in different places.
KPN needs VLAN 6, PPPoE and a larger MTU
KPN delivers internet on VLAN 6 and wants you to log in with PPPoE. On Debian
you need the ppp and vlan packages for that, and ifup fails with unclear
errors until they are installed:
sudo apt install ppp vlan
/etc/network/interfaces.d/1_kpn sets up the physical port, the VLAN on top of
it and the PPPoE session on top of the VLAN:
# WAN - ETH1
# Physical port to the media converter
auto enp2s0
iface enp2s0 inet manual
mtu 1512
# KPN internet VLAN
auto enp2s0.6
iface enp2s0.6 inet manual
vlan-raw-device enp2s0
mtu 1508
# PPPoE session
auto kpn
iface kpn inet ppp
provider kpn
PPPoE adds an 8 byte header to every packet. With a normal MTU of 1500 on the VLAN that leaves 1492 bytes for the PPPoE session. KPN supports “baby jumbo frames” (RFC 4638), so I raised the MTU of the VLAN to 1508 and of the port to 1512 (the VLAN tag needs 4 more bytes). This gives the PPPoE session a full 1500 bytes, just like the Ziggo line.
The session itself is configured in /etc/ppp/peers/kpn:
plugin pppoe.so
nic-enp2s0.6
user "internet"
password "internet"
noauth
hide-password
mtu 1500
mru 1500
noipdefault
#defaultroute
#usepeerdns
+ipv6
persist
maxfail 0
holdoff 5
lcp-echo-interval 10
lcp-echo-failure 6
noaccomp
nopcomp
novj
novjccomp
nodeflate
nobsdcomp
KPN does not check the username and password, “internet” for both is fine. The
options persist, maxfail 0 and holdoff 5 make pppd reconnect forever, with
5 seconds between attempts. The LCP echo options make pppd send a keepalive
every 10 seconds and give up on the session after 6 missed replies. I will get
to the two options that are commented out below.
PPPoE must not take the default route
The first time the KPN session came up, pppd took the default route away from
Ziggo, because of the defaultroute option. I commented it out, as I want to
decide myself which line carries the traffic. The routing setup comes later in
this post.
A router needs to forward packets, so /etc/sysctl.d/99-router.conf contains:
net.ipv4.ip_forward=1
Run sudo sysctl --system to load it without a reboot.
dnsmasq serves DHCP and DNS on the LAN
The LAN port has a static address in /etc/network/interfaces.d/4_lan:
# LAN ETH4
auto enp5s0
iface enp5s0 inet static
address 192.168.1.1/24
dnsmasq hands out addresses on the LAN and answers DNS queries. It forwards
those to Cloudflare and Google instead of to the DNS servers of one of the
providers, so DNS keeps working whichever line is active. This is
/etc/dnsmasq.d/lan.conf:
interface=enp5s0
bind-dynamic
dhcp-range=192.168.1.10,192.168.1.250,255.255.255.0,12h
dhcp-option=option:router,192.168.1.1
dhcp-option=option:dns-server,192.168.1.1
domain-needed
bogus-priv
no-resolv
server=1.1.1.1
server=8.8.8.8
nftables does the firewall and NAT for both lines
This is /etc/nftables.conf:
#!/usr/sbin/nft -f
flush ruleset
define ZIGGO = "enp1s0"
define KPN = "ppp0"
define WAN = { $ZIGGO, $KPN }
define LAN = "enp5s0"
table inet filter {
chain input {
type filter hook input priority filter; policy drop;
ct state established,related accept
ct state invalid drop
iif lo accept
# WAN: DHCP replies from Ziggo
iifname $ZIGGO udp sport 67 udp dport 68 accept
# LAN: DHCP and DNS (dnsmasq)
iifname $LAN udp dport { 53, 67 } accept
iifname $LAN tcp dport 53 accept
# SSH on all interfaces, rate limited per source address
tcp dport 22 ct state new meter ssh4 { ip saddr limit rate 10/minute burst 5 packets } accept
tcp dport 22 ct state new meter ssh6 { ip6 saddr limit rate 10/minute burst 5 packets } accept
# Ping and IPv6 control traffic
meta l4proto { icmp, ipv6-icmp } accept
}
chain forward {
type filter hook forward priority filter; policy drop;
# Clamp TCP MSS to the route MTU (safety net for PPPoE)
tcp flags syn tcp option maxseg size set rt mtu
ct state established,related accept
ct state invalid drop
iifname $LAN oifname $WAN accept
}
chain output {
type filter hook output priority filter; policy accept;
}
}
table ip nat {
chain postrouting {
type nat hook postrouting priority srcnat; policy accept;
oifname $WAN masquerade
}
}
The WAN define is a set of both lines, so forwarding and NAT work for either
of them. Only the DHCP rule uses ZIGGO, as that is the only line that uses
DHCP. The MSS clamping rule normally does nothing, because the PPPoE session has
a full 1500 byte MTU. If KPN ever falls back to 1492, it prevents TCP
connections that hang as soon as a large packet is sent. It has to come before
the accept rules, as an accept ends the evaluation of the chain.
I want to reach the router with SSH from home, so port 22 is open on all interfaces. SSH only accepts keys, and the meter limits every source address to 10 new connections per minute with a burst of 5. Note that every SYN counts, including the retries of your SSH client, so when you reconnect a few times in a row you may get a “Connection timed out” for a minute.
Apply the file with:
sudo nft -c -f /etc/nftables.conf && sudo systemctl restart nftables
The -c checks the file first, so a typo does not leave you with half a
firewall.
When I started I had also installed ufw, and it ran next to nftables. A packet
has to pass both rule sets, so this is confusing at best. When I purged ufw my
SSH connections from outside stopped working. The allow 22 of ufw had been the
only rule that let SSH in from the internet, my nftables configuration only
allowed SSH from the LAN. So if you remove ufw, add the SSH rule to nftables
first. I could still log in from the LAN to fix it.
You can test the KPN path before you send any real traffic over it, by routing a single address over ppp0 and running a traceroute from the LAN:
sudo ip route add 1.0.0.1 dev ppp0
traceroute -n 1.0.0.1
sudo ip route del 1.0.0.1 dev ppp0
The first hop after the router was the KPN peer address, 195.190.228.120, so forwarding and NAT over KPN worked.
Each line gets its own routing table
When the router sends a packet from its KPN address, that packet must leave through KPN, whatever the default route is. Ziggo drops packets that leave with a KPN source address and the other way around. This matters for SSH (the reply to a connection that came in on KPN must go out through KPN) and for the health checks of the failover script, which have to test each line separately.
The solution is a routing table per line, with a rule that selects the table by
source address. On Debian 13 there is no /etc/iproute2 directory by default,
the defaults live in /usr/share/iproute2. Your own table names go in
/etc/iproute2/rt_tables.d:
sudo mkdir -p /etc/iproute2/rt_tables.d
printf '100 ziggo\n200 kpn\n' | sudo tee /etc/iproute2/rt_tables.d/wan.conf
The addresses of both lines are dynamic, so the tables and rules are filled by
hooks. For Ziggo that is /etc/dhcpcd.exit-hook, which dhcpcd runs after every
change in the lease:
# Keep routing table "ziggo" in sync with the DHCP lease on enp1s0
if [ "$interface" = "enp1s0" ]; then
case "$reason" in
BOUND|RENEW|REBIND|REBOOT)
ip route replace "$new_network_number/$new_subnet_cidr" dev enp1s0 table ziggo
ip route replace default via "${new_routers%% *}" dev enp1s0 table ziggo
while ip rule del lookup ziggo 2>/dev/null; do :; done
ip rule add from "$new_ip_address" lookup ziggo priority 100
;;
EXPIRE|FAIL|RELEASE|STOP|NOCARRIER)
ip route flush table ziggo
while ip rule del lookup ziggo 2>/dev/null; do :; done
;;
esac
fi
For KPN pppd runs the scripts in /etc/ppp/ip-up.d and /etc/ppp/ip-down.d.
Here is my /etc/ppp/ip-down.d/kpn-table:
#!/bin/sh
[ "$PPP_IFACE" = "ppp0" ] || exit 0
ip route flush table kpn
while ip rule del lookup kpn 2>/dev/null; do :; done
I show the ip-up.d script in the next section, as it also sets the default
route. Make both executable, and do not put a dot in the file name, because
run-parts skips those files. You can test both lines from the router with:
ping -c3 -I <ziggo-address> 1.1.1.1
ping -c3 -I <kpn-address> 1.1.1.1
curl -4 --interface <kpn-address> ifconfig.me
The last command should print the KPN address.
KPN is the primary line because it is faster
At first the Ziggo line carried all traffic and KPN was the backup. Then I looked at the numbers: a ping to 1.1.1.1 takes 4 ms over KPN and 12 ms over Ziggo. After a reboot the KPN session was up within 13 seconds, while the Ziggo modem needed about 80 seconds to hand out a lease. So I made KPN the primary line.
The switching is done with route metrics. dhcpcd installs the Ziggo default
route with metric 1002. The ip-up.d script adds a default route over ppp0 with
metric 100, and the lowest metric wins. This is the updated
/etc/ppp/ip-up.d/kpn-table:
#!/bin/sh
# Keep routing table "kpn" in sync with the PPPoE session
[ "$PPP_IFACE" = "ppp0" ] || exit 0
ip route replace default dev ppp0 table kpn
while ip rule del lookup kpn 2>/dev/null; do :; done
ip rule add from "$PPP_LOCAL" lookup kpn priority 200
# KPN is the primary line: lower metric than the Ziggo route from dhcpcd (1002)
ip route replace default dev ppp0 metric 100
When the PPPoE session drops, the ppp0 interface disappears and the kernel removes its routes, so traffic falls back to Ziggo without any script. When the session comes back, the hook adds the route again. The main table then looks like this:
default dev ppp0 scope link metric 100
default via <ziggo-gateway> dev enp1s0 proto dhcp src <ziggo-address> metric 1002 mtu 1500
Connections that the LAN had open over Ziggo were set up with the Ziggo address by NAT, and they break when the route changes. You can delete them from the connection tracking table, so the clients reconnect right away:
sudo conntrack -D -s 192.168.1.0 --mask-src 255.255.255.0
This only matches connections from the LAN. I first tried
conntrack -D --src-nat, but that also deleted entries of incoming pings that
were never translated.
The router’s own DNS must not depend on Ziggo
My first switch to KPN failed. The LAN worked, but curl ifconfig.me on the
router itself hung. The cause was /etc/resolv.conf, which dhcpcd had filled
with the DNS servers of Ziggo. Those do not answer queries from a KPN address:
dig +time=3 +tries=1 -b <kpn-address> @62.179.104.196 ifconfig.me
;; communications error to 62.179.104.196#53: timed out
The LAN did not suffer from this, as dnsmasq forwards to 1.1.1.1 and 8.8.8.8. So
the fix was to let the router use dnsmasq as well. dnsmasq also listens on the
loopback interface when you configure an interface, so /etc/resolv.conf
becomes:
nameserver 127.0.0.1
To stop dhcpcd from overwriting this file, add this to /etc/dhcpcd.conf:
nohook resolv.conf
And for the same reason usepeerdns is commented out in the KPN peer file.
Otherwise the Debian script /etc/ppp/ip-up.d/0000usepeerdns adds the KPN DNS
servers to the file on every connect.
A running dhcpcd does not read its configuration again by itself. It rewrites
/etc/resolv.conf on every network event, including the IPv6 router
advertisements of Ziggo, so my edit was gone within minutes. dhcpcd -n enp1s0
reloads the configuration, but it also rebinds the Ziggo lease, and I was
connected through Ziggo. Thanks to the routing tables I could SSH in through the
KPN address instead, which does not depend on Ziggo at all.
A script handles a KPN line that is up but broken
The route metrics cover the case where the PPPoE session drops. They do not
cover the case where ppp0 is up, but no traffic gets through. pppd also needs up
to a minute (6 missed keepalives at 10 second intervals) before it gives up on a
dead session. For these cases I wrote a small script,
/usr/local/sbin/wan-failover:
#!/bin/sh
# KPN (ppp0) is the primary line, Ziggo (enp1s0) the backup.
# If the PPPoE session drops, the kernel removes the KPN route by itself.
# This script covers the other case: ppp0 is up, but KPN has no internet.
TARGETS="1.1.1.1 9.9.9.9 8.8.8.8"
INTERVAL=5 # seconds between checks
FAIL_AFTER=3 # failed KPN checks before switching to Ziggo
BACK_AFTER=6 # good KPN checks before switching back
# IPv4 address of an interface, empty if it has none
wan_ip() {
ip -4 -o addr show "$1" 2>/dev/null | awk '{print $4}' | cut -d/ -f1
}
# A line is healthy when at least one target answers from its address.
# The source address selects the per-line routing table.
healthy() {
src=$(wan_ip "$1")
[ -n "$src" ] || return 1
for t in $TARGETS; do
ping -n -q -c 1 -W 2 -I "$src" "$t" >/dev/null 2>&1 && return 0
done
return 1
}
on_kpn() {
ip route show default | grep -q 'dev ppp0'
}
after_switch() {
# Drop LAN connections that are still NAT-ed to the old line
conntrack -D -s 192.168.1.0 --mask-src 255.255.255.0 >/dev/null 2>&1
}
fails=0
goods=0
ziggo_ok=unknown
while true; do
if healthy ppp0; then
fails=0
goods=$((goods + 1))
else
goods=0
fails=$((fails + 1))
fi
if healthy enp1s0; then z=yes; else z=no; fi
if [ "$z" != "$ziggo_ok" ]; then
echo "Ziggo healthy: $z"
ziggo_ok=$z
fi
if on_kpn; then
if [ "$fails" -ge "$FAIL_AFTER" ]; then
if [ "$ziggo_ok" = yes ]; then
echo "KPN failed $fails checks, switching to Ziggo"
ip route del default dev ppp0 metric 100
after_switch
elif [ "$fails" -eq "$FAIL_AFTER" ]; then
echo "KPN failed $fails checks, but Ziggo is down too, staying on KPN"
fi
fi
elif [ -n "$(wan_ip ppp0)" ] && [ "$goods" -ge "$BACK_AFTER" ]; then
echo "KPN healthy for $goods checks, switching back to KPN"
ip route replace default dev ppp0 metric 100
after_switch
fi
sleep "$INTERVAL"
done
The script pings three targets from the address of each line, and a line is healthy when one of them answers. Because of the routing tables these checks work whether a line carries the default route or not. After three failed rounds it removes the KPN default route, so Ziggo takes over. A failed round takes about 11 seconds (5 seconds of sleep and three pings with a 2 second timeout), so that is about half a minute. It only switches back after KPN has been healthy for six rounds in a row, so a flapping line does not cause constant switching.
The script does not keep track of the active line itself, it reads the routing
table every round. This keeps it in line with the ip-up.d hook, which adds the
KPN route on every reconnect. When both lines are down, it does nothing and logs
that once.
It runs as a systemd service, /etc/systemd/system/wan-failover.service:
[Unit]
Description=WAN failover (KPN primary, Ziggo backup)
After=network-online.target
Wants=network-online.target
[Service]
ExecStart=/usr/local/sbin/wan-failover
Restart=always
RestartSec=5
[Install]
WantedBy=multi-user.target
Enable it with:
sudo apt install conntrack
sudo systemctl daemon-reload
sudo systemctl enable --now wan-failover
The output of the script ends up in the journal, which you can follow with
journalctl -u wan-failover -f.
Ziggo should not get a link-local default route
After a reboot I looked at the boot log. The Ziggo link went down and up twice
while the modem and the network card agreed on the link, and it took 80 seconds
before Ziggo handed out a lease. In the meantime dhcpcd had given enp1s0 a
link-local address (169.254.x.x) and a default route over it. That route leads
nowhere. To prevent this, add this line to /etc/dhcpcd.conf:
noipv4ll
Dynamic DNS follows the active line
To reach the router with SSH I use a DNS name with a short TTL. A cron job on
the router calls dynip.php, a small PHP script on my web server. The script
takes the address that the request came from, and when it changed, updates the A
record through the TransIP API:
<?php
$key = trim(file_get_contents('transip.key'));
$ip = $_SERVER['REMOTE_ADDR'];
$token = $_GET['token'];
if ($token != trim(file_get_contents('token.txt'))) {
die('ko');
}
if (trim(file_get_contents('ip.txt')) == $ip) {
die('ok');
}
function curlReq($method, $url, $body, $headers)
{
$ch = curl_init();
curl_setopt($ch, CURLOPT_URL, $url);
curl_setopt($ch, CURLOPT_POSTFIELDS, json_encode($body));
curl_setopt($ch, CURLOPT_HTTPHEADER, array_merge(['Content-Type: application/json'], $headers));
curl_setopt($ch, CURLOPT_CUSTOMREQUEST, $method);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
$data = json_decode(curl_exec($ch), true);
curl_close($ch);
if ($data['error'] ?? '') {
throw new \Exception($data['error']);
}
return $data;
}
function getToken($key)
{
$method = 'POST';
$url = 'https://api.transip.nl/v6/auth';
$body = [
'login' => 'tip_username',
'nonce' => random_int(0, PHP_INT_MAX),
'read_only' => false,
'expiration_time' => '10 seconds',
'label' => 'dynip',
'global_key' => false,
];
openssl_sign(json_encode($body), $signature, $key, OPENSSL_ALGO_SHA512);
$headers = ['Signature: ' . base64_encode($signature)];
$data = curlReq($method, $url, $body, $headers);
return $data['token'];
}
function setDns($token, $ip)
{
$method = 'PATCH';
$url = 'https://api.transip.nl/v6/domains/domeinnaam.nl/dns';
$body = [
'dnsEntry' => [
'name' => 'router',
'expire' => 300,
'type' => 'A',
'content' => $ip,
]
];
$headers = ['Authorization: Bearer ' . $token];
curlReq($method, $url, $body, $headers);
}
setDns(getToken($key), $ip);
file_put_contents('ip.txt', $ip);
die('ok');
The request leaves the router through the line that has the default route, so the name follows the active line. Thanks to the routing tables SSH works on both addresses, so it does not matter much which one the name points to.
Testing failover without unplugging cables
To test the script without physical access to the router, I blocked the pings of the router over ppp0 in a separate nftables table. LAN traffic and the PPPoE session keep working, but the health checks fail:
sudo nft add table inet failtest
sudo nft add chain inet failtest out '{ type filter hook output priority 0; }'
sudo nft add rule inet failtest out oifname "ppp0" icmp type echo-request drop
The script switched to Ziggo after three failed rounds. After deleting the table
again with sudo nft delete table inet failtest, it switched back to KPN after
six good rounds. Restarting nftables also removes the table, as the
configuration file starts with flush ruleset.
Future work: unplug tests, IPv6 and load balancing
There are a few things left to do. I still want to unplug each line and measure how long the LAN is offline, and to check a full reboot with the new DNS settings. The LAN is IPv4 only for now: Ziggo gives the router a single IPv6 address and KPN only a link-local one. Getting IPv6 on the LAN requires DHCPv6 prefix delegation and router advertisements on enp5s0, and a review of the firewall for IPv6. Load balancing over both lines may come after that, but for now failover does what is needed.
Enjoy!