Author SHA1 Message Date
kurogeek 596ac1f4bb router: make CrowdSec opt-in (crowdsec.enable, default off)
Not every site wants the ban engine (hub sync needs internet at
activation, and it is one more moving part on a small box). Gate
crowdsec.nix on a new `crowdsec.enable` option like `omada.enable`.

gw-cnx-1 keeps it on; the VM test drops its mkForce overrides.
2026-09-21 07:20:17 +00:00
kurogeek e1e18dd9f3 router/tests: cover the staging uplink as NATed fallback exit
Neither regression fixed by the previous commit was observable: the VM
test had no staging port, and pinging the ISP's PPPoE address only needs
ppp0's connected route, not the default route pppd refused to install.

Add a fifth node, `oldlan`: a networkd DHCP server on vlan 4 handing gw
its staging lease, with a second address (203.0.113.1) that gw can only
reach through that lease's default route. It has no route back to the
VLANs, so client pings only work if gw masquerades.

Proven: pppd's metric-0 default wins over the DHCP one while ppp0 is up;
lan reaches the old LAN via the staging port and iot (allowWan = false)
does not; after stopping pppd the staging route carries WAN traffic with
the same allowWan split; restarting pppd makes ppp0 preferred again.
2026-09-21 04:32:14 +00:00
kurogeek 8090ab3e6d router: let allowWan VLANs out through the staging uplink
With stagingPort set, the box itself had internet over the staging DHCP
uplink but LAN/Wi-Fi clients had none: forward and masquerade were scoped
to ppp0 only. Worse, pppd's `defaultroute` refuses to install its route
while the staging DHCP default route (metric 1024) exists ("not replacing
existing default route"), so even a live PPPoE session was never used.

- firewall: forward-allow + masquerade allowWan VLANs -> stagingPort in a
  separate `router-staging-nat` postrouting chain (networking.nat only
  takes one external interface). Same allowWan set as nixos-nat.
- pppoe: `defaultroute-metric 0`, so pppd only checks for a metric-0
  default route, installs ppp0 as the preferred exit and removes it on
  hangup, leaving the staging route as the fallback.
2026-09-21 04:32:07 +00:00
8 changed files with 151 additions and 56 deletions
+2 -2
View File
@@ -28,7 +28,7 @@ mgmt-only trust model, SSH exposure and the Wi-Fi bridge ports. Run it with
| DHCP | Kea, one subnet per VLAN |
| DNS | Blocky (blocklist resolver), metrics on :4000 scraped by control |
| IPv6 | DHCPv6-PD on ppp0, /64 per VLAN via SLAAC |
| Bans | CrowdSec + nftables bouncer (sshd log parsing) |
| Bans | Optional per site (`crowdsec.enable`): CrowdSec + nftables bouncer (sshd log parsing) |
| Omada | Optional per site: TP-Link Omada controller as a podman container |
| Wi-Fi | Optional: hostapd on the router's radios; each SSID (`wifi.networks`) is an untagged access port of its VLAN, passphrases via vars prompts — see `modules/clan/router/README.md` |
| Proxy | Optional: Caddy reverse proxy for internal services under `*.<site><n>.cnx.network` with a real Let's Encrypt wildcard (DNS-01 against ns1) |
@@ -71,7 +71,7 @@ Trust model: mgmt → everything; other VLANs → router DNS/DHCP + internet onl
installer: `ls -l /dev/disk/by-id/`).
2. Add the machine to `inventory.machines` in `clan.nix`, to the `router`
instance in `inventory.nix` (`roles.default.machines.gw-<city>-<n>.settings`: `site`,
`siteId` (next free number), port names, VLANs, `omada.enable`; keep the
`siteId` (next free number), port names, VLANs, `omada.enable`, `crowdsec.enable`; keep the
`mgmt`/`lan` VLANs), and to the machine list in `modules/mesh-hosts.nix`.
Do **not** add it to `modules/hosts.nix` (dynamic PPPoE IP; clan connects
over the mesh).
+2
View File
@@ -77,6 +77,8 @@ in
};
# This site runs the Omada controller for its APs/switches.
omada.enable = true;
# sshd ban engine (was unconditional before the option existed).
crowdsec.enable = true;
# Internal reverse proxy: real wildcard cert *.cnx1.cnx.network; Blocky
# resolves the names to the router's LAN address for mgmt+lan clients.
+5 -4
View File
@@ -3,10 +3,11 @@
Turns a machine with several NICs into a site gateway: PPPoE WAN (ISP
credentials via vars prompts), a VLAN-filtering bridge over the LAN ports with
one L3 interface per VLAN, Kea DHCP and Blocky DNS per VLAN, nftables
firewall/NAT, DHCPv6-PD, CrowdSec, an iperf3 server and a WAN speed-test
timer. Optional: a Wi-Fi access point on the router's own radios (hostapd),
the TP-Link Omada controller (podman) and an internal Caddy reverse proxy
with a real wildcard certificate (ACME DNS-01).
firewall/NAT, DHCPv6-PD, an iperf3 server and a WAN speed-test timer.
Optional: a Wi-Fi access point on the router's own radios (hostapd), CrowdSec
with the nftables bouncer (sshd log parsing), the TP-Link Omada controller
(podman) and an internal Caddy reverse proxy with a real wildcard certificate
(ACME DNS-01).
Addressing convention: a site owns `10.<siteId>.0.0/16`; VLAN `<id>` defaults
to `10.<siteId>.<id>.0/24`, router at `.1`, DHCP pool `.100-.199`. The `mgmt`
+5 -2
View File
@@ -1,12 +1,14 @@
# CrowdSec security engine + nftables bouncer: parses sshd auth attempts from
# the journal and bans offending source IPs at the firewall. Log-based (no
# inline DPI) so it costs the N300 next to nothing.
# inline DPI) so it costs the N300 next to nothing. Opt-in per site
# (`crowdsec.enable`): the hub sync needs internet at activation time.
{ settings }:
{ ... }:
{ lib, ... }:
let
cfg = settings;
in
{
config = lib.mkIf cfg.crowdsec.enable {
services.crowdsec = {
enable = true;
autoUpdateService = true;
@@ -43,4 +45,5 @@ in
registerBouncer.enable = true;
settings.mode = "nftables";
};
};
}
+27 -1
View File
@@ -3,7 +3,8 @@
# other VLANs -> DNS/DHCP on the router + WAN (if allowWan); no inter-VLAN
# WAN (ppp0) -> nothing inbound beyond established/related
# mesh -> admin SSH + metrics scrapes (same trust boundary as the fleet)
# staging -> admin SSH only (pre-cutover uplink into the old LAN)
# staging -> admin SSH only inbound (pre-cutover uplink into the old LAN);
# allowWan VLANs are NATed out through it while ppp0 is down
{ settings }:
{ lib, ... }:
let
@@ -14,6 +15,13 @@ let
lib.filterAttrs (_: vlan: vlan.allowWan) cfg.vlans
);
nonMgmtIfs = lib.filter (i: i != "vlan-mgmt") vlanIfs;
# allowWan VLANs may also leave through the staging uplink. Same set as
# networking.nat.internalInterfaces below, so `allowWan` holds on both
# exits. The kernel picks the exit: ppp0 (metric 0, see pppoe.nix) while
# the session is up, the staging DHCP route (metric 1024) otherwise.
stagingExit = cfg.stagingPort != null && wanVlanIfs != [ ];
wanVlanSet = "{ ${lib.concatMapStringsSep ", " (i: ''"${i}"'') wanVlanIfs} }";
in
{
networking.nftables.enable = true;
@@ -50,6 +58,9 @@ in
extraForwardRules = ''
tcp flags syn tcp option maxseg size set rt mtu comment "MSS clamp for PPPoE mtu 1492"
iifname "vlan-mgmt" accept comment "mgmt reaches all VLANs and the WAN"
''
+ lib.optionalString stagingExit ''
iifname ${wanVlanSet} oifname "${cfg.stagingPort}" accept comment "allowWan VLANs out via the staging uplink"
'';
};
@@ -62,4 +73,19 @@ in
externalInterface = "ppp0";
internalInterfaces = wanVlanIfs;
};
# networking.nat only masquerades on its single externalInterface; the
# staging uplink needs its own postrouting chain (nixos-nat's is
# oifname-scoped to ppp0, so the two never both apply).
networking.nftables.tables = lib.optionalAttrs stagingExit {
router-staging-nat = {
family = "ip";
content = ''
chain post {
type nat hook postrouting priority srcnat;
iifname ${wanVlanSet} oifname "${cfg.stagingPort}" masquerade comment "allowWan VLANs out via the staging uplink"
}
'';
};
};
}
+11 -7
View File
@@ -256,13 +256,15 @@ in
example = "enp3s0";
description = ''
Temporary DHCPv4-client uplink into the existing LAN while the box
runs alongside the router it replaces: gives it internet + mesh
before the WAN port is cabled (PPPoE simply retries until then). The
port is in no VLAN zone; the firewall admits only SSH on it. Do NOT
connect the trunk ports to the production switch while staging
Kea on the mgmt tag would fight the old router's DHCP in one
broadcast domain. Set to null at cutover (and usually hand the port
back to `trunkPorts`).
runs alongside the router it replaces: gives it (and, NATed, the
allowWan VLANs) internet + mesh before the WAN port is cabled; once
the PPPoE session is up its default route wins, and the staging
route only carries traffic again if the session drops (PPPoE simply
retries until then). The port is in no VLAN zone; inbound, the
firewall admits only SSH on it. Do NOT connect the trunk ports to
the production switch while staging Kea on the mgmt tag would
fight the old router's DHCP in one broadcast domain. Set to null at
cutover (and usually hand the port back to `trunkPorts`).
'';
};
@@ -283,6 +285,8 @@ in
omada.enable = lib.mkEnableOption "TP-Link Omada SDN controller (podman container)";
crowdsec.enable = lib.mkEnableOption "CrowdSec (sshd log parsing) with the nftables bouncer";
proxy = {
enable = lib.mkEnableOption "internal reverse proxy (Caddy, wildcard cert via DNS-01)";
+6
View File
@@ -30,6 +30,11 @@ in
'';
};
# defaultroute-metric 0: pppd refuses `defaultroute` while any other
# default route exists (e.g. the staging uplink's DHCP route, metric 1024,
# network.nix) unless given a metric; with 0 it only checks for a metric-0
# route, installs its own as the preferred exit, and removes it again on
# hangup so the staging route takes over.
services.pppd = {
enable = true;
peers.wan = {
@@ -40,6 +45,7 @@ in
file ${creds.files."user-opts".path}
noipdefault
defaultroute
defaultroute-metric 0
noauth
hide-password
persist
+60 -7
View File
@@ -1,14 +1,17 @@
# End-to-end VM test of the router service: a PPPoE access concentrator plays
# the ISP on the WAN port, a trunk carries tagged lan/iot VLANs to `client`,
# and an untagged access port carries mgmt to `admin`.
# an untagged access port carries mgmt to `admin`, and `oldlan` is the DHCP
# network the box is staged in before cutover.
#
# isp ---(vlan 1: PPPoE)--- wan [gw] trunk ---(vlan 2: tagged 20/40)--- client
# access --(vlan 3: untagged mgmt)--- admin
# staging -(vlan 4: DHCP client)--- oldlan
#
# What is proven: PPPoE dial-in with the vars-provided credentials, bridge
# VLAN tagging/untagging, Kea leases and reservations per VLAN, Blocky
# answering on the VLAN with the blocklist active, NAT to the WAN, and the
# firewall trust model (allowWan, mgmt-only SSH, no inter-VLAN forwarding).
# answering on the VLAN with the blocklist active, NAT to the WAN, the
# firewall trust model (allowWan, mgmt-only SSH, no inter-VLAN forwarding),
# and the staging uplink as NATed fallback exit behind ppp0.
{ pkgs, lib, ... }:
let
# The vars mock answers every prompt with "mock-prompt-value-<name>"; the
@@ -20,6 +23,10 @@ let
clientAddress = "10.9.20.50";
adminMac = "02:00:00:00:00:10";
adminAddress = "10.9.10.50";
oldlanAddress = "192.168.88.1";
# Only reachable through oldlan's router role, i.e. via gw's staging
# default route (metric 1024); ppp0's metric-0 default must win while up.
beyondStaging = "203.0.113.1";
in
{
name = "router";
@@ -36,6 +43,7 @@ in
isp = { };
client = { };
admin = { };
oldlan = { };
};
instances.router = {
@@ -48,6 +56,7 @@ in
wan.interface = "wan";
trunkPorts = [ "trunk" ];
accessPorts.access = "mgmt";
stagingPort = "staging";
vlans = {
mgmt = {
id = 10;
@@ -111,6 +120,10 @@ in
vlan = 3;
assignIP = false;
};
staging = {
vlan = 4;
assignIP = false;
};
};
# Something must listen on 22 for the mgmt-only SSH rule to be observable
@@ -118,13 +131,11 @@ in
services.openssh.enable = true;
# The sandbox has no internet: serve the blocklist from a local file
# instead of GitHub, and skip CrowdSec, whose hub sync needs the network
# (it is not what this test exercises).
# instead of GitHub. (CrowdSec, whose hub sync needs the network too,
# is opt-in and stays off.)
services.blocky.settings.blocking.denylists.ads = lib.mkForce [
(toString (pkgs.writeText "ads.hosts" "0.0.0.0 ads.example.com\n"))
];
services.crowdsec.enable = lib.mkForce false;
services.crowdsec-firewall-bouncer.enable = lib.mkForce false;
# Two simulated radios: wlan0 is the AP (settings above), wlan1 plays a
# wireless client. It lives in its own network namespace, like the
@@ -251,6 +262,32 @@ in
};
environment.systemPackages = [ pkgs.netcat ];
};
# The LAN the box is staged in: a DHCP server handing gw its uplink
# lease, plus an address that is only reachable via that uplink's
# default route. No route back to 10.9.0.0/16: replies only reach the
# clients if gw masquerades them.
oldlan = {
virtualisation.interfaces.staging = {
vlan = 4;
assignIP = false;
};
networking.useDHCP = false;
networking.useNetworkd = true;
systemd.network.networks."10-staging" = {
matchConfig.Name = "staging";
address = [
"${oldlanAddress}/24"
"${beyondStaging}/32"
];
networkConfig.DHCPServer = true;
dhcpServerConfig = {
PoolOffset = 100;
PoolSize = 50;
};
};
networking.firewall.allowedUDPPorts = [ 67 ];
};
};
testScript = ''
@@ -302,5 +339,21 @@ in
gw.wait_until_succeeds("ip netns exec sta wpa_cli -i wlan1 status | grep -q wpa_state=COMPLETED")
gw.succeed("timeout 60 sta-dhcp")
gw.succeed("ip netns exec sta ip -4 addr show wlan1 | grep -q 'inet 10.9.20.1[0-9][0-9]/24'")
with subtest("Staging uplink: NATed exit for allowWan VLANs, behind ppp0 while it is up"):
gw.wait_until_succeeds("ip -4 route show default dev staging | grep -q 'via ${oldlanAddress}'")
# pppd installs its default route despite the DHCP one (defaultroute-metric 0).
gw.succeed("ip route get ${beyondStaging} | grep -q 'dev ppp0'")
# On-link old-LAN hosts are reached through the staging port regardless.
client.succeed("ping -c1 -W2 -I lan0 ${oldlanAddress}")
client.fail("ping -c1 -W2 -I iot0 ${oldlanAddress}")
# ppp0 down: the staging route carries the WAN traffic, allowWan still holds.
gw.systemctl("stop pppd-wan.service")
gw.wait_until_succeeds("ip route get ${beyondStaging} | grep -q 'dev staging'")
client.succeed("ping -c1 -W2 -I lan0 ${beyondStaging}")
client.fail("ping -c1 -W2 -I iot0 ${beyondStaging}")
# ppp0 back: preferred again.
gw.systemctl("start pppd-wan.service")
gw.wait_until_succeeds("ip route get ${beyondStaging} | grep -q 'dev ppp0'")
'';
}