feat: refuse silent v3->v2 downgrade; monthly retrain timer with staged promote
train_all() now raises RainUnavailableError when use_rain=True and the
Open-Meteo history cannot be loaded, instead of logging a warning and
writing gauge-only (v2) bundles over the deployed v3 set -- which is what
the 2026-09-01 server retrain did unnoticed. --no-rain remains the explicit
way to get v2. CLI exits 2 with a one-line error. Three tests cover the
guard, the opt-out, and the v3 happy path.
scripts/retrain.sh trains into models/.staging, refuses to promote unless
metrics.json shows hgb-v3+ and >=14 trained stations, then renames bundles
into place (previous generation kept in models/.previous). No API restart:
predict.py reloads by mtime on the hourly precompute.
water-monitor-retrain.{service,timer}: 1st of each month 03:30, Persistent,
OMP_NUM_THREADS=4, Nice=15, same sandbox as the API unit. install.sh now
does `uv sync` into .venv (one env rule; removes a stale venv/) and enables
the timer. water-monitor.service in the repo matched neither the deployed
unit nor the uv env; it now does (run.py --web-api, .venv, EnvironmentFile).
This commit is contained in:
@@ -0,0 +1,17 @@
|
||||
[Unit]
|
||||
Description=Monthly flood-model retrain (docs/FLOOD_FORECASTING.md section 7)
|
||||
|
||||
[Timer]
|
||||
# Policy: at minimum once pre-monsoon (May-June), monthly through the season
|
||||
# (July-November), and after any major flood. A retrain costs ~12 min and RAM
|
||||
# peaks ~300 MB, so running it every month all year is cheaper than remembering
|
||||
# which months matter. 1st of the month, 03:30 server-local -- between the
|
||||
# hourly scrapes and outside Thai daytime traffic.
|
||||
OnCalendar=*-*-01 03:30:00
|
||||
# Catch up if the box was off at the scheduled time.
|
||||
Persistent=true
|
||||
RandomizedDelaySec=20min
|
||||
Unit=water-monitor-retrain.service
|
||||
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
Reference in New Issue
Block a user