Compare commits

...
23 Commits
Author SHA1 Message Date
grabowski 32399f1899 feat: recalibrate P.77/P.75 thresholds; capacity guard on alerts
CI / Format & lint (push) Successful in 10s
CI / Test suite (push) Successful in 17s
Security / Static analysis (push) Successful in 12s
Security / License report (push) Successful in 45s
Security / Dependency vulnerabilities (push) Successful in 49s
The first ntfy cycle announced "Warning level at P.77" at 3.02 m. That
gauge's 2.85 m threshold sat below its own dry-season baseline (2.6-2.7 m
at 8-14 % channel capacity): P.77 had been "above warning" for 761 of the
last 2 146 hours, at 22 % capacity. Across 2018-2024, 75-85 % capacity
reads 3.35-4.57 m and 95-105 % reads 4.27-5.08 m; set 4.30 / 4.90. P.75
moved 2.75/3.50 -> 3.20/3.65 on the same evidence (2024: 3.45 / 3.72).
The predictor already handles changed thresholds (regression-derived
probabilities until the Oct 1 retrain).

Second line of defence in notify.py: a clear->alert transition is only
announced when RID's discharge_percent for the reading is >= 60 %, so a
re-rated or datum-shifted gauge cannot page subscribers again. P.1 is
exempt (its stages come from the inundation map, not capacity); readings
without a capacity figure fall back to level only; the all-clear edge is
never blocked. 3 tests.
2026-09-12 00:34:59 +02:00
grabowski f4d42c90f4 fix: ntfy listens on the Tailscale address; monitor publishes to it directly
Security / Dependency vulnerabilities (push) Successful in 44s
Security / Static analysis (push) Successful in 9s
CI / Format & lint (push) Successful in 10s
CI / Test suite (push) Successful in 26s
Security / License report (push) Successful in 50s
Docs / Validate documentation (push) Successful in 16s
The reverse proxy is a separate VPS on the tailnet, so a loopback-only
ntfy was unreachable from it. install_ntfy.sh now binds the host's Tailscale
IP (NTFY_LISTEN overrides). New NTFY_PUBLISH_URL: where the monitor POSTs,
separate from the public NTFY_SERVER subscribers see, so an alert never
waits on DNS or the proxy (first cycle logged 502s from Cloudflare while
the domain was not yet proxied).
2026-09-12 00:28:29 +02:00
grabowski 039d24a5c3 fix: init ntfy only in the collection leader
CI / Test suite (push) Successful in 29s
Docs / Validate documentation (push) Successful in 13s
CI / Format & lint (push) Successful in 15s
Security / Dependency vulnerabilities (push) Successful in 42s
Security / License report (push) Successful in 48s
Security / Static analysis (push) Successful in 9s
Every uvicorn worker ran the notification_state DDL at startup; on
Postgres the losers of that race get UniqueViolation on pg_type and the
whole init was skipped (notifications off). Only the leader publishes, so
only the leader initialises, after election. The DDL also tolerates a
concurrent creator now: on failure it verifies the table exists instead
of giving up.
2026-09-12 00:22:03 +02:00
grabowski 777b230baf feat: public flood notifications over self-hosted ntfy
CI / Format & lint (push) Successful in 11s
Security / Static analysis (push) Successful in 12s
CI / Test suite (push) Successful in 18s
Docs / Validate documentation (push) Successful in 11s
Security / Dependency vulnerabilities (push) Successful in 1m30s
Security / License report (push) Successful in 47s
Anyone can now get push alerts on their phone without an account: the
monitor publishes to an ntfy server (one Go binary, ~30 MB RSS) and
subscribers pick topics in the free iOS/Android/web app.

Semantics are transitions, never state. One message when a gauge crosses
its warning or danger threshold, one all-clear when it drops back (0.10 m
hysteresis), nothing while it sits above. A three-day flood is two
messages; a quiet season is zero. Topics: ping-warning / ping-danger
(basin digest), ping-<station>-warning / -danger, ping-p1-outlook (opt-in:
model P(warning within 24 h) at P.1 rises through 50 %, clears below 25 %,
message says it is experimental), ping-status (feed stale >= 3 h /
recovered). Priority 5 on danger so it rings through Do Not Disturb.

src/notify.py runs once per collection cycle in the API process (leader
only, after the forecast precompute, same data the dashboard shows).
Last-sent state lives in a notification_state table so a restart never
re-sends; a failed publish leaves state untouched so the crossing is
retried next cycle instead of lost. Off unless NTFY_SERVER is set.

Dashboard: a "Get alerts" button (only when configured) opens a panel
with the server, per-topic cards, ntfy:// deep links and web links, app
store links and a disclaimer. EN + TH. GET /api/notifications feeds it.

scripts/install_ntfy.sh: .deb install, server.yml (loopback listen,
anonymous read, token-only write scoped to ping-*, 72 h cache, signup/
login/metrics off, tight visitor limits), systemd, user + token, .env.
Verified against ntfy 2.28.0: anon publish 403, token publish 200, token
on foreign topic 403, anon read 200, and a seeded crossing through the
real _notify_transitions path arrived in the topic with priority, tags,
click and action button. docs/NOTIFICATIONS.md has the deployment and
reverse-proxy notes. Tests: 10 for the state machine (159 total).
2026-09-12 00:18:38 +02:00
grabowski 0ec675e9c5 ci: license report from a clean venv, not the runner's site-packages
CI / Format & lint (push) Successful in 10s
Security / Dependency vulnerabilities (push) Successful in 53s
Security / Static analysis (push) Successful in 10s
CI / Test suite (push) Successful in 20s
Security / License report (push) Successful in 50s
2026-09-11 23:54:44 +02:00
grabowski 7b31d4d0dd feat: "Is the model getting better?" - live verification per model version
CI / Test suite (push) Successful in 22s
CI / Format & lint (push) Successful in 16s
Docs / Validate documentation (push) Successful in 10s
Security / Dependency vulnerabilities (push) Successful in 1m34s
Security / Static analysis (push) Successful in 10s
Security / License report (push) Successful in 13s
src/ml/skill.py joins forecast_history (what each deployed version
predicted for the 24 h peak, hourly) to water_measurements (what the
river did) and reports per version: verified hours, peak MAE, bias, the
persistence baseline (peak = current level), skill = 1 - MAE/persistence,
and the same MAE restricted to observed peaks >= 2 m. Only forecasts
whose window has elapsed with >= 75 % of hours observed count; a
version needs 24 verified hours before it is compared.

GET /api/forecast/skill?station_code=P.1&horizon=24 returns it (SWR
cached, 15 min). The dashboard's forecast card gains a panel with a
one-line verdict (current vs previous version), the per-version table,
and a caveat that quiet weeks measure quiet-river accuracy only: the
model is judged on flood-onset lead, which the backtests cover. EN + TH.

On today's production data: hgb-v3+28b62e5 (369 h, Aug 13 - Sep 1)
MAE 15.2 cm, skill -0.05; hgb-v2+f6570ac (224 h, Sep 1 - 11) MAE
12.3 cm, skill 0.36 - the "worse" v2 model scores better on a quieter
fortnight, which is exactly why the panel shows the >= 2 m column and
the caveat. Tests: 3, sqlite, synthetic.

scripts/dev_proxy.py: DEV_PROXY_LOCAL lets a not-yet-deployed endpoint be
answered from a local JSON file while everything else goes to prod.
2026-09-11 23:44:44 +02:00
grabowski 2e19974fad docs: README describes the project that exists; CLAUDE.md for agents
The README was the original template: P.1 "in Nakhon Sawan", P.103 "in
Bangkok", VictoriaMetrics as the recommended database, Docker/Grafana
sections, github.com/your-username support links. Rewritten around what
runs: sources, gap fill, the forecast and its 13 h result, the live API
table, systemd deployment with the retrain timer, repository layout,
real docs links, data-source credits. Other DB adapters are mentioned as
supported-but-not-production.

CLAUDE.md replaces the untracked Ruflo boilerplate with project rules:
Python 3.11, format/test gates, Bangkok timestamps, harness-first model
changes judged on lead, the RainUnavailableError guard, no git add -A.
2026-09-11 23:44:43 +02:00
grabowski b03318210c security: pip-audit + bandit gates that can fail; patch 29 known CVEs
security.yml previously ran safety/bandit/semgrep with `|| true` and could
not go red. Now: pip-audit on requirements.txt is a hard gate (dev deps
reported only), bandit HIGH fails (B104 bind-all skipped: intended behind
Cloudflare/Caddy), pip-licenses uploaded as a report. Weekly + on
dependency/source changes.

Running it locally found 29 advisories, all in pinned-and-forgotten
runtime deps: starlette 0.27 (7, incl. Host-header path confusion and
form DoS), fastapi 0.104, requests 2.31 (3), pymysql 1.1. Bumped to
current: fastapi 0.141.1 / starlette 1.6.0, pydantic 2.13.5, uvicorn
0.52.4, requests 2.34.2, pymysql 1.2.0; dev: pytest 9.1.1, black 26.5.1.
pip-audit is now clean. requires-python narrowed to 3.11 (the truth:
psycopg2-binary 2.9.9 fails on 3.13; pandas 2.0.3 has no 3.12 wheels).
Full suite passes; API smoke-tested (health, stations, forecast, history,
stats, docs, openapi) on the new stack. black 26 reformatted 8 files.
2026-09-11 23:44:43 +02:00
grabowski 97a6694ab2 feat(dashboard): light/dark theme
CI / Format & lint (push) Successful in 36s
CI / Test suite (push) Successful in 27s
Docs / Validate documentation (push) Successful in 12s
Sun/moon button next to the language switch. Follows the OS preference
until the user picks one (persisted in localStorage, tracks OS changes
only while unpinned). Every colour that was a literal white / grey is now
a token with a dark counterpart; flood-verdict, live-pill and lang-toggle
states got their own bg/ink/border tokens. OSM tiles are inverted with
a hue rotate so roads and labels stay legible while the river network,
markers and rain dots (SVG, unfiltered) keep their data colours. Chart.js
reads grid/tick/legend colours from the tokens and the open chart is
rebuilt on toggle. Leaflet popups and controls follow the theme.
2026-09-11 23:16:30 +02:00
grabowski 5ad8e4eac3 ci: green pipelines that check what exists; one formatting contract
The Test Suite job failed on every push since the black check was added
because the tree had never been formatted, and pre-commit said 120
columns while CI ran black's default 88. pyproject.toml now carries
[tool.black] / [tool.isort] (88, black profile) as the single source;
pre-commit reads it; `make format` applied it (13 files, whitespace only,
146 insertions / 128 deletions, tests unchanged at 146 passed).

ci.yml: lint (black, isort, flake8 hard errors) + pytest. The Docker
registry push, VictoriaMetrics integration test, staging/production
deploy and Apache-Bench jobs were template scaffolding for hosts and
registries that do not exist; production is a systemd unit updated by
git pull. Removed rather than left permanently skipped.

docs.yml: the "Check markdown links" step curl'd every URL in every .md
and failed on localhost examples and the Tailscale IP, and the Sphinx
jobs built artifacts nobody read. Replaced by two checks that mean
something: relative links/images in README, CONTRIBUTING and docs/
resolve inside the repo, and the FastAPI OpenAPI schema exports with
the documented endpoints present (uploaded as an artifact).
2026-09-11 23:05:37 +02:00
grabowski 6f4a86edbb fix(dashboard): timestamps are Asia/Bangkok everywhere; stale feed says so
The API emits naive ICT timestamps ("2026-09-12T02:00:00"). The page fed
them to new Date(), which applies the BROWSER's zone: a viewer in Europe
parsed a 02:00 ICT reading as 02:00 CEST, five hours in the future, so
"Last updated" clamped to "0 min ago" forever; a viewer in the Americas
saw thousands of minutes. parseTs() now pins +07:00 on naive strings and
every display formats with timeZone: Asia/Bangkok, so the site shows
river time regardless of where it is opened. Daily/hourly chart buckets
key on the Bangkok calendar day instead of UTC getters.

The tile also gets a real stale state: past 3 h (RID is hourly) it turns
red, reads "Feed stale · last reading <day>, N h ago", and the header
pill switches from LIVE DATA to STALE FEED. Age shows hours past 2 h.

scripts/dev_proxy.py serves the working-copy dashboard with API calls
proxied to water.buildfor.life so browser-side changes can be checked
against live data (and any browser timezone) before deploying.
2026-09-11 23:00:39 +02:00
grabowski ce08312c0f docs: public dashboard URL (water.buildfor.life) replaces the Tailscale IP everywhere
CI/CD Pipeline - Northern Thailand Ping River Monitor / Test Suite (3.11) (push) Failing after 32s
Documentation / Documentation Summary (push) Successful in 3s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Build Docker Image (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Integration Test with Services (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Staging (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Production (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Performance Test (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Code Quality (push) Successful in 15s
Documentation / Validate Documentation (push) Failing after 16s
Documentation / Generate API Documentation (push) Successful in 11s
Documentation / Build Sphinx Documentation (push) Successful in 18s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Cleanup (push) Successful in 1s
2026-09-11 22:38:13 +02:00
grabowski d621aa9ce7 eval: quantile heads and fc48 on top of hgb-v3 (rejected/deferred); HII gauge-rain aggregate
CI/CD Pipeline - Northern Thailand Ping River Monitor / Test Suite (3.11) (push) Failing after 41s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Build Docker Image (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Integration Test with Services (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Staging (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Production (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Performance Test (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Code Quality (push) Successful in 17s
Documentation / Validate Documentation (push) Failing after 16s
Documentation / Generate API Documentation (push) Successful in 11s
Documentation / Build Sphinx Documentation (push) Successful in 18s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Cleanup (push) Successful in 1s
Documentation / Documentation Summary (push) Successful in 3s
Rolling-origin harness gains rise_rain_quantile, rise_rain_quantile_uw,
rise_rain_qsigma (L2 point + quantile sigma) and rise_rain_fc48, all
opt-in, plus --from-cache for reproducible offline reruns. Results in
models/eval_2026-09-12*.json, write-up in docs/FLOOD_FORECASTING.md:

- quantile point prediction: better MAE, worse first-alert lead at 5 of
  11 events (P.103 2022-08-14 +6h -> +1h) -> rejected
- quantile sigma only: Brier within noise (0.0031 -> 0.0029) -> not worth 3x heads
- rain_fc48: neutral everywhere except 2024-10-03 P.1 (+21h -> +72h), n=1
  -> deferred to after the 2026 season

src/ml/hii_rain.py: catchment-mean hourly rain from the ~130 HII gauges in
the upper-Ping box and a 24h-sum comparison against Open-Meteo. Not a
training feature (table exists only since 2026-08-11, no archive); exposed
at GET /api/hii/rainfall/catchment so the two sources' agreement is on
record by the time a fold can test it.

data._read_cache now skips non-station files in models/cache/ (the shared
dir also holds rain_openmeteo / dam_* caches, which crashed the reader).
scripts/summarize_eval.py prints per-variant lead/peak-error tables.
2026-09-11 21:55:37 +02:00
grabowski 764764e07e feat: refuse silent v3->v2 downgrade; monthly retrain timer with staged promote
train_all() now raises RainUnavailableError when use_rain=True and the
Open-Meteo history cannot be loaded, instead of logging a warning and
writing gauge-only (v2) bundles over the deployed v3 set -- which is what
the 2026-09-01 server retrain did unnoticed. --no-rain remains the explicit
way to get v2. CLI exits 2 with a one-line error. Three tests cover the
guard, the opt-out, and the v3 happy path.

scripts/retrain.sh trains into models/.staging, refuses to promote unless
metrics.json shows hgb-v3+ and >=14 trained stations, then renames bundles
into place (previous generation kept in models/.previous). No API restart:
predict.py reloads by mtime on the hourly precompute.

water-monitor-retrain.{service,timer}: 1st of each month 03:30, Persistent,
OMP_NUM_THREADS=4, Nice=15, same sandbox as the API unit. install.sh now
does `uv sync` into .venv (one env rule; removes a stale venv/) and enables
the timer. water-monitor.service in the repo matched neither the deployed
unit nor the uv env; it now does (run.py --web-api, .venv, EnvironmentFile).
2026-09-11 21:37:11 +02:00
grabowski 0a4bf843ff chore: remove empty shell-accident files committed at repo root
CI/CD Pipeline - Northern Thailand Ping River Monitor / Test Suite (3.11) (push) Failing after 1m8s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Build Docker Image (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Integration Test with Services (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Staging (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Production (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Performance Test (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Code Quality (push) Successful in 31s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Cleanup (push) Successful in 0s
2026-09-11 20:54:48 +02:00
grabowski f6570ac10f chore: declutter the repo
CI/CD Pipeline - Northern Thailand Ping River Monitor / Test Suite (3.11) (push) Failing after 1m43s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Build Docker Image (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Integration Test with Services (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Staging (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Production (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Performance Test (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Code Quality (push) Successful in 45s
Documentation / Validate Documentation (push) Failing after 17s
Documentation / Generate API Documentation (push) Successful in 12s
Documentation / Build Sphinx Documentation (push) Successful in 19s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Cleanup (push) Successful in 1s
Documentation / Documentation Summary (push) Successful in 4s
Removes 26 tracked files that no longer describe or serve the running
system, verified one by one against the whole repo (source, tests, docs,
README, Makefile, Dockerfile, .gitea workflows, pyproject, packaging spec)
plus dynamic-reference paths, before deletion.

Root (11): one-off launch/setup write-ups from the project's first weeks
that document events which never happened the way they describe — a
github.com publication (the remote is self-hosted Gitea) and a 15-minute
scheduler (the service runs hourly). Also .gitlab-ci.yml (unused, CI is
.gitea/), .env.postgres and setup.py.backup (a placeholder env file and a
backup in version control), and the PyInstaller packaging trio
build_executable.py / build_simple.py / ping-river-monitor.spec, which
bundled docs that no longer exist and is not how this deploys.

docs (9): stale guides superseded by DATABASE_DEPLOYMENT_GUIDE,
GITEA_WORKFLOWS, FLOOD_FORECASTING and DATA_SOURCES, plus two snapshots
(PROJECT_STATUS, PROJECT_STRUCTURE) describing a 4-file src/ that is now 39.

scripts (5): one-shot bootstrap tools already run — init_git.sh/.bat,
generate_badges.py, migrate_geolocation.py, encode_password.py.

Every inbound reference was fixed rather than left dangling: README doc
index and migration section, three Makefile targets, the docs.yml summary
step, and the GITEA_WORKFLOWS resource list.

src/ is deliberately untouched. The audit proposed removing several live
modules; verification showed those proposals were mis-scoped and would
have broken production.

.gitignore now covers the agent tooling dirs, model eval output and
editor/shell leftovers — the working tree had collected 56 zero-byte
files named after fragments of shell commands.

136 tests pass; production modules import clean.
2026-08-14 13:45:03 +07:00
grabowski 7e64e0cf18 fix: one current river level, not two
CI/CD Pipeline - Northern Thailand Ping River Monitor / Test Suite (3.11) (push) Failing after 29s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Build Docker Image (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Integration Test with Services (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Staging (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Production (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Performance Test (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Code Quality (push) Successful in 18s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Cleanup (push) Successful in 1s
The verdict banner and the P.1 outlook each wrote state.p1Now from a
different feed — the banner from the latest measurement, the outlook from
current_level on the forecast rows, which carries whatever the model saw
at its as_of. Forecasts are precomputed hourly, so the two drifted apart:
production showed 1.66 m in the banner and 1.52 m in the outlook directly
below it. Harmless at low water; at flood stage two contradictory river
levels on one screen undermine the warning.

setP1Level() now arbitrates: freshest timestamp wins, and the replay and
demo hooks pass force since they deliberately pin a level that is not the
live one. The outlook renders the arbitrated value and clamps the shown
peak to at least the current level, so a stale forecast can no longer
predict a peak below where the river already is. endReplay drops the
replayed level so live data re-arbitrates cleanly.

Verified against a stub reproducing the exact production conditions
(gauge 1.66 at 08:35 vs forecast 1.52 at 08:00): both now read 1.66; a
2.50 m rise against a stale 1.81 m peak renders 2.50/2.50; the 2024
replay still tracks its frames and returns to live on stop.
2026-08-14 09:25:39 +07:00
grabowski 382daa7d86 perf: backfill one dam per request via the api/dam range endpoint
CI/CD Pipeline - Northern Thailand Ping River Monitor / Test Suite (3.11) (push) Failing after 1m6s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Build Docker Image (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Integration Test with Services (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Staging (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Production (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Performance Test (push) Skipped
Documentation / Generate API Documentation (push) Successful in 12s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Cleanup (push) Successful in 1s
Documentation / Documentation Summary (push) Successful in 4s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Code Quality (push) Successful in 34s
Documentation / Validate Documentation (push) Failing after 10s
Documentation / Build Sphinx Documentation (push) Successful in 22s
GET app.rid.go.th/reservoir/api/dam?dam_id&date_start&date_end returns a
single dam's whole date range in one response — Mae Ngat's 2009-today
archive is ~4 chunked requests instead of the ~2,900 one-day POSTs the
all-dams path needs. Field names differ from api/dams and are mapped in
parse_dam_range_records, verified equal on spot-checked dates; the range
endpoint also carries DMD_ULevel, the reservoir level in m MSL that
api/dams stopped publishing after ~2013.

scripts/backfill_rid_reservoir.py defaults to the fast per-dam path
(--all-dams keeps the full-fleet crawl, --refresh rewrites stored days).
Already-stored dates are still skipped, junk values are still bounded,
and the consecutive-failure abort still applies.
2026-08-13 23:10:36 +07:00
grabowski b02e815d72 feat: Thai localisation and portrait-phone layout for the dashboard
Thai is the default unless the browser prefers English, chosen by first
match in navigator.languages order and remembered in localStorage. A
STRINGS table carries both languages (interpolated strings as functions),
applyTranslations() drives static markup via data-i18n attributes, and
setLang() rebuilds everything the JS renders — including map layers, so
popups render in the current language and sensor markers are replaced
rather than stacked. Thai dates use the Buddhist era, matching the
replay label; station names lead with the reader's language.

Portrait phones: the header overflowed a 412 px Android viewport by
64 px, so the page scrolled sideways and the Refresh button sat off
screen. The header now wraps into two deliberate rows (DOM order matches
visual order, so focus order is unaffected), the map description box is
hidden on phones, map height is capped by viewport — including a
height-gated rule for landscape phones, whose 850-960 px widths never
matched the width breakpoints — and the forecast grid goes single
column. Verified in a real browser at 412x915, 360x800, 915x412 and
1440x900: zero horizontal overflow, no desktop change.

Review-swarm fixes: a language switch no longer relabels a pinned
SIMULATION or the 2024 replay as LIVE DATA (it kept the mode from
state); Thai wording corrected where it asserted a rising trend the code
never checks, labelled every gauge 'critical', or used a malformed
compound; aria-labels, the Leaflet load failure and the flood-stage
chips are translated; the language toggle states its action instead of
an aria-pressed value that contradicted its label; and Thai font
families sit after the Latin stack so they cannot restyle English text.
2026-08-13 23:10:34 +07:00
grabowski 5dc5850df6 docs: source sweep — P.75 already is the Mae Ngat release signal
Verified from a Thai ISP that lsim.rid.go.th is unreachable (not
geo-blocked), then swept for any better-than-daily Mae Ngat source.

Key finding: P.75 sits 3.8 km below the dam, reports hourly, and has
been a model feature since v1 — the model has always read the dam's
actual outflow, hourly and directly. That is the likelier reason the
daily reservoir table adds nothing, beyond its publication lag.

Intraday reservoir feeds do exist and are open (bigdata-api.rid.go.th
SWOC, ThaiWater ridhydro_TUP.16 at the dam) but are snapshot-only with
no archive, so they cannot retrain against past events; the HII
collector accumulates them from 2026-08-11 for a post-monsoon revisit.
Also documents api/dam (whole per-dam range in one request, vs the
~2,900-request per-day loop) and several cross-check mirrors.
2026-08-13 21:58:47 +07:00
grabowski 70e4da07a0 docs: lsim.rid.go.th is unreachable, not geo-blocked
Probed from a Thai consumer ISP (AIS Fibre): DNS resolves but ICMP and
ports 80/443/8080 are filtered, while app.rid.go.th answers in 0.27 s
over the same connection. Correct the earlier note that assumed the
timeout was a foreign-network block, and stop pointing the dam-feature
follow-up at a host that cannot be reached.
2026-08-13 21:00:24 +07:00
grabowski 28b62e5a36 feat: Mae Ngat dam features — built, evaluated, defaulted OFF
CI/CD Pipeline - Northern Thailand Ping River Monitor / Code Quality (push) Successful in 15s
Documentation / Validate Documentation (push) Failing after 8s
Documentation / Generate API Documentation (push) Successful in 9s
Documentation / Build Sphinx Documentation (push) Successful in 15s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Cleanup (push) Successful in 1s
Documentation / Documentation Summary (push) Successful in 2s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Test Suite (3.11) (push) Failing after 27s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Build Docker Image (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Integration Test with Services (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Staging (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Production (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Performance Test (push) Skipped
src/ml/dam.py loads rid_reservoir_daily into a leakage-safe hourly frame
(daily row visible from 07:00 its own date, ffill capped at 48 h) and is
plumbed through features/train/predict/evaluate exactly like rain, gated
to the six mainstem stations below the Mae Ngat confluence.

The experiment concludes as a documented NEGATIVE result: on the 2024
record-flood backtest every dam-feature subset costs 1-3 h of first-alert
lead (13h -> 10-12h) for <=3 cm of peak-error gain, because the daily RID
report lags up to 31 h and describes yesterday's benign absorbing
reservoir during fast onset. Features therefore default OFF (--dam
opt-in on the training and backtest CLIs; rise_rain_dam/rise_dam harness
variants, excluded from the default variant set). The ablation also
isolated the HII gap-fill as lead-neutral: the acceptance gate holds at
13 h with fill enabled, and docs/img charts are regenerated with the
shipping configuration. Full table in docs/FLOOD_FORECASTING.md §5.

Review-swarm fixes: evaluate.py skips variants whose feature family is
absent instead of crashing the run; --dam forwards --db-url and warns
loudly when no dam history loads; an empty DB result can no longer wipe
a good dam cache; run-level metrics version claims v4 only when a dam
station is actually in the set.
2026-08-13 20:42:21 +07:00
grabowski 6af6fbe02c fix: survive junk RID dam values — widen storage_pct, bound inserts
Documentation / Generate API Documentation (push) Successful in 9s
Documentation / Build Sphinx Documentation (push) Successful in 15s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Test Suite (3.11) (push) Failing after 25s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Build Docker Image (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Integration Test with Services (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Staging (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Production (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Performance Test (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Code Quality (push) Successful in 16s
Documentation / Validate Documentation (push) Failing after 8s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Cleanup (push) Successful in 1s
Documentation / Documentation Summary (push) Successful in 3s
Dam 100602 reports 87798% storage on some days, overflowing
NUMERIC(6,2) and discarding the entire 33-dam daily batch. storage_pct
is now NUMERIC(8,2) (auto-migrated on connect for existing Postgres/
MySQL tables) and every measure column is bounds-checked before insert
so out-of-capacity junk becomes NULL instead of a batch-killing error.
Rerunning the backfill repairs the days the overflow skipped.
2026-08-13 10:50:58 +07:00
92 changed files with 6812 additions and 7458 deletions
-87
View File
@@ -1,87 +0,0 @@
# Northern Thailand Ping River Monitor Configuration
# Copy this file to .env and customize for your environment
# Database Configuration
DB_TYPE=postgresql
# Options: sqlite, mysql, postgresql, influxdb, victoriametrics
# SQLite Configuration (default)
WATER_DB_PATH=water_levels.db
# VictoriaMetrics Configuration
VM_HOST=localhost
VM_PORT=8428
VM_URL=
# InfluxDB Configuration
INFLUX_HOST=localhost
INFLUX_PORT=8086
INFLUX_DATABASE=ping_river_monitoring
INFLUX_USERNAME=
INFLUX_PASSWORD=
# PostgreSQL Configuration (Remote Server)
# Option 1: Full connection string (URL encode special characters in password)
#POSTGRES_CONNECTION_STRING=postgresql://username:url_encoded_password@your-postgres-host:5432/water_monitoring
# Option 2: Individual components (password will be automatically URL encoded)
POSTGRES_HOST=10.0.10.201
POSTGRES_PORT=5432
POSTGRES_DB=ping_river
POSTGRES_USER=ping_river
POSTGRES_PASSWORD=3_%m]k:+16"rx?M#`swIA
# Examples for connection string:
# - Local: postgresql://postgres:password@localhost:5432/water_monitoring
# - Remote: postgresql://user:pass@192.168.1.100:5432/water_monitoring
# - With special chars: postgresql://user:my%3Apass%40word@host:5432/db
# - With SSL: postgresql://user:pass@host:port/db?sslmode=require
# - Connection pooling: postgresql://user:pass@host:port/db?pool_size=20&max_overflow=0
# Special character URL encoding:
# : → %3A @ → %40 # → %23 ? → %3F & → %26 / → %2F % → %25
# MySQL Configuration
MYSQL_CONNECTION_STRING=mysql://user:password@localhost:3306/ping_river_monitoring
# API Configuration
API_HOST=0.0.0.0
API_PORT=8000
API_WORKERS=1
# Data Collection Settings
SCRAPING_INTERVAL_HOURS=1
REQUEST_TIMEOUT=30
MAX_RETRIES=3
RETRY_DELAY_SECONDS=60
# Data Retention
DATA_RETENTION_DAYS=365
# Logging Configuration
LOG_LEVEL=INFO
LOG_FILE=water_monitor.log
# Security (for production)
SECRET_KEY=your-secret-key-here
API_KEY=your-api-key-here
# Monitoring
ENABLE_METRICS=true
ENABLE_HEALTH_CHECKS=true
# Geographic Settings
TIMEZONE=Asia/Bangkok
DEFAULT_LATITUDE=18.7875
DEFAULT_LONGITUDE=99.0045
# External Services
NOTIFICATION_EMAIL=
SMTP_SERVER=
SMTP_PORT=587
SMTP_USERNAME=
SMTP_PASSWORD=
# Development Settings
DEBUG=false
DEVELOPMENT_MODE=false
+13
View File
@@ -84,6 +84,19 @@ SMTP_PORT=587
SMTP_USERNAME= SMTP_USERNAME=
SMTP_PASSWORD= SMTP_PASSWORD=
# Public push notifications via self-hosted ntfy (https://ntfy.sh, single binary).
# Leave NTFY_SERVER empty to disable. Topics published: <prefix>-<station>-warning,
# <prefix>-<station>-danger, <prefix>-warning, <prefix>-danger, <prefix>-p1-outlook,
# <prefix>-status. See docs/NOTIFICATIONS.md.
NTFY_SERVER=
# Where the monitor POSTs (defaults to NTFY_SERVER). Use the local ntfy
# address (loopback or Tailscale IP) so publishing does not depend on
# DNS / the reverse proxy being up.
NTFY_PUBLISH_URL=
NTFY_TOPIC_PREFIX=ping
NTFY_TOKEN=
PUBLIC_URL=https://water.buildfor.life/
# Matrix Alerting Configuration # Matrix Alerting Configuration
MATRIX_HOMESERVER=https://matrix.org MATRIX_HOMESERVER=https://matrix.org
MATRIX_ACCESS_TOKEN= MATRIX_ACCESS_TOKEN=
-2
View File
@@ -1,2 +0,0 @@
DB_TYPE=postgresql
POSTGRES_CONNECTION_STRING=postgresql://postgres:password@localhost:5432/water_monitoring
+60 -327
View File
@@ -1,342 +1,75 @@
name: CI/CD Pipeline - Northern Thailand Ping River Monitor name: CI
# What this checks, on every push and PR to master:
# 1. formatting contract (black + isort, config in pyproject.toml)
# 2. flake8 hard-error gate (syntax, undefined names)
# 3. the pytest suite (synthetic data, no DB/network; ~1 min)
# Docker build / staging / production / perf jobs from the original template
# were removed: there is no registry, no staging host, and production is a
# systemd unit deployed by `git pull` on the server (docs/FLOOD_FORECASTING.md
# section 6, scripts/install.sh). Re-add a job when the thing it deploys exists.
on: on:
push: push:
branches: [ master, develop ] branches: [master, develop]
pull_request: pull_request:
branches: [ master ] branches: [master]
schedule: schedule:
# Run tests daily at 2 AM UTC # daily, catches dependency drift / upstream API changes in the tests
- cron: '0 2 * * *' - cron: "0 2 * * *"
workflow_dispatch:
env: env:
PYTHON_VERSION: '3.11' PYTHON_VERSION: "3.11" # pandas 2.0.3 ships no 3.12 wheels; psycopg2-binary 2.9.9 breaks on 3.13
REGISTRY: git.b4l.co.th
IMAGE_NAME: b4l/northern-thailand-ping-river-monitor
# GitHub token for better rate limits and authentication
GH_TOKEN: ${{ secrets.GH_TOKEN }}
jobs: jobs:
# Test job lint:
name: Format & lint
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: ${{ env.PYTHON_VERSION }}
cache: pip
cache-dependency-path: requirements-dev.txt
- name: Install tools
run: |
python -m pip install --upgrade pip --root-user-action=ignore
pip install --root-user-action=ignore black==26.5.1 isort==5.12.0 flake8==6.1.0
- name: black
run: black --check --diff src/ *.py
- name: isort
run: isort --check-only --diff src/ *.py
- name: flake8 (errors only)
run: flake8 src/ --count --select=E9,F63,F7,F82 --show-source --statistics
test: test:
name: Test Suite name: Test suite
runs-on: ubuntu-latest runs-on: ubuntu-latest
strategy:
matrix:
python-version: ['3.11'] # pandas 2.0.3 ships no 3.12 wheels; widen after upgrading pandas
steps: steps:
- name: Checkout code - uses: actions/checkout@v4
uses: actions/checkout@v4
with:
token: ${{ secrets.GITEA_TOKEN }}
- name: Set up Python ${{ matrix.python-version }} - uses: actions/setup-python@v5
uses: actions/setup-python@v4 with:
with: python-version: ${{ env.PYTHON_VERSION }}
python-version: ${{ matrix.python-version }} cache: pip
cache-dependency-path: |
requirements.txt
requirements-dev.txt
- name: Cache pip dependencies - name: Install dependencies
uses: actions/cache@v3 run: |
with: python -m pip install --upgrade pip --root-user-action=ignore
path: ~/.cache/pip pip install --root-user-action=ignore -r requirements.txt
key: ${{ runner.os }}-pip-${{ hashFiles('**/requirements*.txt') }} pip install --root-user-action=ignore pytest==9.1.1 pytest-asyncio==0.21.1
restore-keys: |
${{ runner.os }}-pip-
- name: Install dependencies - name: pytest
run: | env:
python -m pip install --upgrade pip --root-user-action=ignore DB_TYPE: sqlite
pip install --root-user-action=ignore -r requirements.txt run: pytest -q -p no:cacheprovider
pip install --root-user-action=ignore -r requirements-dev.txt
- name: Lint with flake8
run: |
flake8 src/ --count --select=E9,F63,F7,F82 --show-source --statistics
flake8 src/ --count --exit-zero --max-complexity=10 --max-line-length=100 --statistics
- name: Type check with mypy (advisory)
run: |
# 86 pre-existing errors; blocking typing gate deferred until the debt is paid down
mypy src/ --ignore-missing-imports || true
- name: Format check with black
run: |
black --check src/ *.py
- name: Import sort check
run: |
isort --check-only src/ *.py
- name: Run integration tests
run: |
python tests/test_integration.py
- name: Run station management tests
run: |
python tests/test_station_management.py
- name: Test application startup
run: |
timeout 10s python run.py --test || true
- name: Security scan with bandit
run: |
bandit -r src/ -f json -o bandit-report.json || true
- name: Upload test artifacts
uses: actions/upload-artifact@v3
if: always()
with:
name: test-results-${{ matrix.python-version }}
path: |
bandit-report.json
*.log
# Code quality job
code-quality:
name: Code Quality
runs-on: ubuntu-latest
steps:
- name: Checkout code
uses: actions/checkout@v4
with:
token: ${{ secrets.GITEA_TOKEN }}
- name: Set up Python
uses: actions/setup-python@v4
with:
python-version: ${{ env.PYTHON_VERSION }}
- name: Install dependencies
run: |
python -m pip install --upgrade pip --root-user-action=ignore
pip install --root-user-action=ignore -r requirements-dev.txt
- name: Run safety check
run: |
safety check -r requirements.txt --json --output safety-report.json || true
- name: Run bandit security scan
run: |
bandit -r src/ -f json -o bandit-report.json || true
- name: Upload security reports
uses: actions/upload-artifact@v3
with:
name: security-reports
path: |
safety-report.json
bandit-report.json
# Build Docker image
build:
name: Build Docker Image
runs-on: ubuntu-latest
needs: test
steps:
- name: Checkout code
uses: actions/checkout@v4
with:
token: ${{ secrets.GITEA_TOKEN }}
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v3
- name: Log in to Container Registry
uses: docker/login-action@v3
with:
registry: ${{ env.REGISTRY }}
username: ${{ github.actor }}
password: ${{ secrets.GITEA_TOKEN }}
- name: Extract metadata
id: meta
uses: docker/metadata-action@v5
with:
images: ${{ env.REGISTRY }}/${{ env.IMAGE_NAME }}
tags: |
type=ref,event=branch
type=ref,event=pr
type=sha,prefix={{branch}}-
type=raw,value=latest,enable={{is_default_branch}}
- name: Build and push Docker image
uses: docker/build-push-action@v5
with:
context: .
platforms: linux/amd64,linux/arm64
push: true
tags: ${{ steps.meta.outputs.tags }}
labels: ${{ steps.meta.outputs.labels }}
cache-from: type=gha
cache-to: type=gha,mode=max
env:
GITHUB_TOKEN: ${{ secrets.GH_TOKEN }}
- name: Test Docker image
run: |
docker run --rm ${{ env.REGISTRY }}/${{ env.IMAGE_NAME }}:${{ github.sha }} python run.py --test
# Integration test with services
integration-test:
name: Integration Test with Services
runs-on: ubuntu-latest
needs: build
services:
victoriametrics:
image: victoriametrics/victoria-metrics:latest
ports:
- 8428:8428
options: >-
--health-cmd "wget --quiet --tries=1 --spider http://localhost:8428/health"
--health-interval 30s
--health-timeout 10s
--health-retries 3
steps:
- name: Checkout code
uses: actions/checkout@v4
with:
token: ${{ secrets.GITEA_TOKEN }}
- name: Wait for VictoriaMetrics
run: |
timeout 60s bash -c 'until curl -f http://localhost:8428/health; do sleep 2; done'
- name: Set up Python
uses: actions/setup-python@v4
with:
python-version: ${{ env.PYTHON_VERSION }}
- name: Install dependencies
run: |
python -m pip install --upgrade pip --root-user-action=ignore
pip install --root-user-action=ignore -r requirements.txt
- name: Test with VictoriaMetrics
env:
DB_TYPE: victoriametrics
VM_HOST: localhost
VM_PORT: 8428
run: |
python run.py --test
- name: Start API server
env:
DB_TYPE: victoriametrics
VM_HOST: localhost
VM_PORT: 8428
run: |
python run.py --web-api &
sleep 10
- name: Test API endpoints
run: |
curl -f http://localhost:8000/health
curl -f http://localhost:8000/stations
curl -f http://localhost:8000/metrics
# Deploy to staging (only on develop branch)
deploy-staging:
name: Deploy to Staging
runs-on: ubuntu-latest
needs: [test, build, integration-test]
if: github.ref == 'refs/heads/develop'
environment:
name: staging
url: https://staging.ping-river-monitor.b4l.co.th
steps:
- name: Checkout code
uses: actions/checkout@v4
with:
token: ${{ secrets.GITEA_TOKEN }}
- name: Deploy to staging
run: |
echo "Deploying to staging environment..."
# Add your staging deployment commands here
# Example: kubectl, docker-compose, or webhook call
- name: Health check staging
run: |
sleep 30
curl -f https://staging.ping-river-monitor.b4l.co.th/health
# Deploy to production (only on main branch, manual approval)
deploy-production:
name: Deploy to Production
runs-on: ubuntu-latest
needs: [test, build, integration-test]
if: github.ref == 'refs/heads/master'
environment:
name: production
url: https://ping-river-monitor.b4l.co.th
steps:
- name: Checkout code
uses: actions/checkout@v4
with:
token: ${{ secrets.GITEA_TOKEN }}
- name: Deploy to production
run: |
echo "Deploying to production environment..."
# Add your production deployment commands here
- name: Health check production
run: |
sleep 30
curl -f https://ping-river-monitor.b4l.co.th/health
- name: Notify deployment
run: |
echo "✅ Production deployment successful!"
echo "🌐 URL: https://ping-river-monitor.b4l.co.th"
echo "📊 Grafana: https://grafana.ping-river-monitor.b4l.co.th"
# Performance test (only on main branch)
performance-test:
name: Performance Test
runs-on: ubuntu-latest
needs: deploy-production
if: github.ref == 'refs/heads/master'
steps:
- name: Checkout code
uses: actions/checkout@v4
with:
token: ${{ secrets.GITEA_TOKEN }}
- name: Install Apache Bench
run: |
sudo apt-get update
sudo apt-get install -y apache2-utils
- name: Performance test API endpoints
run: |
# Test health endpoint
ab -n 100 -c 10 https://ping-river-monitor.b4l.co.th/health
# Test stations endpoint
ab -n 50 -c 5 https://ping-river-monitor.b4l.co.th/stations
# Test metrics endpoint
ab -n 50 -c 5 https://ping-river-monitor.b4l.co.th/metrics
# Cleanup old artifacts
cleanup:
name: Cleanup
runs-on: ubuntu-latest
if: always()
needs: [test, build, integration-test]
steps:
- name: Clean up old Docker images
run: |
echo "Cleaning up old Docker images..."
# Add cleanup commands for old images/artifacts
+81 -350
View File
@@ -1,368 +1,99 @@
name: Documentation name: Docs
# Checks that the documentation the project actually ships stays consistent:
# - every relative link / image path in docs/*.md and README.md resolves
# inside the repo (external URLs are NOT fetched: localhost examples,
# rate-limited hosts and the Tailscale-era links made that gate permanently
# red, and a 200 on a curl --head proves nothing about a doc anyway)
# - the FastAPI app imports and its OpenAPI schema is exportable (that is
# the reference at https://water.buildfor.life/docs)
# The previous Sphinx/apidoc jobs produced artifacts nobody read and were
# removed. Reference docs live in docs/*.md; the public overview is at
# https://buildfor.life/docs/tooling/ping-river-monitor/.
on: on:
push: push:
branches: [ master, develop ] branches: [master, develop]
paths: paths:
- 'docs/**' - "docs/**"
- 'README.md' - "README.md"
- 'CONTRIBUTING.md' - "CONTRIBUTING.md"
- 'src/**/*.py' - "src/web_api.py"
- "src/schemas.py"
- ".gitea/workflows/docs.yml"
pull_request: pull_request:
paths: paths:
- 'docs/**' - "docs/**"
- 'README.md' - "README.md"
- 'CONTRIBUTING.md' - "CONTRIBUTING.md"
workflow_dispatch: workflow_dispatch:
env: env:
PYTHON_VERSION: '3.11' PYTHON_VERSION: "3.11"
jobs: jobs:
# Validate documentation docs:
validate-docs: name: Validate documentation
name: Validate Documentation
runs-on: ubuntu-latest runs-on: ubuntu-latest
steps: steps:
- name: Checkout code - uses: actions/checkout@v4
uses: actions/checkout@v4
with:
token: ${{ secrets.GITEA_TOKEN }}
- name: Set up Python - name: Relative links and images resolve
uses: actions/setup-python@v4 run: |
with: python3 - <<'PY'
python-version: ${{ env.PYTHON_VERSION }} import re, sys, pathlib
root = pathlib.Path(".")
files = [root / "README.md", root / "CONTRIBUTING.md", *root.glob("docs/**/*.md")]
link = re.compile(r"!?\[[^\]]*\]\(([^)\s]+)(?:\s+\"[^\"]*\")?\)")
bad = []
for md in files:
if not md.exists():
continue
for m in link.finditer(md.read_text(encoding="utf-8")):
target = m.group(1)
if target.startswith(("http://", "https://", "mailto:", "#")):
continue
path = target.split("#", 1)[0]
if not path:
continue
resolved = (md.parent / path).resolve()
if not resolved.exists():
bad.append(f"{md}: {target}")
if bad:
print("Broken relative links:")
print("\n".join(" " + b for b in bad))
sys.exit(1)
print(f"checked {len(files)} files, all relative links resolve")
PY
- name: Install documentation tools - uses: actions/setup-python@v5
run: | with:
python -m pip install --upgrade pip python-version: ${{ env.PYTHON_VERSION }}
pip install -r requirements.txt cache: pip
pip install sphinx sphinx-rtd-theme sphinx-autodoc-typehints cache-dependency-path: requirements.txt
pip install markdown-link-check || true
- name: Check markdown links - name: Install dependencies
run: | run: |
echo "🔗 Checking markdown links..." python -m pip install --upgrade pip --root-user-action=ignore
find . -name "*.md" -not -path "./.git/*" -not -path "./node_modules/*" | while read file; do pip install --root-user-action=ignore -r requirements.txt
echo "Checking $file"
# Basic link validation (you can enhance this)
grep -o 'http[s]*://[^)]*' "$file" | while read url; do
if curl -s --head "$url" | head -n 1 | grep -q "200 OK"; then
echo "✅ $url"
else
echo "❌ $url (in $file)"
fi
done
done
- name: Validate README structure - name: OpenAPI schema exports
run: | env:
echo "📋 Validating README structure..." DB_TYPE: sqlite
run: |
python - <<'PY'
import json
from src.web_api import app
spec = app.openapi()
paths = sorted(spec["paths"])
required = {"/forecast", "/measurements/latest", "/measurements/history/{station_code}", "/stations", "/api/stats", "/health"}
missing = required - set(paths)
assert not missing, f"documented endpoints missing from the app: {missing}"
json.dump(spec, open("openapi.json", "w"), indent=1)
print(f"{len(paths)} paths; schema written to openapi.json")
PY
required_sections=( - uses: actions/upload-artifact@v3
"# Northern Thailand Ping River Monitor" with:
"## Features" name: openapi-${{ github.run_number }}
"## Quick Start" path: openapi.json
"## Installation"
"## Usage"
"## API Endpoints"
"## Docker"
"## Contributing"
"## License"
)
for section in "${required_sections[@]}"; do
if grep -q "$section" README.md; then
echo "✅ Found: $section"
else
echo "❌ Missing: $section"
fi
done
- name: Check documentation completeness
run: |
echo "📚 Checking documentation completeness..."
# Check if all Python modules have docstrings
python -c "
import ast
import os
def check_docstrings(filepath):
with open(filepath, 'r', encoding='utf-8') as f:
tree = ast.parse(f.read())
missing_docstrings = []
for node in ast.walk(tree):
if isinstance(node, (ast.FunctionDef, ast.ClassDef, ast.AsyncFunctionDef)):
if not ast.get_docstring(node):
missing_docstrings.append(f'{node.name} in {filepath}')
return missing_docstrings
all_missing = []
for root, dirs, files in os.walk('src'):
for file in files:
if file.endswith('.py') and not file.startswith('__'):
filepath = os.path.join(root, file)
missing = check_docstrings(filepath)
all_missing.extend(missing)
if all_missing:
print('⚠️ Missing docstrings:')
for item in all_missing[:10]: # Show first 10
print(f' - {item}')
if len(all_missing) > 10:
print(f' ... and {len(all_missing) - 10} more')
else:
print('✅ All functions and classes have docstrings')
"
# Generate API documentation
generate-api-docs:
name: Generate API Documentation
runs-on: ubuntu-latest
steps:
- name: Checkout code
uses: actions/checkout@v4
with:
token: ${{ secrets.GITEA_TOKEN }}
- name: Set up Python
uses: actions/setup-python@v4
with:
python-version: ${{ env.PYTHON_VERSION }}
- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install -r requirements.txt
- name: Generate OpenAPI spec
run: |
echo "📝 Generating OpenAPI specification..."
python -c "
import json
import sys
sys.path.insert(0, 'src')
try:
from web_api import app
openapi_spec = app.openapi()
with open('openapi.json', 'w') as f:
json.dump(openapi_spec, f, indent=2)
print('✅ OpenAPI spec generated: openapi.json')
except Exception as e:
print(f'❌ Failed to generate OpenAPI spec: {e}')
"
- name: Generate API documentation
run: |
echo "📖 Generating API documentation..."
# Create API documentation from OpenAPI spec
if [ -f openapi.json ]; then
cat > api-docs.md << 'EOF'
# API Documentation
This document describes the REST API endpoints for the Northern Thailand Ping River Monitor.
## Base URL
- Production: `https://ping-river-monitor.b4l.co.th`
- Staging: `https://staging.ping-river-monitor.b4l.co.th`
- Development: `http://localhost:8000`
## Authentication
Currently, the API does not require authentication. This may change in future versions.
## Endpoints
EOF
# Extract endpoints from OpenAPI spec
python -c "
import json
with open('openapi.json', 'r') as f:
spec = json.load(f)
for path, methods in spec.get('paths', {}).items():
for method, details in methods.items():
print(f'### {method.upper()} {path}')
print()
print(details.get('summary', 'No description available'))
print()
if 'parameters' in details:
print('**Parameters:**')
for param in details['parameters']:
print(f'- `{param[\"name\"]}` ({param.get(\"in\", \"query\")}): {param.get(\"description\", \"No description\")}')
print()
print('---')
print()
" >> api-docs.md
echo "✅ API documentation generated: api-docs.md"
fi
- name: Upload documentation artifacts
uses: actions/upload-artifact@v3
with:
name: documentation-${{ github.run_number }}
path: |
openapi.json
api-docs.md
# Build Sphinx documentation
build-sphinx-docs:
name: Build Sphinx Documentation
runs-on: ubuntu-latest
steps:
- name: Checkout code
uses: actions/checkout@v4
with:
token: ${{ secrets.GITEA_TOKEN }}
- name: Set up Python
uses: actions/setup-python@v4
with:
python-version: ${{ env.PYTHON_VERSION }}
- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install -r requirements.txt
pip install sphinx sphinx-rtd-theme sphinx-autodoc-typehints
- name: Create Sphinx configuration
run: |
mkdir -p docs/sphinx
cat > docs/sphinx/conf.py << 'EOF'
import os
import sys
sys.path.insert(0, os.path.abspath('../../src'))
project = 'Northern Thailand Ping River Monitor'
copyright = '2025, Ping River Monitor Team'
author = 'Ping River Monitor Team'
version = '3.1.3'
release = '3.1.3'
extensions = [
'sphinx.ext.autodoc',
'sphinx.ext.viewcode',
'sphinx.ext.napoleon',
'sphinx_autodoc_typehints',
]
templates_path = ['_templates']
exclude_patterns = ['_build', 'Thumbs.db', '.DS_Store']
html_theme = 'sphinx_rtd_theme'
html_static_path = ['_static']
autodoc_default_options = {
'members': True,
'member-order': 'bysource',
'special-members': '__init__',
'undoc-members': True,
'exclude-members': '__weakref__'
}
EOF
cat > docs/sphinx/index.rst << 'EOF'
Northern Thailand Ping River Monitor Documentation
================================================
.. toctree::
:maxdepth: 2
:caption: Contents:
modules
Indices and tables
==================
* :ref:`genindex`
* :ref:`modindex`
* :ref:`search`
EOF
- name: Generate module documentation
run: |
cd docs/sphinx
sphinx-apidoc -o . ../../src
- name: Build documentation
run: |
cd docs/sphinx
sphinx-build -b html . _build/html
- name: Upload Sphinx documentation
uses: actions/upload-artifact@v3
with:
name: sphinx-docs-${{ github.run_number }}
path: docs/sphinx/_build/html/
# Documentation summary
docs-summary:
name: Documentation Summary
runs-on: ubuntu-latest
needs: [validate-docs, generate-api-docs, build-sphinx-docs]
if: always()
steps:
- name: Generate documentation summary
run: |
echo "# 📚 Documentation Build Summary" > docs-summary.md
echo "" >> docs-summary.md
echo "**Build Date:** $(date -u)" >> docs-summary.md
echo "**Repository:** ${{ github.repository }}" >> docs-summary.md
echo "**Commit:** ${{ github.sha }}" >> docs-summary.md
echo "" >> docs-summary.md
echo "## 📊 Results" >> docs-summary.md
echo "" >> docs-summary.md
if [ "${{ needs.validate-docs.result }}" = "success" ]; then
echo "- ✅ **Documentation Validation**: Passed" >> docs-summary.md
else
echo "- ❌ **Documentation Validation**: Failed" >> docs-summary.md
fi
if [ "${{ needs.generate-api-docs.result }}" = "success" ]; then
echo "- ✅ **API Documentation**: Generated" >> docs-summary.md
else
echo "- ❌ **API Documentation**: Failed" >> docs-summary.md
fi
if [ "${{ needs.build-sphinx-docs.result }}" = "success" ]; then
echo "- ✅ **Sphinx Documentation**: Built" >> docs-summary.md
else
echo "- ❌ **Sphinx Documentation**: Failed" >> docs-summary.md
fi
echo "" >> docs-summary.md
echo "## 🔗 Available Documentation" >> docs-summary.md
echo "" >> docs-summary.md
echo "- [README.md](../README.md)" >> docs-summary.md
echo "- [API Documentation](../docs/)" >> docs-summary.md
echo "- [Contributing Guide](../CONTRIBUTING.md)" >> docs-summary.md
echo "- [Deployment Checklist](../DEPLOYMENT_CHECKLIST.md)" >> docs-summary.md
cat docs-summary.md
- name: Upload documentation summary
uses: actions/upload-artifact@v3
with:
name: docs-summary-${{ github.run_number }}
path: docs-summary.md
+70 -254
View File
@@ -1,293 +1,109 @@
name: Security & Dependency Updates name: Security
# Two gates that can actually fail, plus one report:
# - pip-audit against requirements.txt: any known vulnerability in a runtime
# dependency fails the job (dev-only tools are reported, not gated)
# - bandit on src/: HIGH severity findings fail; medium/low are listed.
# B104 (bind 0.0.0.0) is skipped: the service is meant to listen on all
# interfaces behind Cloudflare/Caddy.
# - pip-licenses report as an artifact (informational; the project is MIT
# and its runtime deps are MIT/BSD/Apache/PSF)
# The old file ran safety/bandit/semgrep with `|| true` and could not go red.
on: on:
schedule: schedule:
# Run security scans daily at 3 AM UTC - cron: "0 3 * * 1" # weekly, Monday 03:00 UTC
- cron: "0 3 * * *"
workflow_dispatch: workflow_dispatch:
push: push:
paths: paths:
- "requirements*.txt" - "requirements*.txt"
- "Dockerfile" - "pyproject.toml"
- "uv.lock"
- "src/**/*.py"
- ".gitea/workflows/security.yml" - ".gitea/workflows/security.yml"
pull_request:
paths:
- "requirements*.txt"
- "pyproject.toml"
- "src/**/*.py"
env: env:
PYTHON_VERSION: "3.11" PYTHON_VERSION: "3.11"
# GitHub token for better rate limits and authentication
GH_TOKEN: ${{ secrets.GH_TOKEN }}
jobs: jobs:
# Dependency vulnerability scan dependencies:
dependency-scan: name: Dependency vulnerabilities
name: Dependency Security Scan
runs-on: ubuntu-latest runs-on: ubuntu-latest
steps: steps:
- name: Checkout code - uses: actions/checkout@v4
uses: actions/checkout@v4
with:
token: ${{ secrets.GITEA_TOKEN }}
- name: Set up Python - uses: actions/setup-python@v5
uses: actions/setup-python@v4
with: with:
python-version: ${{ env.PYTHON_VERSION }} python-version: ${{ env.PYTHON_VERSION }}
- name: Install dependencies - name: Install pip-audit
run: | run: |
python -m pip install --upgrade pip --root-user-action=ignore python -m pip install --upgrade pip --root-user-action=ignore
pip install --root-user-action=ignore safety bandit semgrep pip install --root-user-action=ignore pip-audit
- name: Run Safety check - name: Runtime dependencies (gate)
run: | run: pip-audit -r requirements.txt --strict --desc on
safety check -r requirements.txt --json --output safety-report.json || true
safety check -r requirements-dev.txt --json --output safety-dev-report.json || true
- name: Run Bandit security scan - name: Dev dependencies (report only)
run: | run: pip-audit -r requirements-dev.txt --desc on || echo "::warning::dev-only dependency advisories above"
bandit -r src/ -f json -o bandit-report.json || true
- name: Run Semgrep security scan code:
run: | name: Static analysis
semgrep --config=auto src/ --json --output=semgrep-report.json || true
- name: Upload security reports
uses: actions/upload-artifact@v3
with:
name: security-reports-${{ github.run_number }}
path: |
safety-report.json
safety-dev-report.json
bandit-report.json
semgrep-report.json
- name: Check for critical vulnerabilities
run: |
echo "Checking for critical vulnerabilities..."
# Check Safety results
if [ -f safety-report.json ]; then
critical_count=$(jq '.vulnerabilities | length' safety-report.json 2>/dev/null || echo "0")
if [ "$critical_count" -gt 0 ]; then
echo "Found $critical_count dependency vulnerabilities"
jq '.vulnerabilities[] | "- \(.package_name) \(.installed_version): \(.vulnerability_id)"' safety-report.json
else
echo "No dependency vulnerabilities found"
fi
fi
# Check Bandit results
if [ -f bandit-report.json ]; then
high_severity=$(jq '.results[] | select(.issue_severity == "HIGH") | length' bandit-report.json 2>/dev/null | wc -l)
if [ "$high_severity" -gt 0 ]; then
echo "Found $high_severity high-severity security issues"
else
echo "No high-severity security issues found"
fi
fi
# License compliance check
license-check:
name: License Compliance
runs-on: ubuntu-latest runs-on: ubuntu-latest
steps: steps:
- name: Checkout code - uses: actions/checkout@v4
uses: actions/checkout@v4
with:
token: ${{ secrets.GITEA_TOKEN }}
- name: Set up Python - uses: actions/setup-python@v5
uses: actions/setup-python@v4
with: with:
python-version: ${{ env.PYTHON_VERSION }} python-version: ${{ env.PYTHON_VERSION }}
- name: Install pip-licenses - name: Install bandit
run: | run: |
python -m pip install --upgrade pip --root-user-action=ignore python -m pip install --upgrade pip --root-user-action=ignore
pip install --root-user-action=ignore pip-licenses pip install --root-user-action=ignore bandit
pip install --root-user-action=ignore -r requirements.txt
- name: Check licenses - name: bandit (HIGH fails; medium/low listed)
run: | run: |
echo "Checking dependency licenses..." bandit -r src/ -q --skip B104 -ll -ii || true
bandit -r src/ -q --skip B104 --severity-level high --confidence-level medium
licenses:
name: License report
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: ${{ env.PYTHON_VERSION }}
cache: pip
cache-dependency-path: requirements.txt
# A fresh venv, not the runner's site-packages: the report must list the
# project's runtime deps, not whatever the runner image or a previous
# workflow happened to leave installed (semgrep once showed up here).
- name: Install into a clean venv
run: |
python -m venv .lic && . .lic/bin/activate
pip install --upgrade pip --root-user-action=ignore
pip install --root-user-action=ignore -r requirements.txt pip-licenses
- name: Report
run: |
. .lic/bin/activate
pip-licenses --format=markdown --with-urls --output-file=licenses.md
pip-licenses --format=json --output-file=licenses.json pip-licenses --format=json --output-file=licenses.json
pip-licenses --format=markdown --output-file=licenses.md echo "Copyleft licenses among runtime deps (informational; LGPL is fine to link from MIT):"
pip-licenses --format=plain --ignore-packages pip-licenses | grep -iE 'GPL|AGPL|LGPL' || echo " none"
# Check for problematic licenses - uses: actions/upload-artifact@v3
problematic_licenses=("GPL" "AGPL" "LGPL")
for license in "${problematic_licenses[@]}"; do
if grep -i "$license" licenses.json; then
echo "Found potentially problematic license: $license"
fi
done
echo "License check completed"
- name: Upload license report
uses: actions/upload-artifact@v3
with: with:
name: license-report-${{ github.run_number }} name: licenses-${{ github.run_number }}
path: | path: |
licenses.json
licenses.md licenses.md
licenses.json
# Dependency update check
dependency-update:
name: Check for Dependency Updates
runs-on: ubuntu-latest
steps:
- name: Checkout code
uses: actions/checkout@v4
with:
token: ${{ secrets.GITEA_TOKEN }}
- name: Set up Python
uses: actions/setup-python@v4
with:
python-version: ${{ env.PYTHON_VERSION }}
- name: Install pip-check-updates equivalent
run: |
python -m pip install --upgrade pip --root-user-action=ignore
pip install --root-user-action=ignore pip-review
- name: Check for outdated packages
run: |
echo "Checking for outdated packages..."
pip install --root-user-action=ignore -r requirements.txt
pip list --outdated --format=json > outdated-packages.json || true
if [ -s outdated-packages.json ]; then
echo "Outdated packages found:"
cat outdated-packages.json | jq -r '.[] | "- \(.name): \(.version) -> \(.latest_version)"'
else
echo "All packages are up to date"
fi
- name: Upload dependency reports
uses: actions/upload-artifact@v3
with:
name: dependency-reports-${{ github.run_number }}
path: |
outdated-packages.json
# Code quality metrics
code-quality:
name: Code Quality Metrics
runs-on: ubuntu-latest
steps:
- name: Checkout code
uses: actions/checkout@v4
with:
token: ${{ secrets.GITEA_TOKEN }}
- name: Set up Python
uses: actions/setup-python@v4
with:
python-version: ${{ env.PYTHON_VERSION }}
- name: Install quality tools
run: |
python -m pip install --upgrade pip --root-user-action=ignore
pip install --root-user-action=ignore radon xenon vulture
pip install --root-user-action=ignore -r requirements.txt
- name: Calculate code complexity
run: |
echo "Calculating code complexity..."
radon cc src/ --json > complexity-report.json
radon mi src/ --json > maintainability-report.json
echo "Complexity Summary:"
radon cc src/ --average
echo "Maintainability Summary:"
radon mi src/
- name: Find dead code
run: |
echo "Checking for dead code..."
vulture src/ --json > dead-code-report.json || true
- name: Check for code smells
run: |
echo "Checking for code smells..."
xenon --max-absolute B --max-modules A --max-average A src/ || true
- name: Upload quality reports
uses: actions/upload-artifact@v3
with:
name: code-quality-reports-${{ github.run_number }}
path: |
complexity-report.json
maintainability-report.json
dead-code-report.json
# Security summary
security-summary:
name: Security Summary
runs-on: ubuntu-latest
needs: [dependency-scan, license-check, code-quality]
if: always()
steps:
- name: Download all artifacts
uses: actions/download-artifact@v3
- name: Generate security summary
run: |
echo "# Security Scan Summary" > security-summary.md
echo "" >> security-summary.md
echo "**Scan Date:** $(date -u)" >> security-summary.md
echo "**Repository:** ${{ github.repository }}" >> security-summary.md
echo "**Commit:** ${{ github.sha }}" >> security-summary.md
echo "" >> security-summary.md
echo "## Results" >> security-summary.md
echo "" >> security-summary.md
# Dependency scan results
if [ -f security-reports-*/safety-report.json ]; then
vuln_count=$(jq '.vulnerabilities | length' security-reports-*/safety-report.json 2>/dev/null || echo "0")
if [ "$vuln_count" -eq 0 ]; then
echo "- Dependency Scan: No vulnerabilities found" >> security-summary.md
else
echo "- Dependency Scan: $vuln_count vulnerabilities found" >> security-summary.md
fi
else
echo "- Dependency Scan: Results not available" >> security-summary.md
fi
# Docker scan results (removed Trivy)
echo "- Docker Scan: Skipped (Trivy removed)" >> security-summary.md
# License check results
if [ -f license-report-*/licenses.json ]; then
echo "- License Check: Completed" >> security-summary.md
else
echo "- License Check: Results not available" >> security-summary.md
fi
# Code quality results
if [ -f code-quality-reports-*/complexity-report.json ]; then
echo "- Code Quality: Analyzed" >> security-summary.md
else
echo "- Code Quality: Results not available" >> security-summary.md
fi
echo "" >> security-summary.md
echo "## Detailed Reports" >> security-summary.md
echo "" >> security-summary.md
echo "Detailed reports are available in the workflow artifacts." >> security-summary.md
cat security-summary.md
- name: Upload security summary
uses: actions/upload-artifact@v3
with:
name: security-summary-${{ github.run_number }}
path: security-summary.md
+22
View File
@@ -148,6 +148,28 @@ grafana_data/
models/*.joblib models/*.joblib
models/cache/ models/cache/
models/metrics.json models/metrics.json
# scripts/retrain.sh working dirs (staging + one rollback generation)
models/.staging/
models/.previous/
# Playwright MCP browser artifacts (screenshots/snapshots from agent sessions) # Playwright MCP browser artifacts (screenshots/snapshots from agent sessions)
.playwright-mcp/ .playwright-mcp/
# Agent tooling state (whole dirs; the entries above only covered subpaths)
.claude/
.claude-flow/
.swarm/
# local MCP server wiring, not project config
.mcp.json
# CLAUDE.md is intentionally NOT ignored — track it if you want the agent
# conventions shared with collaborators; it is untracked today.
# Model evaluation output (regenerate with scripts/evaluate_variants.py)
models/eval_*.json
# Editor/merge leftovers and stray shell-redirect artifacts. The repo root once
# collected 56 zero-byte files named after fragments of shell commands.
*.orig
*.rej
*.bak
*~
-129
View File
@@ -1,129 +0,0 @@
# GitLab CI/CD Pipeline for Northern Thailand Ping River Monitor
stages:
- test
- build
- deploy
variables:
PYTHON_VERSION: "3.11"
PIP_CACHE_DIR: "$CI_PROJECT_DIR/.cache/pip"
cache:
paths:
- .cache/pip
- venv/
# Test stage
test:
stage: test
image: python:${PYTHON_VERSION}-slim
before_script:
- apt-get update && apt-get install -y build-essential
- python -m venv venv
- source venv/bin/activate
- pip install --upgrade pip
- pip install -r requirements-dev.txt
script:
- python test_integration.py
- python test_station_management.py
- flake8 src/ --max-line-length=100
- mypy src/
coverage: '/TOTAL.*\s+(\d+%)$/'
artifacts:
reports:
coverage_report:
coverage_format: cobertura
path: coverage.xml
paths:
- htmlcov/
expire_in: 1 week
# Code quality
code_quality:
stage: test
image: python:${PYTHON_VERSION}-slim
before_script:
- python -m venv venv
- source venv/bin/activate
- pip install black isort flake8 mypy
script:
- black --check src/ *.py
- isort --check-only src/ *.py
- flake8 src/ --max-line-length=100
- mypy src/
allow_failure: true
# Security scan
security_scan:
stage: test
image: python:${PYTHON_VERSION}-slim
before_script:
- pip install safety bandit
script:
- safety check -r requirements.txt
- bandit -r src/
allow_failure: true
# Build Docker image
build:
stage: build
image: docker:latest
services:
- docker:dind
before_script:
- docker login -u $CI_REGISTRY_USER -p $CI_REGISTRY_PASSWORD $CI_REGISTRY
script:
- docker build -t $CI_REGISTRY_IMAGE:$CI_COMMIT_SHA .
- docker build -t $CI_REGISTRY_IMAGE:latest .
- docker push $CI_REGISTRY_IMAGE:$CI_COMMIT_SHA
- docker push $CI_REGISTRY_IMAGE:latest
only:
- main
- develop
# Deploy to staging
deploy_staging:
stage: deploy
image: alpine:latest
before_script:
- apk add --no-cache curl
script:
- echo "Deploying to staging environment"
- curl -X POST "$STAGING_WEBHOOK_URL" -H "Content-Type: application/json" -d '{"image":"'$CI_REGISTRY_IMAGE:$CI_COMMIT_SHA'"}'
environment:
name: staging
url: https://staging.ping-river-monitor.example.com
only:
- develop
# Deploy to production
deploy_production:
stage: deploy
image: alpine:latest
before_script:
- apk add --no-cache curl
script:
- echo "Deploying to production environment"
- curl -X POST "$PRODUCTION_WEBHOOK_URL" -H "Content-Type: application/json" -d '{"image":"'$CI_REGISTRY_IMAGE:$CI_COMMIT_SHA'"}'
environment:
name: production
url: https://ping-river-monitor.example.com
when: manual
only:
- main
# Health check after deployment
health_check:
stage: deploy
image: alpine:latest
before_script:
- apk add --no-cache curl jq
script:
- sleep 30 # Wait for deployment
- curl -f $HEALTH_CHECK_URL/health
- curl -s $HEALTH_CHECK_URL/metrics | jq .
dependencies:
- deploy_production
only:
- main
+2 -4
View File
@@ -19,22 +19,20 @@ repos:
# Python code formatting with Black # Python code formatting with Black
- repo: https://github.com/psf/black - repo: https://github.com/psf/black
rev: 23.11.0 rev: 26.5.1
hooks: hooks:
- id: black - id: black
language_version: python3 language_version: python3
args: ['--line-length=120']
# Import sorting with isort # Import sorting with isort
- repo: https://github.com/pycqa/isort - repo: https://github.com/pycqa/isort
rev: 5.12.0 rev: 5.12.0
hooks: hooks:
- id: isort - id: isort
args: ['--profile', 'black', '--line-length', '120']
# Linting with flake8 # Linting with flake8
- repo: https://github.com/pycqa/flake8 - repo: https://github.com/pycqa/flake8
rev: 6.1.0 rev: 6.1.0
hooks: hooks:
- id: flake8 - id: flake8
args: ['--max-line-length=120', '--extend-ignore=E203,W503'] args: ['--max-line-length=100', '--extend-ignore=E203,W503']
+40
View File
@@ -0,0 +1,40 @@
# CLAUDE.md
Guidance for AI coding agents working in this repository.
## What this is
Flood monitoring and forecasting for the Ping River, Chiang Mai. Public dashboard and
API at https://water.buildfor.life/ (never publish the server's private/Tailscale IP).
Production: one systemd unit on a small VPS, `/opt/thailand-water-monitor`, user
`water-monitor`, interpreter `.venv/bin/python` (uv-managed), updated by `git pull`.
## Rules
- Python 3.11 only. `uv sync --python 3.11`; run everything as `uv run ...`.
- `make format` (black 88 / isort black profile, config in pyproject.toml) before
committing; CI fails on formatting. `make test` must stay green — tests are
synthetic-data only, never add one that needs the DB or network.
- Timestamps everywhere are Asia/Bangkok wall-clock with no offset. The dashboard
parses them with `parseTs()` and renders with `timeZone: TZ`; keep it that way.
- Model changes go through the rolling-origin harness (`scripts/evaluate_variants.py`)
and are judged on first-alert LEAD and false alarms, not MAE. Record results, positive
or negative, in `docs/FLOOD_FORECASTING.md` section 5. Do not change what is deployed
(`rise_rain` / hgb-v3) without a harness result that beats it on lead.
- `train_all()` must never silently produce a gauge-only (v2) model; the guard that
raises `RainUnavailableError` stays.
- No `git add -A`: zero-byte shell-accident files (`#`, `$(wc`, ...) have been committed
before. Stage files by name.
- Do not add Co-Authored-By trailers.
- The dashboard is a single file, `src/static/dashboard.html`, EN + TH via the `t()`
table: every user-visible string needs both languages.
## Where things are
- `src/web_api.py` FastAPI app; `src/water_scraper_v3.py` RID collector;
`src/hii_collector.py` ThaiWater/HII; `src/ml/` features/train/evaluate/predict,
`rain.py` (Open-Meteo), `dam.py`, `hii_rain.py`.
- `scripts/retrain.sh` + `water-monitor-retrain.timer`: monthly retrain with staged
promote. `scripts/dev_proxy.py`: serve the working-copy dashboard against the live API.
- `docs/FLOOD_FORECASTING.md` is the authoritative model write-up; `docs/DATA_SOURCES.md`
the source catalog.
-268
View File
@@ -1,268 +0,0 @@
# 🚀 Deployment Checklist - Northern Thailand Ping River Monitor
## ✅ Pre-Deployment Checklist
### **Code Quality**
- [ ] All tests pass (`make test`)
- [ ] Code formatting applied (`make format`)
- [ ] Linting checks pass (`make lint`)
- [ ] No security vulnerabilities (`safety check`)
- [ ] Documentation updated
- [ ] Version number updated in `setup.py` and `src/__init__.py`
### **Configuration**
- [ ] Environment variables configured (`.env` file)
- [ ] Database connection tested
- [ ] API endpoints tested
- [ ] Log levels appropriate for environment
- [ ] Security settings configured (API keys, secrets)
- [ ] Resource limits set (memory, CPU)
### **Dependencies**
- [ ] All required packages in `requirements.txt`
- [ ] No unused dependencies
- [ ] Security updates applied
- [ ] Compatible Python version (3.9+)
## 🐳 Docker Deployment
### **Pre-Docker Checklist**
- [ ] Dockerfile tested locally
- [ ] Docker Compose configuration verified
- [ ] Volume mounts configured correctly
- [ ] Network settings configured
- [ ] Health checks working
- [ ] Resource limits set
### **Docker Commands**
```bash
# Build and test locally
make docker-build
docker run --rm ping-river-monitor python run.py --test
# Deploy with Docker Compose
make docker-run
# Verify deployment
make health-check
```
### **Post-Docker Checklist**
- [ ] All services running (`docker-compose ps`)
- [ ] Health checks passing
- [ ] Logs showing normal operation
- [ ] API accessible (`curl http://localhost:8000/health`)
- [ ] Database connectivity verified
- [ ] Grafana dashboards loading
## 🌐 Production Deployment
### **Infrastructure Requirements**
- [ ] Server specifications adequate (CPU, RAM, Storage)
- [ ] Network connectivity to external APIs
- [ ] SSL certificates configured (if HTTPS)
- [ ] Firewall rules configured
- [ ] Backup strategy implemented
- [ ] Monitoring alerts configured
### **Security Checklist**
- [ ] API keys secured (environment variables)
- [ ] Database credentials secured
- [ ] HTTPS enabled for web interface
- [ ] Input validation enabled
- [ ] Rate limiting configured
- [ ] Log sanitization enabled
### **Performance Checklist**
- [ ] Database indexes created
- [ ] Connection pooling configured
- [ ] Caching enabled where appropriate
- [ ] Resource monitoring enabled
- [ ] Performance baselines established
## 📊 Monitoring Setup
### **Health Monitoring**
- [ ] Health check endpoints responding
- [ ] Database health monitoring
- [ ] API response time monitoring
- [ ] Memory usage monitoring
- [ ] Disk space monitoring
### **Alerting**
- [ ] Critical error alerts configured
- [ ] Performance degradation alerts
- [ ] Database connectivity alerts
- [ ] Disk space alerts
- [ ] API availability alerts
### **Logging**
- [ ] Log rotation configured
- [ ] Log levels appropriate
- [ ] Structured logging enabled
- [ ] Log aggregation configured (if applicable)
- [ ] Log retention policy set
## 🔄 CI/CD Pipeline
### **GitLab CI/CD**
- [ ] `.gitlab-ci.yml` configured
- [ ] Pipeline variables set
- [ ] Test stage passing
- [ ] Build stage creating artifacts
- [ ] Deploy stage configured
- [ ] Rollback procedure documented
### **Pipeline Stages**
- [ ] **Test**: Unit tests, integration tests, linting
- [ ] **Build**: Docker image creation, artifact generation
- [ ] **Deploy**: Staging deployment, production deployment
- [ ] **Verify**: Health checks, smoke tests
## 🗄️ Database Setup
### **Database Configuration**
- [ ] Database server running and accessible
- [ ] Database created with correct permissions
- [ ] Connection string configured
- [ ] Migration scripts run (if applicable)
- [ ] Backup strategy implemented
- [ ] Performance tuning applied
### **Database-Specific Checklist**
#### **SQLite**
- [ ] Database file permissions set correctly
- [ ] WAL mode enabled for better concurrency
- [ ] Regular backup scheduled
#### **MySQL/PostgreSQL**
- [ ] User accounts created with minimal privileges
- [ ] Connection pooling configured
- [ ] Query performance optimized
- [ ] Replication configured (if applicable)
#### **InfluxDB**
- [ ] Retention policies configured
- [ ] Continuous queries set up (if needed)
- [ ] Backup strategy implemented
#### **VictoriaMetrics**
- [ ] Storage configuration optimized
- [ ] Retention period set
- [ ] Resource limits configured
## 🌐 Web Interface
### **API Deployment**
- [ ] FastAPI server running
- [ ] All endpoints responding correctly
- [ ] API documentation accessible (`/docs`)
- [ ] CORS configured correctly
- [ ] Rate limiting working
- [ ] Authentication configured (if applicable)
### **Frontend Integration**
- [ ] Grafana dashboards configured
- [ ] Data sources connected
- [ ] Visualizations working
- [ ] Alerts configured
- [ ] User access configured
## 📈 Performance Verification
### **Load Testing**
- [ ] API endpoints tested under load
- [ ] Database performance under load
- [ ] Memory usage under load
- [ ] Response times acceptable
- [ ] Error rates acceptable
### **Capacity Planning**
- [ ] Expected data volume calculated
- [ ] Storage growth projected
- [ ] Scaling strategy documented
- [ ] Resource monitoring thresholds set
## 🔧 Operational Procedures
### **Maintenance**
- [ ] Update procedure documented
- [ ] Backup and restore procedures tested
- [ ] Rollback procedure documented
- [ ] Monitoring runbooks created
- [ ] Incident response procedures documented
### **Documentation**
- [ ] Deployment guide updated
- [ ] API documentation current
- [ ] Configuration documentation complete
- [ ] Troubleshooting guide available
- [ ] Contact information updated
## ✅ Post-Deployment Verification
### **Functional Testing**
- [ ] Data collection working
- [ ] API endpoints responding
- [ ] Database writes successful
- [ ] Web interface accessible
- [ ] Station management working
### **Integration Testing**
- [ ] External API connectivity
- [ ] Database integration
- [ ] Monitoring integration
- [ ] Alert system working
- [ ] Backup system working
### **Performance Testing**
- [ ] Response times acceptable
- [ ] Memory usage normal
- [ ] CPU usage normal
- [ ] Disk I/O normal
- [ ] Network usage normal
## 🚨 Rollback Plan
### **Rollback Triggers**
- [ ] Critical errors in production
- [ ] Performance degradation
- [ ] Data corruption
- [ ] Security vulnerabilities
- [ ] Service unavailability
### **Rollback Procedure**
1. [ ] Stop current deployment
2. [ ] Restore previous Docker images
3. [ ] Restore database backup (if needed)
4. [ ] Verify system functionality
5. [ ] Update monitoring and alerts
6. [ ] Document incident and lessons learned
## 📞 Support Information
### **Emergency Contacts**
- [ ] System administrator contact
- [ ] Database administrator contact
- [ ] Network administrator contact
- [ ] Application developer contact
### **Documentation Links**
- [ ] Deployment guide
- [ ] API documentation
- [ ] Troubleshooting guide
- [ ] Configuration reference
- [ ] Monitoring dashboards
---
**Deployment Date**: ___________
**Deployed By**: ___________
**Version**: v3.1.3
**Environment**: ___________
**Sign-off**:
- [ ] Technical Lead: ___________
- [ ] Operations Team: ___________
- [ ] Security Team: ___________
-193
View File
@@ -1,193 +0,0 @@
# Final GitHub Publication Checklist ✅
This checklist ensures the Thailand Water Level Monitor project is ready for GitHub publication.
## 🎯 **Project Preparation Complete**
### ✅ **Core Repository Files**
- [x] **README.md** - Comprehensive project documentation with badges and quick start
- [x] **LICENSE** - MIT License for open source distribution
- [x] **CONTRIBUTING.md** - Detailed contributor guidelines
- [x] **.gitignore** - Comprehensive ignore rules for all file types
- [x] **requirements.txt** - All Python dependencies listed and tested
### ✅ **Source Code Organization**
- [x] **src/** directory created with clean separation
- [x] **scripts/** directory for utility scripts and system files
- [x] **docs/** directory with comprehensive documentation
- [x] **grafana/** directory with visualization configuration
- [x] All temporary files removed (*.db, *.log, __pycache__)
### ✅ **Documentation Quality**
- [x] **Installation guides** for all platforms and databases
- [x] **Configuration examples** for 5 different database types
- [x] **Troubleshooting guides** for common deployment issues
- [x] **Migration guides** for updating existing systems
- [x] **API references** documenting Thai government data sources
- [x] **Notable documents** section with official resources
### ✅ **Production Readiness**
- [x] **Docker support** with Dockerfile and docker-compose
- [x] **Systemd service** configuration for Linux deployment
- [x] **Multi-database support** (SQLite, PostgreSQL, MySQL, InfluxDB, VictoriaMetrics)
- [x] **Geolocation support** for Grafana geomap visualization
- [x] **Migration scripts** for safe database schema updates
- [x] **HTTPS configuration** guide for secure deployment
### ✅ **Code Quality**
- [x] **Modular architecture** with clean separation of concerns
- [x] **Error handling** and comprehensive logging
- [x] **Configuration management** via environment variables
- [x] **Database abstraction** layer for multiple backends
- [x] **Testing utilities** (demo_databases.py)
### ✅ **Features Verified**
- [x] **Real-time data collection** from 16 Thai water stations
- [x] **15-minute scheduling** with intelligent retry logic
- [x] **Gap filling** for missing historical data
- [x] **Data validation** and error recovery
- [x] **Geolocation integration** with sample coordinates
- [x] **Grafana dashboards** with pre-built visualizations
## 🚀 **Ready for GitHub Publication**
### **Repository Information**
- **Name**: `thailand-water-monitor`
- **Description**: "Real-time water level monitoring system for Thailand's Royal Irrigation Department stations with Grafana visualization"
- **Topics**: `water-monitoring`, `thailand`, `grafana`, `timeseries`, `python`, `iot`, `environmental-monitoring`
- **License**: MIT
- **Language**: Python
### **Repository Settings**
- [x] Enable Issues for bug reports and feature requests
- [x] Enable Discussions for community support
- [x] Enable Wiki for extended documentation
- [x] Set up GitHub Pages for documentation hosting
- [x] Configure branch protection for main branch
### **Initial Release (v1.0.0)**
- **Release Title**: "Thailand Water Level Monitor v1.0.0 - Complete Monitoring Solution"
- **Release Notes**:
- Complete real-time monitoring system
- Multi-database backend support
- Grafana geomap integration
- Production-ready deployment
- Comprehensive documentation
## 📊 **Project Statistics**
### **Code Metrics**
- **Total Files**: 25+ files
- **Python Source Files**: 4 main modules
- **Documentation Files**: 12 comprehensive guides
- **Configuration Files**: 6 deployment configurations
- **Lines of Code**: ~2,000+ lines of Python
- **Documentation**: ~15,000+ words
### **Feature Coverage**
- **Database Backends**: 5 different types supported
- **Monitoring Stations**: 16 across Thailand
- **Data Collection**: Every 15 minutes
- **Data Points**: ~300 measurements per collection cycle
- **Geolocation**: GPS coordinates and geohash support
- **Visualization**: Pre-built Grafana dashboards
### **Documentation Coverage**
- **Installation**: Complete setup for all platforms
- **Configuration**: All database types documented
- **Deployment**: Docker, systemd, manual options
- **Troubleshooting**: Common issues and solutions
- **Migration**: Safe upgrade procedures
- **API**: External data source documentation
## 🌟 **Key Selling Points**
### **For Water Management Professionals**
- Real-time monitoring of 16 stations across Thailand
- Historical data analysis and trend visualization
- Alert capabilities for critical water levels
- Integration with official Thai government data sources
### **For Developers**
- Clean, modular Python codebase
- Multiple database backend options
- Docker containerization for easy deployment
- Comprehensive API documentation
### **For System Administrators**
- Production-ready deployment configurations
- Systemd service integration
- HTTPS and security configuration
- Monitoring and logging capabilities
### **For Data Scientists**
- Time-series data with geolocation
- Grafana visualization and analysis tools
- Historical data gap filling
- Export capabilities for further analysis
## 🎯 **Post-Publication Roadmap**
### **Immediate (Week 1)**
- [ ] Create GitHub repository and upload files
- [ ] Set up initial release v1.0.0
- [ ] Configure repository settings and templates
- [ ] Create project documentation website
### **Short-term (Month 1)**
- [ ] Add GitHub Actions for CI/CD
- [ ] Create issue and PR templates
- [ ] Set up automated testing
- [ ] Add code quality badges
### **Medium-term (Quarter 1)**
- [ ] Community feedback integration
- [ ] Additional database backends
- [ ] Mobile app development
- [ ] Advanced alerting system
### **Long-term (Year 1)**
- [ ] Predictive analytics features
- [ ] Machine learning integration
- [ ] Multi-country expansion
- [ ] Commercial support options
## 🏆 **Success Metrics**
### **Community Engagement**
- GitHub stars and forks
- Issue reports and feature requests
- Community contributions
- Documentation feedback
### **Technical Adoption**
- Download and deployment statistics
- Database backend usage patterns
- Performance benchmarks
- User success stories
### **Impact Measurement**
- Water management improvements
- Early warning system effectiveness
- Data accessibility improvements
- Research and academic usage
---
## ✅ **FINAL VERIFICATION**
**All checklist items completed successfully!**
The Thailand Water Level Monitor project is now:
-**Professionally organized** with clean structure
-**Comprehensively documented** with guides for all use cases
-**Production ready** with multiple deployment options
-**Community friendly** with contribution guidelines
-**Feature complete** with real-time monitoring capabilities
**🚀 Ready for GitHub publication and community engagement!** 🌊
---
*Last updated: July 30, 2025*
*Project status: Ready for publication*
-233
View File
@@ -1,233 +0,0 @@
# 🎉 Gitea Actions Setup Complete!
## 🚀 **What's Been Created**
Your **Northern Thailand Ping River Monitor** now has a complete CI/CD pipeline with Gitea Actions! Here's what's been set up:
### **🔄 Gitea Actions Workflows**
```
.gitea/workflows/
├── ci.yml # Main CI/CD pipeline
├── release.yml # Automated releases
├── security.yml # Security & dependency scanning
└── docs.yml # Documentation generation
```
### **📊 Workflow Features**
#### **1. CI/CD Pipeline (`ci.yml`)**
-**Multi-Python Testing** (3.9, 3.10, 3.11, 3.12)
-**Code Quality Checks** (flake8, mypy, black, isort)
-**Docker Multi-Arch Builds** (amd64, arm64)
-**Integration Testing** with VictoriaMetrics
-**Automated Staging Deployment** (develop branch)
-**Manual Production Deployment** (main branch)
-**Performance Testing** after deployment
#### **2. Release Management (`release.yml`)**
- 🏷️ **Tag-Based Releases** (`v*.*.*` pattern)
- 📝 **Automatic Changelog Generation**
- 🐳 **Multi-Architecture Docker Images**
- 🔒 **Security Scanning** before release
-**Comprehensive Validation** after deployment
#### **3. Security Monitoring (`security.yml`)**
- 🔒 **Daily Security Scans** (3 AM UTC)
- 📦 **Dependency Vulnerability Detection**
- 🐳 **Docker Image Security Scanning**
- 📄 **License Compliance Checking**
- 📊 **Code Quality Metrics**
- 🔄 **Automated Update Notifications**
#### **4. Documentation (`docs.yml`)**
- 📚 **API Documentation Generation**
- 🔗 **Link Validation**
- 📖 **Sphinx Documentation Building**
-**Documentation Completeness Checking**
## 🔧 **Setup Instructions**
### **1. Configure Repository Secrets**
In your Gitea repository settings, add these secrets:
```bash
# Required
GITEA_TOKEN # For container registry access
# Optional (for notifications)
SLACK_WEBHOOK_URL # Slack notifications
STAGING_WEBHOOK_URL # Staging deployment webhook
PRODUCTION_WEBHOOK_URL # Production deployment webhook
```
### **2. Enable Actions**
1. Go to your repository settings in Gitea
2. Enable "Actions" if not already enabled
3. Configure runners if using self-hosted runners
### **3. Push to Repository**
```bash
# Initialize and push
git init
git remote add origin https://git.b4l.co.th/grabowski/Northern-Thailand-Ping-River-Monitor.git
git add .
git commit -m "Initial commit with Gitea Actions workflows"
git push -u origin main
```
## 🎯 **Workflow Triggers**
### **Automatic Triggers**
- **Push to main/develop** → CI/CD Pipeline
- **Pull Request to main** → Testing & Validation
- **Daily at 2 AM UTC** → CI/CD Health Check
- **Daily at 3 AM UTC** → Security Scanning
- **Git Tag `v*.*.*`** → Release Pipeline
- **Documentation Changes** → Documentation Build
### **Manual Triggers**
- **Manual Dispatch** → Any workflow can be triggered manually
- **Release Creation** → Manual release with custom version
## 📊 **Monitoring & Status**
### **Status Badges**
Your README now includes comprehensive status badges:
- CI/CD Pipeline Status
- Security Scan Status
- Documentation Build Status
- Python Version Support
- FastAPI Version
- Docker Ready
- License Information
- Current Version
### **Workflow Artifacts**
Each workflow generates useful artifacts:
- **Test Results** and coverage reports
- **Security Scan Reports** (JSON format)
- **Docker Images** (multi-architecture)
- **Documentation** (HTML and PDF)
- **Performance Reports**
## 🚀 **Usage Examples**
### **Development Workflow**
```bash
# Create feature branch
git checkout -b feature/new-station-type
# Make changes
git add .
git commit -m "Add support for new station type"
git push origin feature/new-station-type
# Create PR in Gitea → Triggers testing
```
### **Release Workflow**
```bash
# Create and push release tag
git tag v3.1.1
git push origin v3.1.1
# → Triggers automated release pipeline
```
### **Security Monitoring**
- **Daily scans** run automatically
- **Security reports** available in Actions artifacts
- **Notifications** sent for critical vulnerabilities
## 🔍 **Validation Commands**
Test your setup locally:
```bash
# Validate workflow syntax
make validate-workflows
# Test workflow components
make workflow-test
# Run full test suite
make test
# Build Docker image
make docker-build
```
## 📈 **Performance & Optimization**
### **Caching Strategy**
- **Pip dependencies** cached across runs
- **Docker layers** cached for faster builds
- **Workflow artifacts** retained for analysis
### **Parallel Execution**
- **Matrix builds** for multiple Python versions
- **Independent jobs** for security and testing
- **Conditional execution** to skip unnecessary steps
### **Resource Management**
- **Appropriate timeouts** prevent hanging workflows
- **Artifact cleanup** manages storage usage
- **Efficient Docker builds** with multi-stage approach
## 🔒 **Security Best Practices**
### **Implemented Security**
-**Secret management** via Gitea repository secrets
-**Multi-stage Docker builds** for minimal attack surface
-**Non-root containers** for better security
-**Vulnerability scanning** before deployment
-**Dependency monitoring** with automated alerts
### **Security Scanning Coverage**
- **Python dependencies** (Safety, Bandit)
- **Docker images** (Trivy)
- **Code quality** (Semgrep)
- **License compliance** (pip-licenses)
## 📚 **Documentation**
### **Available Documentation**
- [Gitea Workflows Guide](docs/GITEA_WORKFLOWS.md) - Detailed workflow documentation
- [Contributing Guide](CONTRIBUTING.md) - How to contribute
- [Deployment Checklist](DEPLOYMENT_CHECKLIST.md) - Production deployment
- [Project Structure](docs/PROJECT_STRUCTURE.md) - Architecture overview
### **Generated Documentation**
- **API Documentation** - Auto-generated from OpenAPI spec
- **Code Documentation** - Sphinx-generated from docstrings
- **Security Reports** - Automated vulnerability reports
## 🎉 **Ready for Production!**
Your repository is now equipped with:
- 🔄 **Enterprise-grade CI/CD pipeline**
- 🔒 **Comprehensive security monitoring**
- 📊 **Automated quality assurance**
- 🚀 **Streamlined release management**
- 📚 **Automated documentation**
- 🐳 **Multi-architecture Docker support**
- 📈 **Performance monitoring**
- 🔍 **Comprehensive testing**
## 🚀 **Next Steps**
1. **Push to Gitea** and watch the workflows run
2. **Configure deployment environments** (staging/production)
3. **Set up monitoring dashboards** for workflow metrics
4. **Configure notifications** for team collaboration
5. **Create your first release** with `git tag v3.1.3`
Your **Northern Thailand Ping River Monitor** is now ready for professional development and deployment! 🎊
---
**Workflow Version**: v3.1.3
**Setup Date**: 2025-08-12
**Repository**: https://git.b4l.co.th/grabowski/Northern-Thailand-Ping-River-Monitor
-203
View File
@@ -1,203 +0,0 @@
# GitHub Publication Summary
This document summarizes the Thailand Water Level Monitor project preparation for GitHub publication.
## 📁 **Final Project Structure**
```
thailand-water-monitor/
├── 📄 README.md # Main project documentation
├── 📄 LICENSE # MIT License
├── 📄 CONTRIBUTING.md # Contributor guidelines
├── 📄 requirements.txt # Python dependencies
├── 📄 .gitignore # Git ignore rules
├── 📄 Dockerfile # Container definition
├── 📄 docker-compose.victoriametrics.yml # Complete stack deployment
├── 📂 src/ # Source Code
│ ├── 🐍 water_scraper_v3.py # Main application
│ ├── 🐍 database_adapters.py # Multi-database support
│ ├── 🐍 config.py # Configuration management
│ └── 🐍 demo_databases.py # Database testing utility
├── 📂 scripts/ # Utility Scripts
│ ├── 🐍 migrate_geolocation.py # Database migration script
│ └── ⚙️ water-monitor.service # Systemd service file
├── 📂 docs/ # Documentation
│ ├── 📖 DATABASE_DEPLOYMENT_GUIDE.md # Complete setup guide
│ ├── 📖 ENHANCED_SCHEDULER_GUIDE.md # 15-minute scheduling
│ ├── 📖 GEOLOCATION_GUIDE.md # Grafana geomap integration
│ ├── 📖 GAP_FILLING_GUIDE.md # Data integrity management
│ ├── 📖 MIGRATION_QUICKSTART.md # Quick migration guide
│ ├── 📖 VICTORIAMETRICS_SETUP.md # High-performance deployment
│ ├── 📖 HTTPS_CONFIGURATION.md # Secure deployment
│ ├── 📖 DEBIAN_TROUBLESHOOTING.md # Linux deployment issues
│ ├── 📖 PROJECT_STATUS.md # Development status
│ └── 📂 references/
│ └── 📖 NOTABLE_DOCUMENTS.md # Official Thai government resources
└── 📂 grafana/ # Grafana Configuration
├── 📂 dashboards/
│ └── 📊 water-monitoring-dashboard.json
└── 📂 provisioning/
├── 📂 dashboards/
│ └── ⚙️ dashboard.yml
└── 📂 datasources/
└── ⚙️ victoriametrics.yml
```
## ✅ **GitHub Readiness Checklist**
### **Core Files**
-**README.md** - Comprehensive project documentation with badges, features, quick start
-**LICENSE** - MIT License for open source distribution
-**CONTRIBUTING.md** - Detailed contributor guidelines and development setup
-**.gitignore** - Comprehensive ignore rules for Python, databases, logs, IDE files
-**requirements.txt** - All Python dependencies listed
### **Source Code Organization**
-**src/** directory - Clean separation of source code
-**scripts/** directory - Utility scripts and system files
-**docs/** directory - Comprehensive documentation
-**grafana/** directory - Visualization configuration
### **Documentation Quality**
-**Installation guides** - Multiple deployment options
-**Configuration examples** - All database types covered
-**Troubleshooting guides** - Common issues and solutions
-**Migration guides** - Updating existing systems
-**API references** - External data sources documented
### **Production Readiness**
-**Docker support** - Containerization ready
-**Systemd service** - Linux service configuration
-**Multi-database support** - 5 different database options
-**Geolocation support** - Grafana geomap integration
-**Migration scripts** - Safe database updates
## 🌟 **Key Features for GitHub**
### **Real-time Monitoring**
- 16 water stations across Thailand
- 15-minute data collection frequency
- Automatic gap filling and data validation
- Multi-database backend support
### **Visualization Ready**
- Pre-built Grafana dashboards
- Geomap integration with coordinates
- Real-time alerts and notifications
- Historical trend analysis
### **Production Deployment**
- Docker containerization
- VictoriaMetrics high-performance backend
- HTTPS and security configuration
- Comprehensive logging and monitoring
### **Developer Friendly**
- Clean, modular code structure
- Comprehensive documentation
- Multiple database adapters
- Easy local development setup
## 📊 **Project Statistics**
### **Code Metrics**
- **Python Files**: 4 main source files
- **Documentation**: 10+ comprehensive guides
- **Database Support**: 5 different backends
- **Monitoring Stations**: 16 across Thailand
- **Data Points**: ~300 every 15 minutes
### **Documentation Coverage**
- **Installation**: Complete setup guides for all platforms
- **Configuration**: All database types documented
- **Deployment**: Docker, systemd, and manual options
- **Troubleshooting**: Common issues and solutions
- **Migration**: Safe upgrade procedures
### **Features Implemented**
- ✅ Real-time data collection
- ✅ Multi-database support
- ✅ Geolocation integration
- ✅ Gap filling and data validation
- ✅ Grafana visualization
- ✅ Docker deployment
- ✅ Production monitoring
- ✅ Migration tools
## 🚀 **Ready for GitHub Publication**
### **Repository Setup**
1. **Create GitHub repository** - "thailand-water-monitor"
2. **Upload all files** - Complete project structure
3. **Configure repository settings**:
- Add description: "Real-time water level monitoring for Thailand's RID stations"
- Add topics: `water-monitoring`, `thailand`, `grafana`, `timeseries`, `python`
- Enable Issues and Discussions
- Set up GitHub Pages for documentation
### **Initial Release**
- **Version**: v1.0.0
- **Release Notes**: Complete feature set with multi-database support
- **Assets**: Include sample configuration files
- **Documentation**: Link to comprehensive guides
### **Community Features**
- **Issues Template**: Bug reports and feature requests
- **Pull Request Template**: Contribution guidelines
- **Discussions**: Community support and questions
- **Wiki**: Extended documentation and tutorials
## 🎯 **Post-Publication Tasks**
### **Community Building**
- Create detailed issue templates
- Set up GitHub Actions for CI/CD
- Add code quality badges
- Create project roadmap
### **Documentation Enhancement**
- Add video tutorials
- Create API documentation
- Add performance benchmarks
- Create deployment examples
### **Feature Development**
- Mobile app integration
- Additional database backends
- Advanced alerting system
- Predictive analytics
## 📞 **Support Channels**
- **GitHub Issues**: Bug reports and feature requests
- **GitHub Discussions**: Community support and questions
- **Documentation**: Comprehensive guides in docs/ directory
- **Examples**: Working configurations and deployments
## 🏆 **Project Highlights**
### **Technical Excellence**
- Clean, modular architecture
- Comprehensive error handling
- Production-ready deployment
- Multi-database abstraction
### **Documentation Quality**
- Step-by-step installation guides
- Troubleshooting for common issues
- Migration procedures for updates
- API and configuration references
### **Community Ready**
- Open source MIT license
- Contributor guidelines
- Development setup instructions
- Code quality standards
---
**The Thailand Water Level Monitor project is now fully prepared for GitHub publication with a professional structure, comprehensive documentation, and production-ready features.** 🌊
-114
View File
@@ -1,114 +0,0 @@
# 🔑 GitHub Token Setup Guide
## 🎯 **Why You Need This**
The Gitea Actions workflows use Trivy for security scanning, which needs to download vulnerability databases from GitHub. Without a GitHub token, you'll hit rate limits and the security scans will fail.
## 🚀 **Quick Setup (5 minutes)**
### **Step 1: Create GitHub Personal Access Token**
1. **Go to GitHub**: https://github.com/settings/tokens
2. **Click "Generate new token"** → "Generate new token (classic)"
3. **Configure the token**:
- **Note**: `B4L Ping River Monitor - Gitea Actions`
- **Expiration**: `90 days` (or longer)
- **Scopes**: Select `public_repo` (for public repositories)
4. **Click "Generate token"**
5. **Copy the token** (you won't see it again!)
### **Step 2: Add Token to Gitea Repository**
1. **Go to your repository**: https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor
2. **Click "Settings"** (in the repository)
3. **Click "Secrets"** in the left sidebar
4. **Click "Add Secret"**
5. **Configure the secret**:
- **Name**: `GITHUB_TOKEN`
- **Value**: Paste the token you copied from GitHub
6. **Click "Add Secret"**
### **Step 3: Verify It's Working**
1. **Trigger a workflow** by pushing a commit or manually running the security workflow
2. **Check the Actions tab** in your repository
3. **Look for the message**: `✅ GITHUB_TOKEN is configured`
## 🔒 **Security Best Practices**
### **Token Permissions**
- **Minimum required**: `public_repo` scope
- **Never use**: `repo` scope unless you need private repo access
- **Avoid**: Admin or write permissions
### **Token Management**
- **Set expiration**: Don't create tokens that never expire
- **Regular rotation**: Update tokens every 90 days
- **Monitor usage**: Check GitHub token usage in settings
### **Repository Security**
- **Only trusted contributors**: Should have access to repository secrets
- **Audit regularly**: Review who has access to secrets
- **Use organization secrets**: For multiple repositories
## 🧪 **Testing the Setup**
### **Manual Test**
```bash
# Trigger the security workflow manually
# Go to: Repository → Actions → Security & Dependency Updates → Run workflow
```
### **Automatic Test**
```bash
# Push any change to trigger workflows
git commit --allow-empty -m "Test GitHub token setup"
git push
```
### **Check Workflow Logs**
1. Go to Actions tab in your repository
2. Click on the latest "Security & Dependency Updates" run
3. Click on "Docker Security Scan" job
4. Look for: `✅ GITHUB_TOKEN is configured`
## ❌ **Troubleshooting**
### **"GITHUB_TOKEN not configured" message**
- **Problem**: Token not added to repository secrets
- **Solution**: Follow Step 2 above, ensure exact name `GITHUB_TOKEN`
### **"Bad credentials" error**
- **Problem**: Token is invalid or expired
- **Solution**: Generate a new token and update the secret
### **Rate limit errors**
- **Problem**: Token doesn't have correct permissions
- **Solution**: Ensure token has `public_repo` scope
### **Trivy still failing**
- **Problem**: Network issues or GitHub API problems
- **Solution**: Wait and retry, or check GitHub status page
## 🎉 **Success Indicators**
When everything is working correctly, you'll see:
**In workflow logs**: `✅ GITHUB_TOKEN is configured`
**Security scans**: Complete without authentication errors
**Trivy reports**: Generated and uploaded as artifacts
**No rate limit errors**: In the workflow execution
## 📚 **Additional Resources**
- [GitHub Personal Access Tokens Documentation](https://docs.github.com/en/authentication/keeping-your-account-and-data-secure/creating-a-personal-access-token)
- [Gitea Secrets Documentation](https://docs.gitea.io/en-us/usage/actions/#secrets)
- [Trivy Action Documentation](https://github.com/aquasecurity/trivy-action)
---
**Setup Time**: ~5 minutes
**Token Validity**: 90 days (recommended)
**Security Level**: High (read-only public repo access)
Your workflows will now run smoothly with proper GitHub API authentication! 🚀
-17
View File
@@ -120,10 +120,6 @@ docker-logs:
docs: docs:
cd docs && make html cd docs && make html
# Database management
db-migrate:
uv run python scripts/migrate_geolocation.py
# Monitoring # Monitoring
health-check: health-check:
curl -f http://localhost:8000/health || exit 1 curl -f http://localhost:8000/health || exit 1
@@ -149,9 +145,6 @@ setup-postgres:
test-postgres: test-postgres:
uv run python -c "from scripts.setup_postgres import test_postgres_connection; from src.config import Config; config = Config.get_database_config(); test_postgres_connection(config['connection_string'])" uv run python -c "from scripts.setup_postgres import test_postgres_connection; from src.config import Config; config = Config.get_database_config(); test_postgres_connection(config['connection_string'])"
encode-password:
uv run python scripts/encode_password.py
migrate-sqlite: migrate-sqlite:
uv run python scripts/migrate_sqlite_to_postgres.py uv run python scripts/migrate_sqlite_to_postgres.py
@@ -161,16 +154,6 @@ migrate-fast:
analyze-sqlite: analyze-sqlite:
uv run python scripts/migrate_sqlite_to_postgres.py --dry-run uv run python scripts/migrate_sqlite_to_postgres.py --dry-run
# Distribution
build-exe:
uv run python build_simple.py
package: build-exe
@echo "Creating distribution package..."
@if exist dist\ping-river-monitor-distribution.zip del dist\ping-river-monitor-distribution.zip
@cd dist && powershell -Command "Compress-Archive -Path * -DestinationPath ping-river-monitor-distribution.zip -Force"
@echo "✅ Distribution package created: dist/ping-river-monitor-distribution.zip"
# Git helpers # Git helpers
git-setup: git-setup:
git remote add origin https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor.git git remote add origin https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor.git
+121 -474
View File
@@ -1,510 +1,157 @@
# Northern Thailand Ping River Monitor 🏔️ # Northern Thailand Ping River Monitor
A comprehensive real-time water level monitoring system for the Ping River Basin in Northern Thailand, covering Royal Irrigation Department (RID) stations from Chiang Dao to Nakhon Sawan with advanced data collection, storage, and visualization capabilities. Live water levels, discharge, rainfall and machine-learning flood forecasts for the
Ping River basin around Chiang Mai. Collects hourly gauge data from public sources,
keeps the full history in PostgreSQL, and serves a bilingual dashboard, an open REST
API, and 6/12/24-hour flood-risk forecasts per gauge.
[![CI/CD](https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor/actions/workflows/ci.yml/badge.svg)](https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor/actions) [![Security](https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor/actions/workflows/security.yml/badge.svg)](https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor/actions) [![Documentation](https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor/actions/workflows/docs.yml/badge.svg)](https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor/actions) [![Python](https://img.shields.io/badge/Python-3.9+-blue.svg)](https://python.org) [![FastAPI](https://img.shields.io/badge/FastAPI-0.104+-green.svg)](https://fastapi.tiangolo.com) [![Docker](https://img.shields.io/badge/Docker-Ready-blue.svg)](https://docker.com) [![License](https://img.shields.io/badge/License-MIT-green.svg)](LICENSE) [![Version](https://img.shields.io/badge/Version-v3.1.3-blue.svg)](https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor/releases) **Live: [water.buildfor.life](https://water.buildfor.life/)** · API reference at
[/docs](https://water.buildfor.life/docs) · built by [buildfor.life](https://buildfor.life)
after the [October 2024 flood](https://buildfor.life/blog/chiang-mai-flood-2024/) —
background in [Teaching a Model to See the Ping River Rise 13 Hours Early](https://buildfor.life/blog/ping-river-monitor/).
## 🌟 Features [![CI](https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor/actions/workflows/ci.yml/badge.svg)](https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor/actions)
[![Security](https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor/actions/workflows/security.yml/badge.svg)](https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor/actions)
[![Docs](https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor/actions/workflows/docs.yml/badge.svg)](https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor/actions)
[![Python 3.11](https://img.shields.io/badge/Python-3.11-blue.svg)](https://python.org)
[![License: MIT](https://img.shields.io/badge/License-MIT-green.svg)](LICENSE)
### 📊 **Real-time Data Collection** ## What it does
- **16 Monitoring Stations** across Thailand
- **15-minute Collection Frequency** with intelligent scheduling
- **Automatic Gap Filling** for missing historical data
- **Data Validation** and error recovery mechanisms
- **Rate Limiting** to prevent API abuse
### 🌐 **Web API Interface (NEW!)** - **Collects** hourly water level and discharge from 16 Royal Irrigation Department
- **FastAPI-powered REST API** with interactive documentation (RID) telemetry gauges, Chiang Dao to the southern basin, since 2018-08; hourly
- **Station Management** - Add, update, and remove monitoring stations rainfall and water level from 400+ ThaiWater/HII stations; Open-Meteo catchment
- **Real-time health monitoring** and system status rainfall (archive + 48 h forecast); daily Mae Ngat reservoir state. Every source and
- **Manual data collection triggers** via web interface its quirks: [docs/DATA_SOURCES.md](docs/DATA_SOURCES.md).
- **Comprehensive metrics** and performance monitoring - **Fills gaps.** The raw RID grid had readings for ~56 % of hours; a full-history
- **CORS support** for web applications re-fetch plus HII cross-fill brought it to ~93 %. `GET /api/stats` reports the
current figure.
- **Forecasts.** Per gauge and horizon, a gradient-boosted model predicts the rise
within 6/12/24 h and the probability of crossing the station's warning and danger
levels. Trained on the monitor's own history plus catchment rain; evaluated
rolling-origin, event by event. On the October 2024 record flood, trained only on
data through August 2024, the first alert came **13 hours before** P.1 crossed
3.70 m. Everything about the model, including what did not work:
[docs/FLOOD_FORECASTING.md](docs/FLOOD_FORECASTING.md).
- **Shows it.** A Leaflet map with the river drawn as OSM geometry and styled by live
discharge, rain gauges, the Chiang Mai inundation zones, per-station history, the
forecast card, a replay of the 2024 flood, English/Thai, light/dark.
- **Notifies.** Public push alerts over a self-hosted [ntfy](https://ntfy.sh): one
message when a gauge crosses its warning or danger level, one all-clear on the
way down, an opt-in early-warning topic from the model, nothing in between.
Subscribe from the free app, no account. Matrix room alerts for a team are
also supported.
### 🗄️ **Multi-Database Support** ## Quick start
- **VictoriaMetrics** (Recommended) - High-performance time-series
- **InfluxDB** - Purpose-built time-series database
- **PostgreSQL + TimescaleDB** - Relational with time-series optimization
- **MySQL** - Traditional relational database
- **SQLite** - Local development and testing
### 🗺️ **Geolocation Support** Python **3.11** (3.13 breaks the pinned `psycopg2-binary`), PostgreSQL for anything
- **Grafana Geomap** integration ready beyond a quick look, [uv](https://docs.astral.sh/uv/).
- **GPS coordinates** and geohash support
- **Interactive mapping** of water stations
### 📈 **Visualization & Monitoring**
- **Pre-built Grafana dashboards**
- **Real-time alerts** and notifications
- **Historical trend analysis**
- **Built-in metrics collection** (counters, gauges, histograms)
- **Health checks** for database, API, and system resources
### 🚀 **Production Ready**
- **Docker containerization** with multi-service support
- **Systemd service** configuration
- **HTTPS support** with SSL certificates
- **Comprehensive logging** with rotation and colored output
- **Type safety** with Pydantic models and type hints
- **Custom exception handling** for better error management
## 🚀 Quick Start
### Prerequisites
- Python 3.9 or higher
- Internet connection for data fetching
- Database server (optional - SQLite works out of the box)
### Installation
```bash ```bash
# Clone the repository
git clone https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor.git git clone https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor.git
cd Northern-Thailand-Ping-River-Monitor cd Northern-Thailand-Ping-River-Monitor
uv sync --python 3.11
# Quick setup with Make cp .env.example .env # DB_TYPE, POSTGRES_CONNECTION_STRING, optional MATRIX_*
make dev-setup uv run python run.py --web-api # dashboard + API on http://localhost:8000
# Or manual setup:
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env
``` ```
### Basic Usage `DB_TYPE=sqlite` works for the dashboard and API; the forecasting path expects the
PostgreSQL history.
```bash ```bash
# Test run with SQLite (default) uv run python run.py --status # collector status
make run-test uv run python run.py --test # one collection cycle
# or: python run.py --test uv run python run.py --fill-gaps 7 # re-fetch the last 7 days from RID
uv run python run.py --collect-hii # one ThaiWater/HII collection cycle
# Run continuous monitoring uv run python run.py --alert-check # evaluate thresholds, notify Matrix
make run uv run python scripts/train_flood_model.py --stations all # retrain (~12 min)
# or: python run.py make test # pytest, synthetic data, no network
make format # black + isort (the CI contract)
# Start web API server (NEW!)
make run-api
# or: python run.py --web-api
# Run all tests
make test
# Demo different databases
python src/demo_databases.py
``` ```
### 🌐 Web API Interface (NEW!) ## API
The system now includes a comprehensive FastAPI web interface: Read-only, no key, JSON. Base URL `https://water.buildfor.life`; timestamps are
Asia/Bangkok wall-clock without an offset suffix.
| Endpoint | Returns |
| --- | --- |
| `GET /stations` | The 16 RID gauges: code, Thai/English names, coordinates |
| `GET /measurements/latest?limit=N` | Newest reading per station |
| `GET /measurements/history/{code}?hours=N` | Hourly history; or `?start=YYYY-MM-DD&end=YYYY-MM-DD`; `limit` ≤ 100000 |
| `GET /forecast` | Current flood-risk forecast, every station × horizon, with thresholds and P.1 inundation-stage probabilities |
| `GET /api/forecast/history/{code}?hours=N&horizon=24` | Forecasts as issued, for auditing lead time after the fact |
| `GET /api/hii/rainfall/latest`, `/api/hii/waterlevel/latest` | Latest ThaiWater/HII gauge readings |
| `GET /api/hii/rainfall/catchment?days=N` | HII gauge catchment-mean rain next to the Open-Meteo series the model uses |
| `GET /api/forecast/skill?station_code=P.1` | Issued forecasts vs what happened, per deployed model version |
| `GET /api/notifications` | ntfy server and topic names for the subscribe panel |
| `GET /api/stats` | Row counts per source, date range, coverage |
| `GET /health` | DB / upstream / memory checks |
Interactive reference with schemas: [water.buildfor.life/docs](https://water.buildfor.life/docs).
Responses are cached briefly server-side; poll no faster than once a minute — the data
changes hourly.
## Deployment
Production is a systemd unit on a small VPS behind Cloudflare, updated by `git pull`.
`scripts/install.sh` (run as root from a checkout) creates the `water-monitor` user,
deploys to `/opt/thailand-water-monitor`, runs `uv sync` into `.venv`, installs
`water-monitor.service` and the monthly `water-monitor-retrain.timer`.
```bash ```bash
# Start the web API
python run.py --web-api
# Access the API at:
# - Dashboard: http://localhost:8000
# - Interactive docs: http://localhost:8000/docs
# - Health check: http://localhost:8000/health
# - Latest data: http://localhost:8000/measurements/latest
```
**Key API Endpoints:**
- `GET /` - Web dashboard
- `GET /health` - System health status
- `GET /metrics` - Application metrics
- `GET /stations` - List all monitoring stations
- `POST /stations` - Add new monitoring station
- `PUT /stations/{id}` - Update station information
- `DELETE /stations/{id}` - Remove monitoring station
- `GET /measurements/latest` - Latest measurements
- `GET /measurements/station/{code}` - Station-specific data
- `POST /scrape/trigger` - Trigger manual data collection
## 📊 Station Information
The system monitors **16 water stations** along the Ping River Basin in Northern Thailand:
| Station | Thai Name | English Name | Location |
|---------|-----------|--------------|----------|
| P.1 | สะพานนวรัฐ | Nawarat Bridge | Nakhon Sawan |
| P.5 | สะพานท่านาง | Tha Nang Bridge | - |
| P.20 | บ้านเชียงดาว | Ban Chiang Dao | Chiang Mai |
| P.21 | บ้านริมใต้ | Ban Rim Tai | - |
| P.4A | บ้านแม่แตง | Ban Mae Taeng | Chiang Mai |
| P.67 | บ้านแม่แต | Ban Tae | - |
| P.75 | บ้านช่อแล | Ban Chai Lat | - |
| P.76 | บ้านแม่อีไฮ | Banb Mae I Hai | - |
| P.77 | บ้านสบแม่สะป๊วด | Baan Sop Mae Sapuord | - |
| P.81 | บ้านโป่ง | Ban Pong | - |
| P.82 | บ้านสบวิน | Ban Sob win | - |
| P.84 | บ้านพันตน | Ban Panton | - |
| P.85 | บ้านหล่ายแก้ว | Baan Lai Kaew | - |
| P.87 | บ้านป่าซาง | Ban Pa Sang | - |
| P.92 | บ้านเมืองกึ๊ด | Ban Muang Aut | - |
| P.103 | สะพานวงแหวนรอบ 3 | Ring Bridge 3 | Bangkok |
### Data Metrics
- **Water Level**: Measured in meters (m)
- **Discharge**: Flow rate in cubic meters per second (cms)
- **Discharge Percentage**: Relative to station capacity
- **Timestamp**: Thai time (UTC+7) with Buddhist calendar support
## 🗄️ Database Configuration
### VictoriaMetrics (Recommended)
**High-performance time-series database with excellent compression and query speed.**
```bash
# Environment variables
export DB_TYPE=victoriametrics
export VM_HOST=localhost
export VM_PORT=8428
# Quick start with Docker
docker run -d \
--name victoriametrics \
-p 8428:8428 \
-v victoria-metrics-data:/victoria-metrics-data \
victoriametrics/victoria-metrics:latest \
--storageDataPath=/victoria-metrics-data \
--retentionPeriod=2y \
--httpListenAddr=:8428
```
### Complete Stack with Grafana
```bash
# Start the complete monitoring stack
docker-compose -f docker-compose.victoriametrics.yml up -d
# Access Grafana at http://localhost:3000
# Username: admin, Password: admin_password
```
### Other Database Options
<details>
<summary>InfluxDB Configuration</summary>
```bash
export DB_TYPE=influxdb
export INFLUX_HOST=localhost
export INFLUX_PORT=8086
export INFLUX_DATABASE=water_monitoring
export INFLUX_USERNAME=water_user
export INFLUX_PASSWORD=your_password
```
</details>
<details>
<summary>PostgreSQL Configuration</summary>
```bash
export DB_TYPE=postgresql
export POSTGRES_CONNECTION_STRING=postgresql://user:password@localhost:5432/water_monitoring
```
</details>
<details>
<summary>MySQL Configuration</summary>
```bash
export DB_TYPE=mysql
export MYSQL_CONNECTION_STRING=mysql://user:password@localhost:3306/water_monitoring
```
</details>
## 📈 Grafana Dashboards
### Pre-built Dashboard Features
- **Real-time water levels** across all stations
- **Historical trends** and patterns
- **Discharge monitoring** with percentage indicators
- **Station status** and health monitoring
- **Geomap visualization** of station locations
- **Alert thresholds** for critical water levels
### Sample Queries
**VictoriaMetrics/Prometheus:**
```promql
# Current water levels
water_level
# High discharge alerts
water_discharge_percent > 80
# Station-specific data
water_level{station_code="P.1"}
```
**SQL Databases:**
```sql
-- Latest readings from all stations
SELECT s.station_code, s.english_name, m.water_level, m.discharge
FROM stations s
JOIN water_measurements m ON s.id = m.station_id
WHERE m.timestamp = (SELECT MAX(timestamp) FROM water_measurements WHERE station_id = s.id);
```
## 🚀 Production Deployment
### Docker Deployment
```bash
# Build the image
docker build -t thailand-water-monitor .
# Run with environment variables
docker run -d \
--name water-monitor \
-e DB_TYPE=victoriametrics \
-e VM_HOST=victoriametrics \
thailand-water-monitor
```
### Systemd Service (Linux)
The install script sets everything up: a dedicated `water-monitor` system user,
a deploy to `/opt/thailand-water-monitor`, a uv-managed virtualenv, and the
enabled systemd unit.
```bash
# From a checkout of the repo, as root:
sudo bash scripts/install.sh sudo bash scripts/install.sh
# Then start and check:
sudo systemctl start water-monitor.service sudo systemctl start water-monitor.service
systemctl status water-monitor.service systemctl list-timers water-monitor-retrain.timer
``` ```
Fill in `/opt/thailand-water-monitor/.env` (Matrix token/room, DB settings) The retrain timer runs `scripts/retrain.sh`, which trains into `models/.staging`,
before starting if the script reports it is missing. refuses to promote anything that is not a rain-enabled (`hgb-v3+`) set covering the
expected stations, and renames the bundles into place. Details and the operations
runbook: [docs/FLOOD_FORECASTING.md](docs/FLOOD_FORECASTING.md) sections 68.
<details> ## Repository layout
<summary>Manual setup (if you prefer not to use the script)</summary>
```bash
sudo useradd --system --no-create-home --shell /usr/sbin/nologin water-monitor
sudo cp scripts/water-monitor.service /etc/systemd/system/
sudo systemctl enable water-monitor.service
sudo systemctl start water-monitor.service
```
</details>
### Migration for Existing Systems
If you have an existing installation, use the migration script to add geolocation support:
```bash
# Stop the service
sudo systemctl stop water-monitor
# Run migration
python scripts/migrate_geolocation.py
# Restart the service
sudo systemctl start water-monitor
```
## 🔧 Command Line Tools
### Main Application
```bash
python src/water_scraper_v3.py # Run continuous monitoring
python src/water_scraper_v3.py --test # Single test cycle
python src/water_scraper_v3.py --help # Show help
```
### Data Management
```bash
python src/water_scraper_v3.py --check-gaps 7 # Check for missing data (7 days)
python src/water_scraper_v3.py --fill-gaps 7 # Fill missing data gaps
python src/water_scraper_v3.py --update-data 2 # Update existing data (2 days)
```
### Database Testing
```bash
python src/demo_databases.py # SQLite demo
python src/demo_databases.py victoriametrics # VictoriaMetrics demo
python src/demo_databases.py all # Test all databases
```
## 📚 Documentation
### Core Documentation
- **[Data Sources & API Catalog](docs/DATA_SOURCES.md)** - Every ingested and available data source (RID, ThaiWater/HII, dams, rainfall, forecasts)
- **[Installation Guide](docs/DATABASE_DEPLOYMENT_GUIDE.md)** - Complete setup instructions
- **[Scheduler Guide](docs/ENHANCED_SCHEDULER_GUIDE.md)** - 15-minute scheduling system
- **[Geolocation Guide](docs/GEOLOCATION_GUIDE.md)** - Grafana geomap integration
- **[Gap Filling Guide](docs/GAP_FILLING_GUIDE.md)** - Data integrity management
### Deployment Guides
- **[VictoriaMetrics Setup](docs/VICTORIAMETRICS_SETUP.md)** - High-performance deployment
- **[HTTPS Configuration](docs/HTTPS_CONFIGURATION.md)** - Secure deployment
- **[Debian Troubleshooting](docs/DEBIAN_TROUBLESHOOTING.md)** - Linux deployment issues
### References
- **[Notable Documents](docs/references/NOTABLE_DOCUMENTS.md)** - Official Thai government resources
- **[Migration Guide](docs/MIGRATION_QUICKSTART.md)** - Updating existing systems
## 🔍 Troubleshooting
### Common Issues
**Database Connection Errors:**
```bash
# Check database status
python src/demo_databases.py
# Test specific database
python src/demo_databases.py victoriametrics
```
**Missing Data:**
```bash
# Check for gaps
python src/water_scraper_v3.py --check-gaps 7
# Fill missing data
python src/water_scraper_v3.py --fill-gaps 7
```
**Service Issues:**
```bash
# Check service status
sudo systemctl status water-monitor
# View logs
sudo journalctl -u water-monitor -f
```
### Health Checks
```bash
# VictoriaMetrics health
curl http://localhost:8428/health
# Check latest data
curl "http://localhost:8428/api/v1/query?query=water_level"
# Application logs
tail -f water_monitor.log
```
## 🌐 API Integration
### VictoriaMetrics API Examples
```bash
# Query current water levels
curl "http://localhost:8428/api/v1/query?query=water_level"
# Query discharge rates for last hour
curl "http://localhost:8428/api/v1/query_range?query=water_discharge&start=$(date -d '1 hour ago' +%s)&end=$(date +%s)&step=300"
# Query specific station
curl "http://localhost:8428/api/v1/query?query=water_level{station_code=\"P.1\"}"
# High discharge alerts
curl "http://localhost:8428/api/v1/query?query=water_discharge_percent>80"
```
## 📊 Performance
### System Requirements
- **CPU**: 1-2 cores (minimal load)
- **RAM**: 512MB - 2GB (depending on database)
- **Storage**: 1GB+ (for historical data)
- **Network**: Stable internet connection
### Performance Metrics
- **Data Collection**: ~300 data points every 15 minutes
- **Database Write Speed**: 1000+ points/second (VictoriaMetrics)
- **Query Response**: <100ms for recent data
- **Storage Efficiency**: 70x compression vs. raw data
## 🤝 Contributing
Contributions are welcome! Please:
1. Fork the repository
2. Create a feature branch
3. Make your changes
4. Add tests if applicable
5. Submit a pull request
### Development Setup
```bash
# Clone your fork
git clone https://github.com/your-username/thailand-water-monitor.git
cd thailand-water-monitor
# Install development dependencies
pip install -r requirements.txt
pip install pytest black flake8
# Run tests
pytest
# Format code
black src/
```
## 📄 License
This project is licensed under the MIT License - see the [LICENSE](LICENSE) file for details.
## 🙏 Acknowledgments
- **Royal Irrigation Department (RID)** of Thailand for providing the data API
- **VictoriaMetrics** team for the excellent time-series database
- **Grafana** team for the visualization platform
- **Python community** for the amazing libraries and tools
## 📞 Support
- **Issues**: [GitHub Issues](https://github.com/your-username/thailand-water-monitor/issues)
- **Discussions**: [GitHub Discussions](https://github.com/your-username/thailand-water-monitor/discussions)
- **Documentation**: [Project Wiki](https://github.com/your-username/thailand-water-monitor/wiki)
---
## 📁 Project Structure
``` ```
Northern-Thailand-Ping-River-Monitor/ src/ collector, API (web_api.py), dashboard (static/dashboard.html)
├── src/ # Main application code src/ml/ features, training, evaluation harness, prediction, rain/dam/HII loaders
├── tests/ # Test suite scripts/ train_flood_model.py, retrain.sh, evaluate_variants.py, install.sh, dev_proxy.py
├── docs/ # Documentation tests/ pytest suite (synthetic data; no DB or network)
├── grafana/ # Grafana dashboards docs/ FLOOD_FORECASTING.md, DATA_SOURCES.md, deployment and station guides
├── scripts/ # Utility scripts models/ trained bundles + metrics.json (gitignored) and evaluation results (tracked)
├── docker-compose.yml # Docker deployment .gitea/workflows/ ci (format/lint/tests), security (pip-audit/bandit), docs (link + OpenAPI checks)
├── Makefile # Development tasks
└── requirements.txt # Dependencies
``` ```
See [docs/PROJECT_STRUCTURE.md](docs/PROJECT_STRUCTURE.md) for detailed architecture information. ## Documentation
## 🔄 CI/CD & Automation - [docs/FLOOD_FORECASTING.md](docs/FLOOD_FORECASTING.md) — the model: data, features, evaluation, measured performance, negatives, deployment, retraining
- [docs/DATA_SOURCES.md](docs/DATA_SOURCES.md) — every ingested and candidate source, endpoints, quirks
- [docs/STATION_MANAGEMENT_GUIDE.md](docs/STATION_MANAGEMENT_GUIDE.md) — adding/editing gauges
- [docs/DATABASE_DEPLOYMENT_GUIDE.md](docs/DATABASE_DEPLOYMENT_GUIDE.md), [POSTGRESQL_SETUP.md](POSTGRESQL_SETUP.md) — database setup
- [docs/NOTIFICATIONS.md](docs/NOTIFICATIONS.md) — public push alerts: topics, semantics, ntfy deployment
- [docs/MATRIX_QUICK_START.md](docs/MATRIX_QUICK_START.md) — Matrix room alerts for a team
- [docs/GAP_FILLING_GUIDE.md](docs/GAP_FILLING_GUIDE.md) — data integrity tooling
- [docs/references/NOTABLE_DOCUMENTS.md](docs/references/NOTABLE_DOCUMENTS.md) — official Thai government resources
- Public overview: [buildfor.life/docs/tooling/ping-river-monitor](https://buildfor.life/docs/tooling/ping-river-monitor/)
The project includes comprehensive Gitea Actions workflows: Other database backends (VictoriaMetrics, InfluxDB, MySQL, SQLite) and the Grafana
dashboards under `grafana/` are supported by the adapters but not what production
runs; see [docs/VICTORIAMETRICS_SETUP.md](docs/VICTORIAMETRICS_SETUP.md) if you want them.
- **🧪 CI/CD Pipeline** - Automated testing, building, and deployment ## Contributing
- **🔒 Security Scanning** - Daily vulnerability and dependency checks
- **📚 Documentation** - Automated API docs and validation
- **🚀 Release Management** - Automated releases with multi-arch Docker builds
See [docs/GITEA_WORKFLOWS.md](docs/GITEA_WORKFLOWS.md) for detailed workflow documentation. `make format` before committing (black 88 columns, isort black profile — the CI gate),
`make test` must stay green, tests use synthetic data only. See
[CONTRIBUTING.md](CONTRIBUTING.md). Issues and merge requests on
[git.b4l.co.th](https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor).
## 🔗 Repository ## Data sources and thanks
- **Main Repository**: https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor Royal Irrigation Department (RID) gauge telemetry; Hydro-Informatics Institute (HII) /
- **Issues**: https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor/issues ThaiWater open API; Open-Meteo; OpenStreetMap contributors for the river geometry;
- **Actions**: https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor/actions Chiang Mai Municipality for the inundation map the P.1 stages are keyed to. All
- **Documentation**: [docs/](docs/) instruments are theirs; we aggregate, store, fill gaps and forecast.
**Made with ❤️ for water resource monitoring in Northern Thailand's Ping River Basin** ## License
MIT — see [LICENSE](LICENSE).
-311
View File
@@ -1,311 +0,0 @@
#!/usr/bin/env python3
"""
Build script to create a standalone executable for Northern Thailand Ping River Monitor
"""
import os
import shutil
import sys
from pathlib import Path
def create_spec_file():
"""Create PyInstaller spec file"""
spec_content = """
# -*- mode: python ; coding: utf-8 -*-
block_cipher = None
# Data files to include
data_files = [
('.env', '.'),
('sql/*.sql', 'sql'),
('README.md', '.'),
('POSTGRESQL_SETUP.md', '.'),
('SQLITE_MIGRATION.md', '.'),
]
# Hidden imports that PyInstaller might miss
hidden_imports = [
'psycopg2',
'psycopg2-binary',
'sqlalchemy.dialects.postgresql',
'sqlalchemy.dialects.sqlite',
'sqlalchemy.dialects.mysql',
'influxdb',
'pymysql',
'dotenv',
'pydantic',
'fastapi',
'uvicorn',
'schedule',
'pandas',
'requests',
'psutil',
]
a = Analysis(
['run.py'],
pathex=['.'],
binaries=[],
datas=data_files,
hiddenimports=hidden_imports,
hookspath=[],
hooksconfig={},
runtime_hooks=[],
excludes=[
'tkinter',
'matplotlib',
'PIL',
'jupyter',
'notebook',
'IPython',
],
win_no_prefer_redirects=False,
win_private_assemblies=False,
cipher=block_cipher,
noarchive=False,
)
pyz = PYZ(a.pure, a.zipped_data, cipher=block_cipher)
exe = EXE(
pyz,
a.scripts,
a.binaries,
a.zipfiles,
a.datas,
[],
name='ping-river-monitor',
debug=False,
bootloader_ignore_signals=False,
strip=False,
upx=True,
upx_exclude=[],
runtime_tmpdir=None,
console=True,
disable_windowed_traceback=False,
argv_emulation=False,
target_arch=None,
codesign_identity=None,
entitlements_file=None,
icon='icon.ico' if os.path.exists('icon.ico') else None,
)
"""
with open("ping-river-monitor.spec", "w") as f:
f.write(spec_content.strip())
print("[OK] Created ping-river-monitor.spec")
def install_pyinstaller():
"""Install PyInstaller if not present"""
try:
import PyInstaller
print("[OK] PyInstaller already installed")
except ImportError:
print("Installing PyInstaller...")
os.system("uv add --dev pyinstaller")
print("[OK] PyInstaller installed")
def build_executable():
"""Build the executable"""
print("🔨 Building executable...")
# Clean previous builds
if os.path.exists("dist"):
shutil.rmtree("dist")
if os.path.exists("build"):
shutil.rmtree("build")
# Build with PyInstaller using uv
result = os.system("uv run pyinstaller ping-river-monitor.spec --clean --noconfirm")
if result == 0:
print("✅ Executable built successfully!")
# Copy additional files to dist directory
dist_dir = Path("dist")
if dist_dir.exists():
# Copy .env file if it exists
if os.path.exists(".env"):
shutil.copy2(".env", dist_dir / ".env")
print("✅ Copied .env file")
# Copy documentation
for doc in ["README.md", "POSTGRESQL_SETUP.md", "SQLITE_MIGRATION.md"]:
if os.path.exists(doc):
shutil.copy2(doc, dist_dir / doc)
print(f"✅ Copied {doc}")
# Copy SQL files
if os.path.exists("sql"):
shutil.copytree("sql", dist_dir / "sql", dirs_exist_ok=True)
print("✅ Copied SQL files")
print(f"\n🎉 Executable created: {dist_dir / 'ping-river-monitor.exe'}")
print(f"📁 All files in: {dist_dir.absolute()}")
else:
print("❌ Build failed!")
return False
return True
def create_batch_files():
"""Create convenient batch files"""
batch_files = {
"start.bat": """@echo off
echo Starting Ping River Monitor...
ping-river-monitor.exe
pause
""",
"start-api.bat": """@echo off
echo Starting Ping River Monitor Web API...
ping-river-monitor.exe --web-api
pause
""",
"test.bat": """@echo off
echo Running Ping River Monitor test...
ping-river-monitor.exe --test
pause
""",
"status.bat": """@echo off
echo Checking Ping River Monitor status...
ping-river-monitor.exe --status
pause
""",
}
dist_dir = Path("dist")
for filename, content in batch_files.items():
batch_file = dist_dir / filename
with open(batch_file, "w") as f:
f.write(content)
print(f"✅ Created {filename}")
def create_readme():
"""Create deployment README"""
readme_content = """# Ping River Monitor - Standalone Executable
This is a standalone executable version of the Northern Thailand Ping River Monitor.
## Quick Start
1. **Configure Database**: Edit `.env` file with your PostgreSQL settings
2. **Test Connection**: Double-click `test.bat`
3. **Start Monitoring**: Double-click `start.bat`
4. **Web Interface**: Double-click `start-api.bat`
## Files Included
- `ping-river-monitor.exe` - Main executable
- `.env` - Configuration file (EDIT THIS!)
- `start.bat` - Start continuous monitoring
- `start-api.bat` - Start web API server
- `test.bat` - Run a test cycle
- `status.bat` - Check system status
- `README.md`, `POSTGRESQL_SETUP.md` - Documentation
- `sql/` - Database initialization scripts
## Configuration
Edit `.env` file:
```
DB_TYPE=postgresql
POSTGRES_HOST=your-server-ip
POSTGRES_PORT=5432
POSTGRES_DB=water_monitoring
POSTGRES_USER=your-username
POSTGRES_PASSWORD=your-password
```
## Usage
### Command Line
```cmd
# Continuous monitoring
ping-river-monitor.exe
# Single test run
ping-river-monitor.exe --test
# Web API server
ping-river-monitor.exe --web-api
# Check status
ping-river-monitor.exe --status
```
### Batch Files
- Just double-click the `.bat` files for easy operation
## Troubleshooting
1. **Database Connection Issues**
- Check `.env` file settings
- Verify PostgreSQL server is accessible
- Test with `test.bat`
2. **Permission Issues**
- Run as administrator if needed
- Check firewall settings for API mode
3. **Log Files**
- Check `water_monitor.log` for detailed logs
- Logs are created in the same directory as the executable
## Support
For issues or questions, check the documentation files included.
"""
with open("dist/DEPLOYMENT_README.txt", "w") as f:
f.write(readme_content)
print("✅ Created DEPLOYMENT_README.txt")
def main():
"""Main build process"""
print("Building Ping River Monitor Executable")
print("=" * 50)
# Check if we're in the right directory
if not os.path.exists("run.py"):
print(
"❌ Error: run.py not found. Please run this from the project root directory."
)
return False
# Install PyInstaller
install_pyinstaller()
# Create spec file
create_spec_file()
# Build executable
if not build_executable():
return False
# Create convenience files
create_batch_files()
create_readme()
print("\n" + "=" * 50)
print("🎉 BUILD COMPLETE!")
print("📁 Check the 'dist' folder for your executable")
print("💡 Edit the .env file before distributing")
print("🚀 Ready for deployment!")
return True
if __name__ == "__main__":
success = main()
sys.exit(0 if success else 1)
-112
View File
@@ -1,112 +0,0 @@
#!/usr/bin/env python3
"""
Simple build script for standalone executable
"""
import os
import shutil
import sys
from pathlib import Path
def main():
print("Building Ping River Monitor Executable")
print("=" * 50)
# Check if PyInstaller is installed
try:
import PyInstaller
print("[OK] PyInstaller available")
except ImportError:
print("[INFO] Installing PyInstaller...")
os.system("uv add --dev pyinstaller")
# Clean previous builds
if os.path.exists("dist"):
shutil.rmtree("dist")
print("[CLEAN] Removed old dist directory")
if os.path.exists("build"):
shutil.rmtree("build")
print("[CLEAN] Removed old build directory")
# Build command with all necessary options
cmd = [
"uv",
"run",
"pyinstaller",
"--onefile",
"--console",
"--name=ping-river-monitor",
"--add-data=.env;.",
"--add-data=sql;sql",
"--add-data=README.md;.",
"--add-data=POSTGRESQL_SETUP.md;.",
"--add-data=SQLITE_MIGRATION.md;.",
"--hidden-import=psycopg2",
"--hidden-import=sqlalchemy.dialects.postgresql",
"--hidden-import=sqlalchemy.dialects.sqlite",
"--hidden-import=dotenv",
"--hidden-import=pydantic",
"--hidden-import=fastapi",
"--hidden-import=uvicorn",
"--hidden-import=schedule",
"--hidden-import=pandas",
"--clean",
"--noconfirm",
"run.py",
]
print("[BUILD] Running PyInstaller...")
print("[CMD] " + " ".join(cmd))
result = os.system(" ".join(cmd))
if result == 0:
print("[SUCCESS] Executable built successfully!")
# Copy .env file to dist if it exists
if os.path.exists(".env") and os.path.exists("dist"):
shutil.copy2(".env", "dist/.env")
print("[COPY] .env file copied to dist/")
# Create batch files for easy usage
batch_files = {
"start.bat": """@echo off
echo Starting Ping River Monitor...
ping-river-monitor.exe
pause
""",
"start-api.bat": """@echo off
echo Starting Web API...
ping-river-monitor.exe --web-api
pause
""",
"test.bat": """@echo off
echo Running test...
ping-river-monitor.exe --test
pause
""",
}
for filename, content in batch_files.items():
if os.path.exists("dist"):
with open(f"dist/{filename}", "w") as f:
f.write(content)
print(f"[CREATE] {filename}")
print("\n" + "=" * 50)
print("BUILD COMPLETE!")
print(f"Executable: dist/ping-river-monitor.exe")
print("Batch files: start.bat, start-api.bat, test.bat")
print("Don't forget to edit .env file before using!")
return True
else:
print("[ERROR] Build failed!")
return False
if __name__ == "__main__":
success = main()
sys.exit(0 if success else 1)
+8 -1
View File
@@ -184,7 +184,14 @@ Oct 2024 flood). Mae Kuang Udom Thara is the second upstream reservoir.
| Source | What | Access | | Source | What | Access |
|---|---|---| |---|---|---|
| `https://app.rid.go.th/reservoir/api/dams` | **INGESTED** — daily snapshot of all ~35 large dams (storage/inflow/outflow MCM, % of usable). `POST` with form field `date=YYYY-MM-DD` (empty = today); GET returns 404 "Unknown method." Archive ≥ 2009; `level_msl` (`DMD_Q`) populated in older years only. Mae Ngat = `DAM_ID 200103` — hit 113% usable capacity, ~19 MCM/day inflow, in Oct 2024. Collected daily by `src/rid_reservoir.py` into `rid_dams` + `rid_reservoir_daily`; backfill via `scripts/backfill_rid_reservoir.py` | Open, no auth | | `https://app.rid.go.th/reservoir/api/dams` | **INGESTED** — daily snapshot of all ~35 large dams (storage/inflow/outflow MCM, % of usable). `POST` with form field `date=YYYY-MM-DD` (empty = today); GET returns 404 "Unknown method." Archive ≥ 2009; `level_msl` (`DMD_Q`) populated in older years only. Mae Ngat = `DAM_ID 200103` — hit 113% usable capacity, ~19 MCM/day inflow, in Oct 2024. Collected daily by `src/rid_reservoir.py` into `rid_dams` + `rid_reservoir_daily`; backfill via `scripts/backfill_rid_reservoir.py` | Open, no auth |
| `https://lsim.rid.go.th/ForeCast?reservoirid=22` | Mae Ngat daily status/forecast (RID) | Open, scrape — timed out from outside RID network when probed 2026-08-13 | | `https://lsim.rid.go.th/ForeCast?reservoirid=22` | Mae Ngat daily status/forecast (RID) | **UNREACHABLE — do not plan around it.** Probed 2026-08-13 from a Thai consumer ISP (AIS Fibre, TH) *and* from abroad: DNS resolves (122.154.18.207) but ICMP is 100% loss and ports 80/443/8080 are filtered, while `app.rid.go.th` answers in 0.27 s over the same connection. Down or RID-internal-only — not a geo-block |
| `https://app.rid.go.th/reservoir/api/dam` | **Per-dam daily series in ONE request**`GET` with `dam_id=200103&date_start=YYYY-MM-DD&date_end=YYYY-MM-DD&percent=`. Archive to 2009 (scattered single-day gaps). Far cheaper than the per-day `api/dams` loop the backfill used (one request vs ~2,900); prefer it for gap repair and for adding other dams. Sibling `api/damgraph` takes the same params | Open, no auth |
| `https://bigdata-api.rid.go.th` (SWOC) | **Intraday reservoir state** — RID SWOC telemetry, hourly with an explicit `hourly_time_utc` stamp; includes Mae Ngat (`TUP.16`) reservoir level m MSL and % capacity. **Snapshot-only — no archive**, so it can only be accumulated forward | Open, no auth |
| ThaiWater `public/waterlevel_load` stations `ridhydro_TUP.16` (at the dam) / `ridhydro_TUP.11` (dam outlet) | Hourly Mae Ngat reservoir level and outlet stage/flow. **Snapshot-only — `waterlevel_graph` returns empty grids for these ids at every era** (verified 2026-08-13). Collected hourly by our HII collector since 2026-08-11; accumulating forward | Open, no auth |
| ThaiWater `public/waterlevel_graph` station `P.75` (id 3253) | **The practical dam-release signal**: hourly stage+discharge 3.8 km below the Mae Ngat dam, history to 2019. Already ingested as a core RID station and a model feature since v1 | Open, no auth |
| ThaiWater `public/waterlevel_graph` station `MOU301` "สะพานน้ำแม่งัด" (id 1475118) | 10-minute stage on the Mae Ngat *above* the reservoir (inflow arm). History only from ~mid-2025; level only, no discharge | Open, no auth |
| ThaiWater `api-v3 .../analyst/dam` (dam.id 53) | EGAT-sourced copy of Mae Ngat carrying reservoir **level in m MSL historically** — the field RID's own API stopped populating (`DMD_Q`) after ~2013. Daily, no observation time | Open, no auth |
| `https://tiwrm.hii.or.th/DATA/REPORT/php/rid_bigcm_raw.php?sdate=YYYY-MM-DD` | HII HTML mirror of the RID large-dam daily table. Daily and *intermittent* (2026 YTD publishes ~108 of 225 days) — a cross-check, not a primary source | Open, no auth |
| `https://water.egat.co.th` | EGAT dams (Bhumibol/Sirikit) hourly+daily inflow/outflow/level | Endpoint catalog not public; contact EGAT (0-2436-8186). Only relevant downstream of Bhumibol | | `https://water.egat.co.th` | EGAT dams (Bhumibol/Sirikit) hourly+daily inflow/outflow/level | Endpoint catalog not public; contact EGAT (0-2436-8186). Only relevant downstream of Bhumibol |
| ThaiWater `/v2/large-dam/*`, `dam_rulecurve/graph` | All large/medium dams incl. hourly | Requires HII API key (§2.2) | | ThaiWater `/v2/large-dam/*`, `dam_rulecurve/graph` | All large/medium dams incl. hourly | Requires HII API key (§2.2) |
-293
View File
@@ -1,293 +0,0 @@
# Enhanced Scheduler Guide
This guide explains the new 15-minute scheduling system that runs continuously throughout each hour to ensure comprehensive data coverage.
## ✅ **New Scheduling Behavior**
### **15-Minute Schedule Pattern**
- **Timing**: Runs every 15 minutes: 1:00, 1:15, 1:30, 1:45, 2:00, 2:15, 2:30, 2:45, etc.
- **Hourly Full Checks**: At :00 minutes (includes gap filling and data updates)
- **Quarter-Hour Quick Checks**: At :15, :30, :45 minutes (data fetch only)
- **Continuous Coverage**: Ensures no data is missed throughout each hour
### **Operation Types**
- **Full Operations** (at :00): Data fetching + gap filling + data updates
- **Quick Operations** (at :15, :30, :45): Data fetching only for performance
## 🔧 **Technical Implementation**
### **Scheduler States**
```python
# State tracking variables
self.last_successful_update = None # Timestamp of last successful data update
self.retry_mode = False # Whether in quick check mode (skip gap filling)
self.next_hourly_check = None # Next scheduled hourly check
```
### **Quarter-Hour Check Process**
```python
def quarter_hour_check(self):
"""15-minute check for new data"""
current_time = datetime.datetime.now()
minute = current_time.minute
# Determine if this is a full hourly check (at :00) or a quarter-hour check
if minute == 0:
logging.info("=== HOURLY CHECK (00:00) ===")
self.retry_mode = False # Full check with gap filling and updates
else:
logging.info(f"=== 15-MINUTE CHECK ({minute:02d}:00) ===")
self.retry_mode = True # Skip gap filling and updates on 15-min checks
new_data_found = self.run_scraping_cycle()
if new_data_found:
self.last_successful_update = datetime.datetime.now()
if minute == 0:
logging.info("New data found during hourly check")
else:
logging.info(f"New data found during 15-minute check at :{minute:02d}")
else:
if minute == 0:
logging.info("No new data found during hourly check")
else:
logging.info(f"No new data found during 15-minute check at :{minute:02d}")
```
### **Scheduler Setup**
```python
def start_scheduler(self):
"""Start enhanced scheduler with 15-minute checks"""
# Schedule checks every 15 minutes (at :00, :15, :30, :45)
schedule.every().hour.at(":00").do(self.quarter_hour_check)
schedule.every().hour.at(":15").do(self.quarter_hour_check)
schedule.every().hour.at(":30").do(self.quarter_hour_check)
schedule.every().hour.at(":45").do(self.quarter_hour_check)
while True:
schedule.run_pending()
time.sleep(30) # Check every 30 seconds
```
## 📊 **New Data Detection Logic**
### **Smart Detection Algorithm**
```python
def has_new_data(self) -> bool:
"""Check if there is new data available since last successful update"""
# Get most recent timestamp from database
latest_data = self.get_latest_data(limit=1)
# Check if we should have newer data by now
now = datetime.datetime.now()
expected_latest = now.replace(minute=0, second=0, microsecond=0)
# If current time is past 5 minutes after the hour, we should have data
if now.minute >= 5:
if latest_timestamp < expected_latest:
return True # New data expected
# Check if we have data for the previous hour
previous_hour = expected_latest - datetime.timedelta(hours=1)
if latest_timestamp < previous_hour:
return True # Missing recent data
return False # Data is up to date
```
### **Actual Data Verification**
```python
# Compare timestamps before and after scraping
initial_timestamp = get_latest_timestamp_before_scraping()
# ... perform scraping ...
latest_timestamp = get_latest_timestamp_after_scraping()
if initial_timestamp is None or latest_timestamp > initial_timestamp:
new_data_found = True
self.last_successful_update = datetime.datetime.now()
```
## 🚀 **Operational Modes**
### **Mode 1: Full Hourly Operation (at :00)**
- **Schedule**: Every hour at :00 minutes (1:00, 2:00, 3:00, etc.)
- **Operations**:
- ✅ Fetch current data
- ✅ Fill data gaps (last 7 days)
- ✅ Update existing data (last 2 days)
- **Purpose**: Comprehensive data collection and maintenance
### **Mode 2: Quick 15-Minute Checks (at :15, :30, :45)**
- **Schedule**: Every 15 minutes at quarter-hour marks
- **Operations**:
- ✅ Fetch current data only
- ❌ Skip gap filling (performance optimization)
- ❌ Skip data updates (performance optimization)
- **Purpose**: Ensure no new data is missed between hourly checks
## 📋 **Logging Output Examples**
### **Successful Hourly Check (at :00)**
```
2025-07-26 01:00:00,123 - INFO - === HOURLY CHECK (00:00) ===
2025-07-26 01:00:00,124 - INFO - Starting scraping cycle...
2025-07-26 01:00:01,456 - INFO - Successfully fetched 384 data points from API
2025-07-26 01:00:02,789 - INFO - New data found: 2025-07-26 01:00:00
2025-07-26 01:00:03,012 - INFO - Filled 5 data gaps
2025-07-26 01:00:04,234 - INFO - Updated 2 existing measurements
2025-07-26 01:00:04,235 - INFO - New data found during hourly check
```
### **15-Minute Quick Check (at :15, :30, :45)**
```
2025-07-26 01:15:00,123 - INFO - === 15-MINUTE CHECK (15:00) ===
2025-07-26 01:15:00,124 - INFO - Starting scraping cycle...
2025-07-26 01:15:01,456 - INFO - Successfully fetched 299 data points from API
2025-07-26 01:15:02,789 - INFO - New data found: 2025-07-26 01:00:00
2025-07-26 01:15:02,790 - INFO - New data found during 15-minute check at :15
```
### **Continuous 15-Minute Pattern**
```
2025-07-26 01:00:00,123 - INFO - === HOURLY CHECK (00:00) ===
2025-07-26 01:00:04,235 - INFO - New data found during hourly check
2025-07-26 01:15:00,123 - INFO - === 15-MINUTE CHECK (15:00) ===
2025-07-26 01:15:02,790 - INFO - No new data found during 15-minute check at :15
2025-07-26 01:30:00,123 - INFO - === 15-MINUTE CHECK (30:00) ===
2025-07-26 01:30:02,790 - INFO - No new data found during 15-minute check at :30
2025-07-26 01:45:00,123 - INFO - === 15-MINUTE CHECK (45:00) ===
2025-07-26 01:45:02,790 - INFO - No new data found during 15-minute check at :45
2025-07-26 02:00:00,123 - INFO - === HOURLY CHECK (00:00) ===
2025-07-26 02:00:04,235 - INFO - New data found during hourly check
```
## ⚙️ **Configuration Options**
### **Environment Variables**
```bash
# Retry interval (default: 5 minutes)
export RETRY_INTERVAL_MINUTES=5
# Data availability buffer (default: 5 minutes after hour)
export DATA_BUFFER_MINUTES=5
# Gap filling days (default: 7 days)
export GAP_FILL_DAYS=7
# Update check days (default: 2 days)
export UPDATE_DAYS=2
```
### **Scheduler Timing**
```python
# Hourly checks at top of hour
schedule.every().hour.at(":00").do(self.hourly_check)
# 5-minute retries (dynamically scheduled)
schedule.every(5).minutes.do(self.retry_check).tag('retry')
# Check every 30 seconds for responsive retry scheduling
time.sleep(30)
```
## 🔍 **Performance Optimizations**
### **Retry Mode Optimizations**
- **Skip Gap Filling**: Avoids expensive historical data fetching during retries
- **Skip Data Updates**: Avoids comparison operations during retries
- **Focused API Calls**: Only fetches current day data during retries
- **Reduced Database Queries**: Minimal database operations during retries
### **Resource Management**
- **API Rate Limiting**: 1-second delays between API calls
- **Database Connection Pooling**: Efficient connection reuse
- **Memory Efficiency**: Selective data processing
- **Error Recovery**: Automatic retry with exponential backoff
## 🛠️ **Troubleshooting**
### **Common Scenarios**
#### **Stuck in Retry Mode**
```
# Check if API is returning data
curl -X POST https://hyd-app-db.rid.go.th/webservice/getGroupHourlyWaterLevelReportAllHL.ashx
# Check database connectivity
python water_scraper_v3.py --check-gaps 1
# Manual data fetch test
python water_scraper_v3.py --test
```
#### **Missing Hourly Triggers**
```
# Check system time synchronization
timedatectl status
# Verify scheduler is running
ps aux | grep water_scraper
# Check logs for scheduler activity
tail -f water_monitor.log | grep "HOURLY CHECK"
```
#### **False New Data Detection**
```
# Check latest data in database
sqlite3 water_monitoring.db "SELECT MAX(timestamp) FROM water_measurements;"
# Verify timestamp parsing
python -c "
import datetime
print('Current hour:', datetime.datetime.now().replace(minute=0, second=0, microsecond=0))
"
```
## 📈 **Monitoring and Alerts**
### **Key Metrics to Monitor**
- **Hourly Success Rate**: Percentage of hourly checks that find new data
- **Retry Duration**: How long system stays in retry mode
- **Data Freshness**: Time since last successful data update
- **API Response Time**: Performance of data fetching operations
### **Alert Conditions**
- **Extended Retry Mode**: System in retry mode for > 30 minutes
- **No Data for 2+ Hours**: No new data found for extended period
- **High Error Rate**: Multiple consecutive API failures
- **Database Issues**: Connection or save failures
### **Health Check Script**
```bash
#!/bin/bash
# Check if system is stuck in retry mode
RETRY_COUNT=$(tail -n 100 water_monitor.log | grep -c "RETRY CHECK")
if [ $RETRY_COUNT -gt 6 ]; then
echo "WARNING: System may be stuck in retry mode ($RETRY_COUNT retries in last 100 log entries)"
fi
# Check data freshness
LATEST_DATA=$(sqlite3 water_monitoring.db "SELECT MAX(timestamp) FROM water_measurements;")
echo "Latest data timestamp: $LATEST_DATA"
```
## 🎯 **Best Practices**
### **Production Deployment**
1. **Monitor Logs**: Watch for retry mode patterns
2. **Set Alerts**: Configure notifications for extended retry periods
3. **Regular Maintenance**: Weekly gap filling and data validation
4. **Backup Strategy**: Regular database backups before major operations
### **Performance Tuning**
1. **Adjust Buffer Time**: Modify data availability buffer based on API patterns
2. **Optimize Retry Interval**: Balance between responsiveness and API load
3. **Database Indexing**: Ensure proper indexes for timestamp queries
4. **Connection Pooling**: Configure appropriate database connection limits
This enhanced scheduler ensures reliable, efficient, and intelligent water level monitoring with automatic adaptation to data availability patterns.
-227
View File
@@ -1,227 +0,0 @@
# 🚀 Northern Thailand Ping River Monitor - Enhancement Summary
## 🎯 **What We've Accomplished**
We've successfully transformed your water monitoring system from a simple scraper into a **production-ready, enterprise-grade monitoring platform** focused on the Ping River Basin in Northern Thailand, with modern web interfaces, station management capabilities, and comprehensive observability.
## 🌟 **Major New Features Added**
### 1. **FastAPI Web Interface** 🌐
- **Interactive Dashboard** at `http://localhost:8000`
- **REST API** with comprehensive endpoints
- **Station Management** - Add, update, delete monitoring stations
- **Real-time Health Monitoring**
- **Manual Data Collection Triggers**
- **Interactive API Documentation** at `/docs`
- **CORS Support** for web applications
### 2. **Enhanced Architecture** 🏗️
- **Type Safety** with Pydantic models and comprehensive type hints
- **Data Validation Layer** with range checking and error handling
- **Custom Exception Classes** for better error management
- **Modular Design** with separated concerns
### 3. **Observability & Monitoring** 📊
- **Metrics Collection System** (counters, gauges, histograms)
- **Health Checks** for database, API, and system resources
- **Performance Tracking** with response times and success rates
- **Enhanced Logging** with colors, rotation, and performance logs
### 4. **Production Features** 🚀
- **Rate Limiting** to prevent API abuse
- **Request Tracking** with detailed statistics
- **Configuration Validation** on startup
- **Graceful Error Handling** and recovery
- **Background Task Management**
## 📁 **New Files Created**
```
src/
├── models.py # Data models and type definitions
├── exceptions.py # Custom exception classes
├── validators.py # Data validation layer
├── metrics.py # Metrics collection system
├── health_check.py # Health monitoring system
├── rate_limiter.py # Rate limiting and request tracking
├── logging_config.py # Enhanced logging configuration
├── web_api.py # FastAPI web interface
├── main.py # Enhanced CLI with multiple modes
└── __init__.py # Package initialization
# Root files
├── run.py # Simple startup script
├── test_integration.py # Integration test suite
├── test_api.py # API endpoint tests
└── ENHANCEMENT_SUMMARY.md # This file
```
## 🔧 **Enhanced Existing Files**
- **`src/water_scraper_v3.py`** - Integrated new features, metrics, validation
- **`src/config.py`** - Added configuration validation
- **`requirements.txt`** - Added FastAPI, Pydantic, and monitoring dependencies
- **`docker-compose.victoriametrics.yml`** - Added web API service
- **`Dockerfile`** - Updated for new startup script
- **`README.md`** - Updated with new features and usage instructions
## 🌐 **Web API Endpoints**
| Endpoint | Method | Description |
|----------|--------|-------------|
| `/` | GET | Interactive dashboard |
| `/docs` | GET | API documentation |
| `/health` | GET | System health status |
| `/metrics` | GET | Application metrics |
| `/stations` | GET | List all monitoring stations |
| `/measurements/latest` | GET | Latest measurements |
| `/measurements/station/{code}` | GET | Station-specific data |
| `/scrape/trigger` | POST | Trigger manual data collection |
| `/scraping/status` | GET | Scraping status and statistics |
| `/config` | GET | Current configuration (masked) |
## 🚀 **Usage Examples**
### **Traditional Mode (Enhanced)**
```bash
# Test single cycle
python run.py --test
# Continuous monitoring
python run.py
# Fill data gaps
python run.py --fill-gaps 7
# Show system status
python run.py --status
```
### **Web API Mode (NEW!)**
```bash
# Start web API server
python run.py --web-api
# Access dashboard
open http://localhost:8000
# View API documentation
open http://localhost:8000/docs
```
### **Docker Deployment**
```bash
# Start complete stack
docker-compose -f docker-compose.victoriametrics.yml up -d
# Services available:
# - Water API: http://localhost:8000
# - Grafana: http://localhost:3000
# - VictoriaMetrics: http://localhost:8428
```
## 📊 **Monitoring & Observability**
### **Built-in Metrics**
- API request counts and response times
- Database connection status and save operations
- Scraping cycle success/failure rates
- System resource usage (memory, etc.)
### **Health Checks**
- Database connectivity and data freshness
- External API availability
- Memory usage monitoring
- Overall system health status
### **Enhanced Logging**
- Colored console output for better readability
- File rotation to prevent disk space issues
- Performance logging for optimization
- Structured logging with proper levels
## 🔒 **Production Ready Features**
### **Security & Reliability**
- Rate limiting to prevent API abuse
- Input validation and sanitization
- Graceful error handling and recovery
- Configuration validation on startup
### **Performance**
- Efficient metrics collection with minimal overhead
- Background task management
- Connection pooling and resource management
- Optimized database operations
### **Scalability**
- Modular architecture for easy extension
- Async support for high concurrency
- Configurable resource limits
- Health checks for load balancer integration
## 🧪 **Testing**
### **Integration Tests**
```bash
# Run all integration tests
python test_integration.py
```
### **API Tests**
```bash
# Test API endpoints (server must be running)
python test_api.py
```
## 📈 **Performance Improvements**
1. **Request Tracking** - Monitor API performance and success rates
2. **Rate Limiting** - Prevent API abuse and ensure stability
3. **Data Validation** - Catch errors early and improve data quality
4. **Metrics Collection** - Identify bottlenecks and optimization opportunities
5. **Health Monitoring** - Proactive issue detection and alerting
## 🎉 **Benefits Achieved**
### **For Developers**
- **Better Developer Experience** with type hints and validation
- **Easier Debugging** with enhanced logging and error messages
- **Comprehensive Testing** with integration and API tests
- **Modern Architecture** following best practices
### **For Operations**
- **Web Dashboard** for easy monitoring and management
- **Health Checks** for automated monitoring integration
- **Metrics Collection** for performance analysis
- **Production-Ready** deployment with Docker support
### **For Users**
- **REST API** for integration with other systems
- **Real-time Data Access** via web interface
- **Manual Controls** for triggering data collection
- **Status Monitoring** for system visibility
## 🔮 **Future Enhancement Opportunities**
1. **Authentication & Authorization** - Add user management and API keys
2. **Real-time WebSocket Updates** - Live data streaming to web clients
3. **Advanced Analytics** - Trend analysis and forecasting
4. **Alert System** - Email/SMS notifications for critical conditions
5. **Multi-tenant Support** - Support for multiple organizations
6. **Data Export** - CSV, Excel, and other format exports
7. **Mobile App** - React Native or Flutter mobile interface
## 🏆 **Summary**
Your Thailand Water Monitor has been transformed from a simple data scraper into a **comprehensive, enterprise-grade monitoring platform** that includes:
-**Modern Web Interface** with FastAPI
-**Production-Ready Architecture** with proper error handling
-**Comprehensive Monitoring** with metrics and health checks
-**Type Safety** and data validation
-**Enhanced Logging** and observability
-**Docker Support** for easy deployment
-**Extensive Testing** for reliability
The system is now ready for production deployment and can serve as a foundation for further enhancements and integrations!
+144 -7
View File
@@ -68,7 +68,7 @@ environment variable, then `Config.get_database_config()` when `DB_TYPE` is
1. **PostgreSQL** (`_fetch_from_db`) — the primary path. NULL discharge stays 1. **PostgreSQL** (`_fetch_from_db`) — the primary path. NULL discharge stays
NULL, which matters because the models must learn from the real missingness NULL, which matters because the models must learn from the real missingness
pattern. pattern.
2. **HTTP API** (`_fetch_from_api`, default `http://100.81.167.42:8000`) — a 2. **HTTP API** (`_fetch_from_api`, default `https://water.buildfor.life`) — a
fallback for running off-server. **Caveat:** the public history endpoint fallback for running off-server. **Caveat:** the public history endpoint
backfills missing discharge with a synthetic rating-curve estimate, so this backfills missing discharge with a synthetic rating-curve estimate, so this
path is not equivalent to the DB path. It is flagged as path is not equivalent to the DB path. It is flagged as
@@ -451,8 +451,9 @@ acts before any gauge rises, and `rain_fc24` — a weather *forecast* — acts
before the rain itself falls. **Remaining honest limits:** marginal before the rain itself falls. **Remaining honest limits:** marginal
just-over-threshold crests (2025: +2 h) are intrinsically short-notice; the just-over-threshold crests (2025: +2 h) are intrinsically short-notice; the
rain series only exists from 2021-03, so older training rows are rain-blind; rain series only exists from 2021-03, so older training rows are rain-blind;
forecast-rain quality bounds what the feature can add; and Mae Ngat/Mae Kuang forecast-rain quality bounds what the feature can add; and Mae Ngat reservoir
dam releases remain uningested (see `docs/DATA_SOURCES.md`). state, though now ingested daily (see `docs/DATA_SOURCES.md`), measurably
*hurts* alert lead as a model feature — see the 2026-08-13 experiment below.
**Danger-level skill at P.1 is unproven.** P.1 never crossed 4.5 m in the **Danger-level skill at P.1 is unproven.** P.1 never crossed 4.5 m in the
2025-01-01 → 2026-08-10 test span (`base_rate_danger` is 0.0, so every danger 2025-01-01 → 2026-08-10 test span (`base_rate_danger` is 0.0, so every danger
@@ -475,6 +476,114 @@ P.1 additionally reports `stages`: exceedance probability for each of the seven
official inundation stages (3.704.60 m), computed from the regression head and official inundation stages (3.704.60 m), computed from the regression head and
its calibration sigma, so they need no retrain and no per-stage classifiers. its calibration sigma, so they need no retrain and no per-stage classifiers.
### 2026-08-13: Mae Ngat dam features — a documented negative result
With `rid_reservoir_daily` backfilled to 2018 (daily Mae Ngat storage/inflow/
outflow, `src/ml/dam.py`), the obvious v4 experiment was to feed reservoir
state to the mainstem models: during the Oct 2024 flood the dam hit 113% of
usable capacity with 1922 MCM/day inflow spikes on the crossing days.
**It fails the acceptance gate.** On the 2024 record-flood backtest (train
< 1 Sep 2024, belt-and-braces alerting, identical to the deployed pipeline):
| dam features | first-alert lead | record-peak err (24 h ahead) |
|----------------------------|------------------|------------------------------|
| none (deployed v3 config) | **+13 h** (PASS) | +0.24 m |
| all four | +10 h (FAIL) | +0.22 m |
| storage % + 3-day delta | +12 h | +0.21…+0.27 m |
| inflow + outflow | +10 h (FAIL) | +0.35 m |
| outflow only | +12 h | +0.20 m |
Every subset costs 13 h of warning for at most a ~3 cm peak-error gain. The
mechanism is the publication lag: RID posts the daily report on the morning of
its own date (features apply it from 07:00, `dam.py`'s leakage rule), so at the
04:00 first-alert hour of 24 Sep 2024 the freshest dam row still described
23 Sep — a benign reservoir quietly absorbing inflow (outflow 0.13 MCM/day).
The columns therefore argue *against* imminent flooding exactly when the rain
features are (correctly) raising the alarm. The rolling-origin harness agrees:
`rise_rain_dam` matches `rise_rain` on leads and false alarms, only nudging
event-peak amplitude (0.11 → 0.03 m on the Sep 2024 event), and `rise_dam`
(dam without rain) is strictly worse with alarm-latch artifacts.
**Disposition:** dam features are OFF by default (`train_all(use_dam=False)`;
opt-in via `--dam` on the training CLI, `scripts/backtest_render.py --dam`,
and the `rise_rain_dam` / `rise_dam` harness variants). The collector keeps
accruing daily rows; revisit post-monsoon when the 2026 season adds dam-era
flood events.
**Why the lag is probably not the whole story — P.75 already *is* the dam
signal.** A 2026-08-13 source sweep put the negative result on firmer
ground: **P.75 "บ้านช่อแล" sits 3.8 km downstream of the Mae Ngat dam** on
the Mae Ngat river (nearest other station: P.4A at 10.9 km), it reports
hourly, and it has been a model input since v1 with a 12 h routed lead into
P.1. Whatever the reservoir releases flows past P.75 within the hour and the
model already reads it. The daily reservoir table therefore offers a stale,
coarser proxy of a signal the features capture hourly and directly — which
is the more likely reason it adds nothing and costs alarm responsiveness.
That reframes what a future intraday source would have to beat: not "no dam
information", but "hourly observed dam *outflow*". Genuine intraday
reservoir-state feeds do exist and are open (`bigdata-api.rid.go.th` SWOC
telemetry, and ThaiWater station `ridhydro_TUP.16` *at the dam*), but both
are **snapshot-only — no archive** (verified: the history endpoint returns
empty grids for them at every era, including the current one). They can only
be accumulated forward, so they cannot retrain against 2024/2025 events.
The HII collector already captures both hourly as of 2026-08-11; revisit
after the 2026 monsoon, when a season of true intraday reservoir state
exists alongside its flood events.
**Shipped from the same work:** the HII gap-fill merge in the data loader
(`fill_from_hii`, +9,341 h at P.81, +682 h at P.92, +810 h at P.20) is
lead-neutral — the gate holds at 13 h with fill on — and ships enabled.
### 2026-09-12: three candidates on top of hgb-v3 — two rejected, one deferred
Same rolling-origin harness (`src/ml/evaluate.py`, five monsoon folds
20212025, P.1 and P.103), all variants run from the identical
`models/cache/` snapshot (`--from-cache`), results in
`models/eval_2026-09-12*.json`, tables via `scripts/summarize_eval.py`.
Baseline is `rise_rain`, the deployed configuration.
**Quantile regression heads (`rise_rain_quantile`, `_uw`) — rejected.** The
August result that quantile loss beat L2 on MAE held with rain in the model
(P.1 0.083/0.081 vs 0.087; P.103 0.143/0.124 vs 0.152), and Brier improved a
hair, but the operational numbers went the wrong way: at P.103 the 2022-08-14
crossing dropped from +6 h to +1 h lead, 2022-10-02 from +9 h to +5/+3 h, and
the 2024-09-30 event from +9 h to +4 h; at P.1 2022 dropped +5 → +3/+2 h and
2025 +2 → +1 h, with one false-alarm episode where the baseline had none. A
median predicts the *typical* rise, and on the run-up to a crossing the typical
rise is not the one that matters. MAE is not the objective; lead is.
**Quantile heads for sigma only (`rise_rain_qsigma`) — no effect.** The
hybrid keeps the L2 point prediction (so every lead is identical to the
baseline by construction — p≥0.5 alerts are sigma-independent) and derives a
per-row sigma from q90q50. Brier moved 0.0031 → 0.0029 at P.1 and
0.0061 → 0.0060 at P.103, i.e. within noise, at the cost of three fitted
heads per horizon instead of one. Per-row uncertainty from this family of
models is not informative enough here to be worth the training time; the
0.15 m floor stays.
**Forward-48 h forecast rain (`rise_rain_fc48`) — deferred.** Adding the
`(t, t+48]` Open-Meteo sum alongside `rain_fc24` left MAE, Brier and false
alarms unchanged and every event lead within ±1 h of baseline, *except* the
2024-10-03 P.1 record crossing, which went from +21 h to +72 h (and +55 → +69 h
at P.103). That is one event with the highest stakes in the record, on the
same feature family that already produced the 2024 gain, but n=1 is not
evidence: the P.103 2025-09-26 event lost 2 h in the same run. Rerun after the
2026 season adds events; if the 48 h window still moves only the biggest
onsets, promote it. Serving would need no new data source (`fetch_forecast`
already pulls `forecast_days=2`).
**HII gauge rain — not evaluable yet.** `hii_rainfall` (~130 gauges in the
upper-Ping box, DWR/FOP/HII/RID/TMD) is the obvious independent rain source,
but the table only exists since 2026-08-11 and the api-v3 archive endpoint
ignores its date range (see `docs/DATA_SOURCES.md` §2.1), so every training
row before that is NaN and no fold in the harness has gauge data in its test
span. `src/ml/hii_rain.py` builds the catchment mean and
`GET /api/hii/rainfall/catchment` exposes it next to the Open-Meteo series with
a 24 h-sum bias/MAE/correlation, so the two sources' relationship is on record
by the time the 2027 fold (train ≤ 2027-04-30, test JunNov 2027) can test it.
## 6. Deployment ## 6. Deployment
### API ### API
@@ -672,6 +781,32 @@ timestamp in every bundle. Both are echoed in every `/forecast` row, so you can
tell from the API response alone which code produced a forecast and how old the tell from the API response alone which code produced a forecast and how old the
model is. model is.
**Scheduled retrain (since 2026-09-12).** `scripts/water-monitor-retrain.timer`
fires `water-monitor-retrain.service` on the 1st of every month at 03:30 server
time (`Persistent=true`, so a missed run catches up at boot). The unit runs
`scripts/retrain.sh` as the service user with `OMP_NUM_THREADS=4`, `Nice=15`:
1. trains all stations into `models/.staging/` (the API keeps serving the old
bundles throughout);
2. refuses to promote unless `metrics.json` reports a `hgb-v3+` version and at
least 14 trained stations (exit 3, staging discarded, old models untouched);
3. renames the new bundles into `models/`, moving the previous generation to
`models/.previous/` for rollback.
No API restart: `predict.py` reloads bundles by mtime on the next hourly
precompute. `systemctl list-timers water-monitor-retrain.timer` shows the next
run; `sudo systemctl start water-monitor-retrain.service` runs it now (after a
flood, say); `journalctl -u water-monitor-retrain` has the log. The installer
(`scripts/install.sh`) enables the timer.
**Why the trainer refuses to run without rain (since 2026-09-12).** On
2026-09-01 the server retrain could not reach the Open-Meteo archive on a
checkout with no `models/cache/`, logged a warning, and quietly overwrote the
v3 bundles with gauge-only v2 ones — the 13-hour early warning on the 2024 flood
became an 18-hour late one and nothing on the dashboard said so. `train_all()`
now raises `RainUnavailableError` (CLI exit 2) in that situation. Gauge-only
bundles are still available, but only by asking for them: `--no-rain`.
## 8. Operations runbook ## 8. Operations runbook
All commands assume the project virtualenv is active (`.venv` locally). All commands assume the project virtualenv is active (`.venv` locally).
@@ -725,10 +860,12 @@ that went quiet, or a bad backfill) rather than a modelling one.
python -m pytest tests/test_flood_forecast.py -v python -m pytest tests/test_flood_forecast.py -v
``` ```
Seven tests covering leakage, label alignment, the coverage gate, forward-fill and Tests cover leakage, label alignment, the coverage gate, forward-fill and
staleness, a train/predict round trip, the heuristic fallback, and feature-name staleness, a train/predict round trip, the heuristic fallback, feature-name
stability. The whole suite runs in about 8 seconds, so there is no excuse for stability, and the rain-downgrade guard (no rain series → `RainUnavailableError`,
skipping it before a deploy. nothing written; `--no-rain` → v2; rain present → v3 with the rain columns in
`feature_names`). The file runs in well under a minute, so there is no excuse
for skipping it before a deploy.
**Understanding graceful degradation.** Three things can make a forecast row **Understanding graceful degradation.** Three things can make a forecast row
non-model-backed, and all of them are visible in the payload: non-model-backed, and all of them are visible in the payload:
-475
View File
@@ -1,475 +0,0 @@
# Geolocation Support for Grafana Geomap
This guide explains the geolocation functionality added to the Thailand Water Monitor for use with Grafana's geomap visualization.
## ✅ **Implemented Features**
### **Database Schema Updates**
All database adapters now support geolocation fields:
- **latitude**: Decimal latitude coordinates (DECIMAL(10,8) for SQL, REAL for SQLite)
- **longitude**: Decimal longitude coordinates (DECIMAL(11,8) for SQL, REAL for SQLite)
- **geohash**: Geohash string for efficient spatial indexing (VARCHAR(20)/TEXT)
### **Station Data Enhancement**
Station mapping now includes geolocation fields:
```python
'8': {
'code': 'P.1',
'thai_name': 'สะพานนวรัฐ',
'english_name': 'Nawarat Bridge',
'latitude': 15.6944, # Decimal degrees
'longitude': 100.2028, # Decimal degrees
'geohash': 'w5q6uuhvfcfp25' # Geohash for P.1
}
```
## 🗄️ **Database Schema**
### **Updated Stations Table**
```sql
CREATE TABLE stations (
id INTEGER PRIMARY KEY,
station_code TEXT UNIQUE NOT NULL,
thai_name TEXT NOT NULL,
english_name TEXT NOT NULL,
latitude REAL, -- NEW: Latitude coordinate
longitude REAL, -- NEW: Longitude coordinate
geohash TEXT, -- NEW: Geohash for spatial indexing
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
updated_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP
);
```
### **Database Support**
- ✅ **SQLite**: REAL columns for coordinates, TEXT for geohash
- ✅ **PostgreSQL**: DECIMAL(10,8) and DECIMAL(11,8) for coordinates, VARCHAR(20) for geohash
- ✅ **MySQL**: DECIMAL(10,8) and DECIMAL(11,8) for coordinates, VARCHAR(20) for geohash
- ✅ **VictoriaMetrics**: Geolocation data included in metric labels
## 📊 **Current Station Data**
### **P.1 - Nawarat Bridge (Sample)**
- **Station Code**: P.1
- **Thai Name**: สะพานนวรัฐ
- **English Name**: Nawarat Bridge
- **Latitude**: 15.6944
- **Longitude**: 100.2028
- **Geohash**: w5q6uuhvfcfp25
### **Remaining Stations**
The following stations are ready for geolocation data when coordinates become available:
- P.20 - บ้านเชียงดาว (Ban Chiang Dao)
- P.75 - บ้านช่อแล (Ban Chai Lat)
- P.92 - บ้านเมืองกึ๊ด (Ban Muang Aut)
- P.4A - บ้านแม่แตง (Ban Mae Taeng)
- P.67 - บ้านแม่แต (Ban Tae)
- P.21 - บ้านริมใต้ (Ban Rim Tai)
- P.103 - สะพานวงแหวนรอบ 3 (Ring Bridge 3)
- P.82 - บ้านสบวิน (Ban Sob win)
- P.84 - บ้านพันตน (Ban Panton)
- P.81 - บ้านโป่ง (Ban Pong)
- P.5 - สะพานท่านาง (Tha Nang Bridge)
- P.77 - บ้านสบแม่สะป๊วด (Baan Sop Mae Sapuord)
- P.87 - บ้านป่าซาง (Ban Pa Sang)
- P.76 - บ้านแม่อีไฮ (Banb Mae I Hai)
- P.85 - บ้านหล่ายแก้ว (Baan Lai Kaew)
## 🗺️ **Grafana Geomap Integration**
### **Data Source Configuration**
The geolocation data is automatically included in all database queries and can be used directly in Grafana:
#### **SQLite/PostgreSQL/MySQL Query Example**
```sql
SELECT
m.timestamp,
s.station_code,
s.english_name,
s.thai_name,
s.latitude,
s.longitude,
s.geohash,
m.water_level,
m.discharge,
m.discharge_percent
FROM water_measurements m
JOIN stations s ON m.station_id = s.id
WHERE s.latitude IS NOT NULL
AND s.longitude IS NOT NULL
ORDER BY m.timestamp DESC
```
#### **VictoriaMetrics Query Example**
```promql
water_level{latitude!="",longitude!=""}
```
### **Geomap Panel Configuration**
#### **1. Create Geomap Panel**
1. Add new panel in Grafana
2. Select "Geomap" visualization
3. Configure data source (SQLite/PostgreSQL/MySQL/VictoriaMetrics)
#### **2. Configure Location Fields**
- **Latitude Field**: `latitude`
- **Longitude Field**: `longitude`
- **Alternative**: Use `geohash` field for geohash-based positioning
#### **3. Configure Display Options**
- **Station Labels**: Use `station_code` or `english_name`
- **Tooltip Information**: Include `thai_name`, `water_level`, `discharge`
- **Color Mapping**: Map to `water_level` or `discharge_percent`
#### **4. Sample Geomap Configuration**
```json
{
"type": "geomap",
"title": "Thailand Water Stations",
"targets": [
{
"rawSql": "SELECT latitude, longitude, station_code, english_name, water_level, discharge_percent FROM stations s JOIN water_measurements m ON s.id = m.station_id WHERE s.latitude IS NOT NULL AND m.timestamp = (SELECT MAX(timestamp) FROM water_measurements WHERE station_id = s.id)",
"format": "table"
}
],
"fieldConfig": {
"defaults": {
"custom": {
"hideFrom": {
"legend": false,
"tooltip": false,
"vis": false
}
},
"mappings": [],
"color": {
"mode": "continuous-GrYlRd",
"field": "water_level"
}
}
},
"options": {
"view": {
"id": "coords",
"lat": 15.6944,
"lon": 100.2028,
"zoom": 8
},
"controls": {
"mouseWheelZoom": true,
"showZoom": true,
"showAttribution": true
},
"layers": [
{
"type": "markers",
"config": {
"size": {
"field": "discharge_percent",
"min": 5,
"max": 20
},
"color": {
"field": "water_level"
},
"showLegend": true
}
}
]
}
}
```
## 🔧 **Adding New Station Coordinates**
### **Method 1: Update Station Mapping**
Edit `water_scraper_v3.py` and add coordinates to the station mapping:
```python
'1': {
'code': 'P.20',
'thai_name': 'บ้านเชียงดาว',
'english_name': 'Ban Chiang Dao',
'latitude': 19.3056, # Add actual coordinates
'longitude': 98.9264, # Add actual coordinates
'geohash': 'w4r6...' # Add actual geohash
}
```
### **Method 2: Direct Database Update**
```sql
UPDATE stations
SET latitude = 19.3056, longitude = 98.9264, geohash = 'w4r6uuhvfcfp25'
WHERE station_code = 'P.20';
```
### **Method 3: Bulk Update Script**
```python
import sqlite3
coordinates = {
'P.20': {'lat': 19.3056, 'lon': 98.9264, 'geohash': 'w4r6uuhvfcfp25'},
'P.75': {'lat': 18.7756, 'lon': 99.1234, 'geohash': 'w4r5uuhvfcfp25'},
# Add more stations...
}
conn = sqlite3.connect('water_monitoring.db')
cursor = conn.cursor()
for station_code, coords in coordinates.items():
cursor.execute("""
UPDATE stations
SET latitude = ?, longitude = ?, geohash = ?
WHERE station_code = ?
""", (coords['lat'], coords['lon'], coords['geohash'], station_code))
conn.commit()
conn.close()
```
## 🌐 **Geohash Information**
### **What is Geohash?**
Geohash is a geocoding system that represents geographic coordinates as a short alphanumeric string. It provides:
- **Spatial Indexing**: Efficient spatial queries
- **Proximity**: Similar geohashes indicate nearby locations
- **Hierarchical**: Longer geohashes provide more precision
### **Geohash Precision Levels**
- **5 characters**: ~2.4km precision
- **6 characters**: ~610m precision
- **7 characters**: ~76m precision
- **8 characters**: ~19m precision
- **9+ characters**: <5m precision
### **Example: P.1 Geohash**
- **Geohash**: `w5q6uuhvfcfp25`
- **Length**: 14 characters
- **Precision**: Sub-meter accuracy
- **Location**: Nawarat Bridge, Thailand
## 📈 **Grafana Visualization Examples**
### **1. Station Location Map**
- **Type**: Geomap with markers
- **Data**: Current station locations
- **Color**: Water level or discharge percentage
- **Size**: Discharge volume
### **2. Regional Water Levels**
- **Type**: Geomap with heatmap
- **Data**: Water level data across regions
- **Visualization**: Color-coded intensity map
- **Filters**: Time range, station groups
### **3. Alert Zones**
- **Type**: Geomap with threshold markers
- **Data**: Stations exceeding alert thresholds
- **Visualization**: Red markers for high water levels
- **Alerts**: Automated notifications for critical levels
## 🔄 **Updating a Running System**
### **Automated Migration Script**
Use the provided migration script to safely add geolocation columns to your existing database:
```bash
# Stop the water monitoring service first
sudo systemctl stop water-monitor
# Run the migration script
python migrate_geolocation.py
# Restart the service
sudo systemctl start water-monitor
```
### **Migration Script Features**
- ✅ **Auto-detects database type** from environment variables
- ✅ **Checks existing columns** to avoid conflicts
- ✅ **Supports all database types** (SQLite, PostgreSQL, MySQL)
- ✅ **Adds sample data** for P.1 station
- ✅ **Safe operation** - won't break existing data
### **Step-by-Step Migration Process**
#### **1. Stop the Application**
```bash
# If running as systemd service
sudo systemctl stop water-monitor
# If running in screen/tmux
# Use Ctrl+C to stop the process
# If running as Docker container
docker stop water-monitor
```
#### **2. Backup Your Database**
```bash
# SQLite backup
cp water_monitoring.db water_monitoring.db.backup
# PostgreSQL backup
pg_dump water_monitoring > water_monitoring_backup.sql
# MySQL backup
mysqldump water_monitoring > water_monitoring_backup.sql
```
#### **3. Run Migration Script**
```bash
# Default (uses environment variables)
python migrate_geolocation.py
# Or specify database path for SQLite
SQLITE_DB_PATH=/path/to/water_monitoring.db python migrate_geolocation.py
```
#### **4. Verify Migration**
```bash
# Check SQLite schema
sqlite3 water_monitoring.db ".schema stations"
# Check PostgreSQL schema
psql -d water_monitoring -c "\d stations"
# Check MySQL schema
mysql -e "DESCRIBE water_monitoring.stations"
```
#### **5. Update Application Code**
Ensure you have the latest version of the application with geolocation support:
```bash
# Pull latest code
git pull origin main
# Install any new dependencies
pip install -r requirements.txt
```
#### **6. Restart Application**
```bash
# Systemd service
sudo systemctl start water-monitor
# Docker container
docker start water-monitor
# Manual execution
python water_scraper_v3.py
```
### **Migration Output Example**
```
2025-07-28 17:30:00,123 - INFO - Starting geolocation column migration...
2025-07-28 17:30:00,124 - INFO - Detected database type: SQLITE
2025-07-28 17:30:00,125 - INFO - Migrating SQLite database: water_monitoring.db
2025-07-28 17:30:00,126 - INFO - Current columns in stations table: ['id', 'station_code', 'thai_name', 'english_name', 'created_at', 'updated_at']
2025-07-28 17:30:00,127 - INFO - Added latitude column
2025-07-28 17:30:00,128 - INFO - Added longitude column
2025-07-28 17:30:00,129 - INFO - Added geohash column
2025-07-28 17:30:00,130 - INFO - Successfully added columns: latitude, longitude, geohash
2025-07-28 17:30:00,131 - INFO - Updated P.1 station with sample geolocation data
2025-07-28 17:30:00,132 - INFO - P.1 station geolocation: ('P.1', 15.6944, 100.2028, 'w5q6uuhvfcfp25')
2025-07-28 17:30:00,133 - INFO - ✅ Migration completed successfully!
2025-07-28 17:30:00,134 - INFO - You can now restart your water monitoring application
2025-07-28 17:30:00,135 - INFO - The system will automatically use the new geolocation columns
```
## 🔍 **Troubleshooting**
### **Migration Issues**
#### **Database Locked Error**
```bash
# Stop all processes using the database
sudo systemctl stop water-monitor
pkill -f water_scraper
# Wait a few seconds, then run migration
sleep 5
python migrate_geolocation.py
```
#### **Permission Denied**
```bash
# Check database file permissions
ls -la water_monitoring.db
# Fix permissions if needed
sudo chown $USER:$USER water_monitoring.db
chmod 664 water_monitoring.db
```
#### **Missing Dependencies**
```bash
# For PostgreSQL
pip install psycopg2-binary
# For MySQL
pip install pymysql
# For all databases
pip install -r requirements.txt
```
### **Verification Issues**
#### **Missing Coordinates**
If stations don't appear on the geomap:
1. Check if latitude/longitude are NULL in database
2. Verify geolocation data in station mapping
3. Ensure database schema includes geolocation columns
4. Run migration script if columns are missing
#### **Incorrect Positioning**
If stations appear in wrong locations:
1. Verify coordinate format (decimal degrees)
2. Check latitude/longitude order (lat first, lon second)
3. Validate geohash accuracy
### **Rollback Procedure**
If migration causes issues:
#### **SQLite Rollback**
```bash
# Stop application
sudo systemctl stop water-monitor
# Restore backup
cp water_monitoring.db.backup water_monitoring.db
# Restart with old version
sudo systemctl start water-monitor
```
#### **PostgreSQL Rollback**
```sql
-- Remove added columns
ALTER TABLE stations DROP COLUMN IF EXISTS latitude;
ALTER TABLE stations DROP COLUMN IF EXISTS longitude;
ALTER TABLE stations DROP COLUMN IF EXISTS geohash;
```
#### **MySQL Rollback**
```sql
-- Remove added columns
ALTER TABLE stations DROP COLUMN latitude;
ALTER TABLE stations DROP COLUMN longitude;
ALTER TABLE stations DROP COLUMN geohash;
```
## 🎯 **Next Steps**
### **Immediate Actions**
1. **Gather Coordinates**: Collect GPS coordinates for all 16 stations
2. **Update Database**: Add coordinates to remaining stations
3. **Create Dashboards**: Build Grafana geomap visualizations
### **Future Enhancements**
1. **Automatic Geocoding**: API integration for address-to-coordinate conversion
2. **Mobile GPS**: Mobile app for field coordinate collection
3. **Satellite Integration**: Satellite imagery overlay in Grafana
4. **Geofencing**: Alert zones based on geographic boundaries
The geolocation functionality is now fully implemented and ready for use with Grafana's geomap visualization. Station P.1 (Nawarat Bridge) serves as a working example with complete coordinate data.
-2
View File
@@ -286,8 +286,6 @@ make validate-workflows
### **Project-Specific Resources** ### **Project-Specific Resources**
- [Contributing Guide](../CONTRIBUTING.md) - [Contributing Guide](../CONTRIBUTING.md)
- [Deployment Checklist](../DEPLOYMENT_CHECKLIST.md)
- [Project Structure](PROJECT_STRUCTURE.md)
### **Monitoring and Alerts** ### **Monitoring and Alerts**
- Workflow status badges in README - Workflow status badges in README
-168
View File
@@ -1,168 +0,0 @@
# Grafana Matrix Alerting Setup
## Overview
Configure Grafana to send water level alerts directly to Matrix channels when thresholds are exceeded.
## Prerequisites
- Grafana instance with your PostgreSQL data source
- Matrix account and access token
- Matrix room for alerts
## Step 1: Configure Matrix Contact Point
1. **In Grafana, go to Alerting → Contact Points**
2. **Add new contact point:**
```
Name: matrix-water-alerts
Integration: Webhook
URL: https://matrix.org/_matrix/client/v3/rooms/!ROOM_ID:matrix.org/send/m.room.message
HTTP Method: POST
```
3. **Add Headers:**
```
Authorization: Bearer YOUR_MATRIX_ACCESS_TOKEN
Content-Type: application/json
```
4. **Message Template:**
```json
{
"msgtype": "m.text",
"body": "🌊 WATER ALERT: {{ .CommonLabels.alertname }}\n\nStation: {{ .CommonLabels.station_code }}\nLevel: {{ .CommonAnnotations.water_level }}m\nStatus: {{ .CommonLabels.severity }}\n\nTime: {{ .CommonAnnotations.time }}"
}
```
## Step 2: Create Alert Rules
### High Water Level Alert
```yaml
Rule Name: high-water-level
Query: water_level > 6.0
Condition: IS ABOVE 6.0 FOR 5m
Labels:
- severity: critical
- station_code: {{ .station_code }}
Annotations:
- water_level: {{ .water_level }}
- summary: "Critical water level at {{ .station_code }}"
```
### Low Water Level Alert
```yaml
Rule Name: low-water-level
Query: water_level < 1.0
Condition: IS BELOW 1.0 FOR 10m
Labels:
- severity: warning
- station_code: {{ .station_code }}
```
### Data Gap Alert
```yaml
Rule Name: data-gap
Query: increase(measurements_total[1h]) == 0
Condition: IS EQUAL TO 0 FOR 30m
Labels:
- severity: warning
- issue: data-gap
```
## Step 3: Matrix Setup
### Get Matrix Access Token
```bash
curl -X POST https://matrix.org/_matrix/client/v3/login \
-H "Content-Type: application/json" \
-d '{
"type": "m.login.password",
"user": "your_username",
"password": "your_password"
}'
```
### Create Alert Room
```bash
curl -X POST "https://matrix.org/_matrix/client/v3/createRoom" \
-H "Authorization: Bearer YOUR_ACCESS_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "Water Level Alerts - Northern Thailand",
"topic": "Automated alerts for Ping River water monitoring",
"preset": "trusted_private_chat"
}'
```
## Example Alert Queries
### Critical Water Levels
```promql
# High water alert
water_level{station_code=~"P.1|P.4A|P.20"} > 6.0
# Dangerous discharge
discharge{station_code=~".*"} > 500
# Rapid level change
increase(water_level[15m]) > 0.5
```
### System Health
```promql
# No data received
up{job="water-monitor"} == 0
# Old data
(time() - timestamp) > 7200
```
## Alert Notification Format
Your Matrix messages will look like:
```
🌊 WATER ALERT: High Water Level
Station: P.1 (Chiang Mai)
Level: 6.2m (CRITICAL)
Discharge: 450 cms
Status: DANGER
Time: 2025-09-26 14:30:00
Trend: Rising (+0.3m in 30min)
📍 Location: 18.7883°N, 98.9853°E
```
## Advanced Features
### Escalation Rules
```yaml
# Send to different rooms based on severity
- if: severity == "critical"
receiver: matrix-emergency
- if: severity == "warning"
receiver: matrix-alerts
- if: time_of_day() outside "08:00-20:00"
receiver: matrix-night-duty
```
### Rate Limiting
```yaml
group_wait: 5m
group_interval: 10m
repeat_interval: 30m
```
## Testing Alerts
1. **Test Contact Point** - Use Grafana's test button
2. **Simulate Alert** - Manually trigger with test data
3. **Verify Matrix** - Check message formatting and delivery
## Troubleshooting
### Common Issues
- **403 Forbidden**: Check Matrix access token
- **Room not found**: Verify room ID format
- **No alerts**: Check query syntax and thresholds
- **Spam**: Configure proper grouping and intervals
-351
View File
@@ -1,351 +0,0 @@
# Complete Grafana Matrix Alerting Setup Guide
## Overview
Configure Grafana to send water level alerts directly to Matrix channels when thresholds are exceeded.
## Prerequisites
- Grafana instance running (v8.0+)
- PostgreSQL data source configured in Grafana
- Matrix account
- Matrix room for alerts
## Step 1: Get Matrix Access Token
### Method 1: Using curl
```bash
curl -X POST https://matrix.org/_matrix/client/v3/login \
-H "Content-Type: application/json" \
-d '{
"type": "m.login.password",
"user": "your_username",
"password": "your_password"
}'
```
### Method 2: Using Element Web Client
1. Open Element in browser: https://app.element.io
2. Login to your account
3. Go to Settings → Help & About → Advanced
4. Copy your Access Token
### Method 3: Using Matrix Admin Panel
- If you have admin access to your homeserver, generate token via admin API
## Step 2: Create Alert Room
```bash
curl -X POST "https://matrix.org/_matrix/client/v3/createRoom" \
-H "Authorization: Bearer YOUR_ACCESS_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "Water Level Alerts - Northern Thailand",
"topic": "Automated alerts for Ping River water monitoring",
"preset": "private_chat"
}'
```
Save the `room_id` from the response (format: !roomid:homeserver.com)
## Step 3: Configure Grafana Contact Point
### Navigate to Alerting
1. In Grafana, go to **Alerting → Contact Points**
2. Click **Add contact point**
### Contact Point Settings
```
Name: matrix-water-alerts
Integration: Webhook
URL: https://matrix.org/_matrix/client/v3/rooms/!YOUR_ROOM_ID:matrix.org/send/m.room.message/{{ .GroupLabels.alertname }}_{{ .GroupLabels.severity }}_{{ now.Unix }}
HTTP Method: POST
```
### Headers
```
Authorization: Bearer YOUR_MATRIX_ACCESS_TOKEN
Content-Type: application/json
```
### Message Template (JSON Body)
```json
{
"msgtype": "m.text",
"body": "🌊 **PING RIVER WATER ALERT**\n\n**Alert:** {{ .GroupLabels.alertname }}\n**Severity:** {{ .GroupLabels.severity | toUpper }}\n**Station:** {{ .GroupLabels.station_code }} ({{ .GroupLabels.station_name }})\n\n{{ range .Alerts }}**Status:** {{ .Status | toUpper }}\n**Water Level:** {{ .Annotations.water_level }}m\n**Threshold:** {{ .Annotations.threshold }}m\n**Time:** {{ .StartsAt.Format \"2006-01-02 15:04:05\" }}\n{{ if .Annotations.discharge }}**Discharge:** {{ .Annotations.discharge }} cms\n{{ end }}{{ if .Annotations.message }}**Details:** {{ .Annotations.message }}\n{{ end }}{{ end }}\n📈 **Dashboard:** {{ .ExternalURL }}\n📍 **Location:** Northern Thailand Ping River"
}
```
## Step 4: Create Alert Rules
### High Water Level Alert
```yaml
# Rule Configuration
Rule Name: high-water-level
Evaluation Group: water-level-alerts
Folder: Water Monitoring
# Query A
SELECT
station_code,
station_name_th as station_name,
water_level,
discharge,
timestamp
FROM water_measurements
WHERE
timestamp > now() - interval '5 minutes'
AND water_level > 6.0
# Condition
IS ABOVE 6.0 FOR 5 minutes
# Labels
severity: critical
alertname: High Water Level
station_code: {{ $labels.station_code }}
station_name: {{ $labels.station_name }}
# Annotations
water_level: {{ $values.water_level }}
threshold: 6.0
discharge: {{ $values.discharge }}
summary: Critical water level detected at {{ $labels.station_code }}
```
### Emergency Water Level Alert
```yaml
Rule Name: emergency-water-level
Query: water_level > 8.0
Condition: IS ABOVE 8.0 FOR 2 minutes
Labels:
severity: emergency
alertname: Emergency Water Level
Annotations:
threshold: 8.0
message: IMMEDIATE ACTION REQUIRED - Flood risk imminent
```
### Low Water Level Alert
```yaml
Rule Name: low-water-level
Query: water_level < 1.0
Condition: IS BELOW 1.0 FOR 15 minutes
Labels:
severity: warning
alertname: Low Water Level
Annotations:
threshold: 1.0
message: Drought conditions detected
```
### Data Gap Alert
```yaml
Rule Name: data-gap
Query:
SELECT
station_code,
MAX(timestamp) as last_seen
FROM water_measurements
GROUP BY station_code
HAVING MAX(timestamp) < now() - interval '2 hours'
Condition: HAS NO DATA FOR 30 minutes
Labels:
severity: warning
alertname: Data Gap
issue: missing-data
```
### Rapid Level Change Alert
```yaml
Rule Name: rapid-level-change
Query:
SELECT
station_code,
water_level,
LAG(water_level, 1) OVER (PARTITION BY station_code ORDER BY timestamp) as prev_level
FROM water_measurements
WHERE timestamp > now() - interval '15 minutes'
HAVING ABS(water_level - prev_level) > 0.5
Condition: CHANGE > 0.5m FOR 1 minute
Labels:
severity: warning
alertname: Rapid Water Level Change
```
## Step 5: Configure Notification Policy
### Create Notification Policy
```yaml
# Policy Tree
- receiver: matrix-water-alerts
match:
severity: emergency|critical
group_wait: 10s
group_interval: 5m
repeat_interval: 30m
- receiver: matrix-water-alerts
match:
severity: warning
group_wait: 30s
group_interval: 10m
repeat_interval: 2h
```
### Grouping Rules
```yaml
group_by: [alertname, station_code]
group_wait: 10s
group_interval: 5m
repeat_interval: 1h
```
## Step 6: Station-Specific Thresholds
Create separate rules for each station with appropriate thresholds:
```sql
-- P.1 (Chiang Mai) - Urban area, higher thresholds
SELECT * FROM water_measurements
WHERE station_code = 'P.1' AND water_level > 6.5
-- P.4A (Mae Ping) - Agricultural area
SELECT * FROM water_measurements
WHERE station_code = 'P.4A' AND water_level > 5.0
-- P.20 (Downstream) - Lower threshold
SELECT * FROM water_measurements
WHERE station_code = 'P.20' AND water_level > 4.0
```
## Step 7: Advanced Features
### Time-Based Routing
```yaml
# Different receivers for day/night
time_intervals:
- name: working_hours
time_intervals:
- times:
- start_time: '08:00'
end_time: '20:00'
weekdays: ['monday:friday']
routes:
- receiver: matrix-alerts-day
match:
severity: warning
active_time_intervals: [working_hours]
- receiver: matrix-alerts-night
match:
severity: warning
active_time_intervals: ['!working_hours']
```
### Multi-Channel Alerts
```yaml
# Send critical alerts to multiple rooms
- receiver: matrix-emergency
webhook_configs:
- url: https://matrix.org/_matrix/client/v3/rooms/!emergency:matrix.org/send/m.room.message
http_config:
authorization:
credentials: "Bearer EMERGENCY_TOKEN"
- url: https://matrix.org/_matrix/client/v3/rooms/!general:matrix.org/send/m.room.message
http_config:
authorization:
credentials: "Bearer GENERAL_TOKEN"
```
## Step 8: Testing
### Test Contact Point
1. Go to Contact Points in Grafana
2. Select your Matrix contact point
3. Click "Test" button
4. Check Matrix room for test message
### Test Alert Rules
1. Temporarily lower thresholds
2. Wait for condition to trigger
3. Verify alert appears in Grafana
4. Verify Matrix message received
5. Reset thresholds
### Manual Alert Trigger
```bash
# Simulate high water level in database
INSERT INTO water_measurements (station_code, water_level, timestamp)
VALUES ('P.1', 7.5, NOW());
```
## Troubleshooting
### Common Issues
#### 403 Forbidden
- **Cause**: Invalid Matrix access token
- **Fix**: Regenerate token or check permissions
#### Room Not Found
- **Cause**: Incorrect room ID format
- **Fix**: Ensure room ID starts with ! and includes homeserver
#### No Alerts Firing
- **Cause**: Query returns no results
- **Fix**: Test queries in Grafana Explore, check data availability
#### Alert Spam
- **Cause**: No grouping configured
- **Fix**: Configure proper group_by and intervals
#### Messages Not Formatted
- **Cause**: Template syntax errors
- **Fix**: Validate JSON template, check Grafana template docs
### Debug Steps
1. Check Grafana alert rule status
2. Verify contact point test succeeds
3. Check Grafana logs: `/var/log/grafana/grafana.log`
4. Test Matrix API directly with curl
5. Verify database connectivity and query results
## Environment Variables
Add to your `.env`:
```bash
MATRIX_HOMESERVER=https://matrix.org
MATRIX_ACCESS_TOKEN=your_access_token_here
MATRIX_ROOM_ID=!your_room_id:matrix.org
GRAFANA_URL=http://your-grafana-host:3000
```
## Example Alert Message
Your Matrix messages will appear as:
```
🌊 **PING RIVER WATER ALERT**
**Alert:** High Water Level
**Severity:** CRITICAL
**Station:** P.1 (สถานีเชียงใหม่)
**Status:** FIRING
**Water Level:** 6.75m
**Threshold:** 6.0m
**Time:** 2025-09-26 14:30:00
**Discharge:** 450.2 cms
📈 **Dashboard:** http://grafana:3000
📍 **Location:** Northern Thailand Ping River
```
## Security Notes
- Store Matrix tokens securely (environment variables)
- Use room-specific tokens when possible
- Enable rate limiting to prevent spam
- Consider using dedicated alerting user account
- Regularly rotate access tokens
This setup provides comprehensive water level monitoring with immediate Matrix notifications when thresholds are exceeded.
-389
View File
@@ -1,389 +0,0 @@
# HTTPS VictoriaMetrics Configuration Guide
This guide explains how to configure the Thailand Water Monitor to connect to VictoriaMetrics through HTTPS and reverse proxies.
## Configuration Options
### 1. Environment Variables for HTTPS
```bash
# Option 1: Full HTTPS URL (Recommended)
export DB_TYPE=victoriametrics
export VM_HOST=https://vm.example.com
export VM_PORT=443
# Option 2: Host and port separately
export DB_TYPE=victoriametrics
export VM_HOST=vm.example.com
export VM_PORT=443
# Option 3: Custom port with HTTPS
export DB_TYPE=victoriametrics
export VM_HOST=https://vm.example.com
export VM_PORT=8443
```
### 2. Windows PowerShell Configuration
```powershell
# Set environment variables for HTTPS
$env:DB_TYPE="victoriametrics"
$env:VM_HOST="https://vm.example.com"
$env:VM_PORT="443"
# Run the water monitor
python water_scraper_v3.py
```
### 3. Linux/Mac Configuration
```bash
# Set environment variables for HTTPS
export DB_TYPE=victoriametrics
export VM_HOST=https://vm.example.com
export VM_PORT=443
# Run the water monitor
python water_scraper_v3.py
```
## Reverse Proxy Examples
### 1. Nginx Reverse Proxy
```nginx
server {
listen 443 ssl http2;
server_name vm.example.com;
# SSL Configuration
ssl_certificate /path/to/certificate.crt;
ssl_certificate_key /path/to/private.key;
ssl_protocols TLSv1.2 TLSv1.3;
ssl_ciphers ECDHE-RSA-AES256-GCM-SHA512:DHE-RSA-AES256-GCM-SHA512;
# Security headers
add_header Strict-Transport-Security "max-age=31536000; includeSubDomains" always;
add_header X-Frame-Options DENY always;
add_header X-Content-Type-Options nosniff always;
# Optional: Basic authentication
# auth_basic "VictoriaMetrics";
# auth_basic_user_file /etc/nginx/.htpasswd;
location / {
proxy_pass http://localhost:8428;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
# WebSocket support (if needed)
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
# Timeouts
proxy_connect_timeout 60s;
proxy_send_timeout 60s;
proxy_read_timeout 60s;
}
}
# Redirect HTTP to HTTPS
server {
listen 80;
server_name vm.example.com;
return 301 https://$server_name$request_uri;
}
```
### 2. Apache Reverse Proxy
```apache
<VirtualHost *:443>
ServerName vm.example.com
# SSL Configuration
SSLEngine on
SSLCertificateFile /path/to/certificate.crt
SSLCertificateKeyFile /path/to/private.key
SSLProtocol all -SSLv3 -TLSv1 -TLSv1.1
SSLCipherSuite ECDHE-ECDSA-AES256-GCM-SHA384:ECDHE-RSA-AES256-GCM-SHA384
# Security headers
Header always set Strict-Transport-Security "max-age=31536000; includeSubDomains"
Header always set X-Frame-Options DENY
Header always set X-Content-Type-Options nosniff
# Reverse proxy configuration
ProxyPreserveHost On
ProxyPass / http://localhost:8428/
ProxyPassReverse / http://localhost:8428/
# Optional: Basic authentication
# AuthType Basic
# AuthName "VictoriaMetrics"
# AuthUserFile /etc/apache2/.htpasswd
# Require valid-user
</VirtualHost>
<VirtualHost *:80>
ServerName vm.example.com
Redirect permanent / https://vm.example.com/
</VirtualHost>
```
### 3. Traefik Reverse Proxy
```yaml
# docker-compose.yml with Traefik
version: '3.8'
services:
traefik:
image: traefik:v2.10
command:
- --api.dashboard=true
- --entrypoints.web.address=:80
- --entrypoints.websecure.address=:443
- --providers.docker=true
- --certificatesresolvers.letsencrypt.acme.tlschallenge=true
- --certificatesresolvers.letsencrypt.acme.email=admin@example.com
- --certificatesresolvers.letsencrypt.acme.storage=/letsencrypt/acme.json
ports:
- "80:80"
- "443:443"
volumes:
- /var/run/docker.sock:/var/run/docker.sock
- letsencrypt:/letsencrypt
labels:
- traefik.http.routers.api.rule=Host(`traefik.example.com`)
- traefik.http.routers.api.tls.certresolver=letsencrypt
victoriametrics:
image: victoriametrics/victoria-metrics:latest
command:
- '--storageDataPath=/victoria-metrics-data'
- '--retentionPeriod=2y'
- '--httpListenAddr=:8428'
volumes:
- vm_data:/victoria-metrics-data
labels:
- traefik.enable=true
- traefik.http.routers.vm.rule=Host(`vm.example.com`)
- traefik.http.routers.vm.tls.certresolver=letsencrypt
- traefik.http.services.vm.loadbalancer.server.port=8428
volumes:
vm_data:
letsencrypt:
```
## Testing HTTPS Configuration
### 1. Test Connection
```bash
# Test HTTPS connection
curl -k https://vm.example.com/health
# Test with specific port
curl -k https://vm.example.com:8443/health
# Test API endpoint
curl -k "https://vm.example.com/api/v1/query?query=up"
```
### 2. Test with Water Monitor
```bash
# Set environment variables
export DB_TYPE=victoriametrics
export VM_HOST=https://vm.example.com
export VM_PORT=443
# Test with demo script
python demo_databases.py victoriametrics
# Run full water monitor
python water_scraper_v3.py
```
### 3. Verify SSL Certificate
```bash
# Check SSL certificate
openssl s_client -connect vm.example.com:443 -servername vm.example.com
# Check certificate expiration
echo | openssl s_client -connect vm.example.com:443 2>/dev/null | openssl x509 -noout -dates
```
## Configuration Examples
### 1. Production HTTPS Setup
```bash
# Environment variables for production
export DB_TYPE=victoriametrics
export VM_HOST=https://metrics.company.com
export VM_PORT=443
export LOG_LEVEL=INFO
export SCRAPING_INTERVAL_HOURS=1
# Run water monitor
python water_scraper_v3.py
```
### 2. Development with Self-Signed Certificate
```bash
# For development with self-signed certificates
export DB_TYPE=victoriametrics
export VM_HOST=https://dev-vm.local
export VM_PORT=443
export PYTHONHTTPSVERIFY=0 # Disable SSL verification (dev only)
python water_scraper_v3.py
```
### 3. Custom Port Configuration
```bash
# Custom HTTPS port
export DB_TYPE=victoriametrics
export VM_HOST=https://vm.example.com
export VM_PORT=8443
python water_scraper_v3.py
```
## Troubleshooting HTTPS Issues
### 1. SSL Certificate Errors
```bash
# Error: SSL certificate verify failed
# Solution: Check certificate validity
openssl x509 -in certificate.crt -text -noout
# Temporary workaround (not recommended for production)
export PYTHONHTTPSVERIFY=0
```
### 2. Connection Timeout
```bash
# Error: Connection timeout
# Check firewall and network connectivity
telnet vm.example.com 443
nc -zv vm.example.com 443
```
### 3. DNS Resolution Issues
```bash
# Error: Name resolution failed
# Check DNS resolution
nslookup vm.example.com
dig vm.example.com
```
### 4. Proxy Configuration Issues
```bash
# Check proxy logs
# Nginx
tail -f /var/log/nginx/error.log
# Apache
tail -f /var/log/apache2/error.log
# Test direct connection to backend
curl http://localhost:8428/health
```
## Security Best Practices
### 1. SSL/TLS Configuration
- Use TLS 1.2 or higher
- Disable weak ciphers
- Enable HSTS headers
- Use strong SSL certificates
### 2. Authentication
```nginx
# Basic authentication in Nginx
auth_basic "VictoriaMetrics Access";
auth_basic_user_file /etc/nginx/.htpasswd;
# Create password file
htpasswd -c /etc/nginx/.htpasswd username
```
### 3. Network Security
- Use firewall rules to restrict access
- Consider VPN for internal access
- Implement rate limiting
- Monitor access logs
### 4. Certificate Management
```bash
# Auto-renewal with Let's Encrypt
certbot renew --dry-run
# Certificate monitoring
echo | openssl s_client -connect vm.example.com:443 2>/dev/null | \
openssl x509 -noout -dates | grep notAfter
```
## Docker Configuration for HTTPS
### 1. Docker Compose with HTTPS
```yaml
version: '3.8'
services:
water-monitor:
build: .
environment:
- DB_TYPE=victoriametrics
- VM_HOST=https://vm.example.com
- VM_PORT=443
restart: unless-stopped
depends_on:
- victoriametrics
victoriametrics:
image: victoriametrics/victoria-metrics:latest
ports:
- "8428:8428"
volumes:
- vm_data:/victoria-metrics-data
command:
- '--storageDataPath=/victoria-metrics-data'
- '--retentionPeriod=2y'
- '--httpListenAddr=:8428'
volumes:
vm_data:
```
### 2. Environment File (.env)
```bash
# .env file
DB_TYPE=victoriametrics
VM_HOST=https://vm.example.com
VM_PORT=443
LOG_LEVEL=INFO
SCRAPING_INTERVAL_HOURS=1
```
This configuration guide provides comprehensive instructions for setting up HTTPS connectivity to VictoriaMetrics through reverse proxies, ensuring secure and reliable data transmission for the Thailand Water Monitor.
-136
View File
@@ -1,136 +0,0 @@
# Geolocation Migration Quick Start
This is a quick reference guide for updating a running Thailand Water Monitor system to add geolocation support for Grafana geomap.
## 🚀 **Quick Migration (5 minutes)**
### **Step 1: Stop Application**
```bash
# Stop the service (choose your method)
sudo systemctl stop water-monitor
# OR
docker stop water-monitor
# OR use Ctrl+C if running manually
```
### **Step 2: Backup Database**
```bash
# SQLite backup
cp water_monitoring.db water_monitoring.db.backup
# PostgreSQL backup
pg_dump water_monitoring > backup.sql
# MySQL backup
mysqldump water_monitoring > backup.sql
```
### **Step 3: Run Migration**
```bash
# Run the automated migration script
python migrate_geolocation.py
```
### **Step 4: Restart Application**
```bash
# Restart the service
sudo systemctl start water-monitor
# OR
docker start water-monitor
# OR
python water_scraper_v3.py
```
## ✅ **Expected Output**
```
2025-07-28 17:30:00,123 - INFO - Starting geolocation column migration...
2025-07-28 17:30:00,124 - INFO - Detected database type: SQLITE
2025-07-28 17:30:00,127 - INFO - Added latitude column
2025-07-28 17:30:00,128 - INFO - Added longitude column
2025-07-28 17:30:00,129 - INFO - Added geohash column
2025-07-28 17:30:00,133 - INFO - ✅ Migration completed successfully!
```
## 🗺️ **Verify Geolocation Works**
### **Check Database**
```bash
# SQLite
sqlite3 water_monitoring.db "SELECT station_code, latitude, longitude, geohash FROM stations WHERE station_code = 'P.1';"
# Expected output: P.1|15.6944|100.2028|w5q6uuhvfcfp25
```
### **Test Application**
```bash
# Run a test cycle
python water_scraper_v3.py --test
# Should complete without errors
```
## 🔧 **Grafana Setup**
### **Query for Geomap**
```sql
SELECT
s.latitude, s.longitude, s.station_code, s.english_name,
m.water_level, m.discharge_percent
FROM stations s
JOIN water_measurements m ON s.id = m.station_id
WHERE s.latitude IS NOT NULL
AND m.timestamp = (SELECT MAX(timestamp) FROM water_measurements WHERE station_id = s.id)
```
### **Geomap Configuration**
1. Create new panel → Select "Geomap"
2. Set **Latitude field**: `latitude`
3. Set **Longitude field**: `longitude`
4. Set **Color field**: `water_level`
5. Set **Size field**: `discharge_percent`
## 🚨 **Troubleshooting**
### **Database Locked**
```bash
sudo systemctl stop water-monitor
pkill -f water_scraper
sleep 5
python migrate_geolocation.py
```
### **Permission Error**
```bash
sudo chown $USER:$USER water_monitoring.db
chmod 664 water_monitoring.db
```
### **Missing Dependencies**
```bash
pip install psycopg2-binary pymysql
```
## 🔄 **Rollback (if needed)**
```bash
# Stop application
sudo systemctl stop water-monitor
# Restore backup
cp water_monitoring.db.backup water_monitoring.db
# Restart
sudo systemctl start water-monitor
```
## 📚 **More Information**
- **Full Guide**: See `GEOLOCATION_GUIDE.md`
- **Migration Script**: `migrate_geolocation.py`
- **Database Schema**: Updated with latitude, longitude, geohash columns
## 🎯 **What You Get**
- ✅ **P.1 Station** ready for geomap (Nawarat Bridge)
- ✅ **Database Schema** updated for all 16 stations
- ✅ **Grafana Compatible** data structure
- ✅ **Backward Compatible** - existing data preserved
**Total Time**: ~5 minutes for complete migration
+132
View File
@@ -0,0 +1,132 @@
# Flood notifications (ntfy)
Public push notifications for threshold crossings, without accounts, mailing
lists or app-store review: the monitor publishes to a self-hosted
[ntfy](https://ntfy.sh) server, and anyone subscribes to the topics they care
about from the free ntfy app (iOS, Android, F-Droid) or a browser tab.
ntfy is one Go binary with a sqlite cache: ~30 MB RSS idle, negligible CPU. It
runs on the same VPS as the monitor.
## What subscribers get
Every message is a **transition**, never a state. Crossing up into a level sends
one message; dropping back below it (with 0.10 m hysteresis) sends one
all-clear. A river that sits at 3.9 m for three days produces two messages, not
seventy-two. In a quiet season a subscriber hears nothing.
| Topic | Trigger | Priority |
|---|---|---|
| `ping-warning` | any gauge crosses its warning threshold; levels falling back | 4 (high) / 2 |
| `ping-danger` | any gauge crosses its danger threshold | 5 (max, breaks Do-Not-Disturb) |
| `ping-<station>-warning` | that gauge crosses warning; back to normal | 4 / 2 |
| `ping-<station>-danger` | that gauge crosses danger; back below danger | 5 / 3 |
| `ping-p1-outlook` | model P(warning within 24 h) at P.1 rises through 50 % (clears below 25 %) | 4 / 2 |
| `ping-status` | gauge feed stale ≥ 3 h; feed recovered | 3 / 2 |
Station slugs are the code lowercased without the dot: `p1`, `p103`, `p67`.
Thresholds are the ones in `src/ml/features.py` (`THRESHOLDS`): P.1 3.70 /
4.20 m, P.103 5.95 / 6.75 m, and so on.
The outlook topic is opt-in for a reason: it is model output, and the message
says so. Observed-crossing topics only ever report a gauge reading.
Each message carries a click-through and an "Open dashboard" action button to
the public dashboard.
## How it runs
`src/notify.py` is called once per collection cycle inside the API process
(leader only), right after the forecast precompute, so it sees exactly the
readings and forecasts the dashboard shows. Per-key last-sent state is stored
in the `notification_state` table of the monitor's own database, so a restart
or redeploy never re-sends and never misses a crossing that happened while
the service was down (the next cycle compares against the persisted state).
If ntfy is unreachable the transition is **not** recorded, so it is retried
on the next cycle rather than silently lost. Any other failure in the notify
step is logged and never reaches the collection loop.
The dashboard's "🔔 Get alerts" button appears only when `NTFY_SERVER` is
set; it reads `GET /api/notifications` and renders subscribe links
(`ntfy://` deep links for the app, https links for the web UI).
## Deployment
On the monitor VPS, as root:
```bash
cd /opt/thailand-water-monitor
NTFY_DOMAIN=ntfy.buildfor.life bash scripts/install_ntfy.sh
```
This installs the ntfy .deb, writes `/etc/ntfy/server.yml` (listen on the
host's Tailscale address, port 2586; anonymous read, token-only write, 72 h
message cache, signup/login/metrics off, tight visitor limits), enables the
systemd unit,
creates the `monitor` user with **write-only access to `ping-*`**, mints a
token, and appends `NTFY_SERVER` (public URL for subscribers),
`NTFY_PUBLISH_URL` (loopback, what the monitor POSTs to), `NTFY_TOPIC_PREFIX`
and `NTFY_TOKEN` to `.env` if they are not there yet. Then:
```bash
systemctl restart water-monitor
journalctl -u water-monitor -n 20 | grep ntfy # "ntfy notifications: https://... topics ping-*"
curl -s 'https://ntfy.buildfor.life/ping-status/json?poll=1' # anonymous read works
```
The reverse proxy is a separate VPS on the same tailnet, so ntfy listens on
the monitor host's Tailscale address and nothing is exposed on a public
interface. On the Caddy machine:
```caddyfile
ntfy.buildfor.life {
reverse_proxy <monitor tailscale ip>:2586
}
```
Caddy proxies websockets and keeps long-poll connections open by default;
subscribers hold one open. `behind-proxy: true` makes ntfy rate-limit on
`X-Forwarded-For` rather than treating every subscriber as the proxy.
Publishing does not depend on the domain: `NTFY_PUBLISH_URL` points the
monitor at the Tailscale address directly, so a DNS or proxy problem never
holds back an alert. Test the pipeline before the domain is live with
`curl -s 'http://<tailscale ip>:2586/ping-status/json?poll=1'`.
## Configuration
| Variable | Default | Meaning |
|---|---|---|
| `NTFY_SERVER` | *(empty = off)* | public base URL subscribers use; shown on the dashboard |
| `NTFY_PUBLISH_URL` | = `NTFY_SERVER` | where the monitor POSTs; the local ntfy address (`http://<tailscale ip>:2586`), so publishing never waits on DNS/proxy |
| `NTFY_TOPIC_PREFIX` | `ping` | first segment of every topic |
| `NTFY_TOKEN` | *(empty)* | bearer token if the server requires auth to publish (it does, see above) |
| `PUBLIC_URL` | `https://water.buildfor.life/` | click-through target in messages |
Tunables in `src/notify.py`: `CLEAR_MARGIN_M` (0.10), `OUTLOOK_ON` / `OUTLOOK_OFF`
(0.50 / 0.25), stale feed threshold (3 h, argument to `evaluate`).
## Testing
`tests/test_notify.py` covers the state machine: quiet river sends nothing;
crossing once, then silence while above, then all-clear; hysteresis on the way
down; escalation to danger and back; basin digest grouping; outlook on/off;
heuristic forecasts ignored; stale feed and recovery; state survives a restart
through sqlite; a failed publish is retried next cycle.
To exercise the real path against a real ntfy locally: run `ntfy serve` (any
platform, same binary), set `NTFY_SERVER`/`NTFY_TOKEN`, seed readings, and
poll the topic JSON. `scripts/e2e_notify.py` does exactly that if you want a
template.
## Why ntfy and not …
- **Matrix** (`src/alerting.py`, still there): needs a homeserver account per
subscriber and a room invite; fine for a team, wrong for the public.
- **Gotify**: also self-hosted and light, but Android-only client and one
account per subscriber.
- **Email / SMS**: deliverability work, cost per message, no priority
semantics; ntfy can forward to email per subscription if someone wants it.
- **Telegram / LINE bots**: platform lock-in and a bot token in the loop; can be
added later as ntfy→webhook fan-out without touching the monitor.
-206
View File
@@ -1,206 +0,0 @@
# Thailand Water Monitor - Current Project Status
## 📁 **Clean Project Structure**
The project has been cleaned up and organized with the following structure:
```
water_level_monitor/
├── 📄 .gitignore # Git ignore rules
├── 📄 README.md # Main project documentation
├── 📄 requirements.txt # Python dependencies
├── 📄 config.py # Configuration management
├── 📄 water_scraper_v3.py # Main application (15-min scheduler)
├── 📄 database_adapters.py # Multi-database support
├── 📄 demo_databases.py # Database demonstration
├── 📄 Dockerfile # Container configuration
├── 📄 docker-compose.victoriametrics.yml # VictoriaMetrics stack
├── 📚 Documentation/
│ ├── 📄 DATABASE_DEPLOYMENT_GUIDE.md # Multi-database setup guide
│ ├── 📄 DEBIAN_TROUBLESHOOTING.md # Linux deployment guide
│ ├── 📄 ENHANCED_SCHEDULER_GUIDE.md # 15-minute scheduler guide
│ ├── 📄 GAP_FILLING_GUIDE.md # Data gap filling guide
│ ├── 📄 HTTPS_CONFIGURATION.md # HTTPS setup guide
│ └── 📄 VICTORIAMETRICS_SETUP.md # VictoriaMetrics guide
└── 📁 grafana/ # Grafana configuration
├── 📁 provisioning/
│ ├── 📁 datasources/
│ │ └── 📄 victoriametrics.yml # VictoriaMetrics data source
│ └── 📁 dashboards/
│ └── 📄 dashboard.yml # Dashboard provider config
└── 📁 dashboards/
└── 📄 water-monitoring-dashboard.json # Pre-built dashboard
```
## 🧹 **Files Removed During Cleanup**
### **Old Data Files**
- ❌ `thailand_water_data_v2.csv` - Old CSV export
- ❌ `water_monitor.log` - Log file (regenerated automatically)
- ❌ `water_monitoring.db` - SQLite database (recreated automatically)
### **Outdated Documentation**
- ❌ `FINAL_SUMMARY.md` - Contained references to non-existent v2 files
- ❌ `PROJECT_SUMMARY.md` - Outdated project information
### **System Files**
- ❌ `__pycache__/` - Python compiled files directory
## ✅ **Current Features**
### **Enhanced 15-Minute Scheduler**
- **Timing**: Runs every 15 minutes (1:00, 1:15, 1:30, 1:45, 2:00, etc.)
- **Full Checks**: At :00 minutes (gap filling + data updates)
- **Quick Checks**: At :15, :30, :45 minutes (data fetch only)
- **Gap Filling**: Automatically fills missing historical data
- **Data Updates**: Updates existing records when values change
### **Multi-Database Support**
- **VictoriaMetrics** (Recommended) - High-performance time-series
- **InfluxDB** - Purpose-built time-series database
- **PostgreSQL + TimescaleDB** - Relational with time-series optimization
- **MySQL** - Traditional relational database
- **SQLite** - Local development and testing
### **Production Features**
- **Docker Support**: Complete containerization
- **Grafana Integration**: Pre-built dashboards
- **HTTPS Configuration**: Secure deployment options
- **Health Monitoring**: Comprehensive logging and error handling
- **Gap Detection**: Automatic identification of missing data
- **Retry Logic**: Database lock handling and network error recovery
## 🚀 **Quick Start**
### **1. Basic Setup (SQLite)**
```bash
cd water_level_monitor
pip install -r requirements.txt
python water_scraper_v3.py
```
### **2. VictoriaMetrics Setup**
```bash
# Start VictoriaMetrics + Grafana
docker-compose -f docker-compose.victoriametrics.yml up -d
# Configure environment
export DB_TYPE=victoriametrics
export VM_HOST=localhost
export VM_PORT=8428
# Run monitor
python water_scraper_v3.py
```
### **3. Test Different Databases**
```bash
# Test all supported databases
python demo_databases.py all
# Test specific database
python demo_databases.py victoriametrics
```
## 📊 **Data Collection**
### **Station Coverage**
- **16 Water Monitoring Stations** across Thailand
- **Accurate Station Codes**: P.1, P.20, P.21, P.4A, P.5, P.67, P.75, P.76, P.77, P.81, P.82, P.84, P.85, P.87, P.92, P.103
- **Bilingual Names**: Thai and English station identification
### **Metrics Collected**
- 🌊 **Water Level**: Measured in meters (m)
- 💧 **Discharge**: Measured in cubic meters per second (cms)
- 📊 **Discharge Percentage**: Relative to station capacity
- ⏰ **Timestamp**: Hour 24 handling (midnight = 00:00 next day)
### **Data Frequency**
- **Every 15 Minutes**: Continuous monitoring
- **~300+ Data Points**: Per collection cycle
- **Automatic Gap Filling**: Historical data recovery
- **Data Updates**: Changed values detection and correction
## 🔧 **Command Line Tools**
### **Main Application**
```bash
python water_scraper_v3.py # Run continuous monitoring
python water_scraper_v3.py --test # Single test cycle
python water_scraper_v3.py --help # Show help
```
### **Gap Management**
```bash
python water_scraper_v3.py --check-gaps [days] # Check for missing data
python water_scraper_v3.py --fill-gaps [days] # Fill missing data gaps
python water_scraper_v3.py --update-data [days] # Update existing data
```
### **Database Testing**
```bash
python demo_databases.py # SQLite demo
python demo_databases.py victoriametrics # VictoriaMetrics demo
python demo_databases.py all # Test all databases
```
## 📈 **Monitoring & Visualization**
### **Grafana Dashboard**
- **URL**: http://localhost:3000 (when using docker-compose)
- **Username**: admin
- **Password**: admin_password
- **Features**: Time series charts, status tables, gauges, alerts
### **VictoriaMetrics API**
- **URL**: http://localhost:8428
- **Health**: http://localhost:8428/health
- **Metrics**: http://localhost:8428/metrics
- **Query API**: http://localhost:8428/api/v1/query
## 🛡️ **Security & Production**
### **HTTPS Configuration**
- Complete guide in `HTTPS_CONFIGURATION.md`
- SSL certificate setup
- Reverse proxy configuration
- Security best practices
### **Deployment Options**
- **Docker**: Containerized deployment
- **Systemd**: Linux service configuration
- **Cloud**: AWS, GCP, Azure deployment guides
- **Monitoring**: Health checks and alerting
## 📚 **Documentation**
### **Available Guides**
1. **README.md** - Main project documentation
2. **DATABASE_DEPLOYMENT_GUIDE.md** - Multi-database setup
3. **ENHANCED_SCHEDULER_GUIDE.md** - 15-minute scheduler details
4. **GAP_FILLING_GUIDE.md** - Data integrity and gap filling
5. **DEBIAN_TROUBLESHOOTING.md** - Linux deployment troubleshooting
6. **VICTORIAMETRICS_SETUP.md** - VictoriaMetrics configuration
7. **HTTPS_CONFIGURATION.md** - Secure deployment setup
### **Key Features Documented**
- ✅ Installation and configuration
- ✅ Multi-database support
- ✅ 15-minute scheduling system
- ✅ Gap filling and data integrity
- ✅ Production deployment
- ✅ Monitoring and troubleshooting
- ✅ Security configuration
## 🎯 **Project Status: PRODUCTION READY**
The Thailand Water Monitor is now:
- ✅ **Clean**: All old and redundant files removed
- ✅ **Organized**: Clear project structure with proper documentation
- ✅ **Enhanced**: 15-minute scheduling with gap filling
- ✅ **Scalable**: Multi-database support with VictoriaMetrics
- ✅ **Secure**: HTTPS configuration and security best practices
- ✅ **Monitored**: Comprehensive logging and Grafana dashboards
- ✅ **Documented**: Complete guides for all features and deployment options
The project is ready for production deployment with professional-grade monitoring capabilities.
-272
View File
@@ -1,272 +0,0 @@
# 🏗️ Project Structure - Northern Thailand Ping River Monitor
## 📁 Directory Layout
```
Northern-Thailand-Ping-River-Monitor/
├── 📁 src/ # Main application source code
│ ├── __init__.py # Package initialization
│ ├── main.py # CLI entry point and main application
│ ├── water_scraper_v3.py # Core data collection engine
│ ├── web_api.py # FastAPI web interface
│ ├── config.py # Configuration management
│ ├── database_adapters.py # Database abstraction layer
│ ├── models.py # Data models and type definitions
│ ├── exceptions.py # Custom exception classes
│ ├── validators.py # Data validation layer
│ ├── metrics.py # Metrics collection system
│ ├── health_check.py # Health monitoring system
│ ├── rate_limiter.py # Rate limiting and request tracking
│ └── logging_config.py # Enhanced logging configuration
├── 📁 docs/ # Documentation files
│ ├── STATION_MANAGEMENT_GUIDE.md # Station management documentation
│ ├── ENHANCEMENT_SUMMARY.md # Feature enhancement summary
│ └── PROJECT_STRUCTURE.md # This file
├── 📁 scripts/ # Utility scripts
│ └── migrate_geolocation.py # Database migration script
├── 📁 grafana/ # Grafana configuration
│ ├── dashboards/ # Dashboard definitions
│ └── provisioning/ # Grafana provisioning config
├── 📁 tests/ # Test files
│ ├── test_integration.py # Integration test suite
│ ├── test_station_management.py # Station management tests
│ └── test_api.py # API endpoint tests
├── 📄 run.py # Simple startup script
├── 📄 requirements.txt # Production dependencies
├── 📄 requirements-dev.txt # Development dependencies
├── 📄 setup.py # Package installation script
├── 📄 Dockerfile # Docker container definition
├── 📄 docker-compose.victoriametrics.yml # Complete stack deployment
├── 📄 Makefile # Common development tasks
├── 📄 .env.example # Environment configuration template
├── 📄 .gitignore # Git ignore patterns
├── 📄 .gitlab-ci.yml # CI/CD pipeline configuration
├── 📄 LICENSE # MIT license
├── 📄 README.md # Main project documentation
└── 📄 CONTRIBUTING.md # Contribution guidelines
```
## 🔧 Core Components
### **Application Layer**
- **`src/main.py`** - Command-line interface and application orchestration
- **`src/web_api.py`** - FastAPI web interface with REST endpoints
- **`src/water_scraper_v3.py`** - Core data collection and processing engine
### **Data Layer**
- **`src/database_adapters.py`** - Multi-database support (SQLite, MySQL, PostgreSQL, InfluxDB, VictoriaMetrics)
- **`src/models.py`** - Pydantic data models and type definitions
- **`src/validators.py`** - Data validation and sanitization
### **Infrastructure Layer**
- **`src/config.py`** - Configuration management with environment variable support
- **`src/logging_config.py`** - Structured logging with rotation and colors
- **`src/metrics.py`** - Application metrics collection (counters, gauges, histograms)
- **`src/health_check.py`** - System health monitoring and status checks
### **Utility Layer**
- **`src/exceptions.py`** - Custom exception hierarchy
- **`src/rate_limiter.py`** - API rate limiting and request tracking
## 🌐 Web API Structure
### **Endpoints Organization**
```
/ # Dashboard homepage
├── /health # System health status
├── /metrics # Application metrics
├── /config # Configuration (masked)
├── /stations # Station management
│ ├── GET / # List all stations
│ ├── POST / # Create new station
│ ├── GET /{id} # Get specific station
│ ├── PUT /{id} # Update station
│ └── DELETE /{id} # Delete station
├── /measurements # Data access
│ ├── /latest # Latest measurements
│ └── /station/{code} # Station-specific data
└── /scraping # Data collection control
├── /trigger # Manual data collection
└── /status # Scraping status
```
### **API Models**
- **Request Models**: Station creation/update, query parameters
- **Response Models**: Station info, measurements, health status
- **Error Models**: Standardized error responses
## 🗄️ Database Architecture
### **Supported Databases**
1. **SQLite** - Local development and testing
2. **MySQL** - Traditional relational database
3. **PostgreSQL** - Advanced relational with TimescaleDB support
4. **InfluxDB** - Purpose-built time-series database
5. **VictoriaMetrics** - High-performance metrics storage
### **Schema Design**
```sql
-- Stations table
stations (
id INTEGER PRIMARY KEY,
station_code VARCHAR(10) UNIQUE,
thai_name VARCHAR(255),
english_name VARCHAR(255),
latitude DECIMAL(10,8),
longitude DECIMAL(11,8),
geohash VARCHAR(20),
status VARCHAR(20),
created_at TIMESTAMP,
updated_at TIMESTAMP
)
-- Measurements table
water_measurements (
id BIGINT PRIMARY KEY,
timestamp DATETIME,
station_id INTEGER,
water_level DECIMAL(10,3),
discharge DECIMAL(10,2),
discharge_percent DECIMAL(5,2),
status VARCHAR(20),
created_at TIMESTAMP,
FOREIGN KEY (station_id) REFERENCES stations(id),
UNIQUE(timestamp, station_id)
)
```
## 🐳 Docker Architecture
### **Multi-Stage Build**
1. **Builder Stage** - Compile dependencies and build artifacts
2. **Production Stage** - Minimal runtime environment
### **Service Composition**
- **ping-river-monitor** - Data collection service
- **ping-river-api** - Web API service
- **victoriametrics** - Time-series database
- **grafana** - Visualization dashboard
## 📊 Monitoring Architecture
### **Metrics Collection**
- **Counters** - API requests, database operations, scraping cycles
- **Gauges** - Current values, connection status, resource usage
- **Histograms** - Response times, processing durations
### **Health Checks**
- **Database Health** - Connection status, data freshness
- **API Health** - External API availability, response times
- **System Health** - Memory usage, disk space, CPU load
### **Logging Levels**
- **DEBUG** - Detailed execution information
- **INFO** - General operational messages
- **WARNING** - Potential issues and recoverable errors
- **ERROR** - Serious problems requiring attention
- **CRITICAL** - System-threatening issues
## 🔧 Configuration Management
### **Environment Variables**
```bash
# Database
DB_TYPE=victoriametrics
VM_HOST=localhost
VM_PORT=8428
# Application
SCRAPING_INTERVAL_HOURS=1
LOG_LEVEL=INFO
DATA_RETENTION_DAYS=365
# Security
SECRET_KEY=your-secret-key
API_KEY=your-api-key
```
### **Configuration Hierarchy**
1. Environment variables (highest priority)
2. .env file
3. Default values in config.py (lowest priority)
## 🧪 Testing Architecture
### **Test Categories**
- **Unit Tests** - Individual component testing
- **Integration Tests** - System component interaction
- **API Tests** - Endpoint functionality and responses
- **Performance Tests** - Load and stress testing
### **Test Data**
- **Mock Data** - Simulated API responses
- **Test Database** - Isolated test environment
- **Fixtures** - Reusable test data sets
## 📦 Deployment Architecture
### **Development**
```bash
python run.py --web-api # Local development server
```
### **Production**
```bash
docker-compose up -d # Full stack deployment
```
### **CI/CD Pipeline**
1. **Test Stage** - Run all tests and quality checks
2. **Build Stage** - Create Docker images
3. **Deploy Stage** - Deploy to staging/production
4. **Health Check** - Verify deployment success
## 🔒 Security Architecture
### **Input Validation**
- Pydantic models for API requests
- Data range validation for measurements
- SQL injection prevention through ORM
### **Authentication** (Future)
- API key authentication
- JWT token support
- Role-based access control
### **Data Protection**
- Environment variable configuration
- Sensitive data masking in logs
- HTTPS support for production
## 📈 Performance Architecture
### **Optimization Strategies**
- Database connection pooling
- Query optimization and indexing
- Response caching for static data
- Async processing for I/O operations
### **Scalability Considerations**
- Horizontal scaling with load balancers
- Database read replicas
- Microservice architecture readiness
- Container orchestration support
## 🔄 Data Flow Architecture
### **Collection Flow**
```
External API → Rate Limiter → Data Validator → Database Adapter → Database
```
### **API Flow**
```
HTTP Request → FastAPI → Business Logic → Database Adapter → HTTP Response
```
### **Monitoring Flow**
```
Application Events → Metrics Collector → Health Checks → Monitoring Dashboard
```
This architecture provides a solid foundation for a production-ready water monitoring system with excellent maintainability, scalability, and observability.
Binary file not shown.

Before

Width:  |  Height:  |  Size: 109 KiB

After

Width:  |  Height:  |  Size: 109 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 140 KiB

After

Width:  |  Height:  |  Size: 141 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 131 KiB

After

Width:  |  Height:  |  Size: 130 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 122 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 522 KiB

+723
View File
@@ -0,0 +1,723 @@
[
{
"station": "P.1",
"warn_thr": 3.7,
"folds": [
{
"year": 2021,
"n_train": 20024,
"n_test": 4392,
"events": [],
"variants": {
"rise_rain": {
"mae": 0.07375926701460789,
"mae_above_2p5": null,
"brier_warn": 0.0,
"events": [],
"false_alarm_episodes": 0
},
"rise_rain_fc48": {
"mae": 0.07332598842242062,
"mae_above_2p5": null,
"brier_warn": 0.0,
"events": [],
"false_alarm_episodes": 0
},
"rise_rain_quantile": {
"mae": 0.07383022350644125,
"mae_above_2p5": null,
"brier_warn": 1.0850721383440065e-12,
"events": [],
"false_alarm_episodes": 0
},
"rise_rain_quantile_uw": {
"mae": 0.06977345537342049,
"mae_above_2p5": null,
"brier_warn": 7.128994064266462e-16,
"events": [],
"false_alarm_episodes": 0
}
}
},
{
"year": 2022,
"n_train": 28784,
"n_test": 4392,
"events": [
{
"crossing": "2022-10-02T19:00:00",
"peak_ts": "2022-10-03T15:00:00",
"peak_level": 4.65
}
],
"variants": {
"rise_rain": {
"mae": 0.08452847354970178,
"mae_above_2p5": 0.25082252888260664,
"brier_warn": 0.004300908725927739,
"events": [
{
"crossing": "2022-10-02T19:00:00",
"lead_h": 5.0,
"peak_level": 4.65,
"peak_pred_24h_before": 3.8173954245046406
}
],
"false_alarm_episodes": 0
},
"rise_rain_fc48": {
"mae": 0.08369773912557228,
"mae_above_2p5": 0.24702357745371217,
"brier_warn": 0.004325741658698973,
"events": [
{
"crossing": "2022-10-02T19:00:00",
"lead_h": 5.0,
"peak_level": 4.65,
"peak_pred_24h_before": 3.8114190118860223
}
],
"false_alarm_episodes": 0
},
"rise_rain_quantile": {
"mae": 0.08230165237819607,
"mae_above_2p5": 0.2599241552494541,
"brier_warn": 0.004191550822623699,
"events": [
{
"crossing": "2022-10-02T19:00:00",
"lead_h": 3.0,
"peak_level": 4.65,
"peak_pred_24h_before": 3.693734826616603
}
],
"false_alarm_episodes": 0
},
"rise_rain_quantile_uw": {
"mae": 0.08181109208140971,
"mae_above_2p5": 0.2639358812546371,
"brier_warn": 0.004581181754314636,
"events": [
{
"crossing": "2022-10-02T19:00:00",
"lead_h": 2.0,
"peak_level": 4.65,
"peak_pred_24h_before": 3.620470606696475
}
],
"false_alarm_episodes": 0
}
}
},
{
"year": 2023,
"n_train": 37539,
"n_test": 4392,
"events": [],
"variants": {
"rise_rain": {
"mae": 0.07696637416933236,
"mae_above_2p5": 0.10643655855133666,
"brier_warn": 1.919860722404625e-15,
"events": [],
"false_alarm_episodes": 0
},
"rise_rain_fc48": {
"mae": 0.07708101918232377,
"mae_above_2p5": 0.09974894450126857,
"brier_warn": 1.7028444765622288e-15,
"events": [],
"false_alarm_episodes": 0
},
"rise_rain_quantile": {
"mae": 0.06968496247786034,
"mae_above_2p5": 0.10210094973854984,
"brier_warn": 1.0987601008721297e-09,
"events": [],
"false_alarm_episodes": 0
},
"rise_rain_quantile_uw": {
"mae": 0.06726603345661021,
"mae_above_2p5": 0.09716644572204487,
"brier_warn": 2.7085590803510675e-09,
"events": [],
"false_alarm_episodes": 0
}
}
},
{
"year": 2024,
"n_train": 46323,
"n_test": 4392,
"events": [
{
"crossing": "2024-09-24T17:00:00",
"peak_ts": "2024-09-26T02:00:00",
"peak_level": 4.93
},
{
"crossing": "2024-10-03T09:00:00",
"peak_ts": "2024-10-05T12:00:00",
"peak_level": 5.3
}
],
"variants": {
"rise_rain": {
"mae": 0.0890019157543212,
"mae_above_2p5": 0.2093245927883516,
"brier_warn": 0.005826169840715695,
"events": [
{
"crossing": "2024-09-24T17:00:00",
"lead_h": 10.0,
"peak_level": 4.93,
"peak_pred_24h_before": 4.817519939833057
},
{
"crossing": "2024-10-03T09:00:00",
"lead_h": 21.0,
"peak_level": 5.3,
"peak_pred_24h_before": 5.546588884631041
}
],
"false_alarm_episodes": 0
},
"rise_rain_fc48": {
"mae": 0.08875916089292855,
"mae_above_2p5": 0.20975114218621593,
"brier_warn": 0.006061154593275248,
"events": [
{
"crossing": "2024-09-24T17:00:00",
"lead_h": 11.0,
"peak_level": 4.93,
"peak_pred_24h_before": 4.800924141216692
},
{
"crossing": "2024-10-03T09:00:00",
"lead_h": 72.0,
"peak_level": 5.3,
"peak_pred_24h_before": 5.577164984770105
}
],
"false_alarm_episodes": 0
},
"rise_rain_quantile": {
"mae": 0.08341835741803395,
"mae_above_2p5": 0.18688839374651373,
"brier_warn": 0.003188229521816099,
"events": [
{
"crossing": "2024-09-24T17:00:00",
"lead_h": 17.0,
"peak_level": 4.93,
"peak_pred_24h_before": 5.058029430632501
},
{
"crossing": "2024-10-03T09:00:00",
"lead_h": 21.0,
"peak_level": 5.3,
"peak_pred_24h_before": 5.3311787370709975
}
],
"false_alarm_episodes": 0
},
"rise_rain_quantile_uw": {
"mae": 0.08412336465267245,
"mae_above_2p5": 0.19173218504592401,
"brier_warn": 0.0034722638888286116,
"events": [
{
"crossing": "2024-09-24T17:00:00",
"lead_h": 15.0,
"peak_level": 4.93,
"peak_pred_24h_before": 5.0118284217314
},
{
"crossing": "2024-10-03T09:00:00",
"lead_h": 21.0,
"peak_level": 5.3,
"peak_pred_24h_before": 5.3520431553190155
}
],
"false_alarm_episodes": 0
}
}
},
{
"year": 2025,
"n_train": 55083,
"n_test": 4392,
"events": [
{
"crossing": "2025-09-27T18:00:00",
"peak_ts": "2025-09-27T22:00:00",
"peak_level": 3.93
}
],
"variants": {
"rise_rain": {
"mae": 0.11202835461695825,
"mae_above_2p5": 0.2456782496767171,
"brier_warn": 0.005178356039745929,
"events": [
{
"crossing": "2025-09-27T18:00:00",
"lead_h": 2.0,
"peak_level": 3.93,
"peak_pred_24h_before": 3.23
}
],
"false_alarm_episodes": 0
},
"rise_rain_fc48": {
"mae": 0.11083068059655521,
"mae_above_2p5": 0.23787013498709667,
"brier_warn": 0.005228043825289021,
"events": [
{
"crossing": "2025-09-27T18:00:00",
"lead_h": 2.0,
"peak_level": 3.93,
"peak_pred_24h_before": 3.23
}
],
"false_alarm_episodes": 0
},
"rise_rain_quantile": {
"mae": 0.10372526043263834,
"mae_above_2p5": 0.23846363852091013,
"brier_warn": 0.005898259026876986,
"events": [
{
"crossing": "2025-09-27T18:00:00",
"lead_h": 1.0,
"peak_level": 3.93,
"peak_pred_24h_before": 3.23
}
],
"false_alarm_episodes": 1
},
"rise_rain_quantile_uw": {
"mae": 0.10267267834487567,
"mae_above_2p5": 0.2507951319537275,
"brier_warn": 0.005020467408392599,
"events": [
{
"crossing": "2025-09-27T18:00:00",
"lead_h": 2.0,
"peak_level": 3.93,
"peak_pred_24h_before": 3.2417205711942434
}
],
"false_alarm_episodes": 1
}
}
}
]
},
{
"station": "P.103",
"warn_thr": 5.95,
"folds": [
{
"year": 2021,
"n_train": 20009,
"n_test": 4392,
"events": [],
"variants": {
"rise_rain": {
"mae": 0.14438683486790602,
"mae_above_2p5": null,
"brier_warn": 0.0,
"events": [],
"false_alarm_episodes": 0
},
"rise_rain_fc48": {
"mae": 0.14305613024043953,
"mae_above_2p5": null,
"brier_warn": 0.0,
"events": [],
"false_alarm_episodes": 0
},
"rise_rain_quantile": {
"mae": 0.13512418754716052,
"mae_above_2p5": null,
"brier_warn": 6.937198832251013e-10,
"events": [],
"false_alarm_episodes": 0
},
"rise_rain_quantile_uw": {
"mae": 0.10270057685061172,
"mae_above_2p5": null,
"brier_warn": 6.565945930999852e-14,
"events": [],
"false_alarm_episodes": 0
}
}
},
{
"year": 2022,
"n_train": 28769,
"n_test": 4392,
"events": [
{
"crossing": "2022-08-14T04:00:00",
"peak_ts": "2022-08-14T08:00:00",
"peak_level": 6.09
},
{
"crossing": "2022-10-02T16:00:00",
"peak_ts": "2022-10-03T16:00:00",
"peak_level": 7.54
}
],
"variants": {
"rise_rain": {
"mae": 0.14392334606606189,
"mae_above_2p5": 0.41075382386412596,
"brier_warn": 0.00764679988213082,
"events": [
{
"crossing": "2022-08-14T04:00:00",
"lead_h": 6.0,
"peak_level": 6.09,
"peak_pred_24h_before": 4.825270553204425
},
{
"crossing": "2022-10-02T16:00:00",
"lead_h": 9.0,
"peak_level": 7.54,
"peak_pred_24h_before": 7.194064117976157
}
],
"false_alarm_episodes": 0
},
"rise_rain_fc48": {
"mae": 0.14284722696431765,
"mae_above_2p5": 0.39401132538096667,
"brier_warn": 0.007427375799450148,
"events": [
{
"crossing": "2022-08-14T04:00:00",
"lead_h": 6.0,
"peak_level": 6.09,
"peak_pred_24h_before": 4.845896993132048
},
{
"crossing": "2022-10-02T16:00:00",
"lead_h": 9.0,
"peak_level": 7.54,
"peak_pred_24h_before": 7.194661123502819
}
],
"false_alarm_episodes": 0
},
"rise_rain_quantile": {
"mae": 0.1458954690776336,
"mae_above_2p5": 0.4674051188369517,
"brier_warn": 0.00899991317044724,
"events": [
{
"crossing": "2022-08-14T04:00:00",
"lead_h": 1.0,
"peak_level": 6.09,
"peak_pred_24h_before": 4.82714042795344
},
{
"crossing": "2022-10-02T16:00:00",
"lead_h": 5.0,
"peak_level": 7.54,
"peak_pred_24h_before": 6.59871451008115
}
],
"false_alarm_episodes": 0
},
"rise_rain_quantile_uw": {
"mae": 0.13203958422483053,
"mae_above_2p5": 0.4902773906613983,
"brier_warn": 0.008514527559601331,
"events": [
{
"crossing": "2022-08-14T04:00:00",
"lead_h": 4.0,
"peak_level": 6.09,
"peak_pred_24h_before": 4.895369771408423
},
{
"crossing": "2022-10-02T16:00:00",
"lead_h": 3.0,
"peak_level": 7.54,
"peak_pred_24h_before": 6.294082735853337
}
],
"false_alarm_episodes": 0
}
}
},
{
"year": 2023,
"n_train": 37524,
"n_test": 4392,
"events": [],
"variants": {
"rise_rain": {
"mae": 0.16195301512514512,
"mae_above_2p5": null,
"brier_warn": 3.0208441147560167e-07,
"events": [],
"false_alarm_episodes": 0
},
"rise_rain_fc48": {
"mae": 0.16334224973870387,
"mae_above_2p5": null,
"brier_warn": 1.1626187140381066e-08,
"events": [],
"false_alarm_episodes": 0
},
"rise_rain_quantile": {
"mae": 0.14755024879522577,
"mae_above_2p5": null,
"brier_warn": 3.850104714388072e-05,
"events": [],
"false_alarm_episodes": 0
},
"rise_rain_quantile_uw": {
"mae": 0.10589805327026759,
"mae_above_2p5": null,
"brier_warn": 1.8896757611081001e-07,
"events": [],
"false_alarm_episodes": 0
}
}
},
{
"year": 2024,
"n_train": 46308,
"n_test": 4058,
"events": [
{
"crossing": "2024-09-24T10:00:00",
"peak_ts": "2024-09-26T00:00:00",
"peak_level": 8.27
},
{
"crossing": "2024-09-30T03:00:00",
"peak_ts": "2024-09-30T06:00:00",
"peak_level": 5.99
},
{
"crossing": "2024-10-03T06:00:00",
"peak_ts": "2024-10-05T07:00:00",
"peak_level": 9.93
}
],
"variants": {
"rise_rain": {
"mae": 0.13901842325128816,
"mae_above_2p5": 0.2887416310966684,
"brier_warn": 0.0072040857753137046,
"events": [
{
"crossing": "2024-09-24T10:00:00",
"lead_h": 19.0,
"peak_level": 8.27,
"peak_pred_24h_before": 8.212734363860193
},
{
"crossing": "2024-09-30T03:00:00",
"lead_h": 9.0,
"peak_level": 5.99,
"peak_pred_24h_before": 5.47565489914091
},
{
"crossing": "2024-10-03T06:00:00",
"lead_h": 55.0,
"peak_level": 9.93,
"peak_pred_24h_before": 8.780278464915938
}
],
"false_alarm_episodes": 0
},
"rise_rain_fc48": {
"mae": 0.13850733053801131,
"mae_above_2p5": 0.29817392926900865,
"brier_warn": 0.007928846574280278,
"events": [
{
"crossing": "2024-09-24T10:00:00",
"lead_h": 19.0,
"peak_level": 8.27,
"peak_pred_24h_before": 8.166020251020889
},
{
"crossing": "2024-09-30T03:00:00",
"lead_h": 10.0,
"peak_level": 5.99,
"peak_pred_24h_before": 5.524584037259288
},
{
"crossing": "2024-10-03T06:00:00",
"lead_h": 69.0,
"peak_level": 9.93,
"peak_pred_24h_before": 8.595120917488185
}
],
"false_alarm_episodes": 0
},
"rise_rain_quantile": {
"mae": 0.12640546558686208,
"mae_above_2p5": 0.27944115421093924,
"brier_warn": 0.007730268539879654,
"events": [
{
"crossing": "2024-09-24T10:00:00",
"lead_h": 20.0,
"peak_level": 8.27,
"peak_pred_24h_before": 8.336487732683672
},
{
"crossing": "2024-09-30T03:00:00",
"lead_h": 4.0,
"peak_level": 5.99,
"peak_pred_24h_before": 5.45284915024346
},
{
"crossing": "2024-10-03T06:00:00",
"lead_h": 69.0,
"peak_level": 9.93,
"peak_pred_24h_before": 8.575915226697406
}
],
"false_alarm_episodes": 0
},
"rise_rain_quantile_uw": {
"mae": 0.12480975852168202,
"mae_above_2p5": 0.2857575557415695,
"brier_warn": 0.005894928998647136,
"events": [
{
"crossing": "2024-09-24T10:00:00",
"lead_h": 21.0,
"peak_level": 8.27,
"peak_pred_24h_before": 8.444720326882825
},
{
"crossing": "2024-09-30T03:00:00",
"lead_h": 9.0,
"peak_level": 5.99,
"peak_pred_24h_before": 5.42
},
{
"crossing": "2024-10-03T06:00:00",
"lead_h": 32.0,
"peak_level": 9.93,
"peak_pred_24h_before": 8.781001847167515
}
],
"false_alarm_episodes": 0
}
}
},
{
"year": 2025,
"n_train": 54577,
"n_test": 4392,
"events": [
{
"crossing": "2025-09-26T06:00:00",
"peak_ts": "2025-09-27T21:00:00",
"peak_level": 6.64
},
{
"crossing": "2025-10-03T06:00:00",
"peak_ts": "2025-10-03T12:00:00",
"peak_level": 6.14
}
],
"variants": {
"rise_rain": {
"mae": 0.17206685418009837,
"mae_above_2p5": 0.3840520085986099,
"brier_warn": 0.015824852891785323,
"events": [
{
"crossing": "2025-09-26T06:00:00",
"lead_h": 6.0,
"peak_level": 6.64,
"peak_pred_24h_before": 5.73043665401637
},
{
"crossing": "2025-10-03T06:00:00",
"lead_h": 8.0,
"peak_level": 6.14,
"peak_pred_24h_before": 5.609919760117381
}
],
"false_alarm_episodes": 2
},
"rise_rain_fc48": {
"mae": 0.1722274451362234,
"mae_above_2p5": 0.3872216974630654,
"brier_warn": 0.016245727130063177,
"events": [
{
"crossing": "2025-09-26T06:00:00",
"lead_h": 4.0,
"peak_level": 6.64,
"peak_pred_24h_before": 5.7312263707793685
},
{
"crossing": "2025-10-03T06:00:00",
"lead_h": 8.0,
"peak_level": 6.14,
"peak_pred_24h_before": 5.685511824758637
}
],
"false_alarm_episodes": 2
},
"rise_rain_quantile": {
"mae": 0.15912275848686555,
"mae_above_2p5": 0.3845695612674703,
"brier_warn": 0.014797606329811499,
"events": [
{
"crossing": "2025-09-26T06:00:00",
"lead_h": 11.0,
"peak_level": 6.64,
"peak_pred_24h_before": 5.73
},
{
"crossing": "2025-10-03T06:00:00",
"lead_h": 10.0,
"peak_level": 6.14,
"peak_pred_24h_before": 5.58133012298031
}
],
"false_alarm_episodes": 2
},
"rise_rain_quantile_uw": {
"mae": 0.15385355845103685,
"mae_above_2p5": 0.41314645854725585,
"brier_warn": 0.016528116336109958,
"events": [
{
"crossing": "2025-09-26T06:00:00",
"lead_h": 9.0,
"peak_level": 6.64,
"peak_pred_24h_before": 5.741254234340573
},
{
"crossing": "2025-10-03T06:00:00",
"lead_h": 8.0,
"peak_level": 6.14,
"peak_pred_24h_before": 5.418105405306207
}
],
"false_alarm_episodes": 2
}
}
}
]
}
]
+439
View File
@@ -0,0 +1,439 @@
[
{
"station": "P.1",
"warn_thr": 3.7,
"folds": [
{
"year": 2021,
"n_train": 20024,
"n_test": 4392,
"events": [],
"variants": {
"rise_rain": {
"mae": 0.07375926701460789,
"mae_above_2p5": null,
"brier_warn": 0.0,
"events": [],
"false_alarm_episodes": 0
},
"rise_rain_qsigma": {
"mae": 0.07375926701460789,
"mae_above_2p5": null,
"brier_warn": 7.852802408474157e-14,
"events": [],
"false_alarm_episodes": 0
}
}
},
{
"year": 2022,
"n_train": 28784,
"n_test": 4392,
"events": [
{
"crossing": "2022-10-02T19:00:00",
"peak_ts": "2022-10-03T15:00:00",
"peak_level": 4.65
}
],
"variants": {
"rise_rain": {
"mae": 0.08452847354970178,
"mae_above_2p5": 0.25082252888260664,
"brier_warn": 0.004300908725927739,
"events": [
{
"crossing": "2022-10-02T19:00:00",
"lead_h": 5.0,
"peak_level": 4.65,
"peak_pred_24h_before": 3.8173954245046406
}
],
"false_alarm_episodes": 0
},
"rise_rain_qsigma": {
"mae": 0.08452847354970178,
"mae_above_2p5": 0.25082252888260664,
"brier_warn": 0.0039815091186836665,
"events": [
{
"crossing": "2022-10-02T19:00:00",
"lead_h": 5.0,
"peak_level": 4.65,
"peak_pred_24h_before": 3.8173954245046406
}
],
"false_alarm_episodes": 0
}
}
},
{
"year": 2023,
"n_train": 37539,
"n_test": 4392,
"events": [],
"variants": {
"rise_rain": {
"mae": 0.07696637416933236,
"mae_above_2p5": 0.10643655855133666,
"brier_warn": 1.919860722404625e-15,
"events": [],
"false_alarm_episodes": 0
},
"rise_rain_qsigma": {
"mae": 0.07696637416933236,
"mae_above_2p5": 0.10643655855133666,
"brier_warn": 4.0844243581898366e-08,
"events": [],
"false_alarm_episodes": 0
}
}
},
{
"year": 2024,
"n_train": 46323,
"n_test": 4392,
"events": [
{
"crossing": "2024-09-24T17:00:00",
"peak_ts": "2024-09-26T02:00:00",
"peak_level": 4.93
},
{
"crossing": "2024-10-03T09:00:00",
"peak_ts": "2024-10-05T12:00:00",
"peak_level": 5.3
}
],
"variants": {
"rise_rain": {
"mae": 0.0890019157543212,
"mae_above_2p5": 0.2093245927883516,
"brier_warn": 0.005826169840715695,
"events": [
{
"crossing": "2024-09-24T17:00:00",
"lead_h": 10.0,
"peak_level": 4.93,
"peak_pred_24h_before": 4.817519939833057
},
{
"crossing": "2024-10-03T09:00:00",
"lead_h": 21.0,
"peak_level": 5.3,
"peak_pred_24h_before": 5.546588884631041
}
],
"false_alarm_episodes": 0
},
"rise_rain_qsigma": {
"mae": 0.0890019157543212,
"mae_above_2p5": 0.2093245927883516,
"brier_warn": 0.005475787026695418,
"events": [
{
"crossing": "2024-09-24T17:00:00",
"lead_h": 10.0,
"peak_level": 4.93,
"peak_pred_24h_before": 4.817519939833057
},
{
"crossing": "2024-10-03T09:00:00",
"lead_h": 21.0,
"peak_level": 5.3,
"peak_pred_24h_before": 5.546588884631041
}
],
"false_alarm_episodes": 0
}
}
},
{
"year": 2025,
"n_train": 55083,
"n_test": 4392,
"events": [
{
"crossing": "2025-09-27T18:00:00",
"peak_ts": "2025-09-27T22:00:00",
"peak_level": 3.93
}
],
"variants": {
"rise_rain": {
"mae": 0.11202835461695825,
"mae_above_2p5": 0.2456782496767171,
"brier_warn": 0.005178356039745929,
"events": [
{
"crossing": "2025-09-27T18:00:00",
"lead_h": 2.0,
"peak_level": 3.93,
"peak_pred_24h_before": 3.23
}
],
"false_alarm_episodes": 0
},
"rise_rain_qsigma": {
"mae": 0.11202835461695825,
"mae_above_2p5": 0.2456782496767171,
"brier_warn": 0.00491346243876589,
"events": [
{
"crossing": "2025-09-27T18:00:00",
"lead_h": 2.0,
"peak_level": 3.93,
"peak_pred_24h_before": 3.23
}
],
"false_alarm_episodes": 0
}
}
}
]
},
{
"station": "P.103",
"warn_thr": 5.95,
"folds": [
{
"year": 2021,
"n_train": 20009,
"n_test": 4392,
"events": [],
"variants": {
"rise_rain": {
"mae": 0.14438683486790602,
"mae_above_2p5": null,
"brier_warn": 0.0,
"events": [],
"false_alarm_episodes": 0
},
"rise_rain_qsigma": {
"mae": 0.14438683486790602,
"mae_above_2p5": null,
"brier_warn": 3.5349670949872053e-13,
"events": [],
"false_alarm_episodes": 0
}
}
},
{
"year": 2022,
"n_train": 28769,
"n_test": 4392,
"events": [
{
"crossing": "2022-08-14T04:00:00",
"peak_ts": "2022-08-14T08:00:00",
"peak_level": 6.09
},
{
"crossing": "2022-10-02T16:00:00",
"peak_ts": "2022-10-03T16:00:00",
"peak_level": 7.54
}
],
"variants": {
"rise_rain": {
"mae": 0.14392334606606189,
"mae_above_2p5": 0.41075382386412596,
"brier_warn": 0.00764679988213082,
"events": [
{
"crossing": "2022-08-14T04:00:00",
"lead_h": 6.0,
"peak_level": 6.09,
"peak_pred_24h_before": 4.825270553204425
},
{
"crossing": "2022-10-02T16:00:00",
"lead_h": 9.0,
"peak_level": 7.54,
"peak_pred_24h_before": 7.194064117976157
}
],
"false_alarm_episodes": 0
},
"rise_rain_qsigma": {
"mae": 0.14392334606606189,
"mae_above_2p5": 0.41075382386412596,
"brier_warn": 0.006975493312169236,
"events": [
{
"crossing": "2022-08-14T04:00:00",
"lead_h": 6.0,
"peak_level": 6.09,
"peak_pred_24h_before": 4.825270553204425
},
{
"crossing": "2022-10-02T16:00:00",
"lead_h": 9.0,
"peak_level": 7.54,
"peak_pred_24h_before": 7.194064117976157
}
],
"false_alarm_episodes": 0
}
}
},
{
"year": 2023,
"n_train": 37524,
"n_test": 4392,
"events": [],
"variants": {
"rise_rain": {
"mae": 0.16195301512514512,
"mae_above_2p5": null,
"brier_warn": 3.0208441147560167e-07,
"events": [],
"false_alarm_episodes": 0
},
"rise_rain_qsigma": {
"mae": 0.16195301512514512,
"mae_above_2p5": null,
"brier_warn": 3.2176039819636275e-05,
"events": [],
"false_alarm_episodes": 0
}
}
},
{
"year": 2024,
"n_train": 46308,
"n_test": 4058,
"events": [
{
"crossing": "2024-09-24T10:00:00",
"peak_ts": "2024-09-26T00:00:00",
"peak_level": 8.27
},
{
"crossing": "2024-09-30T03:00:00",
"peak_ts": "2024-09-30T06:00:00",
"peak_level": 5.99
},
{
"crossing": "2024-10-03T06:00:00",
"peak_ts": "2024-10-05T07:00:00",
"peak_level": 9.93
}
],
"variants": {
"rise_rain": {
"mae": 0.13901842325128816,
"mae_above_2p5": 0.2887416310966684,
"brier_warn": 0.0072040857753137046,
"events": [
{
"crossing": "2024-09-24T10:00:00",
"lead_h": 19.0,
"peak_level": 8.27,
"peak_pred_24h_before": 8.212734363860193
},
{
"crossing": "2024-09-30T03:00:00",
"lead_h": 9.0,
"peak_level": 5.99,
"peak_pred_24h_before": 5.47565489914091
},
{
"crossing": "2024-10-03T06:00:00",
"lead_h": 55.0,
"peak_level": 9.93,
"peak_pred_24h_before": 8.780278464915938
}
],
"false_alarm_episodes": 0
},
"rise_rain_qsigma": {
"mae": 0.13901842325128816,
"mae_above_2p5": 0.2887416310966684,
"brier_warn": 0.006910847607098417,
"events": [
{
"crossing": "2024-09-24T10:00:00",
"lead_h": 19.0,
"peak_level": 8.27,
"peak_pred_24h_before": 8.212734363860193
},
{
"crossing": "2024-09-30T03:00:00",
"lead_h": 9.0,
"peak_level": 5.99,
"peak_pred_24h_before": 5.47565489914091
},
{
"crossing": "2024-10-03T06:00:00",
"lead_h": 55.0,
"peak_level": 9.93,
"peak_pred_24h_before": 8.780278464915938
}
],
"false_alarm_episodes": 0
}
}
},
{
"year": 2025,
"n_train": 54577,
"n_test": 4392,
"events": [
{
"crossing": "2025-09-26T06:00:00",
"peak_ts": "2025-09-27T21:00:00",
"peak_level": 6.64
},
{
"crossing": "2025-10-03T06:00:00",
"peak_ts": "2025-10-03T12:00:00",
"peak_level": 6.14
}
],
"variants": {
"rise_rain": {
"mae": 0.17206685418009837,
"mae_above_2p5": 0.3840520085986099,
"brier_warn": 0.015824852891785323,
"events": [
{
"crossing": "2025-09-26T06:00:00",
"lead_h": 6.0,
"peak_level": 6.64,
"peak_pred_24h_before": 5.73043665401637
},
{
"crossing": "2025-10-03T06:00:00",
"lead_h": 8.0,
"peak_level": 6.14,
"peak_pred_24h_before": 5.609919760117381
}
],
"false_alarm_episodes": 2
},
"rise_rain_qsigma": {
"mae": 0.17206685418009837,
"mae_above_2p5": 0.3840520085986099,
"brier_warn": 0.015889933586185904,
"events": [
{
"crossing": "2025-09-26T06:00:00",
"lead_h": 6.0,
"peak_level": 6.64,
"peak_pred_24h_before": 5.73043665401637
},
{
"crossing": "2025-10-03T06:00:00",
"lead_h": 8.0,
"peak_level": 6.14,
"peak_pred_24h_before": 5.609919760117381
}
],
"false_alarm_episodes": 2
}
}
}
]
}
]
-38
View File
@@ -1,38 +0,0 @@
# -*- mode: python ; coding: utf-8 -*-
a = Analysis(
['run.py'],
pathex=[],
binaries=[],
datas=[('.env', '.'), ('sql', 'sql'), ('README.md', '.'), ('POSTGRESQL_SETUP.md', '.'), ('SQLITE_MIGRATION.md', '.')],
hiddenimports=['psycopg2', 'sqlalchemy.dialects.postgresql', 'sqlalchemy.dialects.sqlite', 'dotenv', 'pydantic', 'fastapi', 'uvicorn', 'schedule', 'pandas'],
hookspath=[],
hooksconfig={},
runtime_hooks=[],
excludes=[],
noarchive=False,
optimize=0,
)
pyz = PYZ(a.pure)
exe = EXE(
pyz,
a.scripts,
a.binaries,
a.datas,
[],
name='ping-river-monitor',
debug=False,
bootloader_ignore_signals=False,
strip=False,
upx=True,
upx_exclude=[],
runtime_tmpdir=None,
console=True,
disable_windowed_traceback=False,
argv_emulation=False,
target_arch=None,
codesign_identity=None,
entitlements_file=None,
)
+26 -13
View File
@@ -34,23 +34,23 @@ classifiers = [
"Environment :: Web Environment", "Environment :: Web Environment",
"Framework :: FastAPI" "Framework :: FastAPI"
] ]
requires-python = ">=3.11" requires-python = ">=3.11,<3.12"
dependencies = [ dependencies = [
# Core dependencies # Core dependencies
"requests==2.31.0", "requests==2.34.2",
"schedule==1.2.0", "schedule==1.2.0",
"pandas==2.0.3", "pandas==2.0.3",
"numpy>=1.24,<2", "numpy>=1.24,<2",
# Flood forecasting (ML) # Flood forecasting (ML)
"scikit-learn==1.9.0", "scikit-learn==1.9.0",
# Web API framework # Web API framework
"fastapi==0.104.1", "fastapi==0.141.1",
"uvicorn[standard]==0.24.0", "uvicorn[standard]==0.52.4",
"pydantic==2.5.0", "pydantic==2.13.5",
# Database adapters # Database adapters
"sqlalchemy==2.0.23", "sqlalchemy==2.0.23",
"influxdb==5.3.1", "influxdb==5.3.1",
"pymysql==1.1.0", "pymysql==1.2.0",
"psycopg2-binary==2.9.9", "psycopg2-binary==2.9.9",
# Monitoring and metrics # Monitoring and metrics
"psutil==5.9.6" "psutil==5.9.6"
@@ -59,11 +59,11 @@ dependencies = [
[project.optional-dependencies] [project.optional-dependencies]
dev = [ dev = [
# Testing # Testing
"pytest==7.4.3", "pytest==9.1.1",
"pytest-cov==4.1.0", "pytest-cov==4.1.0",
"pytest-asyncio==0.21.1", "pytest-asyncio==0.21.1",
# Code formatting and linting # Code formatting and linting
"black==23.11.0", "black==26.5.1",
"flake8==6.1.0", "flake8==6.1.0",
"isort==5.12.0", "isort==5.12.0",
"mypy==1.7.1", "mypy==1.7.1",
@@ -73,7 +73,7 @@ dev = [
"ipython==8.17.2", "ipython==8.17.2",
"jupyter==1.0.0", "jupyter==1.0.0",
# Type stubs # Type stubs
"types-requests==2.31.0.10", "types-requests==2.33.0.20260906",
"types-python-dateutil==2.8.19.14" "types-python-dateutil==2.8.19.14"
] ]
docs = [ docs = [
@@ -83,7 +83,7 @@ docs = [
] ]
all = [ all = [
"influxdb==5.3.1", "influxdb==5.3.1",
"pymysql==1.1.0", "pymysql==1.2.0",
"psycopg2-binary==2.9.9" "psycopg2-binary==2.9.9"
] ]
@@ -100,11 +100,11 @@ Documentation = "https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor/
[dependency-groups] [dependency-groups]
dev = [ dev = [
# Testing # Testing
"pytest==7.4.3", "pytest==9.1.1",
"pytest-cov==4.1.0", "pytest-cov==4.1.0",
"pytest-asyncio==0.21.1", "pytest-asyncio==0.21.1",
# Code formatting and linting # Code formatting and linting
"black==23.11.0", "black==26.5.1",
"flake8==6.1.0", "flake8==6.1.0",
"isort==5.12.0", "isort==5.12.0",
"mypy==1.7.1", "mypy==1.7.1",
@@ -114,7 +114,7 @@ dev = [
"ipython==8.17.2", "ipython==8.17.2",
"jupyter==1.0.0", "jupyter==1.0.0",
# Type stubs # Type stubs
"types-requests==2.31.0.10", "types-requests==2.33.0.20260906",
"types-python-dateutil==2.8.19.14", "types-python-dateutil==2.8.19.14",
# Documentation # Documentation
"sphinx==7.2.6", "sphinx==7.2.6",
@@ -128,3 +128,16 @@ where = ["src"]
[tool.setuptools.package-dir] [tool.setuptools.package-dir]
"" = "src" "" = "src"
# One formatting contract for CI, pre-commit and editors. Black's default 88
# columns; isort in black-compatible mode. Run `make format` before committing.
[tool.black]
line-length = 88
target-version = ["py311"]
extend-exclude = '/(\.venv|venv|models|\.claude-flow|\.swarm)/'
[tool.isort]
profile = "black"
line_length = 88
known_first_party = ["src"]
skip_gitignore = true
+3 -3
View File
@@ -2,12 +2,12 @@
-r requirements.txt -r requirements.txt
# Testing # Testing
pytest==7.4.3 pytest==9.1.1
pytest-cov==4.1.0 pytest-cov==4.1.0
pytest-asyncio==0.21.1 pytest-asyncio==0.21.1
# Code formatting and linting # Code formatting and linting
black==23.11.0 black==26.5.1
flake8==6.1.0 flake8==6.1.0
isort==5.12.0 isort==5.12.0
mypy==1.7.1 mypy==1.7.1
@@ -25,5 +25,5 @@ ipython==8.17.2
jupyter==1.0.0 jupyter==1.0.0
# Type stubs # Type stubs
types-requests==2.31.0.10 types-requests==2.33.0.20260906
types-python-dateutil==2.8.19.14 types-python-dateutil==2.8.19.14
+7 -7
View File
@@ -1,5 +1,5 @@
# Core dependencies # Core dependencies
requests==2.31.0 requests==2.34.2
schedule==1.2.0 schedule==1.2.0
pandas==2.0.3 pandas==2.0.3
numpy>=1.24,<2 # pandas 2.0.3 wheels are ABI-incompatible with numpy 2.x numpy>=1.24,<2 # pandas 2.0.3 wheels are ABI-incompatible with numpy 2.x
@@ -8,23 +8,23 @@ numpy>=1.24,<2 # pandas 2.0.3 wheels are ABI-incompatible with numpy 2.x
scikit-learn==1.9.0 scikit-learn==1.9.0
# Web API framework # Web API framework
fastapi==0.104.1 fastapi==0.141.1
uvicorn[standard]==0.24.0 uvicorn[standard]==0.52.4
pydantic==2.5.0 pydantic==2.13.5
# Database adapters # Database adapters
sqlalchemy==2.0.23 sqlalchemy==2.0.23
influxdb==5.3.1 influxdb==5.3.1
pymysql==1.1.0 pymysql==1.2.0
psycopg2-binary==2.9.9 psycopg2-binary==2.9.9
# Monitoring and metrics # Monitoring and metrics
psutil==5.9.6 psutil==5.9.6
# Development dependencies (optional) # Development dependencies (optional)
pytest==7.4.3 pytest==9.1.1
pytest-cov==4.1.0 pytest-cov==4.1.0
black==23.11.0 black==26.5.1
flake8==6.1.0 flake8==6.1.0
mypy==1.7.1 mypy==1.7.1
pre-commit==3.5.0 pre-commit==3.5.0
+68 -11
View File
@@ -1,14 +1,20 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
"""Backfill rid_reservoir_daily with RID large-dam history (Mae Ngat et al.). """Backfill rid_reservoir_daily with RID large-dam history (Mae Ngat et al.).
One request per day against app.rid.go.th/reservoir/api/dams (archive reaches Two paths, both idempotent and both skipping what is already stored, so a
back to at least 2009). Days already stored are skipped, so reruns only fetch rerun repairs holes left by transient failures and is safe alongside the
what is missing safe alongside the hourly live collector, and a rerun hourly live collector:
repairs holes left by transient failures.
--dam-id (default: Mae Ngat) one dam, whole range, via api/dam a handful
of requests for the entire 2009-today archive
--all-dams all ~35 dams, one request per calendar day via
api/dams thousands of requests, ~25 minutes
Usage: Usage:
uv run scripts/backfill_rid_reservoir.py # missing days since 2018-08-01 uv run scripts/backfill_rid_reservoir.py # Mae Ngat since 2018-08-01
uv run scripts/backfill_rid_reservoir.py --start 2015-01-01 uv run scripts/backfill_rid_reservoir.py --start 2009-01-01 # full archive
uv run scripts/backfill_rid_reservoir.py --refresh # rewrite stored days too
uv run scripts/backfill_rid_reservoir.py --all-dams --start 2015-01-01
uv run scripts/backfill_rid_reservoir.py --db-url postgresql://... uv run scripts/backfill_rid_reservoir.py --db-url postgresql://...
""" """
@@ -21,7 +27,12 @@ import sys
sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..")) sys.path.insert(0, os.path.join(os.path.dirname(__file__), ".."))
from src.config import Config from src.config import Config
from src.rid_reservoir import RidReservoirStore, backfill from src.rid_reservoir import (
MAE_NGAT_DAM_ID,
RidReservoirStore,
backfill,
backfill_dam,
)
DEFAULT_START = datetime.date(2018, 8, 1) # start of the water_measurements grid DEFAULT_START = datetime.date(2018, 8, 1) # start of the water_measurements grid
@@ -34,6 +45,27 @@ def main(argv=None) -> int:
parser.add_argument("--end", type=datetime.date.fromisoformat, default=None) parser.add_argument("--end", type=datetime.date.fromisoformat, default=None)
parser.add_argument("--db-url", default=None) parser.add_argument("--db-url", default=None)
parser.add_argument("--throttle", type=float, default=0.4) parser.add_argument("--throttle", type=float, default=0.4)
parser.add_argument(
"--dam-id",
default=MAE_NGAT_DAM_ID,
help="dam to backfill via the fast range endpoint (default Mae Ngat)",
)
parser.add_argument(
"--all-dams",
action="store_true",
help="every dam, one request per calendar day (slow full-fleet path)",
)
parser.add_argument(
"--chunk-days",
type=int,
default=1830,
help="days per range request; the endpoint imposes no limit of its own",
)
parser.add_argument(
"--refresh",
action="store_true",
help="re-fetch days already stored (adds level_msl to api/dams rows)",
)
args = parser.parse_args(argv) args = parser.parse_args(argv)
logging.basicConfig( logging.basicConfig(
@@ -59,10 +91,35 @@ def main(argv=None) -> int:
return 1 return 1
end = args.end or datetime.date.today() end = args.end or datetime.date.today()
span_days = (end - args.start).days + 1 span_days = (end - args.start).days + 1
missing = span_days - len(store.present_dates(args.start, end)) dam_id = None if args.all_dams else args.dam_id
saved = backfill(store, args.start, end, throttle_seconds=args.throttle) missing = span_days - len(store.present_dates(args.start, end, dam_id=dam_id))
print(f"backfilled {saved} dam-day rows ({missing} days were missing)") if args.all_dams:
return 0 if saved or missing == 0 else 1 saved = backfill(store, args.start, end, throttle_seconds=args.throttle)
print(f"backfilled {saved} dam-day rows ({missing} days were missing)")
return 0 if saved or missing == 0 else 1
stats = {}
saved = backfill_dam(
store,
dam_id=args.dam_id,
start=args.start,
end=end,
chunk_days=args.chunk_days,
throttle_seconds=max(args.throttle, 1.0),
skip_present=not args.refresh,
stats=stats,
)
still_missing = span_days - len(
store.present_dates(args.start, end, dam_id=args.dam_id)
)
print(
f"backfilled {saved} dam-day rows for dam {args.dam_id} "
f"(requests: {stats.get('requests', 0)}, {missing} days were missing, "
f"{still_missing} never published by the source)"
)
# A rerun saves nothing once the archive is complete — only a real
# transport/database failure is an error here.
return 1 if stats.get("aborted") or stats.get("failures") else 0
if __name__ == "__main__": if __name__ == "__main__":
+17 -6
View File
@@ -45,19 +45,23 @@ AMBER = "#c07d10"
RED = "#d9534f" RED = "#d9534f"
def fit_backtest_model(df_long: pd.DataFrame, train_end: str): def fit_backtest_model(df_long: pd.DataFrame, train_end: str, use_dam: bool = False):
"""Train the 24 h regression + warning heads on rows <= train_end only. """Train the 24 h regression + warning heads on rows <= train_end only.
Mirrors the deployed hgb-v3 pipeline: the regression head learns the RISE Mirrors the deployed hgb-v3 pipeline: the regression head learns the RISE
over the current level, with Open-Meteo catchment-rain features (trailing over the current level, with Open-Meteo catchment-rain features (trailing
sums + the forward-24h forecast sum); label statistics are bounded to the sums + the forward-24h forecast sum); label statistics are bounded to the
training cutoff. training cutoff. use_dam=True adds the Mae Ngat reservoir columns an
ablation-only configuration (2026-08-13 result: costs 1-3 h of lead).
""" """
from src.ml import dam as dam_mod
from src.ml import rain as rain_mod from src.ml import rain as rain_mod
rain_series = rain_mod.catchment_mean(rain_mod.load_history()) rain_series = rain_mod.catchment_mean(rain_mod.load_history())
dam_frame = dam_mod.load_history() if use_dam else None
X, Y, _meta = features.build_matrix( X, Y, _meta = features.build_matrix(
df_long, STATION, (HORIZON,), stats_end=train_end, rain=rain_series df_long, STATION, (HORIZON,), stats_end=train_end, rain=rain_series,
dam=dam_frame,
) )
train_mask = X.index <= pd.Timestamp(train_end) train_mask = X.index <= pd.Timestamp(train_end)
X_train, Y_train = X.loc[train_mask], Y.loc[train_mask] X_train, Y_train = X.loc[train_mask], Y.loc[train_mask]
@@ -200,16 +204,23 @@ def main(argv=None) -> int:
parser = argparse.ArgumentParser(description=__doc__) parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--db-url", default=None) parser.add_argument("--db-url", default=None)
parser.add_argument("--out-dir", default=os.path.join("docs", "img")) parser.add_argument("--out-dir", default=os.path.join("docs", "img"))
parser.add_argument("--dam", action="store_true",
help="ablation: include Mae Ngat reservoir features "
"(2026-08 result: costs 1-3 h of alert lead)")
parser.add_argument("--no-hii-fill", action="store_true",
help="ablation: load without the HII gap-fill merge")
args = parser.parse_args(argv) args = parser.parse_args(argv)
df = data.load_measurements(db_url=args.db_url) df = data.load_measurements(
db_url=args.db_url, hii_fill=not args.no_hii_fill
)
if df.empty: if df.empty:
print("no measurement data available", file=sys.stderr) print("no measurement data available", file=sys.stderr)
return 1 return 1
os.makedirs(args.out_dir, exist_ok=True) os.makedirs(args.out_dir, exist_ok=True)
# --- October 2024 record flood: trained only on data before 1 Sep 2024 --- # --- October 2024 record flood: trained only on data before 1 Sep 2024 ---
X, reg, clf = fit_backtest_model(df, "2024-08-31") X, reg, clf = fit_backtest_model(df, "2024-08-31", use_dam=args.dam)
obs, fc, flood_start, first_alert = event_series( obs, fc, flood_start, first_alert = event_series(
df, X, reg, clf, "2024-09-10", "2024-10-14 23:00") df, X, reg, clf, "2024-09-10", "2024-10-14 23:00")
peak = float(obs.max()) peak = float(obs.max())
@@ -235,7 +246,7 @@ def main(argv=None) -> int:
detail=True) detail=True)
# --- September 2025 flood: the deployed configuration (trained <= 2024) --- # --- September 2025 flood: the deployed configuration (trained <= 2024) ---
X25, reg25, clf25 = fit_backtest_model(df, "2024-12-31") X25, reg25, clf25 = fit_backtest_model(df, "2024-12-31", use_dam=args.dam)
obs25, fc25, flood25, alert25 = event_series( obs25, fc25, flood25, alert25 = event_series(
df, X25, reg25, clf25, "2025-09-22", "2025-10-02 12:00") df, X25, reg25, clf25, "2025-09-22", "2025-10-02 12:00")
pred_at_alert = float(fc25.loc[alert25:, "pred_max"].iloc[:24].max()) if alert25 is not None else None pred_at_alert = float(fc25.loc[alert25:, "pred_max"].iloc[:24].max()) if alert25 is not None else None
+58
View File
@@ -0,0 +1,58 @@
"""Serve the working-copy dashboard locally with API calls proxied to the
live server, so browser-side changes can be checked against real data
before deploy. Usage: python scripts/dev_proxy.py [port]"""
import http.server
import os
import sys
import urllib.request
from pathlib import Path
UPSTREAM = "https://water.buildfor.life"
STATIC = Path(__file__).resolve().parents[1] / "src" / "static"
class Handler(http.server.BaseHTTPRequestHandler):
def do_GET(self):
if self.path == "/" or self.path.startswith("/?"):
body = (STATIC / "dashboard.html").read_bytes()
self._send(200, "text/html; charset=utf-8", body)
return
# Local overrides for endpoints not yet deployed: DEV_PROXY_LOCAL=/api/x=file.json,...
for pair in filter(None, os.environ.get("DEV_PROXY_LOCAL", "").split(",")):
prefix, file = pair.split("=", 1)
if self.path.split("?")[0] == prefix:
self._send(200, "application/json", Path(file).read_bytes())
return
if self.path.startswith("/static/"):
f = STATIC / self.path[len("/static/"):].split("?")[0]
if f.is_file():
ctype = "application/json" if f.suffix in (".json", ".geojson") else "application/octet-stream"
self._send(200, ctype, f.read_bytes())
return
try:
req = urllib.request.Request(
UPSTREAM + self.path,
headers={"User-Agent": "Mozilla/5.0 (dev_proxy; +https://buildfor.life)", "Accept": "application/json"},
)
with urllib.request.urlopen(req, timeout=60) as r:
self._send(r.status, r.headers.get("Content-Type", "application/json"), r.read())
except urllib.error.HTTPError as e:
self._send(e.code, "application/json", e.read())
def _send(self, code, ctype, body):
self.send_response(code)
self.send_header("Content-Type", ctype)
self.send_header("Content-Length", str(len(body)))
self.send_header("Cache-Control", "no-store")
self.end_headers()
self.wfile.write(body)
def log_message(self, *a):
pass
if __name__ == "__main__":
port = int(sys.argv[1]) if len(sys.argv) > 1 else 8765
print(f"http://localhost:{port}/ (API -> {UPSTREAM})")
http.server.ThreadingHTTPServer(("127.0.0.1", port), Handler).serve_forever()
+132
View File
@@ -0,0 +1,132 @@
"""Drive the production notify path in-process: startup init -> seeded readings
-> forecast cache -> _notify_transitions -> sqlite state -> real ntfy."""
import asyncio
import datetime
import json
import os
import sys
import requests
os.environ.update(
DB_TYPE="sqlite",
WATER_DB_PATH=os.path.join(os.environ["LOCALAPPDATA"], "Temp", "smoke3.db"),
NTFY_SERVER="http://127.0.0.1:2586",
NTFY_TOKEN=os.environ.get("NTFY_TOKEN", ""),
NTFY_TOPIC_PREFIX="ping",
)
for f in ("smoke3.db",):
p = os.path.join(os.environ["LOCALAPPDATA"], "Temp", f)
if os.path.exists(p):
os.remove(p)
from src import web_api # noqa: E402
from src.config import Config # noqa: E402
assert Config.NTFY_SERVER
async def main():
# what the lifespan does at startup, minus the scheduler
from src import notify as notify_mod
from src.forecast_history import ForecastHistoryStore
from src.water_scraper_v3 import EnhancedWaterMonitorScraper
db_config = Config.get_database_config()
web_api.app_state["scraper"] = EnhancedWaterMonitorScraper(db_config)
store = ForecastHistoryStore(db_config["connection_string"], db_config["type"])
store.connect()
web_api.app_state["forecast_store"] = store
state = notify_mod.NotificationState(store.engine, store.db_type)
pub = notify_mod.NtfyPublisher(
Config.NTFY_SERVER, prefix=Config.NTFY_TOPIC_PREFIX, token=Config.NTFY_TOKEN
)
web_api.app_state["notify"] = (pub, state)
scraper = web_api.app_state["scraper"]
now = datetime.datetime.now().replace(minute=0, second=0, microsecond=0)
def seed(level_p1, level_p103, ts):
rows = [
{
"station_code": "P.1",
"station_id": 1,
"timestamp": ts,
"water_level": level_p1,
"discharge": 400.0,
"station_name_en": "Nawarat Bridge",
"station_name_th": "สะพานนวรัฐ",
"discharge_percent": 30.0,
"status": "active",
},
{
"station_code": "P.103",
"station_id": 2,
"timestamp": ts,
"water_level": level_p103,
"discharge": 300.0,
"station_name_en": "Ring Road 3",
"station_name_th": "วงแหวน 3",
"discharge_percent": 20.0,
"status": "active",
},
]
scraper.db_adapter.save_measurements(rows)
def forecast(p):
with web_api.FORECAST_CACHE_LOCK:
web_api.FORECAST_CACHE["all"] = (
0,
[
{
"station_code": "P.1",
"horizon_hours": 24,
"p_warning": p,
"predicted_max_level": 3.9,
"source": "model",
}
],
)
def poll(topic):
out = []
for line in (
requests.get(f"{Config.NTFY_SERVER}/{topic}/json?poll=1", timeout=5)
.text.strip()
.splitlines()
):
m = json.loads(line)
if m.get("event") == "message":
out.append(m.get("title") or m.get("message", "")[:40])
return out
# cycle 1: quiet
seed(1.6, 3.2, now - datetime.timedelta(hours=2))
forecast(0.02)
await web_api._notify_transitions()
# cycle 2: P.1 crosses warning, model outlook on
seed(3.75, 3.3, now - datetime.timedelta(hours=1))
forecast(0.7)
await web_api._notify_transitions()
# cycle 3: same state -> silence
seed(3.80, 3.3, now)
forecast(0.65)
await web_api._notify_transitions()
print("ping-p1-warning:", poll("ping-p1-warning"))
print("ping-warning: ", poll("ping-warning"))
print("ping-p1-outlook:", poll("ping-p1-outlook"))
print("ping-p103-warning:", poll("ping-p103-warning"))
from sqlalchemy import text
with store.engine.connect() as c:
print(
"state table:",
c.execute(
text("SELECT key, state, value FROM notification_state ORDER BY key")
).fetchall(),
)
asyncio.run(main())
-57
View File
@@ -1,57 +0,0 @@
#!/usr/bin/env python3
"""
Password URL encoder for PostgreSQL connection strings
"""
import urllib.parse
import sys
def encode_password(password: str) -> str:
"""URL encode a password for use in connection strings"""
return urllib.parse.quote(password, safe='')
def build_connection_string(username: str, password: str, host: str, port: int, database: str) -> str:
"""Build a properly encoded PostgreSQL connection string"""
encoded_password = encode_password(password)
return f"postgresql://{username}:{encoded_password}@{host}:{port}/{database}"
def main():
print("PostgreSQL Password URL Encoder")
print("=" * 40)
if len(sys.argv) > 1:
# Password provided as argument
password = sys.argv[1]
else:
# Interactive mode
password = input("Enter your password: ")
encoded = encode_password(password)
print(f"\nOriginal password: {password}")
print(f"URL encoded: {encoded}")
# Optional: build full connection string
try:
build_full = input("\nBuild full connection string? (y/N): ").strip().lower() == 'y'
except (EOFError, KeyboardInterrupt):
print("\nDone!")
return
if build_full:
username = input("Username: ").strip()
host = input("Host: ").strip()
port = input("Port [5432]: ").strip() or "5432"
database = input("Database [water_monitoring]: ").strip() or "water_monitoring"
connection_string = build_connection_string(username, password, host, int(port), database)
print(f"\nComplete connection string:")
print(f"POSTGRES_CONNECTION_STRING={connection_string}")
print(f"\nAdd this to your .env file:")
print(f"DB_TYPE=postgresql")
print(f"POSTGRES_CONNECTION_STRING={connection_string}")
if __name__ == "__main__":
main()
-51
View File
@@ -1,51 +0,0 @@
#!/usr/bin/env python3
"""
Generate status badges for README.md
"""
import json
import requests
from datetime import datetime
def generate_badge_url(label, message, color="brightgreen"):
"""Generate a shields.io badge URL"""
return f"https://img.shields.io/badge/{label}-{message}-{color}"
def generate_workflow_badge(repo_url, workflow_name, branch="main"):
"""Generate workflow status badge"""
# For Gitea, you might need to adjust this based on your instance
badge_url = f"{repo_url}/actions/workflows/{workflow_name}/badge.svg?branch={branch}"
return badge_url
def main():
"""Generate badges for the project"""
repo_url = "https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor"
badges = {
"CI/CD": generate_workflow_badge(repo_url, "ci.yml"),
"Security": generate_workflow_badge(repo_url, "security.yml"),
"Documentation": generate_workflow_badge(repo_url, "docs.yml"),
"Python": generate_badge_url("Python", "3.9%2B", "blue"),
"FastAPI": generate_badge_url("FastAPI", "0.104%2B", "green"),
"Docker": generate_badge_url("Docker", "Ready", "blue"),
"License": generate_badge_url("License", "MIT", "green"),
"Version": generate_badge_url("Version", "v3.1.3", "blue"),
}
print("# Status Badges")
print()
print("Add these badges to your README.md:")
print()
for name, url in badges.items():
print(f"[![{name}]({url})]({repo_url})")
print()
print("# Markdown Format")
print()
badge_line = " ".join([f"[![{name}]({url})]({repo_url})" for name, url in badges.items()])
print(badge_line)
if __name__ == "__main__":
main()
-35
View File
@@ -1,35 +0,0 @@
@echo off
REM Git initialization script for Northern Thailand Ping River Monitor
echo 🏔️ Initializing Git repository for Northern Thailand Ping River Monitor
REM Initialize git repository
git init
REM Add remote origin
git remote add origin https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor.git
REM Add all files
git add .
REM Initial commit
git commit -m "Initial commit: Northern Thailand Ping River Monitor v3.1.3
Features:
- Real-time water level monitoring for Ping River Basin
- 16 monitoring stations from Chiang Dao to Nakhon Sawan
- FastAPI web interface with station management
- Multi-database support (SQLite, MySQL, PostgreSQL, InfluxDB, VictoriaMetrics)
- Comprehensive monitoring and health checks
- Docker deployment with Grafana integration
- Production-ready architecture with CI/CD pipeline"
echo ✅ Git repository initialized successfully!
echo.
echo Next steps:
echo 1. Review and edit .env file with your configuration
echo 2. Push to remote repository:
echo git push -u origin main
echo.
echo 3. Start the application:
echo python run.py --web-api
-89
View File
@@ -1,89 +0,0 @@
#!/bin/bash
# Git initialization script for Northern Thailand Ping River Monitor
echo "🏔️ Initializing Git repository for Northern Thailand Ping River Monitor"
# Initialize git repository
git init
# Add remote origin
git remote add origin https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor.git
# Create .gitignore if it doesn't exist
if [ ! -f .gitignore ]; then
echo "Creating .gitignore file..."
cat > .gitignore << 'EOF'
# Python
__pycache__/
*.py[cod]
*.so
.Python
build/
develop-eggs/
dist/
downloads/
eggs/
.eggs/
lib/
lib64/
parts/
sdist/
var/
wheels/
*.egg-info/
.installed.cfg
*.egg
# Virtual environments
.env
.venv
env/
venv/
ENV/
# IDE
.vscode/
.idea/
*.swp
*.swo
# Logs
*.log
logs/
# Database files
*.db
*.sqlite
*.sqlite3
# OS
.DS_Store
Thumbs.db
EOF
fi
# Add all files
git add .
# Initial commit
git commit -m "Initial commit: Northern Thailand Ping River Monitor v3.1.3
Features:
- Real-time water level monitoring for Ping River Basin
- 16 monitoring stations from Chiang Dao to Nakhon Sawan
- FastAPI web interface with station management
- Multi-database support (SQLite, MySQL, PostgreSQL, InfluxDB, VictoriaMetrics)
- Comprehensive monitoring and health checks
- Docker deployment with Grafana integration
- Production-ready architecture with CI/CD pipeline"
echo "✅ Git repository initialized successfully!"
echo ""
echo "Next steps:"
echo "1. Review and edit .env file with your configuration"
echo "2. Push to remote repository:"
echo " git push -u origin main"
echo ""
echo "3. Start the application:"
echo " make run-api"
echo " # or: python run.py --web-api"
+19 -6
View File
@@ -18,6 +18,7 @@ APP_DIR="${APP_DIR:-/opt/thailand-water-monitor}"
SERVICE_USER="${SERVICE_USER:-water-monitor}" SERVICE_USER="${SERVICE_USER:-water-monitor}"
SERVICE_GROUP="${SERVICE_GROUP:-${SERVICE_USER}}" SERVICE_GROUP="${SERVICE_GROUP:-${SERVICE_USER}}"
SERVICE_NAME="water-monitor.service" SERVICE_NAME="water-monitor.service"
RETRAIN_NAME="water-monitor-retrain"
# Resolve the repo root (parent of this scripts/ directory). # Resolve the repo root (parent of this scripts/ directory).
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
@@ -72,11 +73,18 @@ if ! command -v uv >/dev/null 2>&1; then
fi fi
UV="$(command -v uv)" UV="$(command -v uv)"
log "Creating virtualenv at ${APP_DIR}/venv" log "Syncing uv-managed virtualenv at ${APP_DIR}/.venv"
cd "${APP_DIR}" cd "${APP_DIR}"
# Named 'venv' (not uv's default .venv) to match the systemd unit's ExecStart. # ONE environment: uv sync owns .venv/ (from pyproject.toml + uv.lock, so the
"${UV}" venv venv # ML extras such as scikit-learn/joblib are present) and both systemd units
"${UV}" pip install --python venv/bin/python -r requirements.txt # run its interpreter directly. Never create a second env by another name --
# a stale 'venv/' once coexisted here and broke manual retrains with
# ModuleNotFoundError while the service itself ran fine.
"${UV}" sync --python 3.11 --frozen
if [ -d "${APP_DIR}/venv" ]; then
warn "Removing stale ${APP_DIR}/venv (superseded by .venv)"
rm -rf "${APP_DIR}/venv"
fi
# 4. Environment file ---------------------------------------------------------- # 4. Environment file ----------------------------------------------------------
if [ ! -f "${APP_DIR}/.env" ]; then if [ ! -f "${APP_DIR}/.env" ]; then
@@ -100,11 +108,14 @@ if [ -f "${APP_DIR}/.env" ]; then
chmod 0600 "${APP_DIR}/.env" chmod 0600 "${APP_DIR}/.env"
fi fi
# 6. Install and enable the systemd unit -------------------------------------- # 6. Install and enable the systemd units -------------------------------------
log "Installing systemd unit" log "Installing systemd units"
install -m 0644 "${SCRIPT_DIR}/${SERVICE_NAME}" "/etc/systemd/system/${SERVICE_NAME}" install -m 0644 "${SCRIPT_DIR}/${SERVICE_NAME}" "/etc/systemd/system/${SERVICE_NAME}"
install -m 0644 "${SCRIPT_DIR}/${RETRAIN_NAME}.service" "/etc/systemd/system/${RETRAIN_NAME}.service"
install -m 0644 "${SCRIPT_DIR}/${RETRAIN_NAME}.timer" "/etc/systemd/system/${RETRAIN_NAME}.timer"
systemctl daemon-reload systemctl daemon-reload
systemctl enable "${SERVICE_NAME}" systemctl enable "${SERVICE_NAME}"
systemctl enable --now "${RETRAIN_NAME}.timer"
log "Done." log "Done."
echo echo
@@ -112,3 +123,5 @@ echo "Next steps:"
echo " sudo systemctl start ${SERVICE_NAME}" echo " sudo systemctl start ${SERVICE_NAME}"
echo " systemctl status ${SERVICE_NAME}" echo " systemctl status ${SERVICE_NAME}"
echo " sudo journalctl -u ${SERVICE_NAME} -f" echo " sudo journalctl -u ${SERVICE_NAME} -f"
echo " systemctl list-timers ${RETRAIN_NAME}.timer # monthly flood-model retrain"
echo " sudo systemctl start ${RETRAIN_NAME}.service # retrain now"
+102
View File
@@ -0,0 +1,102 @@
#!/usr/bin/env bash
# Install ntfy (https://ntfy.sh) as the public notification server for the
# Ping River Monitor. Run as root on the monitor VPS. Idempotent.
#
# NTFY_DOMAIN=ntfy.buildfor.life bash scripts/install_ntfy.sh
#
# What it does:
# - installs the ntfy .deb from the official GitHub release (single Go
# binary, ~30 MB RSS, sqlite message cache)
# - writes /etc/ntfy/server.yml: listens on the Tailscale address only
# (the reverse proxy is another VPS on the tailnet; nothing is exposed
# on a public interface), anonymous READ on all topics, WRITE only with
# a token. Override with NTFY_LISTEN=host:port.
# - creates the `monitor` publishing user + token, writes NTFY_SERVER /
# NTFY_TOKEN into /opt/thailand-water-monitor/.env if not present
#
# Reverse proxy (on the Caddy VPS, over Tailscale):
# ntfy.buildfor.life {
# reverse_proxy <this host's tailscale ip>:2586
# }
# Caddy passes websockets and keeps long-poll connections open by default;
# subscribers hold one open. ntfy runs with behind-proxy: true so rate
# limits key on X-Forwarded-For, not on the proxy's address.
set -euo pipefail
NTFY_DOMAIN="${NTFY_DOMAIN:?set NTFY_DOMAIN, e.g. ntfy.buildfor.life}"
NTFY_VERSION="${NTFY_VERSION:-2.28.0}"
MONITOR_DIR="${MONITOR_DIR:-/opt/thailand-water-monitor}"
TS_IP="$(tailscale ip -4 2>/dev/null | head -1 || true)"
LISTEN="${NTFY_LISTEN:-${TS_IP:-127.0.0.1}:2586}"
echo "ntfy will listen on ${LISTEN}"
if ! command -v ntfy >/dev/null || [[ "$(ntfy --version 2>/dev/null | awk '{print $3}')" != "$NTFY_VERSION" ]]; then
tmp=$(mktemp -d)
curl -fsSL -o "$tmp/ntfy.deb" \
"https://github.com/binwiederhier/ntfy/releases/download/v${NTFY_VERSION}/ntfy_${NTFY_VERSION}_linux_amd64.deb"
dpkg -i "$tmp/ntfy.deb"
rm -rf "$tmp"
fi
install -d -m 755 /var/cache/ntfy /var/lib/ntfy
cat > /etc/ntfy/server.yml <<EOF
# Ping River Monitor notification server. Managed by scripts/install_ntfy.sh.
base-url: "https://${NTFY_DOMAIN}"
listen-http: "${LISTEN}"
behind-proxy: true
# Messages are kept so a phone that was offline still gets the crossing.
cache-file: "/var/cache/ntfy/cache.db"
cache-duration: "72h"
# Everyone may subscribe; only the monitor (token) may publish.
auth-file: "/var/lib/ntfy/user.db"
auth-default-access: "read-only"
# The monitor publishes a handful of messages per flood; be strict with
# everything else so the box cannot be used as a free relay.
visitor-request-limit-burst: 30
visitor-request-limit-replenish: "10s"
visitor-subscription-limit: 60
visitor-message-daily-limit: 200
attachment-cache-dir: ""
enable-signup: false
enable-login: false
enable-metrics: false
EOF
systemctl enable --now ntfy
systemctl restart ntfy
sleep 1
curl -fsS "http://${LISTEN}/v1/health" >/dev/null && echo "ntfy up on ${LISTEN}"
# Publishing identity for the monitor
if ! ntfy user list 2>/dev/null | grep -q '^user monitor (role'; then
NTFY_PASSWORD="$(openssl rand -base64 24)" ntfy user add --role=user monitor
fi
ntfy access monitor 'ping-*' write-only >/dev/null
# 'ping-*' read stays anonymous via auth-default-access
token=$(ntfy token list monitor 2>/dev/null | awk '/^- tk_/{print $2; exit}') # '- tk_xxx (label), ...'
if [[ -z "$token" ]]; then
token=$(ntfy token add --label "water-monitor" monitor | grep -oE 'tk_[A-Za-z0-9]+' | head -1) # 'token tk_xxx created for user monitor'
fi
env_file="${MONITOR_DIR}/.env"
if [[ -f "$env_file" ]] && ! grep -q '^NTFY_SERVER=' "$env_file"; then
{
echo ""
echo "# ntfy public notifications (scripts/install_ntfy.sh)"
echo "NTFY_SERVER=https://${NTFY_DOMAIN}"
echo "NTFY_PUBLISH_URL=http://${LISTEN}"
echo "NTFY_TOPIC_PREFIX=ping"
echo "NTFY_TOKEN=${token}"
} >> "$env_file"
echo "wrote NTFY_* to ${env_file}; restart water-monitor to enable"
else
echo "NTFY_TOKEN=${token}"
fi
echo
echo "Subscribe test (anonymous read): curl -s 'http://${LISTEN}/ping-status/json?poll=1'"
echo "Publish test (needs token): curl -s -H 'Authorization: Bearer ${token}' -d 'hello' http://${LISTEN}/ping-status"
-294
View File
@@ -1,294 +0,0 @@
#!/usr/bin/env python3
"""
Migration script to add geolocation columns to existing water monitoring database
"""
import os
import sys
import sqlite3
import logging
from typing import Dict, Any
# Configure logging
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s - %(levelname)s - %(message)s'
)
def migrate_sqlite(db_path: str = 'water_monitoring.db') -> bool:
"""Migrate SQLite database to add geolocation columns"""
try:
logging.info(f"Migrating SQLite database: {db_path}")
# Connect to database
conn = sqlite3.connect(db_path)
cursor = conn.cursor()
# Check if columns already exist
cursor.execute("PRAGMA table_info(stations)")
columns = [column[1] for column in cursor.fetchall()]
logging.info(f"Current columns in stations table: {columns}")
# Add columns if they don't exist
columns_added = []
if 'latitude' not in columns:
cursor.execute("ALTER TABLE stations ADD COLUMN latitude REAL")
columns_added.append('latitude')
logging.info("Added latitude column")
if 'longitude' not in columns:
cursor.execute("ALTER TABLE stations ADD COLUMN longitude REAL")
columns_added.append('longitude')
logging.info("Added longitude column")
if 'geohash' not in columns:
cursor.execute("ALTER TABLE stations ADD COLUMN geohash TEXT")
columns_added.append('geohash')
logging.info("Added geohash column")
if columns_added:
# Update P.1 station with sample geolocation data
cursor.execute("""
UPDATE stations
SET latitude = 15.6944, longitude = 100.2028, geohash = 'w5q6uuhvfcfp25'
WHERE station_code = 'P.1'
""")
# Commit changes
conn.commit()
logging.info(f"Successfully added columns: {', '.join(columns_added)}")
logging.info("Updated P.1 station with sample geolocation data")
else:
logging.info("All geolocation columns already exist")
# Verify the changes
cursor.execute("SELECT station_code, latitude, longitude, geohash FROM stations WHERE station_code = 'P.1'")
result = cursor.fetchone()
if result:
logging.info(f"P.1 station geolocation: {result}")
conn.close()
return True
except Exception as e:
logging.error(f"Error migrating SQLite database: {e}")
return False
def migrate_postgresql(connection_string: str) -> bool:
"""Migrate PostgreSQL database to add geolocation columns"""
try:
import psycopg2
from urllib.parse import urlparse
logging.info("Migrating PostgreSQL database")
# Parse connection string
parsed = urlparse(connection_string)
# Connect to database
conn = psycopg2.connect(
host=parsed.hostname,
port=parsed.port or 5432,
database=parsed.path[1:], # Remove leading slash
user=parsed.username,
password=parsed.password
)
cursor = conn.cursor()
# Check if columns exist
cursor.execute("""
SELECT column_name
FROM information_schema.columns
WHERE table_name = 'stations'
""")
columns = [row[0] for row in cursor.fetchall()]
logging.info(f"Current columns in stations table: {columns}")
# Add columns if they don't exist
columns_added = []
if 'latitude' not in columns:
cursor.execute("ALTER TABLE stations ADD COLUMN latitude DECIMAL(10,8)")
columns_added.append('latitude')
logging.info("Added latitude column")
if 'longitude' not in columns:
cursor.execute("ALTER TABLE stations ADD COLUMN longitude DECIMAL(11,8)")
columns_added.append('longitude')
logging.info("Added longitude column")
if 'geohash' not in columns:
cursor.execute("ALTER TABLE stations ADD COLUMN geohash VARCHAR(20)")
columns_added.append('geohash')
logging.info("Added geohash column")
if columns_added:
# Update P.1 station with sample geolocation data
cursor.execute("""
UPDATE stations
SET latitude = 15.6944, longitude = 100.2028, geohash = 'w5q6uuhvfcfp25'
WHERE station_code = 'P.1'
""")
# Commit changes
conn.commit()
logging.info(f"Successfully added columns: {', '.join(columns_added)}")
logging.info("Updated P.1 station with sample geolocation data")
else:
logging.info("All geolocation columns already exist")
conn.close()
return True
except ImportError:
logging.error("psycopg2 not installed. Run: pip install psycopg2-binary")
return False
except Exception as e:
logging.error(f"Error migrating PostgreSQL database: {e}")
return False
def migrate_mysql(connection_string: str) -> bool:
"""Migrate MySQL database to add geolocation columns"""
try:
import pymysql
from urllib.parse import urlparse
logging.info("Migrating MySQL database")
# Parse connection string
parsed = urlparse(connection_string)
# Connect to database
conn = pymysql.connect(
host=parsed.hostname,
port=parsed.port or 3306,
database=parsed.path[1:], # Remove leading slash
user=parsed.username,
password=parsed.password
)
cursor = conn.cursor()
# Check if columns exist
cursor.execute("DESCRIBE stations")
columns = [row[0] for row in cursor.fetchall()]
logging.info(f"Current columns in stations table: {columns}")
# Add columns if they don't exist
columns_added = []
if 'latitude' not in columns:
cursor.execute("ALTER TABLE stations ADD COLUMN latitude DECIMAL(10,8)")
columns_added.append('latitude')
logging.info("Added latitude column")
if 'longitude' not in columns:
cursor.execute("ALTER TABLE stations ADD COLUMN longitude DECIMAL(11,8)")
columns_added.append('longitude')
logging.info("Added longitude column")
if 'geohash' not in columns:
cursor.execute("ALTER TABLE stations ADD COLUMN geohash VARCHAR(20)")
columns_added.append('geohash')
logging.info("Added geohash column")
if columns_added:
# Update P.1 station with sample geolocation data
cursor.execute("""
UPDATE stations
SET latitude = 15.6944, longitude = 100.2028, geohash = 'w5q6uuhvfcfp25'
WHERE station_code = 'P.1'
""")
# Commit changes
conn.commit()
logging.info(f"Successfully added columns: {', '.join(columns_added)}")
logging.info("Updated P.1 station with sample geolocation data")
else:
logging.info("All geolocation columns already exist")
conn.close()
return True
except ImportError:
logging.error("pymysql not installed. Run: pip install pymysql")
return False
except Exception as e:
logging.error(f"Error migrating MySQL database: {e}")
return False
def load_config_from_env() -> Dict[str, Any]:
"""Load database configuration from environment variables"""
db_type = os.getenv('DB_TYPE', 'sqlite').lower()
if db_type == 'postgresql':
return {
'type': 'postgresql',
'connection_string': os.getenv('POSTGRES_CONNECTION_STRING',
'postgresql://postgres:password@localhost/water_monitoring')
}
elif db_type == 'mysql':
return {
'type': 'mysql',
'connection_string': os.getenv('MYSQL_CONNECTION_STRING',
'mysql://root:password@localhost/water_monitoring')
}
elif db_type == 'victoriametrics':
logging.info("VictoriaMetrics doesn't require schema migration")
return {'type': 'victoriametrics'}
elif db_type == 'influxdb':
logging.info("InfluxDB doesn't require schema migration")
return {'type': 'influxdb'}
else:
# Default to SQLite
return {
'type': 'sqlite',
'db_path': os.getenv('SQLITE_DB_PATH', 'water_monitoring.db')
}
def main():
"""Main migration function"""
logging.info("Starting geolocation column migration...")
# Load configuration
config = load_config_from_env()
db_type = config['type']
logging.info(f"Detected database type: {db_type.upper()}")
success = False
if db_type == 'sqlite':
db_path = config.get('db_path', 'water_monitoring.db')
if not os.path.exists(db_path):
logging.error(f"Database file not found: {db_path}")
sys.exit(1)
success = migrate_sqlite(db_path)
elif db_type == 'postgresql':
success = migrate_postgresql(config['connection_string'])
elif db_type == 'mysql':
success = migrate_mysql(config['connection_string'])
elif db_type in ['victoriametrics', 'influxdb']:
logging.info(f"{db_type.upper()} doesn't require schema migration")
success = True
else:
logging.error(f"Unsupported database type: {db_type}")
sys.exit(1)
if success:
logging.info("✅ Migration completed successfully!")
logging.info("You can now restart your water monitoring application")
logging.info("The system will automatically use the new geolocation columns")
else:
logging.error("❌ Migration failed!")
sys.exit(1)
if __name__ == "__main__":
main()
+90
View File
@@ -0,0 +1,90 @@
#!/usr/bin/env bash
#
# Retrain the flood forecast models safely. Run by water-monitor-retrain.timer
# (monthly) or by hand: sudo systemctl start water-monitor-retrain.service
#
# Why a script rather than ExecStart=train_flood_model.py:
# * train.py writes each station's bundle straight into models/ over ~12 min,
# and the API's hourly precompute reloads bundles by mtime. Training into
# a staging dir and mv-ing (atomic on one filesystem) means the API never
# sees a half-written joblib file or a mixed old/new set.
# * A run that produced gauge-only (v2) bundles, or trained too few stations,
# must NOT replace the deployed models. train.py already aborts on a
# missing rain series; this script re-checks the written metrics anyway.
# * No API restart is needed: predict.py reloads changed bundles on the next
# precompute (every scrape cycle, hourly), so the new models are live
# within an hour. Restart manually if you want them live immediately.
#
# Exit codes: 0 ok, 2 training refused (see log), 3 verification failed.
set -euo pipefail
APP_DIR="${APP_DIR:-/opt/thailand-water-monitor}"
PYTHON="${PYTHON:-${APP_DIR}/.venv/bin/python}"
MODELS_DIR="${APP_DIR}/models"
STAGE_DIR="${MODELS_DIR}/.staging"
# P.4A is NOT_TRAINABLE by design (17% fill); 15 of 16 is the normal outcome.
MIN_TRAINED="${MIN_TRAINED:-14}"
EXPECT_VERSION_PREFIX="${EXPECT_VERSION_PREFIX:-hgb-v3+}"
log() { printf '%s retrain: %s\n' "$(date '+%Y-%m-%d %H:%M:%S')" "$*"; }
cd "${APP_DIR}"
[ -x "${PYTHON}" ] || { log "no interpreter at ${PYTHON} (run uv sync)"; exit 3; }
rm -rf "${STAGE_DIR}"
mkdir -p "${STAGE_DIR}"
log "training into ${STAGE_DIR} (python=${PYTHON}, OMP_NUM_THREADS=${OMP_NUM_THREADS:-unset})"
# train_flood_model.py exits 2 on a missing rain series (RainUnavailableError)
# instead of silently writing v2 bundles -- propagate that unchanged.
set +e
"${PYTHON}" scripts/train_flood_model.py --stations all --models-dir "${STAGE_DIR}" "$@"
rc=$?
set -e
if [ "${rc}" -ne 0 ]; then
log "training failed (exit ${rc}); deployed models untouched"
rm -rf "${STAGE_DIR}"
exit "${rc}"
fi
# Verify before promoting. Reads metrics.json from the stage dir.
VERSION="$("${PYTHON}" - "${STAGE_DIR}/metrics.json" <<'PY'
import json, sys
m = json.load(open(sys.argv[1]))
print(m["model_version"])
PY
)"
TRAINED="$("${PYTHON}" - "${STAGE_DIR}/metrics.json" <<'PY'
import json, sys
m = json.load(open(sys.argv[1]))
print(sum(1 for s in m["stations"].values() if s.get("status") == "trained"))
PY
)"
log "staged model_version=${VERSION} trained_stations=${TRAINED}"
case "${VERSION}" in
"${EXPECT_VERSION_PREFIX}"*) ;;
*)
log "REFUSING to deploy: version '${VERSION}' does not start with '${EXPECT_VERSION_PREFIX}'"
rm -rf "${STAGE_DIR}"
exit 3
;;
esac
if [ "${TRAINED}" -lt "${MIN_TRAINED}" ]; then
log "REFUSING to deploy: only ${TRAINED} stations trained (< ${MIN_TRAINED})"
rm -rf "${STAGE_DIR}"
exit 3
fi
# Promote: per-file rename is atomic; readers see either the old or the new
# bundle, never a partial one. Keep one previous generation for rollback.
mkdir -p "${MODELS_DIR}/.previous"
for f in "${STAGE_DIR}"/flood_*.joblib "${STAGE_DIR}/metrics.json"; do
name="$(basename "${f}")"
if [ -f "${MODELS_DIR}/${name}" ]; then
mv -f "${MODELS_DIR}/${name}" "${MODELS_DIR}/.previous/${name}"
fi
mv -f "${f}" "${MODELS_DIR}/${name}"
done
rm -rf "${STAGE_DIR}"
log "deployed ${VERSION} (${TRAINED} stations); previous generation in models/.previous. The API picks it up on its next hourly precompute."
+62
View File
@@ -0,0 +1,62 @@
"""Summarise rolling-origin harness output side by side.
Usage:
uv run python scripts/summarize_eval.py models/eval_2026-09-12.json [more.json ...]
Aggregates each (station, variant) across folds: mean MAE, mean flood-regime
MAE, mean Brier, total false-alarm episodes, and every warning event with its
first-alert lead and the 24 h-ahead peak error -- the operational numbers that
decide whether a variant ships.
"""
import json
import statistics
import sys
from collections import OrderedDict
def summarize(paths):
for path in paths:
results = json.load(open(path, encoding="utf-8"))
print(f"\n##### {path}")
for station in results:
print(f"\n=== {station['station']} (warn {station['warn_thr']:.2f} m) ===")
agg = OrderedDict()
for fold in station["folds"]:
for name, m in fold["variants"].items():
a = agg.setdefault(
name, {"mae": [], "mae_hi": [], "brier": [], "fa": 0, "events": []}
)
a["mae"].append(m["mae"])
if m.get("mae_above_2p5") is not None:
a["mae_hi"].append(m["mae_above_2p5"])
if m.get("brier_warn") is not None:
a["brier"].append(m["brier_warn"])
a["fa"] += m["false_alarm_episodes"]
for e in m["events"]:
err = (
None
if e["peak_pred_24h_before"] is None
else e["peak_pred_24h_before"] - e["peak_level"]
)
a["events"].append((fold["year"], e["crossing"][:10], e["lead_h"], e["peak_level"], err))
print(f"{'variant':22} {'MAE':>6} {'MAE_hi':>7} {'Brier':>7} {'FA':>3} events: year crossing lead_h peak(err24h)")
for name, a in agg.items():
ev = " ".join(
f"{y} {d} {'' if l is None else format(l, '+.0f')}h {p:.2f}({'' if err is None else format(err, '+.2f')})"
for y, d, l, p, err in a["events"]
)
leads = [l for *_, l, _, _ in a["events"] if l is not None]
print(
f"{name:22} {statistics.mean(a['mae']):6.3f} "
f"{statistics.mean(a['mae_hi']) if a['mae_hi'] else float('nan'):7.3f} "
f"{statistics.mean(a['brier']) if a['brier'] else float('nan'):7.4f} "
f"{a['fa']:>3} {ev}"
)
if leads:
print(f"{'':22} lead: mean {statistics.mean(leads):+.1f} h, min {min(leads):+.0f} h, "
f"missed {sum(1 for *_, l, _, _ in a['events'] if l is None)}/{len(a['events'])}")
if __name__ == "__main__":
summarize(sys.argv[1:] or ["models/eval_variants.json"])
+2 -2
View File
@@ -11,7 +11,7 @@ import sys
sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..")) sys.path.insert(0, os.path.join(os.path.dirname(__file__), ".."))
from src.ml.train import main from src.ml.train import cli
if __name__ == "__main__": if __name__ == "__main__":
main() raise SystemExit(cli())
+39
View File
@@ -0,0 +1,39 @@
[Unit]
Description=Retrain the Ping River flood forecast models
Documentation=https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor/-/blob/master/docs/FLOOD_FORECASTING.md
After=network-online.target
Wants=network-online.target
[Service]
Type=oneshot
User=water-monitor
Group=water-monitor
WorkingDirectory=/opt/thailand-water-monitor
EnvironmentFile=/opt/thailand-water-monitor/.env
# Same interpreter as water-monitor.service -- the uv-managed .venv.
# scripts/retrain.sh trains into models/.staging, refuses to promote anything
# that is not a rain-enabled (hgb-v3) set covering the expected stations, then
# renames the bundles into place. The API reloads them on its next hourly
# precompute; no restart, so a failed run leaves the old models serving.
ExecStart=/bin/bash /opt/thailand-water-monitor/scripts/retrain.sh
# HistGradientBoosting is CPU-bound; cap threads so training cannot starve
# the API (docs/FLOOD_FORECASTING.md section 6 measured 4 as the sweet spot).
Environment=OMP_NUM_THREADS=4
Environment=PYTHONPATH=/opt/thailand-water-monitor
Environment=PYTHONUNBUFFERED=1
Nice=15
IOSchedulingClass=idle
# 15 stations at ~50 s each plus data load: 12 min observed on 2026-09-12.
TimeoutStartSec=45min
# Same sandbox as the API unit.
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=strict
ProtectHome=true
ReadWritePaths=/opt/thailand-water-monitor
CapabilityBoundingSet=
StandardOutput=journal
StandardError=journal
SyslogIdentifier=water-monitor-retrain
+17
View File
@@ -0,0 +1,17 @@
[Unit]
Description=Monthly flood-model retrain (docs/FLOOD_FORECASTING.md section 7)
[Timer]
# Policy: at minimum once pre-monsoon (May-June), monthly through the season
# (July-November), and after any major flood. A retrain costs ~12 min and RAM
# peaks ~300 MB, so running it every month all year is cheaper than remembering
# which months matter. 1st of the month, 03:30 server-local -- between the
# hourly scrapes and outside Thai daytime traffic.
OnCalendar=*-*-01 03:30:00
# Catch up if the box was off at the scheduled time.
Persistent=true
RandomizedDelaySec=20min
Unit=water-monitor-retrain.service
[Install]
WantedBy=timers.target
+7 -5
View File
@@ -9,17 +9,19 @@ Type=simple
User=water-monitor User=water-monitor
Group=water-monitor Group=water-monitor
WorkingDirectory=/opt/thailand-water-monitor WorkingDirectory=/opt/thailand-water-monitor
ExecStart=/opt/thailand-water-monitor/venv/bin/python src/water_scraper_v3.py # The uv-managed env (uv sync -> .venv). Same interpreter for water-monitor-retrain.service.
ExecStart=/opt/thailand-water-monitor/.venv/bin/python run.py --web-api
ExecReload=/bin/kill -HUP $MAINPID ExecReload=/bin/kill -HUP $MAINPID
Restart=always Restart=always
RestartSec=60 RestartSec=60
TimeoutStopSec=30 TimeoutStopSec=30
# Environment variables # DB_TYPE / POSTGRES_CONNECTION_STRING / MATRIX_* come from the .env file.
Environment=DB_TYPE=victoriametrics EnvironmentFile=/opt/thailand-water-monitor/.env
Environment=VM_HOST=localhost
Environment=VM_PORT=8428
Environment=PYTHONPATH=/opt/thailand-water-monitor Environment=PYTHONPATH=/opt/thailand-water-monitor
# Serving path is latency-bound; single-threaded BLAS is 2.6x faster per call
# (docs/FLOOD_FORECASTING.md section 6). Training sets its own value.
Environment=OMP_NUM_THREADS=1
Environment=PYTHONUNBUFFERED=1 Environment=PYTHONUNBUFFERED=1
# Security settings # Security settings
-106
View File
@@ -1,106 +0,0 @@
#!/usr/bin/env python3
"""
Setup script for Northern Thailand Ping River Monitor
"""
from setuptools import setup, find_packages
import os
# Read the README file
with open("README.md", "r", encoding="utf-8") as fh:
long_description = fh.read()
# Read requirements
try:
with open("requirements.txt", "r", encoding="utf-8") as fh:
requirements = [line.strip() for line in fh if line.strip() and not line.startswith("#")]
except FileNotFoundError:
# Fallback to minimal requirements if file not found
requirements = [
"requests>=2.31.0",
"schedule>=1.2.0",
"pandas>=2.1.0",
"fastapi>=0.104.0",
"uvicorn>=0.24.0",
]
# Extract core requirements (exclude dev dependencies)
core_requirements = []
for req in requirements:
if not any(dev_keyword in req.lower() for dev_keyword in ['pytest', 'black', 'flake8', 'mypy', 'sphinx']):
core_requirements.append(req)
setup(
name="northern-thailand-ping-river-monitor",
version="3.1.3",
author="Ping River Monitor Team",
author_email="contact@example.com",
description="Real-time water level monitoring system for the Ping River Basin in Northern Thailand",
long_description=long_description,
long_description_content_type="text/markdown",
url="https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor",
project_urls={
"Bug Tracker": "https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor/issues",
"Documentation": "https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor/wiki",
"Source Code": "https://git.b4l.co.th/B4L/Northern-Thailand-Ping-River-Monitor",
},
packages=find_packages(),
classifiers=[
"Development Status :: 4 - Beta",
"Intended Audience :: Science/Research",
"Intended Audience :: System Administrators",
"Topic :: Scientific/Engineering :: Hydrology",
"Topic :: System :: Monitoring",
"License :: OSI Approved :: MIT License",
"Programming Language :: Python :: 3",
"Programming Language :: Python :: 3.9",
"Programming Language :: Python :: 3.10",
"Programming Language :: Python :: 3.11",
"Programming Language :: Python :: 3.12",
"Operating System :: OS Independent",
"Environment :: Web Environment",
"Framework :: FastAPI",
],
python_requires=">=3.9",
install_requires=core_requirements,
extras_require={
"dev": [
"pytest>=7.4.3",
"pytest-cov>=4.1.0",
"black>=23.11.0",
"flake8>=6.1.0",
"mypy>=1.7.1",
"pre-commit>=3.5.0",
],
"docs": [
"sphinx>=7.2.6",
"sphinx-rtd-theme>=1.3.0",
],
"all": [
"influxdb>=5.3.1",
"pymysql>=1.1.0",
"psycopg2-binary>=2.9.9",
],
},
entry_points={
"console_scripts": [
"ping-river-monitor=src.main:main",
"ping-river-api=src.web_api:main",
],
},
include_package_data=True,
package_data={
"src": ["*.py"],
},
keywords=[
"water monitoring",
"hydrology",
"thailand",
"ping river",
"environmental monitoring",
"time series",
"fastapi",
"real-time data",
],
zip_safe=False,
)
+7 -3
View File
@@ -12,9 +12,13 @@ __description__ = "Northern Thailand Ping River Monitoring System"
from .config import Config from .config import Config
from .database_adapters import DatabaseAdapter, create_database_adapter from .database_adapters import DatabaseAdapter, create_database_adapter
from .exceptions import (APIConnectionError, ConfigurationError, from .exceptions import (
DatabaseConnectionError, DataValidationError, APIConnectionError,
WaterMonitorException) ConfigurationError,
DatabaseConnectionError,
DataValidationError,
WaterMonitorException,
)
from .models import DatabaseConfig, StationInfo, WaterMeasurement from .models import DatabaseConfig, StationInfo, WaterMeasurement
from .water_scraper_v3 import EnhancedWaterMonitorScraper from .water_scraper_v3 import EnhancedWaterMonitorScraper
+11
View File
@@ -38,6 +38,17 @@ class Config:
TARGET_URL = "https://hyd-app-db.rid.go.th/hydro1h.html" TARGET_URL = "https://hyd-app-db.rid.go.th/hydro1h.html"
API_URL = "https://hyd-app-db.rid.go.th/webservice/getGroupHourlyWaterLevelReportAllHL.ashx" API_URL = "https://hyd-app-db.rid.go.th/webservice/getGroupHourlyWaterLevelReportAllHL.ashx"
THAIWATER_API_KEY = os.getenv("THAIWATER_API_KEY") THAIWATER_API_KEY = os.getenv("THAIWATER_API_KEY")
# Public flood notifications (ntfy). Off unless NTFY_SERVER is set.
# NTFY_SERVER is what subscribers use (public https URL, shown on the
# dashboard). NTFY_PUBLISH_URL is where the monitor POSTs; defaults to
# NTFY_SERVER, set it to http://127.0.0.1:2586 when ntfy runs on the same
# host so publishing never depends on DNS/proxy/tunnel being up.
NTFY_SERVER = os.getenv("NTFY_SERVER", "").strip()
NTFY_PUBLISH_URL = os.getenv("NTFY_PUBLISH_URL", "").strip() or NTFY_SERVER
NTFY_TOPIC_PREFIX = os.getenv("NTFY_TOPIC_PREFIX", "ping").strip()
NTFY_TOKEN = os.getenv("NTFY_TOKEN", "").strip() # publish token if ACL enabled
PUBLIC_URL = os.getenv("PUBLIC_URL", "https://water.buildfor.life/").strip()
REQUEST_TIMEOUT = int(os.getenv("REQUEST_TIMEOUT", "30")) REQUEST_TIMEOUT = int(os.getenv("REQUEST_TIMEOUT", "30"))
USER_AGENT = ( USER_AGENT = (
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 " "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
+30 -30
View File
@@ -139,12 +139,16 @@ class InfluxDBAdapter(DatabaseAdapter):
"time": measurement["timestamp"].isoformat(), "time": measurement["timestamp"].isoformat(),
"fields": { "fields": {
"water_level": float(measurement["water_level"]), "water_level": float(measurement["water_level"]),
"discharge": float(measurement["discharge"]) "discharge": (
if measurement.get("discharge") is not None float(measurement["discharge"])
else None, if measurement.get("discharge") is not None
"discharge_percent": float(measurement["discharge_percent"]) else None
if measurement.get("discharge_percent") ),
else None, "discharge_percent": (
float(measurement["discharge_percent"])
if measurement.get("discharge_percent")
else None
),
}, },
} }
points.append(point) points.append(point)
@@ -551,13 +555,13 @@ class SQLAdapter(DatabaseAdapter):
"station_code": row[1], "station_code": row[1],
"station_name_en": row[2], "station_name_en": row[2],
"station_name_th": row[3], "station_name_th": row[3],
"water_level": float(row[4]) "water_level": (
if row[4] is not None float(row[4]) if row[4] is not None else None
else None, ),
"discharge": float(row[5]) if row[5] is not None else None, "discharge": float(row[5]) if row[5] is not None else None,
"discharge_percent": float(row[6]) "discharge_percent": (
if row[6] is not None float(row[6]) if row[6] is not None else None
else None, ),
"status": row[7], "status": row[7],
} }
) )
@@ -611,13 +615,13 @@ class SQLAdapter(DatabaseAdapter):
"station_code": row[1], "station_code": row[1],
"station_name_en": row[2], "station_name_en": row[2],
"station_name_th": row[3], "station_name_th": row[3],
"water_level": float(row[4]) "water_level": (
if row[4] is not None float(row[4]) if row[4] is not None else None
else None, ),
"discharge": float(row[5]) if row[5] is not None else None, "discharge": float(row[5]) if row[5] is not None else None,
"discharge_percent": float(row[6]) "discharge_percent": (
if row[6] is not None float(row[6]) if row[6] is not None else None
else None, ),
"status": row[7], "status": row[7],
} }
) )
@@ -666,13 +670,13 @@ class SQLAdapter(DatabaseAdapter):
"station_id": row[1], "station_id": row[1],
"station_code": row[2] or f"Station_{row[1]}", "station_code": row[2] or f"Station_{row[1]}",
"station_name_th": row[3] or f"Station {row[1]}", "station_name_th": row[3] or f"Station {row[1]}",
"water_level": float(row[4]) "water_level": (
if row[4] is not None float(row[4]) if row[4] is not None else None
else None, ),
"discharge": float(row[5]) if row[5] is not None else None, "discharge": float(row[5]) if row[5] is not None else None,
"discharge_percent": float(row[6]) "discharge_percent": (
if row[6] is not None float(row[6]) if row[6] is not None else None
else None, ),
"status": row[7], "status": row[7],
} }
) )
@@ -767,9 +771,7 @@ class SQLAdapter(DatabaseAdapter):
return hours_by_day return hours_by_day
except Exception as e: except Exception as e:
logging.error( logging.error(f"Error querying {self.db_type.upper()} recorded hours: {e}")
f"Error querying {self.db_type.upper()} recorded hours: {e}"
)
return None return None
def get_database_stats(self) -> Optional[Dict]: def get_database_stats(self) -> Optional[Dict]:
@@ -814,9 +816,7 @@ class SQLAdapter(DatabaseAdapter):
# the DISTINCT day-hour slots and coverage cannot exceed 100% # the DISTINCT day-hour slots and coverage cannot exceed 100%
first_slot = first_ts.replace(minute=0, second=0, microsecond=0) first_slot = first_ts.replace(minute=0, second=0, microsecond=0)
last_slot = last_ts.replace(minute=0, second=0, microsecond=0) last_slot = last_ts.replace(minute=0, second=0, microsecond=0)
expected_hours = ( expected_hours = int((last_slot - first_slot).total_seconds() // 3600) + 1
int((last_slot - first_slot).total_seconds() // 3600) + 1
)
recorded_hours = int(row[4]) recorded_hours = int(row[4])
coverage_percent = round(100.0 * recorded_hours / expected_hours, 1) coverage_percent = round(100.0 * recorded_hours / expected_hours, 1)
+3 -3
View File
@@ -125,9 +125,9 @@ class DatabaseHealthCheck(HealthCheck):
"message": "Database connection OK", "message": "Database connection OK",
"details": { "details": {
"latest_data_count": len(latest_data), "latest_data_count": len(latest_data),
"latest_timestamp": str(latest_data[0].get("timestamp")) "latest_timestamp": (
if latest_data str(latest_data[0].get("timestamp")) if latest_data else None
else None, ),
}, },
} }
+2 -4
View File
@@ -17,9 +17,9 @@ import time
from typing import Dict, List, Optional from typing import Dict, List, Optional
from .hii_collector import ( from .hii_collector import (
PING_BASIN_CODE,
HiiClient, HiiClient,
HiiStore, HiiStore,
PING_BASIN_CODE,
_parse_datetime, _parse_datetime,
_to_float, _to_float,
) )
@@ -145,9 +145,7 @@ def backfill(
station_rows += store.save_waterlevel_history(sid, rows) station_rows += store.save_waterlevel_history(sid, rows)
except Exception as e: except Exception as e:
totals["errors"] += 1 totals["errors"] += 1
logger.warning( logger.warning(f"{label}: {chunk_start}..{chunk_end} failed: {e}")
f"{label}: {chunk_start}..{chunk_end} failed: {e}"
)
time.sleep(sleep_seconds) time.sleep(sleep_seconds)
totals["rows"] += station_rows totals["rows"] += station_rows
logger.info(f"{label} (id {sid}): {station_rows} rows saved") logger.info(f"{label} (id {sid}): {station_rows} rows saved")
+3 -9
View File
@@ -66,9 +66,7 @@ def rid_code_from_oldcode(oldcode: Optional[str]) -> Optional[str]:
return match.group(1) if match else None return match.group(1) if match else None
def parse_rain_records( def parse_rain_records(payload: Dict, basin_code: int = PING_BASIN_CODE) -> List[Dict]:
payload: Dict, basin_code: int = PING_BASIN_CODE
) -> List[Dict]:
"""Extract per-station rainfall rows from a rain_24h payload.""" """Extract per-station rainfall rows from a rain_24h payload."""
records = [] records = []
for row in payload.get("data") or []: for row in payload.get("data") or []:
@@ -186,9 +184,7 @@ class HiiStore:
def __init__(self, connection_string: str, db_type: str): def __init__(self, connection_string: str, db_type: str):
self.db_type = db_type.lower() self.db_type = db_type.lower()
if self.db_type not in ("sqlite", "postgresql", "mysql"): if self.db_type not in ("sqlite", "postgresql", "mysql"):
raise ValueError( raise ValueError(f"HII collection requires a SQL database, got '{db_type}'")
f"HII collection requires a SQL database, got '{db_type}'"
)
self.connection_string = connection_string self.connection_string = connection_string
self.engine = None self.engine = None
@@ -400,9 +396,7 @@ class HiiStore:
from sqlalchemy import text from sqlalchemy import text
now = datetime.datetime.now() now = datetime.datetime.now()
station_sql = self._upsert( station_sql = self._upsert(station_table, ["id"], station_cols + ["updated_at"])
station_table, ["id"], station_cols + ["updated_at"]
)
measurement_sql = self._upsert( measurement_sql = self._upsert(
measurement_table, ["station_id", "timestamp"], measurement_cols measurement_table, ["station_id", "timestamp"], measurement_cols
) )
+109
View File
@@ -0,0 +1,109 @@
"""Mae Ngat reservoir series for the flood models.
rid_reservoir_daily (collected hourly by src/rid_reservoir.py, backfilled to
2018) holds daily storage/inflow/outflow for every RID large dam. Mae Ngat
Somboon Chon (DAM_ID 200103) is the only large dam upstream of Chiang Mai:
in Oct 2024 its inflow hit 19-22 MCM/day and storage 114% of usable capacity
days around the P.1 crossing upstream state no river gauge carries.
Leakage rule: RID publishes the daily report for date D on the morning of D,
so the row becomes visible to features at D 07:00 local time, never earlier.
Forward-fill is capped at FFILL_LIMIT_H so a stalled collector degrades to
NaN (HGB-native) instead of silently serving stale reservoir state.
Known residual optimism: the collector upserts keep-last (and re-fetches
yesterday), so the stored row for date D is RID's FINAL revision, which
training then back-dates to D 07:00 values live serving may not have had
that morning. This bias works IN FAVOR of dam features, so the 2026-08-13
negative result (they cost 1-3 h of alert lead) holds a fortiori; but any
future POSITIVE result must first validate intraday row stability or shift
the flow columns to D+1 07:00.
"""
import datetime
import logging
from pathlib import Path
from typing import Optional
import pandas as pd
from ..rid_reservoir import MAE_NGAT_DAM_ID
from .data import CACHE_DIR, resolve_db_url
logger = logging.getLogger(__name__)
REPORT_HOUR = 7 # daily value valid from 07:00 local on its own date
FFILL_LIMIT_H = 48 # two missed daily reports -> NaN, not stale state
DAM_COLUMNS = ("storage_pct", "inflow_mcm", "outflow_mcm")
CACHE_FILE = f"dam_{MAE_NGAT_DAM_ID}.csv.gz"
def load_daily(
db_url: Optional[str] = None,
dam_id: str = MAE_NGAT_DAM_ID,
start: Optional[datetime.date] = None,
cache_dir: Path = CACHE_DIR,
) -> Optional[pd.DataFrame]:
"""Daily dam rows indexed by date. DB first, on-disk cache as fallback."""
cache_path = Path(cache_dir) / CACHE_FILE
resolved = resolve_db_url(db_url)
if resolved:
try:
from sqlalchemy import create_engine, text
query = (
"SELECT date, storage_pct, inflow_mcm, outflow_mcm "
"FROM rid_reservoir_daily WHERE dam_id = :dam_id"
)
params = {"dam_id": dam_id}
if start is not None:
query += " AND date >= :start"
params["start"] = start
engine = create_engine(resolved, pool_pre_ping=True)
with engine.connect() as conn:
daily = pd.read_sql(text(query + " ORDER BY date"), conn, params=params)
daily["date"] = pd.to_datetime(daily["date"])
daily = daily.set_index("date")
for col in DAM_COLUMNS:
daily[col] = pd.to_numeric(daily[col], errors="coerce")
# Only full, NON-EMPTY loads refresh the cache: a truncated or
# freshly-recreated table must not wipe a good fallback archive.
if start is None and not daily.empty:
cache_path.parent.mkdir(parents=True, exist_ok=True)
daily.to_csv(cache_path, compression="gzip")
return daily
except Exception as error:
logger.warning(f"dam series DB load failed: {error}")
if cache_path.exists():
logger.warning("falling back to on-disk cache for the dam series")
return pd.read_csv(cache_path, index_col=0, parse_dates=True)
return None
def hourly_frame(daily: Optional[pd.DataFrame]) -> Optional[pd.DataFrame]:
"""Step the daily rows onto an hourly grid, each valid from D 07:00."""
if daily is None or daily.empty:
return None
frame = daily.copy()
frame.index = pd.to_datetime(frame.index) + pd.Timedelta(hours=REPORT_HOUR)
frame = frame[~frame.index.duplicated(keep="last")].sort_index()
hourly_index = pd.date_range(
frame.index.min(),
frame.index.max() + pd.Timedelta(hours=FFILL_LIMIT_H),
freq="h",
)
return frame.reindex(hourly_index).ffill(limit=FFILL_LIMIT_H)
def load_history(db_url: Optional[str] = None) -> Optional[pd.DataFrame]:
"""Full hourly Mae Ngat history for training; None when unavailable."""
return hourly_frame(load_daily(db_url))
def serving_frame(
db_url: Optional[str] = None, days: int = 21
) -> Optional[pd.DataFrame]:
"""Recent hourly dam state for inference (covers the 336 h feature window
plus the 72 h storage-delta lag)."""
start = datetime.date.today() - datetime.timedelta(days=days)
return hourly_frame(load_daily(db_url, start=start))
+11 -7
View File
@@ -23,7 +23,9 @@ from .features import UPSTREAM_LEADS
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
DEFAULT_API_URL = "http://100.81.167.42:8000" # Public dashboard. Override with --api-url for a local instance; the
# Tailscale address of the server is deliberately not the default here.
DEFAULT_API_URL = "https://water.buildfor.life"
# Anchored to the repo root so training/prediction work from any CWD; a relative # Anchored to the repo root so training/prediction work from any CWD; a relative
# path here silently produced 0 rows when the CLI ran outside the repo root. # path here silently produced 0 rows when the CLI ran outside the repo root.
CACHE_DIR = Path(__file__).resolve().parents[2] / "models" / "cache" CACHE_DIR = Path(__file__).resolve().parents[2] / "models" / "cache"
@@ -205,15 +207,13 @@ def fill_from_hii(
"timestamp": missing["timestamp"], "timestamp": missing["timestamp"],
"station_code": code, "station_code": code,
"water_level": missing["wl_msl"] - offset, "water_level": missing["wl_msl"] - offset,
"discharge": missing["discharge"] "discharge": (
if code in _HII_EXACT_MIRRORS missing["discharge"] if code in _HII_EXACT_MIRRORS else float("nan")
else float("nan"), ),
} }
) )
fills.append(fill) fills.append(fill)
logger.info( logger.info(f"HII gap-fill {code}: +{len(fill)} hours (offset {offset:.3f} m)")
f"HII gap-fill {code}: +{len(fill)} hours (offset {offset:.3f} m)"
)
if not fills: if not fills:
return df return df
return _normalize_long(pd.concat([df] + fills, ignore_index=True)) return _normalize_long(pd.concat([df] + fills, ignore_index=True))
@@ -282,6 +282,10 @@ def _read_cache(cache_dir: Path, stations: Optional[List[str]]) -> pd.DataFrame:
frames = [] frames = []
for path in sorted(cache_dir.glob("*.csv.gz")): for path in sorted(cache_dir.glob("*.csv.gz")):
code = path.name[: -len(".csv.gz")] code = path.name[: -len(".csv.gz")]
# The dir is shared with rain.py / dam.py caches (rain_openmeteo,
# dam_<id>): only station files (P.<n>) are measurements.
if not code.startswith("P."):
continue
if stations and code not in stations: if stations and code not in stations:
continue continue
with gzip.open(path, "rt", encoding="utf-8") as handle: with gzip.open(path, "rt", encoding="utf-8") as handle:
+164 -39
View File
@@ -52,20 +52,40 @@ def _flood_weights(y_abs: pd.Series) -> np.ndarray:
return 1.0 + 4.0 * np.clip((y_abs.to_numpy() - 2.5) / 1.2, 0.0, 1.0) return 1.0 + 4.0 * np.clip((y_abs.to_numpy() - 2.5) / 1.2, 0.0, 1.0)
# Experimental forward-48h rain sum, built in evaluate_station (not in
# features.build_features) so the served feature set is untouched until the
# harness says it helps. Serving could supply it: fetch_forecast() already
# pulls forecast_days=2.
EXTRA_RAIN_FEATURES = ("rain_fc48",)
class Variant: class Variant:
"""A trainable candidate producing (pred_abs, sigma_per_row) on test rows.""" """A trainable candidate producing (pred_abs, sigma_per_row) on test rows."""
def __init__(self, name: str, target: str, weighted: bool = False, def __init__(
quantile: bool = False, use_rain: bool = False): self,
name: str,
target: str,
weighted: bool = False,
quantile: bool = False,
use_rain: bool = False,
use_dam: bool = False,
use_fc48: bool = False,
qsigma: bool = False,
):
self.name = name self.name = name
self.target = target # 'abs' or 'rise' self.target = target # 'abs' or 'rise'
self.weighted = weighted self.weighted = weighted
self.quantile = quantile self.quantile = quantile
self.use_rain = use_rain self.use_rain = use_rain
self.use_dam = use_dam
self.use_fc48 = use_fc48
# Hybrid: L2 head for the point prediction (keeps the lead-time
# behaviour of the deployed model exactly, since p>=0.5 alerts are
# sigma-independent) and quantile heads ONLY for a per-row sigma.
self.qsigma = qsigma
def fit_predict( def fit_predict(self, X_tr, y_abs_tr, X_te) -> Tuple[np.ndarray, np.ndarray]:
self, X_tr, y_abs_tr, X_te
) -> Tuple[np.ndarray, np.ndarray]:
if not self.use_rain: if not self.use_rain:
drop = [c for c in features.RAIN_FEATURES if c in X_tr.columns] drop = [c for c in features.RAIN_FEATURES if c in X_tr.columns]
X_tr = X_tr.drop(columns=drop) X_tr = X_tr.drop(columns=drop)
@@ -74,6 +94,20 @@ class Variant:
raise ValueError( raise ValueError(
f"{self.name} requires the rain series (run without --no-rain)" f"{self.name} requires the rain series (run without --no-rain)"
) )
if not self.use_dam:
drop = [c for c in features.DAM_FEATURES if c in X_tr.columns]
X_tr = X_tr.drop(columns=drop)
X_te = X_te.drop(columns=drop)
elif "dam_storage_pct" not in X_tr.columns:
raise ValueError(
f"{self.name} requires the dam series (rid_reservoir_daily backfilled)"
)
if not self.use_fc48:
drop = [c for c in EXTRA_RAIN_FEATURES if c in X_tr.columns]
X_tr = X_tr.drop(columns=drop)
X_te = X_te.drop(columns=drop)
elif "rain_fc48" not in X_tr.columns:
raise ValueError(f"{self.name} requires the rain series")
level_tr = X_tr["level"] level_tr = X_tr["level"]
level_te = X_te["level"].to_numpy() level_te = X_te["level"].to_numpy()
y_tr = (y_abs_tr - level_tr) if self.target == "rise" else y_abs_tr y_tr = (y_abs_tr - level_tr) if self.target == "rise" else y_abs_tr
@@ -89,7 +123,13 @@ class Variant:
else: else:
reg = _make_regressor().fit(X_tr, y_tr, sample_weight=weights) reg = _make_regressor().fit(X_tr, y_tr, sample_weight=weights)
pred = reg.predict(X_te) pred = reg.predict(X_te)
sigma = np.full(len(X_te), FIXED_SIGMA) if self.qsigma:
q50 = _quantile_regressor(0.5).fit(X_tr, y_tr, sample_weight=weights)
q90 = _quantile_regressor(0.9).fit(X_tr, y_tr, sample_weight=weights)
spread = np.maximum(q90.predict(X_te) - q50.predict(X_te), 0.0)
sigma = np.maximum(spread / 1.2816, 0.05)
else:
sigma = np.full(len(X_te), FIXED_SIGMA)
pred_abs = pred + level_te if self.target == "rise" else pred pred_abs = pred + level_te if self.target == "rise" else pred
pred_abs = np.maximum(pred_abs, level_te) # peak >= current, as served pred_abs = np.maximum(pred_abs, level_te) # peak >= current, as served
@@ -100,11 +140,45 @@ VARIANTS: Dict[str, Variant] = {
"baseline_abs": Variant("baseline_abs", target="abs"), "baseline_abs": Variant("baseline_abs", target="abs"),
"rise": Variant("rise", target="rise"), "rise": Variant("rise", target="rise"),
"rise_weighted": Variant("rise_weighted", target="rise", weighted=True), "rise_weighted": Variant("rise_weighted", target="rise", weighted=True),
"rise_quantile": Variant("rise_quantile", target="rise", weighted=True, "rise_quantile": Variant(
quantile=True), "rise_quantile", target="rise", weighted=True, quantile=True
),
"rise_rain": Variant("rise_rain", target="rise", use_rain=True), "rise_rain": Variant("rise_rain", target="rise", use_rain=True),
"rise_rain_dam": Variant(
"rise_rain_dam", target="rise", use_rain=True, use_dam=True
),
"rise_dam": Variant("rise_dam", target="rise", use_dam=True),
# 2026-09-12 experiments on top of the deployed rise_rain configuration:
# per-row sigma from quantile heads (the served sigma sits on the 0.15
# floor at every P.1 horizon, so stage probabilities are constant-
# calibrated), and a longer forecast-rain window for the 24 h horizon.
"rise_rain_quantile": Variant(
"rise_rain_quantile", target="rise", weighted=True, quantile=True, use_rain=True
),
"rise_rain_quantile_uw": Variant(
"rise_rain_quantile_uw", target="rise", quantile=True, use_rain=True
),
"rise_rain_fc48": Variant(
"rise_rain_fc48", target="rise", use_rain=True, use_fc48=True
),
"rise_rain_qsigma": Variant(
"rise_rain_qsigma", target="rise", use_rain=True, qsigma=True
),
} }
# Dam variants are opt-in by name: they require dam columns that only exist
# for features.DAM_STATIONS and only when the reservoir series loaded, and
# the 2026-08-13 ablation concluded them a negative result. The 2026-09-12
# experiments are opt-in too (see their results in docs/FLOOD_FORECASTING.md).
DEFAULT_VARIANTS = [
k
for k, v in VARIANTS.items()
if not v.use_dam
and not v.use_fc48
and not v.qsigma
and not (v.quantile and v.use_rain)
]
def _find_events(observed: pd.Series, thr: float) -> List[dict]: def _find_events(observed: pd.Series, thr: float) -> List[dict]:
"""Contiguous >=thr episodes (gaps under EVENT_GAP_H merged).""" """Contiguous >=thr episodes (gaps under EVENT_GAP_H merged)."""
@@ -149,12 +223,16 @@ def _first_alert_lead(
start = crossing - pd.Timedelta(hours=72) start = crossing - pd.Timedelta(hours=72)
if window_start_floor is not None and window_start_floor > start: if window_start_floor is not None and window_start_floor > start:
start = window_start_floor start = window_start_floor
window = p.loc[start: crossing + pd.Timedelta(hours=24)] window = p.loc[start : crossing + pd.Timedelta(hours=24)]
if len(window) < 2: if len(window) < 2:
return None return None
alert = (window >= ALERT_P) & (window.shift(-1) >= ALERT_P) & ( alert = (
(window.index.to_series().shift(-1) - window.index.to_series()) (window >= ALERT_P)
<= pd.Timedelta(hours=2) & (window.shift(-1) >= ALERT_P)
& (
(window.index.to_series().shift(-1) - window.index.to_series())
<= pd.Timedelta(hours=2)
)
) )
hits = window.index[alert.fillna(False)] hits = window.index[alert.fillna(False)]
if len(hits) == 0: if len(hits) == 0:
@@ -162,9 +240,7 @@ def _first_alert_lead(
return float((crossing - hits[0]).total_seconds() / 3600.0) return float((crossing - hits[0]).total_seconds() / 3600.0)
def _false_alarm_episodes( def _false_alarm_episodes(p: pd.Series, observed: pd.Series, thr: float) -> int:
p: pd.Series, observed: pd.Series, thr: float
) -> int:
"""Alert episodes with no observed >=thr within +/- FALSE_ALARM_GRACE_H.""" """Alert episodes with no observed >=thr within +/- FALSE_ALARM_GRACE_H."""
alert_hours = p[p >= ALERT_P].index alert_hours = p[p >= ALERT_P].index
if len(alert_hours) == 0: if len(alert_hours) == 0:
@@ -198,11 +274,18 @@ def evaluate_station(
variants: Optional[List[str]] = None, variants: Optional[List[str]] = None,
seasons: Tuple[int, ...] = SEASONS, seasons: Tuple[int, ...] = SEASONS,
rain: Optional[pd.Series] = None, rain: Optional[pd.Series] = None,
dam: Optional[pd.DataFrame] = None,
) -> Dict: ) -> Dict:
"""Run every fold x variant for one station; returns the results tree.""" """Run every fold x variant for one station; returns the results tree."""
warn_thr, _ = features.get_thresholds(station) warn_thr, _ = features.get_thresholds(station)
grid = features.make_hourly_grid(df_long) grid = features.make_hourly_grid(df_long)
X_all = features.build_features(grid, station, rain=rain) X_all = features.build_features(grid, station, rain=rain, dam=dam)
if rain is not None:
# forward sum over (t, t+48]; same construction as rain_fc24
r = rain.reindex(X_all.index)
X_all["rain_fc48"] = (
r.shift(-1).iloc[::-1].rolling(48, min_periods=1).sum().iloc[::-1]
)
observed = grid.observed[(station, "water_level")] observed = grid.observed[(station, "water_level")]
keep = X_all["obs_age_h"].notna() keep = X_all["obs_age_h"].notna()
@@ -211,7 +294,7 @@ def evaluate_station(
keep &= X_all.index >= pd.Timestamp(train_start) keep &= X_all.index >= pd.Timestamp(train_start)
X_all = X_all.loc[keep] X_all = X_all.loc[keep]
chosen = {k: VARIANTS[k] for k in (variants or VARIANTS)} chosen = {k: VARIANTS[k] for k in (variants or DEFAULT_VARIANTS)}
results: Dict = {"station": station, "warn_thr": warn_thr, "folds": []} results: Dict = {"station": station, "warn_thr": warn_thr, "folds": []}
for year in seasons: for year in seasons:
@@ -228,7 +311,9 @@ def evaluate_station(
tr = (X_all.index <= train_end) & y_abs.notna() tr = (X_all.index <= train_end) & y_abs.notna()
te = (X_all.index >= test_lo) & (X_all.index <= test_hi) te = (X_all.index >= test_lo) & (X_all.index <= test_hi)
if tr.sum() < 5000 or te.sum() < 500: if tr.sum() < 5000 or te.sum() < 500:
logger.info(f"{station} {year}: skipped (train {tr.sum()}, test {te.sum()})") logger.info(
f"{station} {year}: skipped (train {tr.sum()}, test {te.sum()})"
)
continue continue
X_tr, X_te = X_all.loc[tr], X_all.loc[te] X_tr, X_te = X_all.loc[tr], X_all.loc[te]
@@ -253,7 +338,14 @@ def evaluate_station(
} }
for name, variant in chosen.items(): for name, variant in chosen.items():
pred_abs, sigma = variant.fit_predict(X_tr, y_tr, X_te) try:
pred_abs, sigma = variant.fit_predict(X_tr, y_tr, X_te)
except ValueError as error:
# A variant whose required feature family is absent (e.g. a
# dam variant on a non-DAM_STATIONS target) skips this fold
# instead of killing the whole run and its finished results.
logger.warning(f"{station} {year} {name}: skipped ({error})")
continue
pred_series = pd.Series(pred_abs, index=X_te.index) pred_series = pd.Series(pred_abs, index=X_te.index)
p_warn = pd.Series( p_warn = pd.Series(
1.0 - _phi((warn_thr - pred_abs) / sigma), index=X_te.index 1.0 - _phi((warn_thr - pred_abs) / sigma), index=X_te.index
@@ -297,9 +389,7 @@ def evaluate_station(
) )
fold["variants"][name] = { fold["variants"][name] = {
"mae": float(errors.mean()) if len(errors) else None, "mae": float(errors.mean()) if len(errors) else None,
"mae_above_2p5": ( "mae_above_2p5": (float(errors[high].mean()) if high.any() else None),
float(errors[high].mean()) if high.any() else None
),
"brier_warn": brier, "brier_warn": brier,
"events": event_rows, "events": event_rows,
"false_alarm_episodes": _false_alarm_episodes( "false_alarm_episodes": _false_alarm_episodes(
@@ -320,17 +410,20 @@ def summarize(results: Dict) -> str:
lines.append(header) lines.append(header)
for fold in results["folds"]: for fold in results["folds"]:
for name, m in fold["variants"].items(): for name, m in fold["variants"].items():
events = " ".join( events = (
f"[{e['crossing'][:10]}: " " ".join(
f"{'' if e['lead_h'] is None else format(e['lead_h'], '+.0f')}h" f"[{e['crossing'][:10]}: "
+ ( f"{'' if e['lead_h'] is None else format(e['lead_h'], '+.0f')}h"
f" | {e['peak_pred_24h_before'] - e['peak_level']:+.2f}" + (
if e["peak_pred_24h_before"] is not None f" | {e['peak_pred_24h_before'] - e['peak_level']:+.2f}"
else "" if e["peak_pred_24h_before"] is not None
else ""
)
+ "]"
for e in m["events"]
) )
+ "]" or "no events"
for e in m["events"] )
) or "no events"
lines.append( lines.append(
f"{name:16} {fold['year']:>5} " f"{name:16} {fold['year']:>5} "
f"{m['mae'] if m['mae'] is not None else float('nan'):6.3f} " f"{m['mae'] if m['mae'] is not None else float('nan'):6.3f} "
@@ -347,17 +440,32 @@ def main(argv=None) -> int:
parser = argparse.ArgumentParser(description=__doc__) parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--stations", default="P.1") parser.add_argument("--stations", default="P.1")
parser.add_argument("--db-url", default=None) parser.add_argument("--db-url", default=None)
parser.add_argument("--variants", default=None, parser.add_argument("--variants", default=None, help="comma list; default all")
help="comma list; default all")
parser.add_argument("--out", default="models/eval_variants.json") parser.add_argument("--out", default="models/eval_variants.json")
parser.add_argument("--no-rain", action="store_true", parser.add_argument(
help="skip loading the Open-Meteo rain series") "--no-rain", action="store_true", help="skip loading the Open-Meteo rain series"
)
parser.add_argument(
"--no-dam",
action="store_true",
help="skip loading the Mae Ngat reservoir series",
)
parser.add_argument(
"--from-cache",
action="store_true",
help="offline: read models/cache/ only (no DB, no API, "
"no Open-Meteo refresh) -- reproducible reruns",
)
args = parser.parse_args(argv) args = parser.parse_args(argv)
logging.basicConfig( logging.basicConfig(
level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s" level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s"
) )
df = data.load_measurements(db_url=args.db_url) if args.from_cache:
df = data._read_cache(data.CACHE_DIR, None)
logger.info(f"measurements from cache: {len(df)} rows")
else:
df = data.load_measurements(db_url=args.db_url)
if df.empty: if df.empty:
logger.error("no measurement data") logger.error("no measurement data")
return 1 return 1
@@ -366,7 +474,9 @@ def main(argv=None) -> int:
if not args.no_rain: if not args.no_rain:
from . import rain as rain_mod from . import rain as rain_mod
rain_series = rain_mod.catchment_mean(rain_mod.load_history()) rain_series = rain_mod.catchment_mean(
rain_mod.load_history(refresh=not args.from_cache)
)
if rain_series is None: if rain_series is None:
logger.warning("rain history unavailable; rain features will be NaN") logger.warning("rain history unavailable; rain features will be NaN")
else: else:
@@ -375,12 +485,27 @@ def main(argv=None) -> int:
f"{rain_series.index.max()}" f"{rain_series.index.max()}"
) )
dam_frame = None
if not args.no_dam and not args.from_cache:
from . import dam as dam_mod
dam_frame = dam_mod.load_history(db_url=args.db_url)
if dam_frame is None:
logger.warning("dam history unavailable; dam features will be absent")
else:
logger.info(
f"dam series loaded: {dam_frame.index.min()} .. "
f"{dam_frame.index.max()}"
)
variant_names = args.variants.split(",") if args.variants else None variant_names = args.variants.split(",") if args.variants else None
all_results = [] all_results = []
for station in args.stations.split(","): for station in args.stations.split(","):
station = station.strip() station = station.strip()
logger.info(f"Evaluating {station}...") logger.info(f"Evaluating {station}...")
results = evaluate_station(df, station, variant_names, rain=rain_series) results = evaluate_station(
df, station, variant_names, rain=rain_series, dam=dam_frame
)
all_results.append(results) all_results.append(results)
print(summarize(results)) print(summarize(results))
+39 -4
View File
@@ -37,9 +37,17 @@ THRESHOLDS: Dict[str, Tuple[float, float]] = {
"P.4A": (3.40, 3.90), "P.4A": (3.40, 3.90),
"P.5": (4.55, 4.95), "P.5": (4.55, 4.95),
"P.67": (2.45, 2.90), "P.67": (2.45, 2.90),
"P.75": (2.75, 3.50), # P.75: 2024 (the only year with a full flood record, 191% capacity peak)
# puts 75-85% at 3.45 m and 95-105% at 3.72 m; 2018/2022 agree within
# 0.15 m. The 2026-08 value (2.75) alerted on 15 quiet-season hours.
"P.75": (3.20, 3.65),
"P.76": (5.35, 5.45), "P.76": (5.35, 5.45),
"P.77": (2.85, 3.35), # P.77: recalibrated 2026-09-12. The 2026-08 value (2.85) sat below the
# gauge's own dry-season baseline (2.6-2.7 m at 8-14% capacity), so the
# first ntfy cycle fired a "warning" at 22% capacity. Across 2018-2024,
# 75-85% capacity reads 3.35-4.57 m and 95-105% 4.27-5.08 m; 2024 (the
# best-sampled flood year) gives 4.57 / 5.08. Slightly conservative:
"P.77": (4.30, 4.90),
"P.81": (5.15, 6.30), "P.81": (5.15, 6.30),
# P.82 never reached 100% capacity in the record (max level 3.78, max 96.4%); # P.82 never reached 100% capacity in the record (max level 3.78, max 96.4%);
# danger sits just below the observed maximum so the head can actually train. # danger sits just below the observed maximum so the head can actually train.
@@ -202,9 +210,18 @@ def _hours_since_observed(mask_col: pd.Series) -> pd.Series:
RAIN_FEATURES = ("rain_6h", "rain_24h", "rain_72h", "rain_fc24") RAIN_FEATURES = ("rain_6h", "rain_24h", "rain_72h", "rain_fc24")
DAM_FEATURES = ("dam_storage_pct", "dam_storage_pct_d3", "dam_inflow", "dam_outflow")
# Stations hydrologically downstream of the Mae Ngat confluence (Ping mainstem
# at/below Mae Taeng) — the only ones where reservoir state is causal. West-
# tributary and upper-mainstem stations never receive dam columns.
DAM_STATIONS = frozenset({"P.1", "P.103", "P.67", "P.21", "P.5", "P.81"})
def build_features( def build_features(
grid: HourlyGrid, station: str, rain: Optional[pd.Series] = None grid: HourlyGrid,
station: str,
rain: Optional[pd.Series] = None,
dam: Optional[pd.DataFrame] = None,
) -> pd.DataFrame: ) -> pd.DataFrame:
"""Build the deterministic-order feature matrix for one target station. """Build the deterministic-order feature matrix for one target station.
@@ -217,6 +234,11 @@ def build_features(
bundles even when the live fetch fails. rain_fc24 is the forward 24 h bundles even when the live fetch fails. rain_fc24 is the forward 24 h
sum: the archived forecast series at training time, a real weather sum: the archived forecast series at training time, a real weather
forecast at serving time; it never contains river data. forecast at serving time; it never contains river data.
``dam`` is the hourly Mae Ngat reservoir frame (src/ml/dam.py; columns
storage_pct/inflow_mcm/outflow_mcm, already leakage-shifted to 07:00
report time). Same contract as rain: None omits the columns, an empty
frame yields NaN columns; only DAM_STATIONS receive them.
""" """
idx = grid.observed.index idx = grid.observed.index
cols: Dict[str, pd.Series] = {} cols: Dict[str, pd.Series] = {}
@@ -280,6 +302,18 @@ def build_features(
r.shift(-1).iloc[::-1].rolling(24, min_periods=1).sum().iloc[::-1] r.shift(-1).iloc[::-1].rolling(24, min_periods=1).sum().iloc[::-1]
) )
if dam is not None and station in DAM_STATIONS:
d = dam.reindex(idx)
def _dam_col(name: str) -> pd.Series:
return d[name] if name in d.columns else pd.Series(np.nan, index=idx)
storage = _dam_col("storage_pct")
cols["dam_storage_pct"] = storage
cols["dam_storage_pct_d3"] = storage - storage.shift(72)
cols["dam_inflow"] = _dam_col("inflow_mcm")
cols["dam_outflow"] = _dam_col("outflow_mcm")
return pd.DataFrame(cols, index=idx) return pd.DataFrame(cols, index=idx)
@@ -358,10 +392,11 @@ def build_matrix(
horizons: Tuple[int, ...] = (6, 12, 24), horizons: Tuple[int, ...] = (6, 12, 24),
stats_end: Optional[str] = None, stats_end: Optional[str] = None,
rain: Optional[pd.Series] = None, rain: Optional[pd.Series] = None,
dam: Optional[pd.DataFrame] = None,
) -> Tuple[pd.DataFrame, pd.DataFrame, dict]: ) -> Tuple[pd.DataFrame, pd.DataFrame, dict]:
"""Build (X, Y, meta) training/inference matrices for one station.""" """Build (X, Y, meta) training/inference matrices for one station."""
grid = make_hourly_grid(df_long) grid = make_hourly_grid(df_long)
X = build_features(grid, station, rain=rain) X = build_features(grid, station, rain=rain, dam=dam)
Y = build_labels(grid, station, horizons, stats_end=stats_end) Y = build_labels(grid, station, horizons, stats_end=stats_end)
keep = X["obs_age_h"].notna() keep = X["obs_age_h"].notna()
+122
View File
@@ -0,0 +1,122 @@
"""Catchment-mean hourly rain from the HII/ThaiWater gauge network.
Independent of Open-Meteo (src/ml/rain.py): those are model-analysis values,
these are what the gauges measured. The `hii_rainfall` table has been filled
by the hourly collector since 2026-08-11 and there is NO archive behind it
(the api-v3 rain_24h_graph endpoint ignores its date range, see
docs/DATA_SOURCES.md 2.1), so this series cannot yet be a training feature:
every training row before 2026-08 would be NaN and HistGradientBoosting
would learn nothing from the column. It becomes a candidate once a full
monsoon season of gauge rows exists in the rolling-origin harness's test
span -- the 2027 fold (train through 2027-04-30, test Jun-Nov 2027) is the
first that could show anything.
Until then it serves two purposes:
* a live cross-check of the Open-Meteo catchment mean (/api/hii/rainfall
already exposes the raw gauges; this gives the comparable aggregate);
* accumulating the comparison so the eventual feature evaluation has a
documented bias/variance relationship between the two sources.
"""
import logging
from typing import Optional, Sequence, Tuple
import pandas as pd
from .data import resolve_db_url
logger = logging.getLogger(__name__)
# Same footprint as rain.CATCHMENT_POINTS: the upper Ping above P.1. Gauges
# inside this box are averaged; there are ~130 with recent data (DWR, FOP,
# HII, RID, TMD), far denser than the five Open-Meteo points.
CATCHMENT_BOX: Tuple[float, float, float, float] = (18.75, 19.60, 98.60, 99.30)
# A gauge that reports the same rain_24h for many hours is stuck; drop hours
# where fewer than this many gauges reported at all.
MIN_GAUGES_PER_HOUR = 5
def load_gauge_mean(
db_url: Optional[str] = None,
start: Optional[pd.Timestamp] = None,
end: Optional[pd.Timestamp] = None,
box: Sequence[float] = CATCHMENT_BOX,
engine=None,
) -> Optional[pd.Series]:
"""Hourly catchment-mean rain_1h (mm) across HII gauges in `box`.
Pass `engine` (the API's HII store engine) to reuse a pool; otherwise a
connection is resolved from db_url / config. Returns None if the DB is
unavailable or the table is empty. Hours with fewer than
MIN_GAUGES_PER_HOUR reporting gauges are NaN.
"""
if engine is None:
resolved = resolve_db_url(db_url)
if not resolved:
return None
lat_lo, lat_hi, lon_lo, lon_hi = box
try:
from sqlalchemy import create_engine, text
query = (
"SELECT m.timestamp, COUNT(m.rain_1h) AS n, AVG(m.rain_1h) AS rain_1h "
"FROM hii_rainfall m JOIN hii_rain_stations s ON s.id = m.station_id "
"WHERE s.latitude BETWEEN :lat_lo AND :lat_hi "
"AND s.longitude BETWEEN :lon_lo AND :lon_hi "
"AND m.rain_1h IS NOT NULL"
)
params = {
"lat_lo": lat_lo,
"lat_hi": lat_hi,
"lon_lo": lon_lo,
"lon_hi": lon_hi,
}
if start is not None:
query += " AND m.timestamp >= :start"
params["start"] = pd.Timestamp(start).to_pydatetime()
if end is not None:
query += " AND m.timestamp <= :end"
params["end"] = pd.Timestamp(end).to_pydatetime()
query += " GROUP BY m.timestamp ORDER BY m.timestamp"
if engine is None:
engine = create_engine(resolved, pool_pre_ping=True)
with engine.connect() as conn:
frame = pd.read_sql(text(query), conn, params=params)
except Exception as error:
logger.warning(f"HII gauge rain load failed: {error}")
return None
if frame.empty:
return None
frame["timestamp"] = pd.to_datetime(frame["timestamp"]).dt.floor("h")
frame = frame.groupby("timestamp").agg(n=("n", "sum"), rain_1h=("rain_1h", "mean"))
series = pd.to_numeric(frame["rain_1h"], errors="coerce")
series[frame["n"] < MIN_GAUGES_PER_HOUR] = float("nan")
series.name = "hii_gauge_mean"
return series
def compare_with_openmeteo(
gauge: pd.Series, openmeteo: pd.Series, window_h: int = 24
) -> dict:
"""Bias/correlation of Open-Meteo against the gauges over the overlap.
Both are summed over trailing `window_h` so single-hour timing offsets
(gauges report at :00, the model's hour is an interval) do not dominate.
"""
joined = pd.concat({"gauge": gauge, "openmeteo": openmeteo}, axis=1).dropna()
if joined.empty:
return {"overlap_hours": 0}
g = joined["gauge"].rolling(window_h, min_periods=window_h).sum()
o = joined["openmeteo"].rolling(window_h, min_periods=window_h).sum()
both = pd.concat({"g": g, "o": o}, axis=1).dropna()
if both.empty:
return {"overlap_hours": int(len(joined))}
return {
"overlap_hours": int(len(joined)),
"window_h": window_h,
"gauge_mean_mm": float(both["g"].mean()),
"openmeteo_mean_mm": float(both["o"].mean()),
"bias_mm": float((both["o"] - both["g"]).mean()),
"mae_mm": float((both["o"] - both["g"]).abs().mean()),
"corr": float(both["g"].corr(both["o"])),
}
+34 -6
View File
@@ -128,6 +128,7 @@ def _model_forecast(
as_of: pd.Timestamp, as_of: pd.Timestamp,
current_level: float, current_level: float,
rain: Optional[pd.Series] = None, rain: Optional[pd.Series] = None,
dam: Optional[pd.DataFrame] = None,
) -> List[dict]: ) -> List[dict]:
warn_thr = bundle["thresholds"]["warning"] warn_thr = bundle["thresholds"]["warning"]
danger_thr = bundle["thresholds"]["danger"] danger_thr = bundle["thresholds"]["danger"]
@@ -146,7 +147,9 @@ def _model_forecast(
) )
warn_thr, danger_thr = cfg_warn, cfg_danger warn_thr, danger_thr = cfg_warn, cfg_danger
feature_row = features.build_features(grid, station_code, rain=rain).loc[[as_of]] feature_row = features.build_features(grid, station_code, rain=rain, dam=dam).loc[
[as_of]
]
expected_columns = bundle["feature_names"] expected_columns = bundle["feature_names"]
missing = [c for c in expected_columns if c not in feature_row.columns] missing = [c for c in expected_columns if c not in feature_row.columns]
if missing: if missing:
@@ -179,14 +182,18 @@ def _model_forecast(
) )
p_warning = _sigmoid_probability(predicted_max, warn_thr, sigma_h) p_warning = _sigmoid_probability(predicted_max, warn_thr, sigma_h)
if warn_head is not None: if warn_head is not None:
p_warning = max(p_warning, float(warn_head.predict_proba(feature_row)[0][1])) p_warning = max(
p_warning, float(warn_head.predict_proba(feature_row)[0][1])
)
danger_head = ( danger_head = (
None if thresholds_stale else bundle["heads"].get(f"danger_{horizon_h}") None if thresholds_stale else bundle["heads"].get(f"danger_{horizon_h}")
) )
p_danger = _sigmoid_probability(predicted_max, danger_thr, sigma_h) p_danger = _sigmoid_probability(predicted_max, danger_thr, sigma_h)
if danger_head is not None: if danger_head is not None:
p_danger = max(p_danger, float(danger_head.predict_proba(feature_row)[0][1])) p_danger = max(
p_danger, float(danger_head.predict_proba(feature_row)[0][1])
)
p_warning = _clip_probability(p_warning) p_warning = _clip_probability(p_warning)
p_danger = min(_clip_probability(p_danger), p_warning) p_danger = min(_clip_probability(p_danger), p_warning)
@@ -231,6 +238,7 @@ def _forecast_station(
now: pd.Timestamp, now: pd.Timestamp,
horizons: Tuple[int, ...], horizons: Tuple[int, ...],
rain: Optional[pd.Series] = None, rain: Optional[pd.Series] = None,
dam: Optional[pd.DataFrame] = None,
) -> List[dict]: ) -> List[dict]:
level_col = (station_code, "water_level") level_col = (station_code, "water_level")
if level_col not in grid.observed.columns: if level_col not in grid.observed.columns:
@@ -267,7 +275,7 @@ def _forecast_station(
bundle = _load_bundle(bundle_path) bundle = _load_bundle(bundle_path)
model_results = _model_forecast( model_results = _model_forecast(
station_code, grid, bundle, as_of, current_level, rain=rain station_code, grid, bundle, as_of, current_level, rain=rain, dam=dam
) )
if model_results is None: if model_results is None:
return _heuristic_forecast( return _heuristic_forecast(
@@ -306,6 +314,7 @@ def get_forecasts(
models_dir: Union[str, Path] = DEFAULT_MODELS_DIR, models_dir: Union[str, Path] = DEFAULT_MODELS_DIR,
now: Optional[Union[datetime.datetime, str]] = None, now: Optional[Union[datetime.datetime, str]] = None,
rain: Optional[pd.Series] = None, rain: Optional[pd.Series] = None,
dam: Optional[pd.DataFrame] = None,
) -> List[dict]: ) -> List[dict]:
"""Produce flood forecasts for every station present in `readings_by_station`. """Produce flood forecasts for every station present in `readings_by_station`.
@@ -329,7 +338,13 @@ def get_forecasts(
try: try:
results.extend( results.extend(
_forecast_station( _forecast_station(
station_code, grid, models_dir, now, DEFAULT_HORIZONS, rain=rain station_code,
grid,
models_dir,
now,
DEFAULT_HORIZONS,
rain=rain,
dam=dam,
) )
) )
except Exception as error: except Exception as error:
@@ -376,4 +391,17 @@ def get_latest_forecasts(
logger.warning("live rain unavailable; rain features will be NaN") logger.warning("live rain unavailable; rain features will be NaN")
rain = pd.Series(dtype=float) rain = pd.Series(dtype=float)
return get_forecasts(readings_by_station, models_dir=models_dir, rain=rain) # Recent Mae Ngat reservoir state; same empty-not-None contract so
# dam-trained bundles keep their columns (NaN) when the DB read fails.
from . import dam as dam_mod
try:
dam = dam_mod.serving_frame(db_url=db_url)
except Exception as error:
logger.warning(f"dam serving frame failed: {error}")
dam = None
if dam is None:
logger.warning("dam state unavailable; dam features will be NaN")
dam = pd.DataFrame()
return get_forecasts(readings_by_station, models_dir=models_dir, rain=rain, dam=dam)
+4 -9
View File
@@ -121,12 +121,8 @@ def load_history(
cursor = fetch_from.date() cursor = fetch_from.date()
try: try:
while cursor <= end: while cursor <= end:
chunk_end = min( chunk_end = min(datetime.date(cursor.year, 12, 31), end)
datetime.date(cursor.year, 12, 31), end chunks.append(fetch_history(cursor.isoformat(), chunk_end.isoformat()))
)
chunks.append(
fetch_history(cursor.isoformat(), chunk_end.isoformat())
)
cursor = datetime.date(cursor.year + 1, 1, 1) cursor = datetime.date(cursor.year + 1, 1, 1)
except Exception as error: except Exception as error:
logger.warning(f"Open-Meteo history fetch failed: {error}") logger.warning(f"Open-Meteo history fetch failed: {error}")
@@ -174,7 +170,7 @@ def backfill_db(engine, db_type: str, chunk_rows: int = 5000) -> int:
return 0 return 0
total = 0 total = 0
for start in range(0, len(history), chunk_rows): for start in range(0, len(history), chunk_rows):
part = history.iloc[start: start + chunk_rows] part = history.iloc[start : start + chunk_rows]
total += save_to_db(part, engine, db_type) total += save_to_db(part, engine, db_type)
logger.info(f"openmeteo_rain backfill: {total}/{len(history)} rows") logger.info(f"openmeteo_rain backfill: {total}/{len(history)} rows")
return total return total
@@ -204,8 +200,7 @@ def save_to_db(df: pd.DataFrame, engine, db_type: str) -> int:
cols = ["timestamp"] + point_cols + ["catchment_mean"] cols = ["timestamp"] + point_cols + ["catchment_mean"]
placeholders = ", ".join(f":{c}" for c in cols) placeholders = ", ".join(f":{c}" for c in cols)
updates = ", ".join( updates = ", ".join(
f"{c} = " f"{c} = " + (f"VALUES({c})" if db_type == "mysql" else f"EXCLUDED.{c}")
+ (f"VALUES({c})" if db_type == "mysql" else f"EXCLUDED.{c}")
for c in cols[1:] for c in cols[1:]
) )
if db_type == "mysql": if db_type == "mysql":
+170
View File
@@ -0,0 +1,170 @@
"""Live forecast skill: what the deployed model said versus what the river did.
Every hour the precompute stores the issued 24 h forecast (forecast_history);
water_measurements holds what actually happened. Joining the two gives a
verification that needs no retraining and answers the question the dashboard
is asked most: "is the model getting better?" per model version, on the
hours that version was actually serving.
Metrics per version and horizon:
n verified forecasts (issued, and the horizon has since elapsed)
mae |predicted_max - observed_max| over the horizon window, metres
bias mean(predicted - observed): >0 over-predicts the peak
persistence MAE of the trivial "peak = current level" forecast on the
same rows; a model is only useful if it beats this
skill 1 - mae/persistence (0 = no better than persistence, 1 = perfect)
above_2m same MAE restricted to rows where the observed peak >= 2 m,
i.e. the flood-relevant regime
Only the P.1 gauge is verified by default: it is the one the city threshold
is keyed to, and one station keeps the query cheap enough to run on request.
"""
import datetime
import logging
from typing import Dict, List, Optional
logger = logging.getLogger(__name__)
DEFAULT_STATION = "P.1"
DEFAULT_HORIZON = 24
MIN_VERIFIED = 24 # fewer than a day of verified hours is not a number
def _sql_for(db_type: str) -> str:
"""Join each issued forecast to the observed max over (as_of, as_of + h]."""
if db_type == "postgresql":
window_end = "f.as_of + (f.horizon_hours || ' hours')::interval"
elif db_type == "mysql":
window_end = "DATE_ADD(f.as_of, INTERVAL f.horizon_hours HOUR)"
else: # sqlite
window_end = "datetime(f.as_of, '+' || f.horizon_hours || ' hours')"
return f"""
SELECT f.as_of, f.model_version, f.predicted_max_level, f.current_level,
(SELECT MAX(m.water_level) FROM water_measurements m
JOIN stations s ON s.id = m.station_id
WHERE s.station_code = f.station_code
AND m.timestamp > f.as_of AND m.timestamp <= {window_end}) AS observed_max,
(SELECT COUNT(m.water_level) FROM water_measurements m
JOIN stations s ON s.id = m.station_id
WHERE s.station_code = f.station_code
AND m.timestamp > f.as_of AND m.timestamp <= {window_end}) AS observed_n
FROM forecast_history f
WHERE f.station_code = :code AND f.horizon_hours = :horizon
AND f.source = 'model' AND f.predicted_max_level IS NOT NULL
AND f.as_of <= :verifiable_before
ORDER BY f.as_of
"""
def compute_skill(
engine,
db_type: str,
station_code: str = DEFAULT_STATION,
horizon_hours: int = DEFAULT_HORIZON,
now: Optional[datetime.datetime] = None,
) -> Dict:
"""Per-model-version verification of issued forecasts against observations."""
from sqlalchemy import text
now = now or datetime.datetime.now()
verifiable_before = now - datetime.timedelta(hours=horizon_hours)
with engine.connect() as conn:
rows = [
dict(r._mapping)
for r in conn.execute(
text(_sql_for(db_type)),
{
"code": station_code,
"horizon": horizon_hours,
"verifiable_before": verifiable_before,
},
)
]
def _ts(value):
# sqlite hands back strings; postgres/mysql give datetimes
if isinstance(value, datetime.datetime):
return value
return datetime.datetime.fromisoformat(str(value).replace(" ", "T"))
by_version: Dict[str, List[dict]] = {}
for r in rows:
r["as_of"] = _ts(r["as_of"])
# need most of the window observed, or the "max" is not the peak
if r["observed_max"] is None or (r["observed_n"] or 0) < horizon_hours * 0.75:
continue
by_version.setdefault(r["model_version"] or "unknown", []).append(r)
versions = []
for version, vrows in by_version.items():
pred = [float(r["predicted_max_level"]) for r in vrows]
obs = [float(r["observed_max"]) for r in vrows]
cur = [
float(r["current_level"]) if r["current_level"] is not None else None
for r in vrows
]
err = [p - o for p, o in zip(pred, obs)]
mae = sum(abs(e) for e in err) / len(err)
bias = sum(err) / len(err)
pers_rows = [(c, o) for c, o in zip(cur, obs) if c is not None]
persistence = (
sum(abs(c - o) for c, o in pers_rows) / len(pers_rows)
if pers_rows
else None
)
high = [(p, o) for p, o in zip(pred, obs) if o >= 2.0]
versions.append(
{
"model_version": version,
"first_issued": min(r["as_of"] for r in vrows).isoformat(),
"last_issued": max(r["as_of"] for r in vrows).isoformat(),
"n": len(vrows),
"mae_m": round(mae, 3),
"bias_m": round(bias, 3),
"persistence_mae_m": (
None if persistence is None else round(persistence, 3)
),
"skill": (
None if not persistence else round(1.0 - mae / persistence, 3)
),
"above_2m_n": len(high),
"above_2m_mae_m": (
round(sum(abs(p - o) for p, o in high) / len(high), 3)
if high
else None
),
"enough_data": len(vrows) >= MIN_VERIFIED,
}
)
versions.sort(key=lambda v: v["first_issued"])
# Headline: current version vs the previous one that had enough data
current = versions[-1] if versions else None
previous = (
next((v for v in reversed(versions[:-1]) if v["enough_data"]), None)
if versions
else None
)
trend = None
if current and previous and current["enough_data"]:
trend = {
"previous_version": previous["model_version"],
"mae_delta_m": round(current["mae_m"] - previous["mae_m"], 3),
"skill_delta": (
None
if current["skill"] is None or previous["skill"] is None
else round(current["skill"] - previous["skill"], 3)
),
"better": current["mae_m"] < previous["mae_m"],
}
return {
"station_code": station_code,
"horizon_hours": horizon_hours,
"verified_until": verifiable_before.isoformat(),
"min_verified": MIN_VERIFIED,
"versions": versions,
"current": current,
"trend": trend,
}
+106 -16
View File
@@ -45,6 +45,15 @@ MIN_SIGMA = 0.15
MIN_ROWS_TO_TRAIN = 200 MIN_ROWS_TO_TRAIN = 200
MIN_ROWS_FOR_HEAD = 50 MIN_ROWS_FOR_HEAD = 50
class RainUnavailableError(RuntimeError):
"""Raised when a rain-enabled training run cannot obtain the rain series.
Training would otherwise fall through to gauge-only (v2) bundles and
overwrite the deployed v3 artifacts without anyone noticing.
"""
HGB_PARAMS = { HGB_PARAMS = {
"max_iter": 300, "max_iter": 300,
"learning_rate": 0.06, "learning_rate": 0.06,
@@ -204,9 +213,10 @@ def train_station(
split_test_start: str = SPLIT_B_TEST_START, split_test_start: str = SPLIT_B_TEST_START,
split_test_end: str = SPLIT_B_TEST_END, split_test_end: str = SPLIT_B_TEST_END,
rain: Optional[pd.Series] = None, rain: Optional[pd.Series] = None,
dam: Optional[pd.DataFrame] = None,
) -> Tuple[Optional[dict], dict]: ) -> Tuple[Optional[dict], dict]:
"""Train every head for one station. Returns (bundle_or_None, station_metrics).""" """Train every head for one station. Returns (bundle_or_None, station_metrics)."""
X, Y, meta = features.build_matrix(df_long, station, horizons, rain=rain) X, Y, meta = features.build_matrix(df_long, station, horizons, rain=rain, dam=dam)
if meta["n_rows"] < MIN_ROWS_TO_TRAIN: if meta["n_rows"] < MIN_ROWS_TO_TRAIN:
return None, { return None, {
"status": "failed", "status": "failed",
@@ -310,9 +320,9 @@ def train_station(
skipped_heads, skipped_heads,
) )
else: else:
skipped_heads[ skipped_heads[head_key] = (
head_key f"only {n_pos} positives in train span (< {MIN_POSITIVES_FOR_CLASSIFIER})"
] = f"only {n_pos} positives in train span (< {MIN_POSITIVES_FOR_CLASSIFIER})" )
heads[head_key] = clf heads[head_key] = clf
if not skip_eval: if not skip_eval:
@@ -407,13 +417,18 @@ def train_station(
if clf is not None: if clf is not None:
skipped_heads.pop(head_key, None) skipped_heads.pop(head_key, None)
else: else:
skipped_heads[ skipped_heads[head_key] = (
head_key f"only {n_pos} positives in train span (< {MIN_POSITIVES_FOR_CLASSIFIER})"
] = f"only {n_pos} positives in train span (< {MIN_POSITIVES_FOR_CLASSIFIER})" )
final_heads[head_key] = None final_heads[head_key] = None
# v3 = rise target + Open-Meteo rain features; v2 = rise target only # v4 = + Mae Ngat dam features; v3 = rise + rain; v2 = rise target only
version_prefix = "hgb-v3" if "rain_24h" in feature_names else "hgb-v2" if "dam_storage_pct" in feature_names:
version_prefix = "hgb-v4"
elif "rain_24h" in feature_names:
version_prefix = "hgb-v3"
else:
version_prefix = "hgb-v2"
bundle = { bundle = {
"station_code": station, "station_code": station,
"model_version": f"{version_prefix}+{_git_short_sha()}", "model_version": f"{version_prefix}+{_git_short_sha()}",
@@ -443,13 +458,20 @@ def train_all(
skip_eval: bool = False, skip_eval: bool = False,
hgb_overrides: Optional[dict] = None, hgb_overrides: Optional[dict] = None,
use_rain: bool = True, use_rain: bool = True,
use_dam: bool = False,
db_url: Optional[str] = None,
) -> dict: ) -> dict:
"""Train and save every requested station's models. Returns the metrics.json payload.""" """Train and save every requested station's models. Returns the metrics.json payload."""
models_dir = Path(models_dir) models_dir = Path(models_dir)
models_dir.mkdir(parents=True, exist_ok=True) models_dir.mkdir(parents=True, exist_ok=True)
# Catchment rain (Open-Meteo archive, 2021+). Optional: without it the # Catchment rain (Open-Meteo archive, 2021+). A rain-less run produces v2
# models train as v2 (no rain columns) and still serve correctly. # bundles that serve fine but have measurably less flood lead (the 2024
# record flood: 13 h early with rain vs 18 h late without). The 2026-09-01
# server retrain hit exactly that -- the archive fetch failed on a checkout
# with no models/cache/ and the run quietly wrote v2 over v3. So the
# downgrade is now an error unless the caller opts out with use_rain=False
# (the --no-rain flag), which is the only way to get v2 deliberately.
rain_series = None rain_series = None
if use_rain: if use_rain:
try: try:
@@ -457,12 +479,57 @@ def train_all(
rain_series = rain_mod.catchment_mean(rain_mod.load_history()) rain_series = rain_mod.catchment_mean(rain_mod.load_history())
except Exception as error: except Exception as error:
logger.warning(f"rain history unavailable, training without it: {error}") raise RainUnavailableError(
f"rain history unavailable ({error}); refusing to silently "
"downgrade to v2 bundles -- fix Open-Meteo access or restore "
"models/cache/rain_openmeteo.csv.gz, or pass --no-rain to "
"train gauge-only bundles on purpose"
) from error
if rain_series is None:
raise RainUnavailableError(
"rain history unavailable (Open-Meteo archive unreachable and "
"no models/cache/rain_openmeteo.csv.gz); refusing to silently "
"downgrade to v2 bundles -- fix access, restore the cache file, "
"or pass --no-rain to train gauge-only bundles on purpose"
)
if rain_series is not None: if rain_series is not None:
logger.info( logger.info(
f"rain series: {rain_series.index.min()} .. {rain_series.index.max()}" f"rain series: {rain_series.index.min()} .. {rain_series.index.max()}"
) )
version_prefix = "hgb-v3" if rain_series is not None else "hgb-v2"
# Mae Ngat reservoir state (rid_reservoir_daily, 2018+). OFF by default:
# the 2026-08-13 backtest ablation showed every dam-feature subset COSTS
# 1-3 h of first-alert lead on the 2024 record flood (the daily report
# lags up to 31 h, so during fast onset the columns describe yesterday's
# benign reservoir and damp the alarm). Kept as an opt-in for post-monsoon
# re-evaluation once the 2026 season adds dam-era flood events.
dam_frame = None
if use_dam:
try:
from . import dam as dam_mod
dam_frame = dam_mod.load_history(db_url=db_url)
except Exception as error:
logger.warning(f"dam history unavailable, training without it: {error}")
if dam_frame is None:
# load_history returns None (no raise) when both DB and cache
# miss — an explicitly requested experiment must say so loudly.
logger.warning(
"--dam requested but no dam history available; "
"training v3-style bundles WITHOUT dam features"
)
if dam_frame is not None:
logger.info(f"dam series: {dam_frame.index.min()} .. {dam_frame.index.max()}")
# Run-level version: v4 only if some requested station actually receives
# dam columns (they are gated to DAM_STATIONS; per-bundle versions are
# derived from each station's own feature_names and remain authoritative).
if dam_frame is not None and any(s in features.DAM_STATIONS for s in stations):
version_prefix = "hgb-v4"
elif rain_series is not None:
version_prefix = "hgb-v3"
else:
version_prefix = "hgb-v2"
model_version = f"{version_prefix}+{_git_short_sha()}" model_version = f"{version_prefix}+{_git_short_sha()}"
station_results: Dict[str, dict] = {} station_results: Dict[str, dict] = {}
@@ -480,6 +547,7 @@ def train_all(
skip_eval=skip_eval, skip_eval=skip_eval,
hgb_overrides=hgb_overrides, hgb_overrides=hgb_overrides,
rain=rain_series, rain=rain_series,
dam=dam_frame,
) )
if bundle is None: if bundle is None:
logger.warning(f"{station}: failed ({station_metrics.get('reason')})") logger.warning(f"{station}: failed ({station_metrics.get('reason')})")
@@ -540,7 +608,15 @@ def main(argv: Optional[List[str]] = None) -> None:
parser.add_argument( parser.add_argument(
"--no-rain", "--no-rain",
action="store_true", action="store_true",
help="train without the Open-Meteo rain features (v2-style bundles)", help="DELIBERATELY train without the Open-Meteo rain features "
"(v2-style bundles). Without this flag a missing rain series aborts "
"the run instead of quietly downgrading the deployed model",
)
parser.add_argument(
"--dam",
action="store_true",
help="EXPERIMENTAL: include Mae Ngat reservoir features (v4 bundles); "
"the 2026-08 ablation showed they cost 1-3 h of alert lead",
) )
args = parser.parse_args(argv) args = parser.parse_args(argv)
@@ -570,14 +646,28 @@ def main(argv: Optional[List[str]] = None) -> None:
models_dir=Path(args.models_dir), models_dir=Path(args.models_dir),
skip_eval=args.skip_eval, skip_eval=args.skip_eval,
use_rain=not args.no_rain, use_rain=not args.no_rain,
use_dam=args.dam,
db_url=resolve_db_url(args.db_url),
) )
trained = sum( trained = sum(
1 for s in metrics_payload["stations"].values() if s["status"] == "trained" 1 for s in metrics_payload["stations"].values() if s["status"] == "trained"
) )
logger.info( logger.info(
f"Done: {trained}/{len(stations)} stations trained. metrics.json written to {args.models_dir}" f"Done: {trained}/{len(stations)} stations trained "
f"({metrics_payload['model_version']}). "
f"metrics.json written to {args.models_dir}"
) )
def cli() -> int:
"""Console entry: RainUnavailableError becomes a one-line error, exit 2."""
try:
main()
except RainUnavailableError as error:
logger.error(str(error))
return 2
return 0
if __name__ == "__main__": if __name__ == "__main__":
main() raise SystemExit(cli())
+454
View File
@@ -0,0 +1,454 @@
"""Public flood notifications over ntfy.
Runs once per collection cycle inside the API process (leader only), right
after the forecast precompute, so it sees the same readings and forecasts the
dashboard shows. Publishes to a self-hosted ntfy server; anyone subscribes to
a topic from the free app or a browser, no account needed.
Topics (all under one configurable prefix, default "ping"):
{prefix}-{station}-warning observed level crossed the station's warning threshold
{prefix}-{station}-danger observed level crossed the danger threshold
{prefix}-warning any station crossed warning (basin-wide digest)
{prefix}-danger any station crossed danger
{prefix}-p1-outlook model early warning for Chiang Mai city: P.1's 24 h
warning probability crossed the alert level (opt-in;
the forecast is experimental and says so)
{prefix}-status feed/monitor health: data stale, recovered
Each notification is a TRANSITION, not a state: crossing UP into a level sends
one message; dropping back below (with hysteresis) sends an all-clear. While
the river sits above a threshold nothing is repeated, so a subscriber in a
flood gets a handful of messages, not one an hour. The per-topic state is
persisted (notification_state table) so a restart never re-sends.
Everything is fail-safe: ntfy unreachable, table missing, malformed
reading -> a logged warning, never an exception into the collection loop.
"""
import datetime
import logging
from dataclasses import dataclass
from typing import Dict, Iterable, List, Optional
import requests
from .ml import features
logger = logging.getLogger(__name__)
# Hysteresis: an all-clear needs the level this far BELOW the threshold, so a
# river bobbing around 3.70 m does not toggle warning/clear every hour.
CLEAR_MARGIN_M = 0.10
# Capacity guard. The level thresholds in features.THRESHOLDS were calibrated
# from RID's discharge_percent (% of channel capacity); if RID re-rates a
# gauge or moves its datum, the level crosses while capacity says the channel
# is nearly empty (P.77, 2026-09: 3.0 m "warning" at 22 %). A crossing is
# only announced when the reported capacity agrees that the river is high.
# P.1 is exempt: its stages come from the municipal inundation map, not from
# capacity. Readings without a capacity figure fall back to level only.
CAPACITY_GUARD_MIN_PCT = 60.0
CAPACITY_GUARD_EXEMPT = {"P.1"}
# Outlook alert fires when p_warning(24h) rises through ON, clears below OFF.
OUTLOOK_ON = 0.50
OUTLOOK_OFF = 0.25
# Below this the outlook is not announced at all (avoid "5 % chance" noise).
OUTLOOK_HORIZON = 24
STATION_NAMES: Dict[str, str] = {
"P.1": "Nawarat Bridge, Chiang Mai city",
"P.103": "Ring Road Bridge 3, Chiang Mai",
"P.67": "Ban Tae (Mae Taeng)",
"P.21": "Ban Rim Tai (Mae Rim)",
"P.75": "Ban Chai Lat",
"P.92": "Ban Muang Aut",
"P.20": "Ban Chiang Dao",
"P.4A": "Ban Mae Taeng",
"P.5": "Tha Nang Bridge (downstream)",
"P.81": "Ban Pong (downstream)",
"P.82": "Ban Sob Win",
"P.84": "Ban Panton",
"P.87": "Ban Pa Sang",
"P.77": "Ban Sop Mae Sapuat",
"P.85": "Ban Lai Kaew",
"P.76": "Ban Mae I Hai",
}
def _slug(code: str) -> str:
return code.lower().replace(".", "")
@dataclass
class Notification:
topic: str
title: str
message: str
priority: int = 3 # ntfy: 1 min .. 5 max
tags: Optional[List[str]] = None
click: Optional[str] = None
class NtfyPublisher:
def __init__(
self,
server: str,
prefix: str = "ping",
token: Optional[str] = None,
dashboard_url: str = "https://water.buildfor.life/",
timeout: int = 10,
):
self.server = server.rstrip("/")
self.prefix = prefix
self.token = token
self.dashboard_url = dashboard_url
self.timeout = timeout
def topic(self, *parts: str) -> str:
return "-".join([self.prefix, *parts])
def publish(self, n: Notification) -> bool:
headers = {
"Title": n.title,
"Priority": str(n.priority),
"Click": n.click or self.dashboard_url,
"Actions": f"view, Open dashboard, {n.click or self.dashboard_url}",
}
if n.tags:
headers["Tags"] = ",".join(n.tags)
if self.token:
headers["Authorization"] = f"Bearer {self.token}"
try:
r = requests.post(
f"{self.server}/{n.topic}",
data=n.message.encode("utf-8"),
headers=headers,
timeout=self.timeout,
)
if r.status_code >= 300:
logger.warning(f"ntfy {n.topic}: HTTP {r.status_code} {r.text[:120]}")
return False
return True
except Exception as error:
logger.warning(f"ntfy {n.topic}: {error}")
return False
class NotificationState:
"""Per-key last-sent state, in the monitor's own SQL database."""
def __init__(self, engine, db_type: str):
self.engine = engine
self.db_type = db_type
self._ensure()
def _ensure(self) -> None:
from sqlalchemy import text
ddl = (
"CREATE TABLE IF NOT EXISTS notification_state ("
"key VARCHAR(64) PRIMARY KEY, state VARCHAR(16) NOT NULL, "
"value NUMERIC(8,3), updated_at TIMESTAMP NOT NULL)"
)
try:
with self.engine.begin() as conn:
conn.execute(text(ddl))
except Exception as error:
# Postgres: two sessions racing CREATE TABLE IF NOT EXISTS can
# both pass the existence check; the loser fails with a unique
# violation on pg_type. The table exists either way; verify.
with self.engine.connect() as conn:
conn.execute(text("SELECT 1 FROM notification_state WHERE 1=0"))
logger.debug(f"notification_state DDL raced, table present: {error}")
def get(self, key: str) -> Optional[str]:
from sqlalchemy import text
with self.engine.connect() as conn:
row = conn.execute(
text("SELECT state FROM notification_state WHERE key = :k"), {"k": key}
).fetchone()
return row[0] if row else None
def set(self, key: str, state: str, value: Optional[float] = None) -> None:
from sqlalchemy import text
now = datetime.datetime.now()
with self.engine.begin() as conn:
if self.db_type == "mysql":
sql = (
"INSERT INTO notification_state (key, state, value, updated_at) "
"VALUES (:k, :s, :v, :t) ON DUPLICATE KEY UPDATE "
"state = VALUES(state), value = VALUES(value), updated_at = VALUES(updated_at)"
)
else:
sql = (
"INSERT INTO notification_state (key, state, value, updated_at) "
"VALUES (:k, :s, :v, :t) ON CONFLICT (key) DO UPDATE SET "
"state = EXCLUDED.state, value = EXCLUDED.value, updated_at = EXCLUDED.updated_at"
)
conn.execute(text(sql), {"k": key, "s": state, "v": value, "t": now})
class InMemoryState(NotificationState):
"""For tests and when no SQL engine is available (loses state on restart)."""
def __init__(self): # noqa: D107 - intentionally skips the SQL parent
self._d: Dict[str, str] = {}
def get(self, key: str) -> Optional[str]:
return self._d.get(key)
def set(self, key: str, state: str, value: Optional[float] = None) -> None:
self._d[key] = state
def _level_state(level: float, warn: float, danger: float, prev: Optional[str]) -> str:
"""'clear' | 'warning' | 'danger', with hysteresis on the way down."""
if level >= danger:
return "danger"
if level >= warn:
# from danger: stay 'danger' until below danger - margin
if prev == "danger" and level >= danger - CLEAR_MARGIN_M:
return "danger"
return "warning"
if prev in ("warning", "danger") and level >= warn - CLEAR_MARGIN_M:
return "warning"
return "clear"
def evaluate(
readings: Iterable[dict],
forecasts: Iterable[dict],
state: NotificationState,
publisher: NtfyPublisher,
stale_after_h: float = 3.0,
now: Optional[datetime.datetime] = None,
) -> List[Notification]:
"""Compare current readings/forecasts with last-sent state; publish transitions.
readings: rows with station_code, water_level, timestamp (latest per station)
forecasts: /forecast rows (station_code, horizon_hours, p_warning, predicted_max_level)
Returns the notifications that were published (for logs/tests).
"""
now = now or datetime.datetime.now()
sent: List[Notification] = []
def emit(n: Notification) -> bool:
ok = publisher.publish(n)
if ok:
sent.append(n)
return ok
# ---- observed levels, per station, plus basin-wide fan-out
basin_changes: Dict[str, List[str]] = {"warning": [], "danger": [], "clear": []}
latest_ts: Optional[datetime.datetime] = None
for r in readings:
code = r.get("station_code")
level = r.get("water_level")
if not code or level is None:
continue
try:
level = float(level)
except (TypeError, ValueError):
continue
ts = r.get("timestamp")
if isinstance(ts, str):
try:
ts = datetime.datetime.fromisoformat(ts)
except ValueError:
ts = None
if isinstance(ts, datetime.datetime) and (latest_ts is None or ts > latest_ts):
latest_ts = ts
warn, danger = features.get_thresholds(code)
key = f"level:{code}"
prev = state.get(key) or "clear"
cur = _level_state(level, warn, danger, prev)
pct = r.get("discharge_percent")
if (
cur != "clear"
and prev == "clear"
and code not in CAPACITY_GUARD_EXEMPT
and pct is not None
):
try:
if float(pct) < CAPACITY_GUARD_MIN_PCT:
logger.info(
f"{code}: level {level:.2f} m >= {warn:.2f} but only "
f"{float(pct):.0f}% capacity; threshold looks stale, not alerting"
)
continue
except (TypeError, ValueError):
pass
if cur == prev:
continue
name = STATION_NAMES.get(code, code)
slug = _slug(code)
when = (
ts.strftime("%d %b %H:%M") if isinstance(ts, datetime.datetime) else "now"
)
if cur == "danger":
ok = emit(
Notification(
publisher.topic(slug, "danger"),
f"DANGER level at {code}",
f"{name}: {level:.2f} m at {when}, above the danger level of {danger:.2f} m.",
priority=5,
tags=["rotating_light", code],
)
)
basin_changes["danger"].append(f"{code} {level:.2f} m")
elif cur == "warning":
if prev == "danger":
ok = emit(
Notification(
publisher.topic(slug, "danger"),
f"{code} back below danger level",
f"{name}: {level:.2f} m at {when}; still above the warning level of {warn:.2f} m.",
priority=3,
tags=["arrow_down", code],
)
)
basin_changes["clear"].append(f"{code} below danger ({level:.2f} m)")
else:
ok = emit(
Notification(
publisher.topic(slug, "warning"),
f"Warning level at {code}",
f"{name}: {level:.2f} m at {when}, above the warning level of {warn:.2f} m.",
priority=4,
tags=["warning", code],
)
)
basin_changes["warning"].append(f"{code} {level:.2f} m")
else: # clear
ok = emit(
Notification(
publisher.topic(slug, "warning"),
f"{code} back to normal",
f"{name}: {level:.2f} m at {when}, below the warning level of {warn:.2f} m.",
priority=2,
tags=["white_check_mark", code],
)
)
basin_changes["clear"].append(f"{code} normal ({level:.2f} m)")
# Only remember the transition once it was actually delivered: if ntfy
# was down, the next cycle retries instead of silently swallowing a
# flood crossing.
if ok:
state.set(key, cur, level)
if basin_changes["danger"]:
emit(
Notification(
publisher.topic("danger"),
"Ping River: danger level reached",
"; ".join(basin_changes["danger"]),
priority=5,
tags=["rotating_light"],
)
)
if basin_changes["warning"]:
emit(
Notification(
publisher.topic("warning"),
"Ping River: warning level reached",
"; ".join(basin_changes["warning"]),
priority=4,
tags=["warning"],
)
)
if basin_changes["clear"]:
emit(
Notification(
publisher.topic("warning"),
"Ping River: levels falling",
"; ".join(basin_changes["clear"]),
priority=2,
tags=["white_check_mark"],
)
)
# ---- model outlook for the city gauge (opt-in topic, experimental)
p1 = next(
(
f
for f in forecasts
if f.get("station_code") == "P.1"
and f.get("horizon_hours") == OUTLOOK_HORIZON
and f.get("source") == "model"
),
None,
)
if p1 and p1.get("p_warning") is not None:
p = float(p1["p_warning"])
key = "outlook:P.1"
prev = state.get(key) or "off"
cur = (
"on" if (p >= OUTLOOK_ON or (prev == "on" and p >= OUTLOOK_OFF)) else "off"
)
if cur != prev:
peak = p1.get("predicted_max_level")
warn, _ = features.get_thresholds("P.1")
if cur == "on":
ok = emit(
Notification(
publisher.topic("p1-outlook"),
"Early warning: Chiang Mai flood risk rising",
f"The forecast model gives a {p * 100:.0f}% chance that Nawarat Bridge (P.1) "
f"reaches {warn:.2f} m within 24 h"
+ (
f" (expected peak {float(peak):.2f} m)"
if peak is not None
else ""
)
+ ". Experimental model output, not an official warning; "
"follow ThaiWater/TMD for official alerts.",
priority=4,
tags=["crystal_ball"],
)
)
else:
ok = emit(
Notification(
publisher.topic("p1-outlook"),
"Chiang Mai flood risk easing",
f"The model's 24 h probability of reaching {warn:.2f} m at P.1 has dropped to {p * 100:.0f}%.",
priority=2,
tags=["crystal_ball"],
)
)
if ok:
state.set(key, cur, p)
# ---- feed health
if latest_ts is not None:
age_h = (now - latest_ts).total_seconds() / 3600.0
key = "feed"
prev = state.get(key) or "ok"
cur = "stale" if age_h >= stale_after_h else "ok"
if cur != prev:
if cur == "stale":
ok = emit(
Notification(
publisher.topic("status"),
"Ping River monitor: gauge feed stale",
f"No new readings for {age_h:.0f} h (last {latest_ts:%d %b %H:%M}). "
"Levels and forecasts on the dashboard are not current.",
priority=3,
tags=["hourglass"],
)
)
else:
ok = emit(
Notification(
publisher.topic("status"),
"Ping River monitor: feed recovered",
f"Readings are current again (latest {latest_ts:%d %b %H:%M}).",
priority=2,
tags=["white_check_mark"],
)
)
if ok:
state.set(key, cur, age_h)
return sent
+5 -7
View File
@@ -51,8 +51,7 @@ class PostgresHistory:
if start >= end: if start >= end:
raise ValueError("start must be before end") raise ValueError("start must be before end")
query = text( query = text("""
"""
SELECT m.timestamp, s.station_code, m.water_level, SELECT m.timestamp, s.station_code, m.water_level,
m.discharge, m.discharge_percent m.discharge, m.discharge_percent
FROM water_measurements m FROM water_measurements m
@@ -62,8 +61,7 @@ class PostgresHistory:
AND m.timestamp <= :end_time AND m.timestamp <= :end_time
ORDER BY m.timestamp ASC ORDER BY m.timestamp ASC
LIMIT :limit LIMIT :limit
""" """)
)
with self.engine.connect() as connection: with self.engine.connect() as connection:
rows = connection.execute( rows = connection.execute(
query, query,
@@ -91,9 +89,9 @@ class PostgresHistory:
"station_code": station_code, "station_code": station_code,
"water_level": water_level, "water_level": water_level,
"discharge": discharge, "discharge": discharge,
"discharge_percent": float(row[4]) "discharge_percent": (
if row[4] is not None float(row[4]) if row[4] is not None else None
else None, ),
} }
) )
return result return result
+5 -3
View File
@@ -173,8 +173,10 @@ class RequestTracker:
"failed_requests": self.failed_requests, "failed_requests": self.failed_requests,
"success_rate": self.successful_requests / self.total_requests, "success_rate": self.successful_requests / self.total_requests,
"average_response_time": self.total_response_time / self.total_requests, "average_response_time": self.total_response_time / self.total_requests,
"last_request_time": self.last_request_time.isoformat() "last_request_time": (
if self.last_request_time self.last_request_time.isoformat()
else None, if self.last_request_time
else None
),
"error_breakdown": dict(self.error_count_by_type), "error_breakdown": dict(self.error_count_by_type),
} }
+308 -26
View File
@@ -7,6 +7,13 @@ dam in Thailand — storage, inflow and outflow in MCM — with archive depth
back to at least 2009. GET returns 404 ("Unknown method."); the POST body back to at least 2009. GET returns 404 ("Unknown method."); the POST body
may be empty but must carry a Content-Length. may be empty but must carry a Content-Length.
A sibling endpoint transposes that axis: ``GET .../api/dam`` with
``dam_id``/``date_start``/``date_end`` returns ONE dam over a whole date
range. It is GET-only (POST answers 404 "Unknown method.") and served the
full 2009-01-01..today archive 6,362 rows, 4.7 MB in a single ~6 s
response, so backfilling one dam costs one request instead of one per
calendar day. Field names differ from api/dams; see parse_dam_range_records.
The reservoir that matters for P.1 flood forecasting is Mae Ngat Somboon The reservoir that matters for P.1 flood forecasting is Mae Ngat Somboon
Chon (DAM_ID 200103), the only large dam upstream of Chiang Mai: during the Chon (DAM_ID 200103), the only large dam upstream of Chiang Mai: during the
Oct 2024 record flood it reached 113% of usable capacity with inflow spikes Oct 2024 record flood it reached 113% of usable capacity with inflow spikes
@@ -26,7 +33,25 @@ import requests
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
RID_DAMS_URL = "https://app.rid.go.th/reservoir/api/dams" RID_DAMS_URL = "https://app.rid.go.th/reservoir/api/dams"
RID_DAM_RANGE_URL = "https://app.rid.go.th/reservoir/api/dam"
MAE_NGAT_DAM_ID = "200103" MAE_NGAT_DAM_ID = "200103"
RID_ARCHIVE_START = datetime.date(2009, 1, 1) # earliest date api/dam serves
# Per-column NUMERIC capacity; source junk beyond these becomes NULL instead
# of overflowing the insert and discarding the whole daily batch.
_MEASURE_BOUNDS = {
"storage_mcm": 1e8,
"storage_pct": 1e6,
"inflow_mcm": 1e8,
"outflow_mcm": 1e8,
"level_msl": 1e6,
}
def _bounded(value: Optional[float], limit: float) -> Optional[float]:
if value is not None and abs(value) >= limit:
return None
return value
def _to_float(value: Any) -> Optional[float]: def _to_float(value: Any) -> Optional[float]:
@@ -75,6 +100,62 @@ def parse_dam_records(payload: Dict) -> List[Dict]:
return records return records
def parse_dam_range_records(payload: Dict) -> List[Dict]:
"""Flatten one api/dam single-dam range payload into the same rows as
parse_dam_records, so both endpoints feed one store.
api/dam names its columns differently and suffixes each measurement
``_curr`` / ``_prev``; ``_prev`` is the SAME calendar date one year
earlier (confirmed against api/dams' DMD_Date_prev) and is dropped —
those days are rows of their own. Mapping, verified equal to api/dams
on 2019-01-05, 2024-09-24, 2024-10-05 and 2026-08-08:
DMD_QUse_curr -> storage_mcm (identical)
PERCENT_DMD_QUse_curr -> storage_pct (2 dp; api/dams rounds to
whole percent, 112.62 vs 113)
DMD_Inflow_curr -> inflow_mcm (identical)
DMD_Outflow_curr -> outflow_mcm (identical)
DMD_ULevel_curr -> level_msl (only source of the level:
api/dams' DMD_Q is ' - ' for
all 35 dams, and api/dam's
DMD_Q_curr is a constant 0.00)
Dam metadata (name, region, capacities) and coordinates are carried too,
so the range path never has to blank rid_dams.
"""
coords = payload.get("dam_coordinates") or {}
payload_dam_id = payload.get("dam_id")
records = []
for row in payload.get("dam_data") or []:
dam_id = row.get("DAM_ID") or payload_dam_id
try:
date = datetime.date.fromisoformat(row.get("DATE_curr") or "")
except ValueError:
continue
if not dam_id:
continue
records.append(
{
"dam_id": dam_id,
"region": row.get("DAM_Region") or payload.get("dam_region"),
"name_th": row.get("DAM_Name") or payload.get("dam_name"),
"latitude": _to_float(coords.get("lat")),
"longitude": _to_float(coords.get("lng")),
"capacity_max_mcm": _to_float(row.get("DAM_QMax")),
"capacity_normal_mcm": _to_float(row.get("DAM_QStore")),
"date": date,
"storage_mcm": _to_float(row.get("DMD_QUse_curr")),
"storage_pct": _to_float(row.get("PERCENT_DMD_QUse_curr")),
"inflow_mcm": _to_float(row.get("DMD_Inflow_curr")),
"outflow_mcm": _to_float(row.get("DMD_Outflow_curr")),
# An MSL elevation of exactly 0 is "not published", not a
# reading — these dams sit between 45 and 400 m.
"level_msl": _to_float(row.get("DMD_ULevel_curr")) or None,
}
)
return records
class RidReservoirClient: class RidReservoirClient:
"""HTTP client for the RID reservoir daily-status API.""" """HTTP client for the RID reservoir daily-status API."""
@@ -82,9 +163,11 @@ class RidReservoirClient:
self, self,
url: str = RID_DAMS_URL, url: str = RID_DAMS_URL,
session: Optional[requests.Session] = None, session: Optional[requests.Session] = None,
timeout: int = 60, timeout: int = 120,
range_url: str = RID_DAM_RANGE_URL,
): ):
self.url = url self.url = url
self.range_url = range_url
self.session = session or requests.Session() self.session = session or requests.Session()
self.timeout = timeout self.timeout = timeout
@@ -100,6 +183,36 @@ class RidReservoirClient:
response.raise_for_status() response.raise_for_status()
return parse_dam_records(response.json()) return parse_dam_records(response.json())
def fetch_dam_range(
self, dam_id: str, start: datetime.date, end: datetime.date
) -> List[Dict]:
"""Every published day in [start, end] for one dam, in one request.
GET only api/dam answers POST with 404 "Unknown method.", the exact
opposite of api/dams. An unrecognised dam_id still returns HTTP 200
but with a PHP notice page instead of JSON, so a decode failure is
reported as a bad request rather than a transport error.
"""
response = self.session.get(
self.range_url,
params={
"dam_id": dam_id,
"date_start": start.isoformat(),
"date_end": end.isoformat(),
"percent": "",
},
timeout=self.timeout,
)
response.raise_for_status()
try:
payload = response.json()
except ValueError as e:
raise ValueError(
f"api/dam returned non-JSON for dam_id '{dam_id}' "
f"({start}..{end}) — unknown dam_id?"
) from e
return parse_dam_range_records(payload)
class RidReservoirStore: class RidReservoirStore:
"""SQL persistence for dam metadata + daily measurements. """SQL persistence for dam metadata + daily measurements.
@@ -151,7 +264,7 @@ class RidReservoirStore:
dam_id VARCHAR(10) NOT NULL, dam_id VARCHAR(10) NOT NULL,
date DATE NOT NULL, date DATE NOT NULL,
storage_mcm NUMERIC(10,2), storage_mcm NUMERIC(10,2),
storage_pct NUMERIC(6,2), storage_pct NUMERIC(8,2),
inflow_mcm NUMERIC(10,2), inflow_mcm NUMERIC(10,2),
outflow_mcm NUMERIC(10,2), outflow_mcm NUMERIC(10,2),
level_msl NUMERIC(8,2), level_msl NUMERIC(8,2),
@@ -168,21 +281,80 @@ class RidReservoirStore:
with self.engine.begin() as conn: with self.engine.begin() as conn:
for statement in ddl: for statement in ddl:
conn.execute(text(statement)) conn.execute(text(statement))
# Widen storage_pct on tables created before 2026-08-13: the source
# publishes junk percents (dam 100602 reports 87798%) that overflowed
# NUMERIC(6,2) and discarded whole daily batches.
if self.db_type == "postgresql":
migrations = (
"ALTER TABLE rid_reservoir_daily "
"ALTER COLUMN storage_pct TYPE NUMERIC(8,2)",
)
elif self.db_type == "mysql":
migrations = (
"ALTER TABLE rid_reservoir_daily MODIFY storage_pct NUMERIC(8,2)",
)
else: # sqlite: NUMERIC is affinity only, nothing to widen
migrations = ()
for statement in migrations:
try:
with self.engine.begin() as conn:
conn.execute(text(statement))
except Exception as e:
logger.warning(f"rid_reservoir_daily migration skipped: {e}")
def _upsert(self, table: str, key_cols: List[str], value_cols: List[str]) -> str: def _upsert(
self,
table: str,
key_cols: List[str],
value_cols: List[str],
preserve_cols: "tuple[str, ...]" = (),
) -> str:
"""Build an upsert; columns in `preserve_cols` keep their stored value
when the incoming one is NULL.
Dam metadata needs that: any payload that omits a name or coordinate
would otherwise blank a good rid_dams row on every later write.
"""
cols = key_cols + value_cols cols = key_cols + value_cols
col_list = ", ".join(cols) col_list = ", ".join(cols)
params = ", ".join(f":{c}" for c in cols) params = ", ".join(f":{c}" for c in cols)
conflict = ", ".join(key_cols)
if self.db_type == "sqlite": if self.db_type == "sqlite":
return f"INSERT OR REPLACE INTO {table} ({col_list}) VALUES ({params})" if not preserve_cols:
if self.db_type == "postgresql": return f"INSERT OR REPLACE INTO {table} ({col_list}) VALUES ({params})"
updates = ", ".join(f"{c} = EXCLUDED.{c}" for c in value_cols) updates = ", ".join(
conflict = ", ".join(key_cols) (
f"{c} = COALESCE(excluded.{c}, {table}.{c})"
if c in preserve_cols
else f"{c} = excluded.{c}"
)
for c in value_cols
)
return ( return (
f"INSERT INTO {table} ({col_list}) VALUES ({params}) " f"INSERT INTO {table} ({col_list}) VALUES ({params}) "
f"ON CONFLICT ({conflict}) DO UPDATE SET {updates}" f"ON CONFLICT ({conflict}) DO UPDATE SET {updates}"
) )
updates = ", ".join(f"{c} = VALUES({c})" for c in value_cols) if self.db_type == "postgresql":
updates = ", ".join(
(
f"{c} = COALESCE(EXCLUDED.{c}, {table}.{c})"
if c in preserve_cols
else f"{c} = EXCLUDED.{c}"
)
for c in value_cols
)
return (
f"INSERT INTO {table} ({col_list}) VALUES ({params}) "
f"ON CONFLICT ({conflict}) DO UPDATE SET {updates}"
)
updates = ", ".join(
(
f"{c} = COALESCE(VALUES({c}), {c})"
if c in preserve_cols
else f"{c} = VALUES({c})"
)
for c in value_cols
)
return ( return (
f"INSERT INTO {table} ({col_list}) VALUES ({params}) " f"INSERT INTO {table} ({col_list}) VALUES ({params}) "
f"ON DUPLICATE KEY UPDATE {updates}" f"ON DUPLICATE KEY UPDATE {updates}"
@@ -211,9 +383,20 @@ class RidReservoirStore:
"outflow_mcm", "outflow_mcm",
"level_msl", "level_msl",
] ]
dam_sql = self._upsert("rid_dams", ["dam_id"], dam_cols + ["updated_at"]) dam_sql = self._upsert(
"rid_dams",
["dam_id"],
dam_cols + ["updated_at"],
preserve_cols=tuple(dam_cols), # never blank metadata we already have
)
measure_sql = self._upsert( measure_sql = self._upsert(
"rid_reservoir_daily", ["dam_id", "date"], measure_cols "rid_reservoir_daily",
["dam_id", "date"],
measure_cols,
# Only api/dam carries a level (api/dams' DMD_Q is ' - ' for every
# dam), so the hourly collector would blank the backfilled level
# of today and yesterday on every cycle.
preserve_cols=("level_msl",),
) )
now = datetime.datetime.now() now = datetime.datetime.now()
dams = {} dams = {}
@@ -222,10 +405,10 @@ class RidReservoirStore:
dam_row = {c: record.get(c) for c in dam_cols} dam_row = {c: record.get(c) for c in dam_cols}
dam_row.update({"dam_id": record["dam_id"], "updated_at": now}) dam_row.update({"dam_id": record["dam_id"], "updated_at": now})
dams[record["dam_id"]] = dam_row dams[record["dam_id"]] = dam_row
measure_row = {c: record.get(c) for c in measure_cols} measure_row = {
measure_row.update( c: _bounded(record.get(c), _MEASURE_BOUNDS[c]) for c in measure_cols
{"dam_id": record["dam_id"], "date": record["date"]} }
) measure_row.update({"dam_id": record["dam_id"], "date": record["date"]})
measurements.append(measure_row) measurements.append(measure_row)
try: try:
with self.engine.begin() as conn: with self.engine.begin() as conn:
@@ -237,21 +420,31 @@ class RidReservoirStore:
return 0 return 0
def present_dates( def present_dates(
self, start: datetime.date, end: datetime.date self,
start: datetime.date,
end: datetime.date,
dam_id: Optional[str] = None,
) -> "set[datetime.date]": ) -> "set[datetime.date]":
"""Dates in [start, end] that already have rows, for backfill skipping.""" """Dates in [start, end] that already have rows, for backfill skipping.
`dam_id` narrows the answer to one dam: the per-day fleet backfill can
treat any stored date as done, but a per-dam backfill must not skip a
date merely because some other dam published it.
"""
if not self.engine and not self.connect(): if not self.engine and not self.connect():
return set() return set()
from sqlalchemy import text from sqlalchemy import text
sql = (
"SELECT DISTINCT date FROM rid_reservoir_daily "
"WHERE date >= :start AND date <= :end"
)
params = {"start": start, "end": end}
if dam_id is not None:
sql += " AND dam_id = :dam_id"
params["dam_id"] = dam_id
with self.engine.begin() as conn: with self.engine.begin() as conn:
values = conn.execute( values = conn.execute(text(sql), params).fetchall()
text(
"SELECT DISTINCT date FROM rid_reservoir_daily "
"WHERE date >= :start AND date <= :end"
),
{"start": start, "end": end},
).fetchall()
dates = set() dates = set()
for (value,) in values: for (value,) in values:
if isinstance(value, str): # sqlite returns ISO strings if isinstance(value, str): # sqlite returns ISO strings
@@ -306,9 +499,7 @@ def backfill(
if not store.engine and not store.connect(): if not store.engine and not store.connect():
logger.error("backfill aborted: database connection failed") logger.error("backfill aborted: database connection failed")
return 0 return 0
span = [ span = [start + datetime.timedelta(days=i) for i in range((end - start).days + 1)]
start + datetime.timedelta(days=i) for i in range((end - start).days + 1)
]
present = store.present_dates(start, end) present = store.present_dates(start, end)
targets = [d for d in span if d not in present] targets = [d for d in span if d not in present]
logger.info( logger.info(
@@ -336,6 +527,97 @@ def backfill(
return total return total
def backfill_dam(
store: RidReservoirStore,
dam_id: str = MAE_NGAT_DAM_ID,
start: datetime.date = RID_ARCHIVE_START,
end: Optional[datetime.date] = None,
client: Optional[RidReservoirClient] = None,
chunk_days: int = 1830,
throttle_seconds: float = 1.0,
skip_present: bool = True,
stats: Optional[Dict] = None,
) -> int:
"""Backfill ONE dam over [start, end] using the range endpoint.
Costs one request per chunk instead of one per calendar day: Mae Ngat's
whole 2009-today archive is ~4 requests here versus ~6,400 with
`backfill`. No server-side range limit was observed (2009-01-01..today
answered in full), so `chunk_days` exists only to bound the response size
and the time a single request can hang, not to satisfy the API.
Chunks whose dates are already stored are skipped without a request, and
returned rows are filtered to the missing dates, so a rerun repairs holes
rather than rewriting the archive. `skip_present=False` re-fetches
everything, which is how rows first written by the api/dams collector
gain a level_msl and two-decimal storage_pct.
Returns rows saved. Some dates stay missing however often this runs
the source simply never published them (72 days of Mae Ngat's archive,
absent from api/dams too) so a 0-row rerun is normal and callers must
not read it as failure; pass `stats` to get the request/failure counts
that actually distinguish an outage.
"""
client = client or RidReservoirClient()
end = end or datetime.date.today()
counters = {"requests": 0, "failures": 0, "aborted": False}
if stats is not None:
stats.update(counters)
counters = stats
if not store.engine and not store.connect():
logger.error("backfill_dam aborted: database connection failed")
counters["aborted"] = True
return 0
span_days = (end - start).days + 1
if span_days <= 0:
return 0
present = store.present_dates(start, end, dam_id=dam_id) if skip_present else set()
logger.info(
f"backfill_dam {dam_id}: {span_days - len(present)} of {span_days} "
f"days missing in [{start}, {end}]"
)
total = 0
failures = 0
chunk_start = start
while chunk_start <= end:
chunk_end = min(chunk_start + datetime.timedelta(days=chunk_days - 1), end)
wanted = {
chunk_start + datetime.timedelta(days=i)
for i in range((chunk_end - chunk_start).days + 1)
} - present
if not wanted:
chunk_start = chunk_end + datetime.timedelta(days=1)
continue
try:
counters["requests"] += 1
records = client.fetch_dam_range(dam_id, chunk_start, chunk_end)
records = [r for r in records if r["date"] in wanted]
saved = store.save(records)
if records and not saved:
raise RuntimeError("database save persisted 0 rows")
total += saved
failures = 0
logger.info(
f"backfill_dam {dam_id} [{chunk_start}, {chunk_end}]: "
f"{saved} rows ({total} total)"
)
except Exception as e:
failures += 1
counters["failures"] += 1
logger.warning(
f"backfill_dam {dam_id} [{chunk_start}, {chunk_end}] failed "
f"({failures} in a row): {e}"
)
if failures >= 5:
logger.error("5 consecutive failures — aborting backfill_dam")
counters["aborted"] = True
break
chunk_start = chunk_end + datetime.timedelta(days=1)
if chunk_start <= end:
time.sleep(throttle_seconds)
return total
def create_collector_from_config() -> Optional[RidReservoirCollector]: def create_collector_from_config() -> Optional[RidReservoirCollector]:
"""Build a collector from app Config; None when disabled or non-SQL DB.""" """Build a collector from app Config; None when disabled or non-SQL DB."""
from .config import Config from .config import Config
+1174 -212
View File
File diff suppressed because it is too large Load Diff
+1 -3
View File
@@ -632,9 +632,7 @@ class EnhancedWaterMonitorScraper:
if data: if data:
if self.save_to_database(data): if self.save_to_database(data):
filled_count += len(data) filled_count += len(data)
logger.info( logger.info(f"Filled {len(data)} measurements for {fetch_date}")
f"Filled {len(data)} measurements for {fetch_date}"
)
else: else:
logger.warning(f"Failed to save data for {fetch_date}") logger.warning(f"Failed to save data for {fetch_date}")
else: else:
+239 -19
View File
@@ -210,7 +210,9 @@ async def lifespan(app: FastAPI):
app_state["leader_lock"] = _acquire_collection_leadership( app_state["leader_lock"] = _acquire_collection_leadership(
Config.COLLECTION_LEADER_PORT Config.COLLECTION_LEADER_PORT
) )
app_state["notify"] = None
if app_state["leader_lock"]: if app_state["leader_lock"]:
app_state["notify"] = _init_notifications()
app_state["scraping_task"] = asyncio.create_task(background_scraping_task()) app_state["scraping_task"] = asyncio.create_task(background_scraping_task())
logger.info("This worker is the background-collection leader") logger.info("This worker is the background-collection leader")
else: else:
@@ -296,6 +298,73 @@ async def _persist_rain():
logger.warning(f"rain persistence failed: {e}") logger.warning(f"rain persistence failed: {e}")
def _init_notifications():
"""Publisher + persisted state for ntfy, or None if off/unavailable.
Called only by the collection leader: it is the one process that
publishes, so the notification_state DDL runs exactly once per host.
"""
if not Config.NTFY_SERVER:
return None
try:
from . import notify as notify_mod
store = app_state.get("forecast_store")
if store and not store.engine:
store.connect()
state = (
notify_mod.NotificationState(store.engine, store.db_type)
if store and store.engine
else notify_mod.InMemoryState()
)
if isinstance(state, notify_mod.InMemoryState):
logger.warning(
"ntfy: no SQL store; notification state is in-memory "
"(a restart may re-send the current level)"
)
publisher = notify_mod.NtfyPublisher(
Config.NTFY_PUBLISH_URL,
prefix=Config.NTFY_TOPIC_PREFIX,
token=Config.NTFY_TOKEN or None,
dashboard_url=Config.PUBLIC_URL,
)
logger.info(
f"ntfy notifications: publish to {Config.NTFY_PUBLISH_URL}, "
f"subscribers use {Config.NTFY_SERVER}, topics {Config.NTFY_TOPIC_PREFIX}-*"
)
return publisher, state
except Exception as e:
logger.error(f"ntfy init failed (notifications off): {e}")
return None
async def _notify_transitions():
"""Publish flood/outlook/feed transitions to ntfy (leader only, fail-safe)."""
cfg = app_state.get("notify")
if not cfg:
return
publisher, state = cfg
try:
from . import notify as notify_mod
scraper = app_state["scraper"]
readings = await asyncio.to_thread(
scraper.db_adapter.get_latest_measurements, 200
)
with FORECAST_CACHE_LOCK:
cached = FORECAST_CACHE.get("all")
forecasts = cached[1] if cached else []
sent = await asyncio.to_thread(
notify_mod.evaluate, readings, forecasts, state, publisher
)
if sent:
logger.info(
"ntfy: published " + ", ".join(f"{n.topic}: {n.title}" for n in sent)
)
except Exception as e:
logger.warning(f"ntfy notify cycle failed: {e}")
async def _precompute_forecasts(): async def _precompute_forecasts():
"""Refresh the forecast cache and persist the issued forecasts (leader only).""" """Refresh the forecast cache and persist the issued forecasts (leader only)."""
try: try:
@@ -374,9 +443,7 @@ async def background_scraping_task():
hii_counts = await asyncio.get_event_loop().run_in_executor( hii_counts = await asyncio.get_event_loop().run_in_executor(
None, hii_collector.run_cycle None, hii_collector.run_cycle
) )
set_gauge( set_gauge("hii_rainfall_rows_saved", hii_counts["rainfall"])
"hii_rainfall_rows_saved", hii_counts["rainfall"]
)
set_gauge( set_gauge(
"hii_waterlevel_rows_saved", hii_counts["waterlevel"] "hii_waterlevel_rows_saved", hii_counts["waterlevel"]
) )
@@ -404,6 +471,10 @@ async def background_scraping_task():
# evaluation. # evaluation.
await _precompute_forecasts() await _precompute_forecasts()
# Push notifications for threshold crossings (uses the
# forecasts just computed; no-op unless NTFY_SERVER set).
await _notify_transitions()
app_state["is_scraping"] = False app_state["is_scraping"] = False
# Calculate next run time # Calculate next run time
@@ -521,7 +592,9 @@ _STATIC_DIR = os.path.dirname(_DASHBOARD_HTML_PATH)
@app.get("/robots.txt", include_in_schema=False) @app.get("/robots.txt", include_in_schema=False)
async def robots_txt(): async def robots_txt():
return FileResponse(os.path.join(_STATIC_DIR, "robots.txt"), media_type="text/plain") return FileResponse(
os.path.join(_STATIC_DIR, "robots.txt"), media_type="text/plain"
)
@app.get("/llms.txt", include_in_schema=False) @app.get("/llms.txt", include_in_schema=False)
@@ -864,7 +937,7 @@ def _hii_rows(sql: str, params: Dict[str, Any]) -> List[Dict[str, Any]]:
# flood, slightly stale readings with a visible timestamp beat an error page. # flood, slightly stale readings with a visible timestamp beat an error page.
HII_CACHE: Dict[str, Any] = {} HII_CACHE: Dict[str, Any] = {}
HII_CACHE_LOCK = Lock() HII_CACHE_LOCK = Lock()
_HII_COMPUTE_LOCKS = {"rain": Lock(), "waterlevel": Lock()} _HII_COMPUTE_LOCKS = {"rain": Lock(), "waterlevel": Lock(), "skill": Lock()}
LATEST_CACHE: Dict[str, Any] = {} LATEST_CACHE: Dict[str, Any] = {}
LATEST_CACHE_LOCK = Lock() LATEST_CACHE_LOCK = Lock()
_LATEST_COMPUTE_LOCK = Lock() _LATEST_COMPUTE_LOCK = Lock()
@@ -1001,6 +1074,91 @@ async def get_hii_rainfall_latest(
return rows return rows
@app.get("/api/hii/rainfall/catchment")
async def get_hii_rainfall_catchment(
response: Response, days: int = Query(14, ge=1, le=60)
):
"""Upper-Ping catchment-mean hourly rain: HII gauges vs the Open-Meteo
series the flood model actually uses, plus their agreement over the window.
Evidence-gathering endpoint (docs/FLOOD_FORECASTING.md, HII gauge rain):
the gauge table only exists since 2026-08 so it cannot be a training
feature yet; this makes the two sources' relationship observable meanwhile.
"""
increment_counter("api_requests", labels={"endpoint": "hii_rain_catchment"})
start = datetime.now() - timedelta(days=days)
def compute():
import pandas as pd
from .ml import hii_rain
engine = _hii_engine()
if engine is None:
return {
"box": hii_rain.CATCHMENT_BOX,
"gauge": [],
"openmeteo": [],
"comparison_24h_sums": {"overlap_hours": 0},
}
gauge = hii_rain.load_gauge_mean(start=pd.Timestamp(start), engine=engine)
openmeteo = None
try:
from sqlalchemy import text
with engine.connect() as conn:
frame = pd.read_sql(
text(
"SELECT timestamp, catchment_mean FROM openmeteo_rain "
"WHERE timestamp >= :start ORDER BY timestamp"
),
conn,
params={"start": start},
)
if not frame.empty:
frame["timestamp"] = pd.to_datetime(frame["timestamp"])
openmeteo = pd.to_numeric(
frame.set_index("timestamp")["catchment_mean"], errors="coerce"
)
except Exception as error: # openmeteo_rain may not exist yet
logger.warning(f"openmeteo_rain read failed: {error}")
def series_rows(s):
if s is None:
return []
return [
{
"timestamp": ts.isoformat(),
"rain_mm": None if pd.isna(v) else round(float(v), 2),
}
for ts, v in s.items()
]
comparison = (
hii_rain.compare_with_openmeteo(gauge, openmeteo)
if gauge is not None and openmeteo is not None
else {"overlap_hours": 0}
)
return {
"box": hii_rain.CATCHMENT_BOX,
"gauge": series_rows(gauge),
"openmeteo": series_rows(openmeteo),
"comparison_24h_sums": comparison,
}
payload, stale = await _cached_swr(
HII_CACHE,
HII_CACHE_LOCK,
_HII_COMPUTE_LOCKS["rain"],
f"rain_catchment:{days}",
Config.HII_CACHE_TTL_SECONDS,
compute,
)
if stale:
response.headers["X-Data-Stale"] = "true"
return payload
@app.get("/api/hii/waterlevel/latest") @app.get("/api/hii/waterlevel/latest")
async def get_hii_waterlevel_latest( async def get_hii_waterlevel_latest(
response: Response, hours: int = Query(26, ge=1, le=168) response: Response, hours: int = Query(26, ge=1, le=168)
@@ -1043,9 +1201,7 @@ async def get_postgres_history(
return cached[1] return cached[1]
try: try:
db_config = Config.get_database_config() db_config = Config.get_database_config()
end_time = ( end_time = datetime.combine(end, datetime.max.time()) if end else datetime.now()
datetime.combine(end, datetime.max.time()) if end else datetime.now()
)
start_time = ( start_time = (
datetime.combine(start, datetime.min.time()) datetime.combine(start, datetime.min.time())
if start if start
@@ -1138,17 +1294,85 @@ async def get_forecast_history(
store = app_state.get("forecast_store") store = app_state.get("forecast_store")
if not store: if not store:
return [] return []
end_dt = ( end_dt = datetime.combine(end, datetime.max.time()) if end else datetime.now()
datetime.combine(end, datetime.max.time()) if end else datetime.now()
)
start_dt = ( start_dt = (
datetime.combine(start, datetime.min.time()) datetime.combine(start, datetime.min.time())
if start if start
else end_dt - timedelta(hours=hours) else end_dt - timedelta(hours=hours)
) )
return await asyncio.to_thread( return await asyncio.to_thread(store.fetch, station_code, start_dt, end_dt, horizon)
store.fetch, station_code, start_dt, end_dt, horizon
@app.get("/api/notifications")
async def get_notifications_config():
"""Public ntfy settings so the dashboard can offer subscribe links."""
server = Config.NTFY_SERVER
if not server:
return {"enabled": False}
prefix = Config.NTFY_TOPIC_PREFIX
return {
"enabled": True,
"server": server,
"prefix": prefix,
"topics": {
"warning": f"{prefix}-warning",
"danger": f"{prefix}-danger",
"p1_outlook": f"{prefix}-p1-outlook",
"status": f"{prefix}-status",
"station_pattern": f"{prefix}-<station>-warning | {prefix}-<station>-danger (station code lowercase, no dot: p1, p103)",
},
"semantics": "transitions only: one message on crossing up, one all-clear on the way down (0.10 m hysteresis)",
}
@app.get("/api/forecast/skill")
async def get_forecast_skill(
response: Response,
station_code: str = Query("P.1"),
horizon: int = Query(24, ge=1, le=48),
):
"""Is the model getting better? Issued forecasts verified against what the
river then did, per model version, with a persistence baseline.
Read from forecast_history (what each deployed version predicted, hourly)
joined to water_measurements; no retraining involved. Cached like the HII
feeds because the join is a few hundred correlated subqueries.
"""
increment_counter("api_requests", labels={"endpoint": "forecast_skill"})
store = app_state.get("forecast_store")
if not store:
return {
"station_code": station_code,
"horizon_hours": horizon,
"versions": [],
"current": None,
"trend": None,
}
def compute():
from .ml import skill
if not store.engine and not store.connect():
return {
"station_code": station_code,
"horizon_hours": horizon,
"versions": [],
"current": None,
"trend": None,
}
return skill.compute_skill(store.engine, store.db_type, station_code, horizon)
payload, stale = await _cached_swr(
HII_CACHE,
HII_CACHE_LOCK,
_HII_COMPUTE_LOCKS["skill"],
f"skill:{station_code}:{horizon}",
max(Config.HII_CACHE_TTL_SECONDS, 900),
compute,
) )
if stale:
response.headers["X-Data-Stale"] = "true"
return payload
@app.get("/measurements/latest", response_model=List[MeasurementResponse]) @app.get("/measurements/latest", response_model=List[MeasurementResponse])
@@ -1234,9 +1458,7 @@ async def get_database_stats():
from sqlalchemy import text from sqlalchemy import text
with engine.connect() as conn: with engine.connect() as conn:
return conn.execute( return conn.execute(text("""
text(
"""
SELECT (SELECT COUNT(*) FROM hii_rainfall) AS rain_n, SELECT (SELECT COUNT(*) FROM hii_rainfall) AS rain_n,
(SELECT COUNT(*) FROM hii_waterlevel) AS wl_n, (SELECT COUNT(*) FROM hii_waterlevel) AS wl_n,
(SELECT COUNT(*) FROM hii_rain_stations) AS rain_s, (SELECT COUNT(*) FROM hii_rain_stations) AS rain_s,
@@ -1245,9 +1467,7 @@ async def get_database_stats():
(SELECT MAX(timestamp) FROM hii_rainfall) AS rain_hi, (SELECT MAX(timestamp) FROM hii_rainfall) AS rain_hi,
(SELECT MIN(timestamp) FROM hii_waterlevel) AS wl_lo, (SELECT MIN(timestamp) FROM hii_waterlevel) AS wl_lo,
(SELECT MAX(timestamp) FROM hii_waterlevel) AS wl_hi (SELECT MAX(timestamp) FROM hii_waterlevel) AS wl_hi
""" """)).one()
)
).one()
def compute(): def compute():
# Heavy: full-table counts and coverage over ~1.7M rows. Runs at most # Heavy: full-table counts and coverage over ~1.7M rows. Runs at most
+128
View File
@@ -0,0 +1,128 @@
"""Tests for the Mae Ngat dam series (src/ml/dam.py) and its feature gating."""
import datetime
import numpy as np
import pandas as pd
from src.ml import features
from src.ml.dam import FFILL_LIMIT_H, REPORT_HOUR, hourly_frame
def _daily(days=5, start="2024-09-20"):
idx = pd.date_range(start, periods=days, freq="D")
return pd.DataFrame(
{
"storage_pct": np.linspace(90, 110, days),
"inflow_mcm": np.linspace(2, 20, days),
"outflow_mcm": np.linspace(0.5, 5, days),
},
index=idx,
)
class TestHourlyFrame:
def test_daily_value_visible_from_report_hour_only(self):
hourly = hourly_frame(_daily())
day0 = pd.Timestamp("2024-09-20")
# Nothing before the first report hour
assert hourly.index.min() == day0 + pd.Timedelta(hours=REPORT_HOUR)
# The day's value holds from 07:00 through the next morning
assert hourly.loc[day0 + pd.Timedelta(hours=7), "storage_pct"] == 90.0
assert hourly.loc[day0 + pd.Timedelta(hours=23), "storage_pct"] == 90.0
next_6am = day0 + pd.Timedelta(days=1, hours=6)
next_7am = day0 + pd.Timedelta(days=1, hours=7)
assert hourly.loc[next_6am, "storage_pct"] == 90.0 # yesterday's value
assert hourly.loc[next_7am, "storage_pct"] == 95.0 # today's report
def test_ffill_capped_after_missing_days(self):
daily = _daily(days=2).drop(index=pd.Timestamp("2024-09-21"))
# extend with a far-later row so the gap sits mid-frame
late = _daily(days=1, start="2024-09-28")
hourly = hourly_frame(pd.concat([daily, late]))
gap_ts = pd.Timestamp("2024-09-20") + pd.Timedelta(
hours=REPORT_HOUR + FFILL_LIMIT_H + 1
)
assert np.isnan(hourly.loc[gap_ts, "storage_pct"])
def test_none_and_empty(self):
assert hourly_frame(None) is None
assert hourly_frame(pd.DataFrame()) is None
class TestLoadDaily:
def test_empty_db_result_does_not_wipe_cache(self, tmp_path):
from sqlalchemy import create_engine, text
from src.ml.dam import CACHE_FILE, load_daily
# Good cache from a previous run
cache_path = tmp_path / CACHE_FILE
_daily(3).rename_axis("date").to_csv(cache_path, compression="gzip")
# Reachable DB whose table exists but is empty
db = f"sqlite:///{tmp_path}/empty_dam.db"
with create_engine(db).begin() as conn:
conn.execute(
text(
"CREATE TABLE rid_reservoir_daily (dam_id TEXT, date DATE, "
"storage_pct REAL, inflow_mcm REAL, outflow_mcm REAL)"
)
)
result = load_daily(db_url=db, cache_dir=tmp_path)
assert result.empty # honest empty result...
cached = pd.read_csv(cache_path, index_col=0)
assert len(cached) == 3 # ...but the good cache survives
def _grid(hours=400, start="2024-09-15"):
idx = pd.date_range(start, periods=hours, freq="h")
frames = []
for code in ("P.1", "P.82"):
frames.append(
pd.DataFrame(
{
"timestamp": idx,
"station_code": code,
"water_level": 2.0,
"discharge": 100.0,
}
)
)
return features.make_hourly_grid(pd.concat(frames, ignore_index=True))
class TestFeatureGating:
def test_dam_columns_only_for_dam_stations(self):
grid = _grid()
dam = hourly_frame(_daily(days=20, start="2024-09-10"))
X_p1 = features.build_features(grid, "P.1", dam=dam)
X_p82 = features.build_features(grid, "P.82", dam=dam)
for col in features.DAM_FEATURES:
assert col in X_p1.columns
assert col not in X_p82.columns
# values actually aligned, not all-NaN
assert X_p1["dam_storage_pct"].notna().any()
assert X_p1["dam_inflow"].notna().any()
def test_none_dam_omits_columns(self):
X = features.build_features(_grid(), "P.1", dam=None)
for col in features.DAM_FEATURES:
assert col not in X.columns
def test_empty_dam_frame_yields_nan_columns(self):
# Serving contract: empty frame -> columns exist as NaN so dam-trained
# bundles pass the feature guard when the DB read fails.
X = features.build_features(_grid(), "P.1", dam=pd.DataFrame())
for col in features.DAM_FEATURES:
assert col in X.columns
assert X[col].isna().all()
def test_storage_delta_72h(self):
grid = _grid(hours=24 * 12, start="2024-09-15")
dam = hourly_frame(_daily(days=20, start="2024-09-10"))
X = features.build_features(grid, "P.1", dam=dam)
ts = pd.Timestamp("2024-09-24 12:00")
expected = X.loc[ts, "dam_storage_pct"] - X.loc[
ts - pd.Timedelta(hours=72), "dam_storage_pct"
]
assert abs(X.loc[ts, "dam_storage_pct_d3"] - expected) < 1e-9
+123 -13
View File
@@ -14,12 +14,18 @@ def test_dashboard_contains_live_map_and_flow_visualization():
assert "leaflet" in html.lower() assert "leaflet" in html.lower()
def test_dashboard_explains_flow_legend_and_refresh(): def _body(html: str) -> str:
html = DASHBOARD_PATH.read_text(encoding="utf-8") """Markup only. The STRINGS table repeats every English phrase, so a
whole-file search would pass even if an element were deleted."""
return html[html.index("<body>"):html.index("<script src=")]
assert "Flow status" in html
assert "Last updated" in html def test_dashboard_explains_flow_legend_and_refresh():
assert "Refresh" in html body = _body(DASHBOARD_PATH.read_text(encoding="utf-8"))
assert 'data-i18n="legend.flow"' in body
assert 'data-i18n="stat.updated"' in body
assert 'data-i18n="action.refresh"' in body
def test_dashboard_uses_mapped_river_network_instead_of_station_connections(): def test_dashboard_uses_mapped_river_network_instead_of_station_connections():
@@ -35,7 +41,7 @@ def test_dashboard_loads_additional_thaiwater_sensors():
html = DASHBOARD_PATH.read_text(encoding="utf-8") html = DASHBOARD_PATH.read_text(encoding="utf-8")
assert "fetch('/sensors/thaiwater')" in html assert "fetch('/sensors/thaiwater')" in html
assert "Additional basin stations" in html assert 'data-i18n="sensors.title"' in _body(html)
assert 'id="station-search"' in html assert 'id="station-search"' in html
assert "applyStationSearch" in html assert "applyStationSearch" in html
@@ -67,23 +73,127 @@ def test_dashboard_shows_hii_rainfall_layer():
assert "fetch('/api/hii/waterlevel/latest')" in html assert "fetch('/api/hii/waterlevel/latest')" in html
assert "renderRainLayer" in html assert "renderRainLayer" in html
assert "rain-toggle" in html assert "rain-toggle" in html
assert "Rainfall · 24 h" in html body = _body(html)
assert 'data-i18n="legend.rain"' in body
# TMD rain classes on the legend # TMD rain classes on the legend
assert "Heavy 3590" in html assert 'data-i18n="legend.rain.heavy"' in body
assert "Extreme &gt; 150" in html assert 'data-i18n="legend.rain.extreme"' in body
def test_dashboard_loads_station_history_chart(): def test_dashboard_loads_station_history_chart():
html = DASHBOARD_PATH.read_text(encoding="utf-8") html = DASHBOARD_PATH.read_text(encoding="utf-8")
assert "Station history" in html assert 'id="history-card"' in html
assert "/api/forecast/history/" in html assert "/api/forecast/history/" in html
assert "Model 24 h peak (as issued)" in html assert "'chart.model'" in html
assert "PostgreSQL" not in html assert "PostgreSQL" not in html
assert "/measurements/history/" in html assert "/measurements/history/" in html
assert "history-chart" in html assert "history-chart" in html
# Date-range picker alongside the quick-range dropdown # Date-range picker alongside the quick-range dropdown
assert 'id="history-start"' in html assert 'id="history-start"' in html
assert 'id="history-end"' in html assert 'id="history-end"' in html
assert "Last 7 days" in html assert 'data-i18n="range.7d"' in _body(html)
assert "Last 90 days" in html assert 'data-i18n="range.90d"' in _body(html)
def _extract_lang_tables(html: str) -> dict:
"""Pull the `en:` / `th:` key sets out of the STRINGS literal.
Parsing the JS with a regex is crude, but it is enough to catch the failure
that matters: a key added to one language and forgotten in the other, which
silently falls back to English for Thai readers.
"""
import re
start = html.index("const STRINGS = {")
end = html.index("// Thai unless the visitor's browser", start)
block = html[start:end]
tables = {}
for lang in ("en", "th"):
section = re.search(rf"\n {lang}: {{\n(.*?)\n }},\n", block, re.S)
assert section, f"{lang} table not found in STRINGS"
tables[lang] = set(re.findall(r"^\s{12}'([^']+)':", section.group(1), re.M))
return tables
def test_dashboard_translations_cover_both_languages():
html = DASHBOARD_PATH.read_text(encoding="utf-8")
tables = _extract_lang_tables(html)
assert len(tables["en"]) > 100, "expected the full English string table"
missing_th = tables["en"] - tables["th"]
missing_en = tables["th"] - tables["en"]
assert not missing_th, f"keys missing a Thai translation: {sorted(missing_th)}"
assert not missing_en, f"Thai-only keys with no English fallback: {sorted(missing_en)}"
def test_dashboard_i18n_markup_keys_exist():
"""Every data-i18n attribute must resolve to a real string key."""
import re
html = DASHBOARD_PATH.read_text(encoding="utf-8")
keys = _extract_lang_tables(html)["en"]
used = set(re.findall(r'data-i18n(?:-placeholder|-title|-aria)?="([^"]+)"', html))
unknown = used - keys
assert not unknown, f"markup references undefined string keys: {sorted(unknown)}"
def test_dashboard_is_mobile_portrait_safe():
html = DASHBOARD_PATH.read_text(encoding="utf-8")
# The header row overflowed a 412 px Android viewport by 64 px until it wrapped
assert "flex-wrap: wrap" in html
assert "overflow-x: hidden" in html
assert 'id="lang-toggle"' in html
assert "ping-monitor-lang" in html # remembered language choice
def test_dashboard_translates_user_visible_aria_labels():
"""A Thai page must not hand screen-reader users English landmarks."""
import re
html = DASHBOARD_PATH.read_text(encoding="utf-8")
body = _body(html)
for match in re.finditer(r'<[^>]*\saria-label="[^"]+"[^>]*>', body):
tag = match.group(0)
if "data-i18n-aria" in tag or "id=\"lang-toggle\"" in tag:
continue # the toggle sets its own label per language in JS
raise AssertionError(f"aria-label without a translation key: {tag[:120]}")
def test_dashboard_keeps_simulation_and_replay_labels_on_language_switch():
"""Relabelling a pinned simulation as LIVE would present fake flood data as real."""
html = DASHBOARD_PATH.read_text(encoding="utf-8")
assert "setLiveIndicator(state.liveMode || 'live', state.liveLabelKey)" in html
assert "if (!state.replayTimer) setLiveIndicator('live');" not in html
def test_dashboard_default_language_respects_browser_order():
"""navigator.languages = ['th-TH','en-US'] must resolve to Thai, not English."""
html = DASHBOARD_PATH.read_text(encoding="utf-8")
assert "langs.some" not in html # the old any-English-wins test
assert "if (code.startsWith('th')) return 'th';" in html
def test_dashboard_shows_one_current_river_level():
"""The verdict banner and the P.1 outlook must not disagree about "now".
Production served 1.66 m in the banner and 1.52 m in the outlook at the
same moment: the banner used the latest measurement, the outlook used the
forecast payload's current_level from an older as_of.
"""
html = DASHBOARD_PATH.read_text(encoding="utf-8")
# One arbitrated writer, freshest-wins
assert "function setP1Level(" in html
assert "state.p1NowAt" in html
# Neither feed may assign the level directly any more
assert "state.p1Now = Number(p1.water_level)" not in html
assert "state.p1Now = row.current_level" not in html
# The outlook renders the arbitrated level, and never a peak below it
assert "const shownNow = state.p1Now != null" in html
assert "Math.max(Number(row.predicted_max_level), shownNow)" in html
+71 -2
View File
@@ -197,7 +197,7 @@ def test_train_smoke_and_roundtrip(tmp_path):
df = make_synth(n, data_stations, seed=7, pulses=pulses) df = make_synth(n, data_stations, seed=7, pulses=pulses)
metrics = train.train_all( metrics = train.train_all(
df, target_stations, models_dir=tmp_path, skip_eval=True, hgb_overrides={"max_iter": 20}, use_rain=False df, target_stations, models_dir=tmp_path, skip_eval=True, hgb_overrides={"max_iter": 20}, use_rain=False, use_dam=False
) )
assert metrics["stations"]["P.1"]["status"] == "trained" assert metrics["stations"]["P.1"]["status"] == "trained"
assert metrics["stations"]["P.20"]["status"] == "trained" assert metrics["stations"]["P.20"]["status"] == "trained"
@@ -233,11 +233,80 @@ def test_heuristic_fallback(tmp_path):
assert row["trained_at"] is None assert row["trained_at"] is None
def _p1_synth(n: int = 300, seed: int = 11) -> pd.DataFrame:
upstream = [code for code, _lead in features.UPSTREAM_LEADS["P.1"]]
return make_synth(n, ["P.1"] + upstream, seed=seed, pulses={"P.1": [(100, 20, 2.0)]})
def test_train_refuses_silent_rain_downgrade(tmp_path, monkeypatch):
"""use_rain=True with no rain series must abort, not write v2 bundles.
Regression for the 2026-09-01 server retrain that overwrote v3 with v2
because the Open-Meteo archive fetch failed on a cache-less checkout.
"""
from src.ml import rain as rain_mod
df = _p1_synth()
overrides = {"max_iter": 10}
# Case 1: the loader returns None (archive unreachable, no cache file)
monkeypatch.setattr(rain_mod, "load_history", lambda *a, **k: None)
with pytest.raises(train.RainUnavailableError, match="--no-rain"):
train.train_all(df, ["P.1"], models_dir=tmp_path, skip_eval=True, hgb_overrides=overrides, use_rain=True, use_dam=False)
assert not (tmp_path / "flood_P.1.joblib").exists()
assert not (tmp_path / "metrics.json").exists()
# Case 2: the loader raises (network / parse error)
def boom(*a, **k):
raise ConnectionError("simulated Open-Meteo outage")
monkeypatch.setattr(rain_mod, "load_history", boom)
with pytest.raises(train.RainUnavailableError, match="simulated Open-Meteo outage"):
train.train_all(df, ["P.1"], models_dir=tmp_path, skip_eval=True, hgb_overrides=overrides, use_rain=True, use_dam=False)
assert not (tmp_path / "flood_P.1.joblib").exists()
# Explicit opt-out still produces v2 bundles as before
metrics = train.train_all(df, ["P.1"], models_dir=tmp_path, skip_eval=True, hgb_overrides=overrides, use_rain=False, use_dam=False)
assert metrics["model_version"].startswith("hgb-v2+")
assert (tmp_path / "flood_P.1.joblib").exists()
def test_train_with_rain_series_yields_v3(tmp_path, monkeypatch):
from src.ml import rain as rain_mod
df = _p1_synth()
idx = pd.date_range(df["timestamp"].min(), df["timestamp"].max(), freq="h")
fake_rain = pd.DataFrame({"a": np.linspace(0, 1, len(idx)), "b": 0.5}, index=idx)
monkeypatch.setattr(rain_mod, "load_history", lambda *a, **k: fake_rain)
metrics = train.train_all(df, ["P.1"], models_dir=tmp_path, skip_eval=True, hgb_overrides={"max_iter": 10}, use_rain=True, use_dam=False)
assert metrics["model_version"].startswith("hgb-v3+")
bundle = joblib.load(tmp_path / "flood_P.1.joblib")
assert set(features.RAIN_FEATURES) <= set(bundle["feature_names"])
def test_cli_exit_code_on_rain_failure(tmp_path, monkeypatch, caplog):
"""The console entry turns the guard into a one-line error and exit 2."""
from src.ml import rain as rain_mod
df = _p1_synth()
monkeypatch.setattr(rain_mod, "load_history", lambda *a, **k: None)
monkeypatch.setattr(train, "load_measurements", lambda *a, **k: df)
monkeypatch.setattr(train, "resolve_db_url", lambda *a, **k: None)
monkeypatch.setattr(
"sys.argv",
["train", "--stations", "P.1", "--models-dir", str(tmp_path), "--skip-eval"],
)
assert train.cli() == 2
assert "refusing to silently downgrade" in caplog.text
assert not (tmp_path / "metrics.json").exists()
def test_feature_name_stability(tmp_path): def test_feature_name_stability(tmp_path):
upstream = [code for code, _lead in features.UPSTREAM_LEADS["P.1"]] upstream = [code for code, _lead in features.UPSTREAM_LEADS["P.1"]]
data_stations = ["P.1"] + upstream data_stations = ["P.1"] + upstream
df = make_synth(300, data_stations, seed=11, pulses={"P.1": [(100, 20, 2.0)]}) df = make_synth(300, data_stations, seed=11, pulses={"P.1": [(100, 20, 2.0)]})
train.train_all(df, ["P.1"], models_dir=tmp_path, skip_eval=True, hgb_overrides={"max_iter": 10}, use_rain=False) train.train_all(df, ["P.1"], models_dir=tmp_path, skip_eval=True, hgb_overrides={"max_iter": 10}, use_rain=False, use_dam=False)
# Safe: loading the bundle this same test just wrote to tmp_path, not an external file. # Safe: loading the bundle this same test just wrote to tmp_path, not an external file.
bundle = joblib.load(tmp_path / "flood_P.1.joblib") bundle = joblib.load(tmp_path / "flood_P.1.joblib")
+93
View File
@@ -0,0 +1,93 @@
"""Forecast skill verification: issued forecasts vs observed peaks (sqlite)."""
import datetime
import pytest
from sqlalchemy import create_engine, text
from src.ml import skill
@pytest.fixture
def engine(tmp_path):
eng = create_engine(f"sqlite:///{tmp_path / 'skill.db'}")
with eng.begin() as c:
c.execute(text("CREATE TABLE stations (id INTEGER PRIMARY KEY, station_code TEXT)"))
c.execute(text("INSERT INTO stations VALUES (1, 'P.1')"))
c.execute(
text(
"CREATE TABLE water_measurements (timestamp DATETIME, station_id INTEGER, water_level REAL)"
)
)
c.execute(
text(
"CREATE TABLE forecast_history (as_of TIMESTAMP, station_code TEXT, horizon_hours INTEGER, "
"predicted_max_level REAL, p_warning REAL, p_danger REAL, current_level REAL, "
"model_version TEXT, source TEXT)"
)
)
return eng
def _fill(engine, start, hours, level_fn, forecasts):
"""hours of hourly observations from `start`, plus (as_of_offset_h, version, pred) rows."""
with engine.begin() as c:
for h in range(hours):
ts = start + datetime.timedelta(hours=h)
c.execute(
text("INSERT INTO water_measurements VALUES (:t, 1, :l)"),
{"t": ts, "l": level_fn(h)},
)
for off, version, pred in forecasts:
ts = start + datetime.timedelta(hours=off)
c.execute(
text(
"INSERT INTO forecast_history VALUES (:t, 'P.1', 24, :p, 0, 0, :cur, :v, 'model')"
),
{"t": ts, "p": pred, "cur": level_fn(off), "v": version},
)
def test_skill_per_version_and_trend(engine):
start = datetime.datetime(2026, 8, 1)
# river: flat 1.5 m, with a bump to 2.4 m around hour 100
level = lambda h: 2.4 if 96 <= h <= 104 else 1.5
forecasts = []
# old version: always predicts 1.5 (persistence-like, misses the bump)
for off in range(0, 60):
forecasts.append((off, "hgb-v2+aaaaaaa", 1.5))
# new version: predicts 1.5 normally and 2.3 ahead of the bump
for off in range(60, 200):
pred = 2.3 if 72 <= off <= 104 else 1.5
forecasts.append((off, "hgb-v3+bbbbbbb", pred))
_fill(engine, start, 260, level, forecasts)
out = skill.compute_skill(engine, "sqlite", "P.1", 24, now=start + datetime.timedelta(hours=300))
assert [v["model_version"] for v in out["versions"]] == ["hgb-v2+aaaaaaa", "hgb-v3+bbbbbbb"]
old, new = out["versions"]
assert old["n"] == 60 and old["enough_data"]
assert new["n"] == 140 and new["enough_data"]
# the old version issued only on flat hours: perfect there, no bump rows
assert old["mae_m"] == 0.0 and old["above_2m_n"] == 0
# the new version saw the bump: nonzero MAE but positive skill vs persistence
assert new["above_2m_n"] > 0
assert new["skill"] is not None and new["skill"] > 0
assert out["current"]["model_version"] == "hgb-v3+bbbbbbb"
assert out["trend"]["previous_version"] == "hgb-v2+aaaaaaa"
assert out["trend"]["better"] is False # honest: old had an easier period
def test_skill_requires_full_window(engine):
start = datetime.datetime(2026, 8, 1)
# forecasts issued at the very end have no observed window yet
_fill(engine, start, 30, lambda h: 1.5, [(o, "hgb-v3+ccccccc", 1.5) for o in range(0, 30)])
out = skill.compute_skill(engine, "sqlite", "P.1", 24, now=start + datetime.timedelta(hours=30))
# only as_of <= now-24h AND with >= 18 observed hours in the window count
assert out["versions"] and out["versions"][0]["n"] == 7 # as_of 0..6 h: <= now-24h with >= 18 observed hours
assert out["versions"][0]["enough_data"] is False
assert out["trend"] is None
def test_skill_empty(engine):
out = skill.compute_skill(engine, "sqlite", "P.1", 24)
assert out["versions"] == [] and out["current"] is None and out["trend"] is None
+23
View File
@@ -353,6 +353,29 @@ class TestHiiApiEndpoints:
self._get(web_api, "get_hii_rainfall_latest", hours=48) self._get(web_api, "get_hii_rainfall_latest", hours=48)
assert calls["n"] == 2 assert calls["n"] == 2
def test_rainfall_catchment(self, web_api):
"""One gauge in the box (CHM005, 19.12N 98.94E) is below the
MIN_GAUGES_PER_HOUR floor, so the catchment mean is NaN -> null, the
openmeteo_rain table does not exist in this store, and the comparison
reports no overlap. Shape is what matters: the endpoint must not 500
on a fresh database."""
payload, response = self._get(web_api, "get_hii_rainfall_catchment", days=7)
assert "x-data-stale" not in response.headers
assert list(payload) == ["box", "gauge", "openmeteo", "comparison_24h_sums"]
assert payload["openmeteo"] == []
assert payload["comparison_24h_sums"] == {"overlap_hours": 0}
assert len(payload["gauge"]) == 1
assert payload["gauge"][0]["rain_mm"] is None # < MIN_GAUGES_PER_HOUR
def test_rainfall_catchment_disabled(self, monkeypatch):
from src import web_api
monkeypatch.setitem(web_api.app_state, "hii_collector", None)
web_api.HII_CACHE.clear()
web_api._REFRESH_IN_FLIGHT.clear()
payload, _ = self._get(web_api, "get_hii_rainfall_catchment", days=7)
assert payload["gauge"] == [] and payload["openmeteo"] == []
def test_stale_served_on_recompute_failure(self, web_api, monkeypatch): def test_stale_served_on_recompute_failure(self, web_api, monkeypatch):
# Prime the cache, expire it, break the DB: the stale copy is served # Prime the cache, expire it, break the DB: the stale copy is served
# and flagged via the X-Data-Stale header. # and flagged via the X-Data-Stale header.
+47
View File
@@ -0,0 +1,47 @@
"""HII gauge-rain aggregate: pure-function tests (no DB)."""
import numpy as np
import pandas as pd
from src.ml import hii_rain
def _hourly(start, n):
return pd.date_range(start, periods=n, freq="h")
def test_compare_identical_series_has_zero_bias():
idx = _hourly("2026-08-12", 200)
rng = np.random.default_rng(1)
rain = pd.Series(rng.exponential(0.5, len(idx)), index=idx)
out = hii_rain.compare_with_openmeteo(rain, rain.copy(), window_h=24)
assert out["overlap_hours"] == 200
assert out["bias_mm"] == 0.0
assert out["mae_mm"] == 0.0
assert out["corr"] > 0.999
def test_compare_reports_constant_bias():
idx = _hourly("2026-08-12", 100)
gauge = pd.Series(1.0, index=idx)
model = pd.Series(1.5, index=idx) # model wetter by 0.5 mm/h
out = hii_rain.compare_with_openmeteo(gauge, model, window_h=24)
assert abs(out["bias_mm"] - 12.0) < 1e-9 # 0.5 mm/h x 24 h
def test_compare_uses_overlap_only():
gauge = pd.Series(1.0, index=_hourly("2026-08-12", 100))
model = pd.Series(1.0, index=_hourly("2026-08-14", 100)) # 52 h overlap
out = hii_rain.compare_with_openmeteo(gauge, model, window_h=24)
assert out["overlap_hours"] == 52
def test_compare_no_overlap():
gauge = pd.Series(1.0, index=_hourly("2026-01-01", 10))
model = pd.Series(1.0, index=_hourly("2026-06-01", 10))
assert hii_rain.compare_with_openmeteo(gauge, model) == {"overlap_hours": 0}
def test_load_gauge_mean_without_db_returns_none(monkeypatch):
monkeypatch.setattr(hii_rain, "resolve_db_url", lambda *a, **k: None)
assert hii_rain.load_gauge_mean() is None
+285
View File
@@ -0,0 +1,285 @@
"""ntfy notification state machine: transitions only, hysteresis, restart-safe."""
import datetime
import pytest
from src import notify
class FakePublisher(notify.NtfyPublisher):
def __init__(self):
super().__init__("http://ntfy.test", prefix="ping")
self.sent = []
def publish(self, n):
self.sent.append(n)
return True
@pytest.fixture
def pub():
return FakePublisher()
def _reading(code, level, ts="2026-09-24T12:00:00"):
return {"station_code": code, "water_level": level, "timestamp": ts}
def _fc(p, peak=None):
return [
{
"station_code": "P.1",
"horizon_hours": 24,
"p_warning": p,
"predicted_max_level": peak,
"source": "model",
}
]
NOW = datetime.datetime(2026, 9, 24, 12, 30)
def topics(pub):
return [n.topic for n in pub.sent]
def test_quiet_river_sends_nothing(pub):
state = notify.InMemoryState()
for h in range(48):
notify.evaluate(
[_reading("P.1", 1.6), _reading("P.103", 3.2)],
_fc(0.01),
state,
pub,
now=NOW,
)
assert pub.sent == []
def test_warning_crossing_once_then_silence_then_clear(pub):
state = notify.InMemoryState()
# rising through 3.70 (P.1 warning)
notify.evaluate([_reading("P.1", 3.65)], [], state, pub, now=NOW)
assert pub.sent == []
notify.evaluate([_reading("P.1", 3.72)], [], state, pub, now=NOW)
assert topics(pub) == ["ping-p1-warning", "ping-warning"]
assert pub.sent[0].priority == 4 and "3.72 m" in pub.sent[0].message
# stays above: no repeats for many hours
for level in (3.80, 3.95, 4.05, 3.90, 3.75):
notify.evaluate([_reading("P.1", level)], [], state, pub, now=NOW)
assert len(pub.sent) == 2
# dips to 3.65: within hysteresis, still no message
notify.evaluate([_reading("P.1", 3.65)], [], state, pub, now=NOW)
assert len(pub.sent) == 2
# 3.55: clear
notify.evaluate([_reading("P.1", 3.55)], [], state, pub, now=NOW)
assert topics(pub)[2:] == ["ping-p1-warning", "ping-warning"]
assert "back to normal" in pub.sent[2].title
def test_danger_escalation_and_deescalation(pub):
state = notify.InMemoryState()
notify.evaluate([_reading("P.1", 3.9)], [], state, pub, now=NOW) # warning
notify.evaluate(
[_reading("P.1", 4.25)], [], state, pub, now=NOW
) # danger (>= 4.20)
assert topics(pub) == [
"ping-p1-warning",
"ping-warning",
"ping-p1-danger",
"ping-danger",
]
assert pub.sent[2].priority == 5
notify.evaluate(
[_reading("P.1", 4.15)], [], state, pub, now=NOW
) # hysteresis: still danger
assert len(pub.sent) == 4
notify.evaluate([_reading("P.1", 4.05)], [], state, pub, now=NOW) # back to warning
assert topics(pub)[4:] == ["ping-p1-danger", "ping-warning"]
assert "below danger" in pub.sent[4].title
def test_jump_straight_to_danger(pub):
state = notify.InMemoryState()
notify.evaluate(
[_reading("P.103", 7.0)], [], state, pub, now=NOW
) # P.103 danger 6.75
assert topics(pub) == ["ping-p103-danger", "ping-danger"]
def test_basin_digest_groups_stations(pub):
state = notify.InMemoryState()
notify.evaluate(
[_reading("P.1", 3.8), _reading("P.103", 6.0), _reading("P.67", 1.0)],
[],
state,
pub,
now=NOW,
)
basin = [n for n in pub.sent if n.topic == "ping-warning"]
assert len(basin) == 1 and "P.1" in basin[0].message and "P.103" in basin[0].message
def test_outlook_on_off_with_hysteresis(pub):
state = notify.InMemoryState()
r = [_reading("P.1", 2.9)]
notify.evaluate(r, _fc(0.30), state, pub, now=NOW)
assert pub.sent == []
notify.evaluate(r, _fc(0.55, 3.9), state, pub, now=NOW)
assert topics(pub) == ["ping-p1-outlook"]
assert "55%" in pub.sent[0].message and "3.90 m" in pub.sent[0].message
assert "not an official warning" in pub.sent[0].message
notify.evaluate(
r, _fc(0.40), state, pub, now=NOW
) # between OFF and ON: stays on, silent
assert len(pub.sent) == 1
notify.evaluate(r, _fc(0.20), state, pub, now=NOW)
assert len(pub.sent) == 2 and "easing" in pub.sent[1].title
def test_heuristic_forecast_ignored(pub):
state = notify.InMemoryState()
fc = [
{
"station_code": "P.1",
"horizon_hours": 24,
"p_warning": 0.9,
"source": "heuristic",
}
]
notify.evaluate([_reading("P.1", 2.0)], fc, state, pub, now=NOW)
assert pub.sent == []
def test_stale_feed_and_recovery(pub):
state = notify.InMemoryState()
notify.evaluate(
[_reading("P.1", 1.6, "2026-09-24T12:00:00")], [], state, pub, now=NOW
)
assert pub.sent == []
later = NOW + datetime.timedelta(hours=4)
notify.evaluate(
[_reading("P.1", 1.6, "2026-09-24T12:00:00")], [], state, pub, now=later
)
assert topics(pub) == ["ping-status"] and "stale" in pub.sent[0].title
notify.evaluate(
[_reading("P.1", 1.6, "2026-09-24T12:00:00")],
[],
state,
pub,
now=later + datetime.timedelta(hours=1),
)
assert len(pub.sent) == 1 # still stale, no repeat
notify.evaluate(
[_reading("P.1", 1.6, "2026-09-24T17:00:00")],
[],
state,
pub,
now=later + datetime.timedelta(hours=1),
)
assert len(pub.sent) == 2 and "recovered" in pub.sent[1].title
def test_capacity_guard_blocks_stale_threshold(pub):
"""P.77 2026-09: 3.02 m >= 2.85 m 'warning' at 22 % capacity -> not a flood."""
state = notify.InMemoryState()
r = {
"station_code": "P.77",
"water_level": 4.40,
"timestamp": "2026-09-24T12:00:00",
"discharge_percent": 10.3,
}
notify.evaluate([r], [], state, pub, now=NOW)
assert pub.sent == [] and state.get("level:P.77") is None
# same level with capacity agreeing -> alert
r["discharge_percent"] = 82.0
notify.evaluate([r], [], state, pub, now=NOW)
assert topics(pub) == ["ping-p77-warning", "ping-warning"]
def test_capacity_guard_exempts_p1_and_missing_pct(pub):
state = notify.InMemoryState()
notify.evaluate(
[
{
"station_code": "P.1",
"water_level": 3.75,
"timestamp": "2026-09-24T12:00:00",
"discharge_percent": 40.0,
}
],
[],
state,
pub,
now=NOW,
)
assert topics(pub) == ["ping-p1-warning", "ping-warning"]
pub.sent.clear()
notify.evaluate(
[
{
"station_code": "P.103",
"water_level": 6.0,
"timestamp": "2026-09-24T12:00:00",
}
],
[],
state,
pub,
now=NOW,
)
assert topics(pub) == ["ping-p103-warning", "ping-warning"]
def test_capacity_guard_does_not_block_clearing(pub):
"""Guard applies only to the clear->alert edge; the all-clear always goes out."""
state = notify.InMemoryState()
r = {
"station_code": "P.67",
"water_level": 2.6,
"timestamp": "2026-09-24T12:00:00",
"discharge_percent": 90.0,
}
notify.evaluate([r], [], state, pub, now=NOW)
assert len(pub.sent) == 2
r.update(water_level=2.2, discharge_percent=30.0)
notify.evaluate([r], [], state, pub, now=NOW)
assert "back to normal" in pub.sent[2].title
def test_state_survives_restart_via_sql(tmp_path, pub):
from sqlalchemy import create_engine
eng = create_engine(f"sqlite:///{tmp_path / 'n.db'}")
state = notify.NotificationState(eng, "sqlite")
notify.evaluate([_reading("P.1", 3.8)], [], state, pub, now=NOW)
assert len(pub.sent) == 2
# "restart": new state object on the same DB, same reading -> nothing re-sent
state2 = notify.NotificationState(eng, "sqlite")
notify.evaluate([_reading("P.1", 3.8)], [], state2, pub, now=NOW)
assert len(pub.sent) == 2
def test_publish_failure_does_not_advance_state():
"""If ntfy is down the transition must be retried next cycle, not lost."""
class Down(notify.NtfyPublisher):
def __init__(self):
super().__init__("http://ntfy.test")
self.calls = 0
def publish(self, n):
self.calls += 1
return False
pub = Down()
state = notify.InMemoryState()
notify.evaluate([_reading("P.1", 3.8)], [], state, pub, now=NOW)
assert pub.calls == 2 and state.get("level:P.1") is None
# next cycle, ntfy back: the crossing is delivered
good = FakePublisher()
notify.evaluate([_reading("P.1", 3.8)], [], state, good, now=NOW)
assert topics(good) == ["ping-p1-warning", "ping-warning"]
assert state.get("level:P.1") == "warning"
+406
View File
@@ -9,6 +9,8 @@ from src.rid_reservoir import (
RidReservoirCollector, RidReservoirCollector,
RidReservoirStore, RidReservoirStore,
backfill, backfill,
backfill_dam,
parse_dam_range_records,
parse_dam_records, parse_dam_records,
) )
@@ -56,6 +58,47 @@ def _dams_payload(date="2026-08-13"):
} }
def _range_row(date, **overrides):
"""One api/dam row, shaped like the real 2024-09-24 Mae Ngat response."""
row = {
"DAM_ID": "200103",
"DATE_curr": date,
"DAM_Name": "เขื่อนแม่งัดสมบูรณ์ชล",
"DAM_Region": "เหนือ",
"DAM_QMax": "323.00",
"DAM_QStore": "265.00",
"DAM_QUsage": "253.00",
"DUL_Useless": "12.00",
"DMD_ULevel_curr": "395.91",
"DMD_Q_curr": "0.00",
"DMD_Inflow_curr": "19.06",
"DMD_Outflow_curr": "0.13",
"VAL_DMD_Q_curr": "242.87",
"DMD_QUse_curr": "254.87",
"PERCENT_DMD_QUse_curr": "96.18",
# Same calendar date one year earlier -> must be ignored, not stored
"DMD_ULevel_prev": "391.52",
"DMD_QUse_prev": "195.64",
"PERCENT_DMD_QUse_prev": "73.83",
"DMD_Inflow_prev": "1.64",
"DMD_Outflow_prev": "3.79",
}
row.update(overrides)
return row
def _range_payload(dates=("2024-09-24",), rows=None):
return {
"dam_id": "200103",
"dam_name": "เขื่อนแม่งัดสมบูรณ์ชล",
"dam_region": "เหนือ",
"dam_coordinates": {"lat": 19.16138, "lng": 99.04011},
"dam_data": rows if rows is not None else [_range_row(d) for d in dates],
"max": {"max_DMD_QUse_curr": "254.87"},
"min": {"min_DMD_QUse_curr": "254.87"},
}
class FakeClient: class FakeClient:
def __init__(self, payload=None, fail_dates=()): def __init__(self, payload=None, fail_dates=()):
self.payload = payload or _dams_payload() self.payload = payload or _dams_payload()
@@ -71,6 +114,28 @@ class FakeClient:
return parse_dam_records(self.payload) return parse_dam_records(self.payload)
class FakeRangeClient:
"""Serves every day in the requested window, minus `gaps` (the real
endpoint omits scattered days rather than returning empty rows)."""
def __init__(self, gaps=(), fail_chunks=()):
self.gaps = {datetime.date.fromisoformat(d) for d in gaps}
self.fail_chunks = set(fail_chunks) # (start, end) tuples that raise
self.calls = []
def fetch_dam_range(self, dam_id, start, end):
self.calls.append((dam_id, start, end))
if (start, end) in self.fail_chunks:
raise ConnectionError("boom")
days = [
start + datetime.timedelta(days=i) for i in range((end - start).days + 1)
]
return parse_dam_range_records(
_range_payload(rows=[_range_row(d.isoformat()) for d in days
if d not in self.gaps])
)
class TestParsing: class TestParsing:
def test_parse_dam_records(self): def test_parse_dam_records(self):
records = parse_dam_records(_dams_payload()) records = parse_dam_records(_dams_payload())
@@ -88,6 +153,46 @@ class TestParsing:
def test_parse_empty_payload(self): def test_parse_empty_payload(self):
assert parse_dam_records({}) == [] assert parse_dam_records({}) == []
def test_parse_dam_range_records(self):
records = parse_dam_range_records(_range_payload(("2024-09-24",)))
assert len(records) == 1 # the _prev columns are last year, not a row
row = records[0]
assert row["dam_id"] == MAE_NGAT_DAM_ID
assert row["date"] == datetime.date(2024, 9, 24)
assert row["region"] == "เหนือ"
assert row["latitude"] == 19.16138
assert row["longitude"] == 99.04011
assert row["capacity_max_mcm"] == 323.0
assert row["capacity_normal_mcm"] == 265.0
# Agrees with api/dams for this date, at higher percent precision
assert row["storage_mcm"] == 254.87
assert row["storage_pct"] == 96.18
assert row["inflow_mcm"] == 19.06
assert row["outflow_mcm"] == 0.13
assert row["level_msl"] == 395.91 # DMD_ULevel, absent from api/dams
def test_parse_dam_range_records_produces_same_keys_as_dams(self):
assert set(parse_dam_range_records(_range_payload())[0]) == set(
parse_dam_records(_dams_payload())[0]
)
def test_parse_dam_range_zero_level_is_missing_not_a_reading(self):
payload = _range_payload(rows=[_range_row("2024-09-24", DMD_ULevel_curr="0.00")])
assert parse_dam_range_records(payload)[0]["level_msl"] is None
def test_parse_dam_range_skips_broken_rows(self):
payload = _range_payload(
rows=[
_range_row("n/a"), # unparseable date
_range_row("2024-09-24"),
]
)
assert len(parse_dam_range_records(payload)) == 1
def test_parse_dam_range_empty_payload(self):
assert parse_dam_range_records({}) == []
assert parse_dam_range_records(_range_payload(rows=[])) == []
class TestStore: class TestStore:
@pytest.fixture @pytest.fixture
@@ -123,6 +228,35 @@ class TestStore:
def test_save_empty(self, store): def test_save_empty(self, store):
assert store.save([]) == 0 assert store.save([]) == 0
def test_junk_source_values_are_nulled_not_fatal(self, store):
# Real junk from 2019-01-05: dam 100602 reported 87798% storage,
# which overflowed NUMERIC(6,2) and discarded the whole 33-dam batch.
payload = _dams_payload()
payload["regions"][0]["dams"].append(
{
"DAM_ID": "100602",
"DAM_Name": "junk",
"DMD_Date": "2026-08-13",
"DMD_QUse": "343292.00",
"PERCENT_DMD_QUse": "87798.00",
"DMD_Inflow": "0.59",
"DMD_Outflow": "1e12", # beyond NUMERIC(10,2) -> NULL
}
)
assert store.save(parse_dam_records(payload)) == 3
from sqlalchemy import text
with store.engine.begin() as conn:
row = conn.execute(
text(
"SELECT storage_mcm, storage_pct, outflow_mcm "
"FROM rid_reservoir_daily WHERE dam_id = '100602'"
)
).fetchone()
assert float(row[0]) == 343292.00 # fits NUMERIC(10,2), kept raw
assert float(row[1]) == 87798.00 # fits widened NUMERIC(8,2)
assert row[2] is None # beyond capacity -> NULL, batch survives
def test_present_dates(self, store): def test_present_dates(self, store):
lo, hi = datetime.date(2026, 8, 1), datetime.date(2026, 8, 31) lo, hi = datetime.date(2026, 8, 1), datetime.date(2026, 8, 31)
assert store.present_dates(lo, hi) == set() assert store.present_dates(lo, hi) == set()
@@ -131,6 +265,55 @@ class TestStore:
# Outside the window -> excluded # Outside the window -> excluded
assert store.present_dates(lo, datetime.date(2026, 8, 12)) == set() assert store.present_dates(lo, datetime.date(2026, 8, 12)) == set()
def test_daily_collector_does_not_blank_a_backfilled_level(self, store):
date = "2026-08-13"
assert store.save(parse_dam_range_records(_range_payload((date,)))) == 1
# api/dams has no level column at all; its rewrite must not clear one
dams_row = [r for r in parse_dam_records(_dams_payload(date))
if r["dam_id"] == MAE_NGAT_DAM_ID]
assert dams_row[0]["level_msl"] is None
store.save(dams_row)
from sqlalchemy import text
with store.engine.begin() as conn:
row = conn.execute(
text(
"SELECT level_msl, storage_mcm FROM rid_reservoir_daily "
"WHERE dam_id = :d"
),
{"d": MAE_NGAT_DAM_ID},
).fetchone()
assert float(row[0]) == 395.91 # kept
assert float(row[1]) == 222.01 # published columns still overwritten
def test_present_dates_per_dam(self, store):
lo, hi = datetime.date(2026, 8, 1), datetime.date(2026, 8, 31)
store.save(parse_dam_records(_dams_payload()))
day = datetime.date(2026, 8, 13)
assert store.present_dates(lo, hi, dam_id=MAE_NGAT_DAM_ID) == {day}
# Another dam having the date must not mark this one done
assert store.present_dates(lo, hi, dam_id="999999") == set()
def test_null_metadata_does_not_wipe_known_dam_details(self, store):
store.save(parse_dam_records(_dams_payload()))
blank = parse_dam_records(_dams_payload("2026-08-14"))
for record in blank: # a payload that omits metadata
record.update({"name_th": None, "latitude": None, "region": None})
assert store.save(blank) == 2
from sqlalchemy import text
with store.engine.begin() as conn:
row = conn.execute(
text(
"SELECT name_th, latitude, region FROM rid_dams "
"WHERE dam_id = :d"
),
{"d": MAE_NGAT_DAM_ID},
).fetchone()
assert row[0] == "เขื่อนแม่งัดสมบูรณ์ชล"
assert float(row[1]) == 19.16138
assert row[2] == "เหนือ"
class TestCollectorAndBackfill: class TestCollectorAndBackfill:
def test_run_cycle_today_and_yesterday(self, tmp_path): def test_run_cycle_today_and_yesterday(self, tmp_path):
@@ -216,3 +399,226 @@ class TestCollectorAndBackfill:
) )
assert saved == 4 # first 2 days succeeded, then 5 failures -> abort assert saved == 4 # first 2 days succeeded, then 5 failures -> abort
assert len(client.calls) == 7 assert len(client.calls) == 7
class TestBackfillDam:
@pytest.fixture
def store(self, tmp_path):
store = RidReservoirStore(f"sqlite:///{tmp_path}/range.db", "sqlite")
assert store.connect()
return store
def test_whole_range_in_one_request(self, store):
client = FakeRangeClient()
saved = backfill_dam(
store,
start=datetime.date(2024, 9, 24),
end=datetime.date(2024, 10, 6),
client=client,
throttle_seconds=0,
)
assert saved == 13
assert client.calls == [
(MAE_NGAT_DAM_ID, datetime.date(2024, 9, 24), datetime.date(2024, 10, 6))
]
def test_chunks_long_ranges(self, store):
client = FakeRangeClient()
saved = backfill_dam(
store,
start=datetime.date(2024, 1, 1),
end=datetime.date(2024, 1, 10),
client=client,
chunk_days=4,
throttle_seconds=0,
)
assert saved == 10
assert client.calls == [
(MAE_NGAT_DAM_ID, datetime.date(2024, 1, 1), datetime.date(2024, 1, 4)),
(MAE_NGAT_DAM_ID, datetime.date(2024, 1, 5), datetime.date(2024, 1, 8)),
(MAE_NGAT_DAM_ID, datetime.date(2024, 1, 9), datetime.date(2024, 1, 10)),
]
def test_source_gaps_are_tolerated(self, store):
client = FakeRangeClient(gaps=("2024-01-03",))
saved = backfill_dam(
store,
start=datetime.date(2024, 1, 1),
end=datetime.date(2024, 1, 5),
client=client,
throttle_seconds=0,
)
assert saved == 4 # the day the source never published stays absent
assert store.present_dates(
datetime.date(2024, 1, 1), datetime.date(2024, 1, 5)
) == {
datetime.date(2024, 1, d) for d in (1, 2, 4, 5)
}
def test_stored_days_are_skipped_and_whole_chunks_cost_no_request(self, store):
client = FakeRangeClient()
backfill_dam(
store,
start=datetime.date(2024, 1, 1),
end=datetime.date(2024, 1, 4),
client=client,
throttle_seconds=0,
)
# Rerun over a wider window: the stored chunk is not re-requested and
# only the missing days are written
saved = backfill_dam(
store,
start=datetime.date(2024, 1, 1),
end=datetime.date(2024, 1, 8),
client=client,
chunk_days=4,
throttle_seconds=0,
)
assert saved == 4
assert client.calls[1:] == [
(MAE_NGAT_DAM_ID, datetime.date(2024, 1, 5), datetime.date(2024, 1, 8))
]
def test_partly_stored_chunk_saves_only_missing_days(self, store):
client = FakeRangeClient()
backfill_dam(
store,
start=datetime.date(2024, 1, 3),
end=datetime.date(2024, 1, 3),
client=client,
throttle_seconds=0,
)
saved = backfill_dam(
store,
start=datetime.date(2024, 1, 1),
end=datetime.date(2024, 1, 5),
client=client,
throttle_seconds=0,
)
assert saved == 4 # day 3 already present, requested but not rewritten
def test_refresh_rewrites_stored_days(self, store):
client = FakeRangeClient()
window = dict(
start=datetime.date(2024, 1, 1),
end=datetime.date(2024, 1, 3),
client=client,
throttle_seconds=0,
)
assert backfill_dam(store, **window) == 3
assert backfill_dam(store, skip_present=False, **window) == 3
assert len(client.calls) == 2
def test_junk_values_are_nulled_not_fatal(self, store):
client = FakeRangeClient()
client.fetch_dam_range = lambda dam_id, start, end: parse_dam_range_records(
_range_payload(
rows=[
_range_row(
"2019-01-05",
DMD_QUse_curr="343292.00",
PERCENT_DMD_QUse_curr="87798.47",
DMD_Outflow_curr="1e12", # beyond NUMERIC(10,2) -> NULL
)
]
)
)
assert (
backfill_dam(
store,
start=datetime.date(2019, 1, 5),
end=datetime.date(2019, 1, 5),
client=client,
throttle_seconds=0,
)
== 1
)
from sqlalchemy import text
with store.engine.begin() as conn:
row = conn.execute(
text(
"SELECT storage_pct, outflow_mcm FROM rid_reservoir_daily "
"WHERE date = '2019-01-05'"
)
).fetchone()
assert float(row[0]) == 87798.47
assert row[1] is None
def test_stats_separate_an_empty_rerun_from_an_outage(self, store):
window = dict(
start=datetime.date(2024, 1, 1),
end=datetime.date(2024, 1, 3),
throttle_seconds=0,
)
backfill_dam(store, client=FakeRangeClient(), **window)
# Everything already stored: no request, no failure, still a success
stats = {}
assert backfill_dam(store, client=FakeRangeClient(), stats=stats, **window) == 0
assert stats == {"requests": 0, "failures": 0, "aborted": False}
# A source gap keeps requesting, but still reports no failure
gapped = FakeRangeClient(gaps=("2024-01-05",))
stats = {}
assert (
backfill_dam(
store,
client=gapped,
stats=stats,
start=datetime.date(2024, 1, 1),
end=datetime.date(2024, 1, 5),
throttle_seconds=0,
)
== 1
)
assert stats["requests"] == 1 and not stats["failures"]
def test_stats_record_failures(self, store):
chunk = (datetime.date(2024, 1, 1), datetime.date(2024, 1, 3))
client = FakeRangeClient(fail_chunks=[chunk])
stats = {}
assert (
backfill_dam(
store,
client=client,
stats=stats,
start=chunk[0],
end=chunk[1],
throttle_seconds=0,
)
== 0
)
assert stats["failures"] == 1 and not stats["aborted"]
def test_aborts_after_consecutive_failures(self, store):
start = datetime.date(2024, 1, 1)
chunks = [
(start + datetime.timedelta(days=i), start + datetime.timedelta(days=i))
for i in range(30)
]
client = FakeRangeClient(fail_chunks=chunks[1:])
stats = {}
saved = backfill_dam(
store,
start=start,
end=start + datetime.timedelta(days=29),
client=client,
chunk_days=1,
throttle_seconds=0,
stats=stats,
)
assert saved == 1 # first chunk succeeded, then 5 failures -> abort
assert len(client.calls) == 6
assert stats["aborted"] and stats["failures"] == 5
def test_aborts_when_db_saves_nothing(self, store, monkeypatch):
monkeypatch.setattr(store, "save", lambda records: 0) # broken DB
client = FakeRangeClient()
backfill_dam(
store,
start=datetime.date(2024, 1, 1),
end=datetime.date(2024, 3, 1),
client=client,
chunk_days=1,
throttle_seconds=0,
)
assert len(client.calls) == 5 # aborted, not one request per chunk
Generated
+129 -713
View File
File diff suppressed because it is too large Load Diff