src/static/custom-markers.json holds user points of interest
({name, lat, lon, optional note}); the dashboard renders them as pin
markers whose popups show which inundation zone the point sits in, its
P.1 trigger level, and the live 24 h exceedance probability (ray-cast
point-in-polygon against the zone geojson; smallest matching zone wins).
Seeded with the owner's four properties.
The 7 Chiang Mai inundation zones are now hand-traced by the project
owner against the real basemap (cnx_flood.geojson), replacing the
scan-digitized approximation and its georeferencing error entirely.
Features are ordered zone 7 -> 1 so the earlier-flooding (smaller)
zones render on top of the wider extents in Leaflet.
The P.1 marker (Nawarat Bridge, 18.7875) sat inside zone 2 near its top
edge, but the bridge IS Chang Khlan's northern boundary - the layer was
~0.0095 deg too far north (the district-centroid correction in c4e0fb6
overshot). Zone 2's north edge now lands at the bridge.
The previous affine fit slid along the mostly north-south river (an
ill-conditioned direction) leaving the zones ~1 km south-west of their
true position. New approach: extract the scanned river band by color,
match it to the OSM Ping mainstem centerline by arc length (which pins
the along-river position), then correct residual translation using
known district locations cross-checked against the river residual
(533 m -> ~300 m median; Chang Khlan zone centroid now within 200 m).
Mountain-area artifacts are clipped away via the municipality hull.
Zones were also nearly invisible at fillOpacity 0.16 - base opacity is
now 0.38, rising with the live 24 h exceedance probability, and 0.72
with a solid red border once the river is at or above a zone's level.
Digitize the 7 Chiang Mai inundation zones from the municipal map into
geojson polygons: pixels classified by legend color, vectorized via
marching squares, and georeferenced by a 6-parameter affine fitted
least-squares to the OSM Ping centerline (109 m median residual after
outlier-filtered refits - the scanned image overlay could never scale
correctly and is removed).
The dashboard renders the zones as a Leaflet geoJSON layer styled live
from the forecast: fill opacity scales with each zone's 24 h exceedance
probability, zones whose trigger level the river has already reached get
a solid red border, and each polygon's popup shows its trigger level and
live probability. Zone styles refresh with every forecast load.
From the adversarial review and threshold backtest (swarm verification):
- predict.py: when a bundle's trained thresholds differ from the current
config (deploy before retrain), skip its stale classifier heads and
derive p_warning/p_danger from the regression + sigma against the
CURRENT thresholds - the dashboard can no longer show contradictory
old-threshold classifier output next to new-threshold stages
- features.py: decouple the low-coverage regression-label rescue from
the warning threshold (now the station's own p97.5 level); the old
coupling silently dropped 34% of P.5's regression training rows and
cost +46% MAE when its threshold rose
- features.py: P.82 danger 3.80 -> 3.75 (3.80 was above the station's
8-year maximum of 3.78, so danger could never train or fire)
- data.py / predict.py: anchor models/cache paths to the repo root; the
relative paths silently returned zero rows when run from another CWD
- annotate P.4A thresholds as low-confidence (11 supporting readings)
47 tests pass. Retrain required for the label-rescue and P.82 changes
to reach the classifier heads.
The push-CI gates (black/isort/mypy) had never actually run before the
branch-trigger fix, and the codebase predates them. Formatting is now
black/isort clean repo-wide. mypy keeps running but non-blocking: 86
pre-existing errors are a separate cleanup, not a gate to hold hostage.
Workflows listened on 'main'/'develop' but the repo's default branch is
master, so push events never started a job (every historical run is a
schedule event). Point push/PR triggers and the deploy-gate refs at
master, and trim the test matrix to 3.11/3.12 to match
requires-python >=3.11.
Replace the network-wide (3.0, 4.5) m thresholds with per-station values
calibrated from the DB's discharge_percent (RID % of channel capacity):
warning = median level at 75-85% capacity, danger = median at 95-105%.
Fixes P.103 over-alerting (bank-full ~6.75 m, not 4.5) and P.67
under-alerting (overflow ~2.9 m). Requires a retrain to take effect in
the classifier heads.
P.1 uses the official Chiang Mai municipal inundation map instead:
warning 3.70 m (stage 1, city flooding begins), danger 4.20 m (stage 5),
with the full 7-stage table (3.70-4.60 m + discharge) in
features.P1_FLOOD_STAGES. Forecast rows for P.1 now include per-stage
exceedance probabilities computed from the regression head + calibration
sigma - available immediately without retraining.
Dashboard: "Chiang Mai city flood outlook" block above the forecast grid
(predicted peak + 7 stage-probability chips) and a toggleable
georeferenced overlay of the official flood-zone map
(static/flood-zones-p1.jpg, bounds tunable in FLOOD_ZONE_BOUNDS).
scikit-learn 1.9.0 supports only Python >=3.11, so uv could not resolve
the >=3.9 range. The deployment runs 3.11. Also pins numpy back to
1.26.4 in the lockfile (pandas 2.0.3 ABI).
Add src/ml/ package predicting, per station and per 6/12/24 h horizon,
the probability of exceeding warning (3.0 m) and danger (4.5 m) levels
plus expected peak level, trained on the 592k-row PostgreSQL history:
- features.py: hourly grid with coverage gating and no future leakage;
upstream stations enter at empirically measured travel-time lags
(P.20 +17h ... P.103 +1h vs P.1); hour-of-day deliberately excluded
(it encodes the scrape schedule, not hydrology)
- train.py: HistGradientBoosting regression + warn/danger classifier
heads per station x horizon, >=30-positives gate with calibrated
sigmoid-on-regression fallback, strict temporal splits, per-event
lead-time evaluation; guards against sklearn 1.9.0 crash on
degenerate feature columns
- predict.py: bundle loading with feature-name checks, heuristic
fallback tier, get_latest_forecasts() for the API; raises when no
models are trained so the endpoint 503s instead of serving
persistence output as forecasts
- data.py: Postgres-first loader (FLOOD_ML_DB_URL override), HTTP API
fallback (flagged: that path backfills synthetic discharge), csv.gz
cache
- /forecast endpoint (15-min TTL cache) + dashboard flood-risk panel
(hidden until models exist)
- docs/FLOOD_FORECASTING.md: full system doc with measured deployment
numbers (~335 MB RSS, CPU negligible, ~6 min full retrain) and
retraining policy
Validation: out-of-sample backtest of the record 2024 flood season
(train <= Aug 2024) alerted 24-48 h ahead of the Oct 5 peak; 2025-26
test split: P.1 6h PR-AUC 0.974, recall 98.3% at 1% false-alarm rate.
Also: fix P.81 station coordinates (was Ban Pong/Ratchaburi, 493 km
out of basin; now 18.6936 N 99.0819 E per RID station page), pin
scikit-learn==1.9.0 and numpy<2, gitignore model artifacts (~100 MB,
train on the server via scripts/train_flood_model.py).
Station selection showed no history since 21ca844: Chart.js v4 datasets
had parsing:false with plain number arrays, drawing empty axes. Remove
the flag so the chart parses values again.
Backend hardening for the same flow:
- /measurements/history/{code} no longer 503s on non-Postgres configs;
it falls back to the configured adapter (reversed to ascending order)
- DB_TYPE defaults to postgresql when POSTGRES_CONNECTION_STRING is set
and DB_TYPE is unset, so the .env psql wins over the sqlite default
- zero readings (0.0) are no longer coerced to None, which would fail
MeasurementResponse validation and 500 /measurements/latest
Map visualization:
- river segments are now colored, widened and dash-speed-animated by
the discharge at the nearest gauge (same scale as the marker legend)
- fix z-order bug that drew the animated flow line behind its casing
- legend entries for river lines, reduced-motion fallback
River geometry: rebuild ping-river-network.geojson from Overpass
(110 -> 202 features), restoring missing Ping mainstem reaches through
the Bhumibol reservoir and the Tak-Kamphaeng Phet braided section
(unnamed waterway=river ways in OSM), with short synthetic connectors
(connector: true) bridging remaining sub-8 km holes.
Extract the inline root() dashboard markup into src/static/dashboard.html,
loaded once at import. Keeps HTML out of the Python module (web_api 616 -> 576
lines) with a defensive fallback if the file is missing.
Full <500 compliance for web_api still needs the endpoints split into
APIRouter modules; tracked as remaining #4 work.
Cover the previously-untested, highest-risk parsing in
fetch_water_data_for_date by mocking the HTTP call:
- hour 1-23 map to the same day; hour 24 rolls to next-day midnight
- qvalues "***" / None yield discharge None (no crash)
- None water level is skipped
- out-of-range (0, 25) and empty hourlytime rows are skipped
- missing "rows" key returns an empty list
This gives the scraper a safety net before its module is split.
Move the seven request/response models out of web_api.py into a dedicated
src/schemas.py (separation of concerns; first step of the file-size cleanup).
web_api.py imports them back, so behaviour is unchanged.
Note: web_api.py is still over the 500-line guideline; the remaining bulk is
the inline HTML dashboard in root(), to be extracted in a follow-up.
- tests/conftest.py: put repo root on sys.path so `import src...` resolves
under pytest regardless of invocation directory.
- test_matrix_formatting.py: lock in HTML formatted_body + plain-text fallback,
URL linkification, HTML escaping, and send_alert field rendering.
- test_station_persistence.py: cover default-load, save/reload round-trip
(incl. Thai text), runtime-file precedence, and atomic-write cleanup.
These are real assert-based tests (unlike the existing print-style scripts) so
CI can gate on them. 13 tests, all passing.
- .env now chmod 0600 and APP_DIR chmod 0750 after chown, so the Matrix token
and DB credentials are not world-readable.
- uv auto-install (curl | sh as root) is now opt-in via AUTO_INSTALL_UV=1 and
pins a specific uv version; otherwise the script requires uv to be
pre-installed and fails with instructions, avoiding unattended remote code
execution as root.
- scripts/install.sh: one-command hardened deploy (creates the water-monitor
system user, deploys to /opt, builds a uv-managed venv, installs and enables
the systemd unit). Idempotent; excludes .env/*.db/stations.json from sync so
runtime state is preserved.
- Fix placeholder Documentation= URL in water-monitor.service.
- README: document the script as the primary systemd install path, with manual
steps kept as a fallback.
Station CRUD via the API previously mutated the scraper's in-memory
station_mapping only, so changes were lost on restart (and the systemd
service auto-restarts).
- Extract the 130-line hardcoded station_mapping into bundled defaults at
src/data/stations.json; the scraper loads from a runtime-writable config
file (STATION_CONFIG_PATH, default stations.json) and falls back to the
bundled defaults to seed it.
- Add scraper.save_stations() with an atomic temp-file + os.replace write.
- create/update/delete station endpoints now persist and roll back the
in-memory change if the write fails; re-raise HTTPException so persistence
errors surface as real 500s instead of being swallowed.
- Backend-agnostic (works for the VictoriaMetrics deployment, which has no
relational stations table). Runtime stations.json is gitignored.
Also clears pre-existing flake8 debt in water_scraper_v3.py (unused imports,
long lines, duplicate logging import) and dedupes the User-Agent to
Config.USER_AGENT.
- MeasurementResponse.discharge is now Optional[float]; measurements with a
null discharge no longer raise a Pydantic ValidationError (HTTP 500) on the
/measurements/latest and /measurements/station endpoints.
- InfluxDB save_measurements guards float(discharge) against None instead of
crashing with TypeError.
- Extract the duplicated measurement->response mapping into a single
_to_measurement_response helper used by both measurement endpoints.
Matrix alerts:
- Send HTML formatted_body (org.matrix.custom.html) so **bold** and URLs
render instead of showing literal Markdown; add plain-text body fallback.
Add dependency-free markdown_to_matrix_html/strip_markdown helpers with
HTML escaping of station/message data.
Security:
- InfluxDB: bind untrusted station_codes as query params and cast limit to
int (was f-string interpolation / injection risk).
- VictoriaMetrics: escape Prometheus label values and coerce metric values
to float, preventing exposition-format injection and None crashes.
- web_api: run blocking scrape cycle via run_in_executor so it no longer
freezes the event loop; make CORS origins configurable and only allow
credentials with explicit origins ("*" + credentials is invalid/unsafe).
- config: remove hardcoded root/postgres password fallbacks (raise instead)
and stop defaulting VM_HOST to a real infrastructure hostname.
Also remove unused imports and wrap long lines to satisfy flake8.
- Stale data alerts now only trigger after 12 hours without new data
- Reduces false alerts during expected data gaps
Co-Authored-By: Claude <noreply@anthropic.com>
- Replaced deprecated tool.uv.dev-dependencies with dependency-groups.dev
- Follows new uv standard for dependency group declaration
Co-Authored-By: Claude <noreply@anthropic.com>
- Added Black for code formatting (line-length 120)
- Added isort for import sorting
- Added flake8 for linting
- Added standard pre-commit hooks for file checks
Co-Authored-By: Claude <noreply@anthropic.com>
- Alerts now run automatically after every successful new data fetch
- Works for both hourly fetches and retry mode exits
- Alert check runs when fresh data is saved to database
- Logs alert results (total generated and sent count)
Co-Authored-By: Claude <noreply@anthropic.com>
- Created test suite for zone-based water level alerts (9 test cases)
- Created test suite for rate-of-change alerts (5 test cases)
- Created combined alert scenario test
- Fixed rate-of-change detection to use station_code instead of station_id
- All 3 test suites passing (14 total test cases)
Test coverage:
- Zone alerts: P.1 zones 1-8 with INFO/WARNING/CRITICAL/EMERGENCY levels
- Rate-of-change: 0.15/0.25/0.40 m/h thresholds for WARNING/CRITICAL/EMERGENCY
- Combined: Simultaneous zone and rate-of-change alert triggering
Co-Authored-By: Claude <noreply@anthropic.com>
- Implement check_rate_of_change() to detect rapid water level rises
- Monitor water level changes over configurable lookback period (default 3 hours)
- Define rate-of-change thresholds for P.1 and other stations
- Alert on moderate (15cm/h), rapid (25cm/h), and very rapid (40cm/h) rises
- Only alert on rising water levels (positive rate of change)
- Integrate rate-of-change checks into run_alert_check() cycle
- Support both SQLite and PostgreSQL database adapters with fallback
Rate thresholds for P.1 (Nawarat Bridge):
- Warning: 0.15 m/h (15 cm/hour) - moderate rise
- Critical: 0.25 m/h (25 cm/hour) - rapid rise
- Emergency: 0.40 m/h (40 cm/hour) - very rapid rise
Default thresholds for other stations:
- Warning: 0.20 m/h, Critical: 0.35 m/h, Emergency: 0.50 m/h
Alert messages include:
- Rate of change in m/h and cm/h
- Total level change over period
- Time period analyzed
This early warning system detects dangerous trends before absolute
thresholds are reached, allowing for earlier response to flooding events.
🤖 Generated with [Claude Code](https://claude.ai/code)
Co-Authored-By: Claude <noreply@anthropic.com>
- Change HTTP method from POST to PUT for Matrix API v3
- Matrix API requires PUT when transaction ID is included in URL path
- Move transaction ID construction before URL building for clarity
- Fixes "405 Method Not Allowed" error when sending notifications
The Matrix API v3 endpoint structure:
PUT /_matrix/client/v3/rooms/{roomId}/send/{eventType}/{txnId}
Previous error:
POST request was being rejected with 405 Method Not Allowed
Now working:
PUT request successfully sends messages to Matrix rooms
🤖 Generated with [Claude Code](https://claude.ai/code)
Co-Authored-By: Claude <noreply@anthropic.com>
- Add 8 water level zones plus NewEdge threshold for P.1 station
- Zone 1: 3.7m (Info), Zone 2: 3.9m (Info)
- Zone 3-5: 4.0-4.2m (Warning levels)
- Zone 6-7: 4.3-4.6m (Critical levels)
- Zone 8/NewEdge: 4.8m (Emergency level)
- Implement special zone-based checking logic for P.1
- Maintain backward compatibility with standard warning/critical/emergency thresholds
- Keep standard threshold checking for other stations
Zone progression for P.1:
- 3.7m: Zone 1 alert (Info)
- 3.9m: Zone 2 alert (Info)
- 4.0m: Zone 3 alert (Warning)
- 4.1m: Zone 4 alert (Warning)
- 4.2m: Zone 5 alert (Warning)
- 4.3m: Zone 6 alert (Critical)
- 4.6m: Zone 7 alert (Critical)
- 4.8m: Zone 8/NewEdge alert (Emergency)
🤖 Generated with [Claude Code](https://claude.ai/code)
Co-Authored-By: Claude <noreply@anthropic.com>