Files
grabowski a0086086a2
CI/CD Pipeline - Northern Thailand Ping River Monitor / Test Suite (3.11) (push) Failing after 24s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Build Docker Image (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Integration Test with Services (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Staging (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Performance Test (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Code Quality (push) Successful in 13s
Documentation / Validate Documentation (push) Failing after 8s
Documentation / Generate API Documentation (push) Successful in 8s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Production (push) Skipped
Documentation / Build Sphinx Documentation (push) Successful in 15s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Cleanup (push) Successful in 1s
Documentation / Documentation Summary (push) Successful in 2s
feat: rolling-origin event-aware evaluation harness for model variants
One fold per monsoon season (train <= 30 Apr, test Jun-Nov, 2021-2025)
replaces the single fixed holdout that contained only ~4 warning events.
Metrics are what matters operationally: sustained first-alert lead vs
each observed 3.70m crossing (two consecutive alerting samples required;
lookback floored at the previous event's end so multi-peak floods can't
launder lead credit), peak error from the prediction actually issued 24h
before the peak (3h match tolerance, null on outages), false-alarm
episodes (12h gap tolerance), MAE / flood-regime MAE, and a Brier score
on warning exceedance — included because sigma cancels algebraically in
any p>=0.5 alert metric, so lead times compare predictors while Brier
compares uncertainty models.

Variants: baseline_abs (current), rise (target = future max - current
level), rise_weighted (flood-regime sample weights 1x->5x), and
rise_quantile (q50/q90 heads, spread-implied sigma). Harness verified by
a 3-agent adversarial review (features bit-identical across fold
cutoffs; three metric flaws found and fixed before first use).

Also: features.build_labels/build_matrix gain stats_end so the rescue
quantile is computed from pre-cutoff data only, closing the label-
construction leak flagged in the earlier ML review.
2026-08-12 15:19:26 +07:00

19 lines
485 B
Python

#!/usr/bin/env python3
"""CLI for the rolling-origin model-variant evaluation.
Usage:
uv run scripts/evaluate_variants.py # P.1, all variants
uv run scripts/evaluate_variants.py --stations P.1,P.103
uv run scripts/evaluate_variants.py --variants baseline_abs,rise_quantile
"""
import os
import sys
sys.path.insert(0, os.path.join(os.path.dirname(__file__), ".."))
from src.ml.evaluate import main
if __name__ == "__main__":
sys.exit(main())