perf: single-flight /api/stats; fast-fail health probe
CI/CD Pipeline - Northern Thailand Ping River Monitor / Test Suite (3.11) (push) Failing after 25s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Build Docker Image (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Integration Test with Services (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Staging (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Production (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Performance Test (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Code Quality (push) Successful in 13s
Documentation / Generate API Documentation (push) Successful in 10s
Documentation / Build Sphinx Documentation (push) Successful in 16s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Cleanup (push) Successful in 1s
Documentation / Documentation Summary (push) Successful in 4s
Documentation / Validate Documentation (push) Failing after 8s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Test Suite (3.11) (push) Failing after 25s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Build Docker Image (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Integration Test with Services (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Staging (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Deploy to Production (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Performance Test (push) Skipped
CI/CD Pipeline - Northern Thailand Ping River Monitor / Code Quality (push) Successful in 13s
Documentation / Generate API Documentation (push) Successful in 10s
Documentation / Build Sphinx Documentation (push) Successful in 16s
CI/CD Pipeline - Northern Thailand Ping River Monitor / Cleanup (push) Successful in 1s
Documentation / Documentation Summary (push) Successful in 4s
Documentation / Validate Documentation (push) Failing after 8s
The on-box rerun after the inline fast path showed the cached endpoints healthy (latest p50 72-160ms, HII ~100ms) but /api/stats at 81% timeouts and /health at 33%: stats had a cache but NO single-flight, so every concurrent miss ran the heavy whole-DB counts (~1.7M rows) in parallel, re-jamming Postgres and the executor — which also dragged uncached history windows into 60s timeouts. /api/stats now computes through _ttl_cached_stale (one computation per 5min TTL, stale served on failure). The health API probe fails fast (5s instead of 30s) and its cache TTL rises to 30s, so a slow upstream can no longer pin executor threads longer than the cache lifetime.
This commit is contained in:
@@ -140,6 +140,10 @@ class APIHealthCheck(HealthCheck):
|
||||
|
||||
def __init__(self, api_url: str, session, name: str = "api"):
|
||||
super().__init__(name)
|
||||
# A liveness probe should fail fast: the default 30s timeout meant a
|
||||
# slow upstream pinned executor threads for longer than the /health
|
||||
# cache TTL, so the pool never drained under load.
|
||||
self.timeout_seconds = 5
|
||||
self.api_url = api_url
|
||||
self.session = session
|
||||
|
||||
|
||||
Reference in New Issue
Block a user