Benchmarks
Every run loads the same randomized data into two writers and a reader on MinIO, next to Prometheus or Mimir, then sends both the same randomized PromQL queries. Each answer is checked for correctness and timed. The newest run is shown by default.
| Query family | Cases | p99 latency | p50 ratio | Non-match |
|---|---|---|---|---|
| promql.instant | 1,688 | 32.4 ms 247 ms | 0.08× | 25 |
| promql.range | 1,640 | 76.2 ms 460 ms | 0.10× | 41 |
| promql.series | 153 | 63.5 ms 76.1 ms | 0.26× | 0 |
| promql.label_values | 150 | 26.5 ms 145 ms | 0.16× | 0 |
| promql.labels | 16 | 20.3 ms 46.7 ms | 0.14× | 0 |
- match3,581
- both error43
- inconclusive21
- mismatch1
- oracle error1
How runs are measured
The differential fuzzer in tests/regression/harness/fuzz runs both systems on one Docker host, each pinned to its own CPU set. A run fails on any mismatch, implementation error or timeout, unstable implementation answer, or rejected write. The recent scenario queries data still being ingested; historical queries older data that has been flushed to object storage, with artificial latency added to MinIO.
To record a new run from the repository root:
PYTHONPATH=tests/regression mise exec -- python -m harness.fuzz.bench --products metrics --duration 30mResults land in documentation/benchmarks/fuzz and appear here on the next docs build.