Reference

Benchmarks

Every run loads the same randomized data into two writers and a reader on MinIO, next to Prometheus or Mimir, then sends both the same randomized PromQL queries. Each answer is checked for correctness and timed. The newest run is shown by default.

Runhistoricalvs Mimir store-gateway 3.2.1s3Full report ↗
Failed· 1 mismatch case(s)2026-10-08 03:54Z · 97bdc82e99 (dirty) · simd-perf · 30 min · OrbStack · aarch64 · 12 CPUs · 8.8 GiB
3,647
queries compared
1
mismatches
66 non-match total
2.33 ms
p50 query
27.3 ms Mimir
55.2 ms
p99 query
342 ms Mimir
0.10×
p50 latency ratio
per query, impl / oracle
0.03
mean CPU cores
0.53 Mimir
Latency by percentile
Plural MetricsMimir store-gateway 3.2.1
Latency of every query both sides answered. Each query runs against both systems with the same data, so the percentiles compare like with like.
By query family
Query familyCasesp99 latencyp50 ratioNon-match
promql.instant1,688
32.4 ms
247 ms
0.08×25
promql.range1,640
76.2 ms
460 ms
0.10×41
promql.series153
63.5 ms
76.1 ms
0.26×0
promql.label_values150
26.5 ms
145 ms
0.16×0
promql.labels16
20.3 ms
46.7 ms
0.14×0
Bars share one linear scale across families. The ratio is the median of per-query implementation / oracle latency, so values below 1× mean Plural answered faster.
Correctness
  • match3,581
  • both error43
  • inconclusive21
  • mismatch1
  • oracle error1
Inconclusive results are differences the API contract allows, such as both sides truncating at a limit.
Resources
CPU, mean cores
Plural
0.03
Mimir
0.53
Memory, p95 working set
Plural
379 MiB
Mimir
1,271 MiB
Sampled from the Docker Engine API every two seconds over the run. Object storage is excluded from both sides.
History · historical · Mimir store-gateway 3.2.1
Plural MetricsMimir store-gateway 3.2.1
Only one run has been recorded on this scenario and oracle so far. Later runs will chart here.

How runs are measured

The differential fuzzer in tests/regression/harness/fuzz runs both systems on one Docker host, each pinned to its own CPU set. A run fails on any mismatch, implementation error or timeout, unstable implementation answer, or rejected write. The recent scenario queries data still being ingested; historical queries older data that has been flushed to object storage, with artificial latency added to MinIO.

To record a new run from the repository root:

bash
PYTHONPATH=tests/regression mise exec -- python -m harness.fuzz.bench --products metrics --duration 30m

Results land in documentation/benchmarks/fuzz and appear here on the next docs build.