Reference

Benchmarks

Every run loads the same randomized data into a single Logs process on MinIO, next to Loki, then sends both the same randomized LogQL queries. Each answer is checked for correctness and timed. The newest run is shown by default.

Runhistoricalvs Loki 3.5.5s3Full report ↗
Failed· 2 mismatch case(s)2026-10-08 20:32Z · 0b54fe965a (dirty) · simd-perf · 30 min · OrbStack · arm64 · 12 CPUs · 8.8 GiB
9,154
queries compared
2
mismatches
353 non-match total
10.1 ms
p50 query
11.9 ms Loki
252 ms
p99 query
539 ms Loki
0.54×
p50 latency ratio
per query, impl / oracle
0.23
mean CPU cores
0.40 Loki
Latency by percentile
Plural LogsLoki 3.5.5
Latency of every query both sides answered. Each query runs against both systems with the same data, so the percentiles compare like with like.
By query family
Query familyCasesp99 latencyp50 ratioNon-match
logql.metric.range3,231
341 ms
1.1 s
0.47×154
logql.metric.instant1,922
151 ms
195 ms
0.67×88
logql.log.parsed1,497
191 ms
402 ms
0.50×17
logql.log1,191
179 ms
385 ms
0.60×26
logql.log.parsed.limited348
126 ms
120 ms
0.67×7
logql.labels341
39.8 ms
28.2 ms
0.49×0
logql.label_values314
40.1 ms
28.6 ms
0.55×47
logql.log.limited310
210 ms
115 ms
0.86×14
Bars share one linear scale across families. The ratio is the median of per-query implementation / oracle latency, so values below 1× mean Plural answered faster.
Correctness
  • match8,801
  • both error223
  • inconclusive69
  • oracle unstable47
  • oracle error6
  • oracle timeout6
  • mismatch2
Inconclusive results are differences the API contract allows, such as both sides truncating at a limit.
Resources
CPU, mean cores
Plural
0.23
Loki
0.40
Memory, p95 working set
Plural
743 MiB
Loki
695 MiB
Sampled from the Docker Engine API every two seconds over the run. Object storage is excluded from both sides.
History · historical · Loki 3.5.5
Plural LogsLoki 3.5.5
Query latency for every run on this scenario and oracle. Click a point to open that run.

How runs are measured

The differential fuzzer in tests/regression/harness/fuzz runs both systems on one Docker host, each pinned to its own CPU set. A run fails on any mismatch, implementation error or timeout, unstable implementation answer, or rejected write. The recent scenario queries data still being ingested; historical queries older data that has been flushed to object storage, with artificial latency added to MinIO.

To record a new run from the repository root:

bash
PYTHONPATH=tests/regression mise exec -- python -m harness.fuzz.bench --products logs --duration 30m

Results land in documentation/benchmarks/fuzz and appear here on the next docs build.