Architecture
A Prometheus-compatible time-series database on SlateDB. It takes remote write and OTLP, answers PromQL, and runs as two StatefulSets and a bucket.
- 2
- component kinds
- 6
- pods
- 0
- replicated disks
- 0
- hash rings
More writers means more shards from the next hour on. Nothing is copied or rebalanced, and readers scale on their own.
- 10
- component kinds
- 48
- pods
- 15 PVCs
- replicated disks
- 4
- hash rings
Each series lives in memory on three ingesters until a 2-hour block is cut and uploaded. Adding ingesters reshards series via the ring; store-gateways reshard blocks via their own ring.
Write path
- Prometheus, Alloy or an OTel Collector sends a batch to any writer, over Remote Write 1.0 or 2.0 or OTLP.
- Each series is hashed by its namespace and sorted labels to a storage shard, using the routing epoch in effect at each sample's timestamp. Samples for other shards are forwarded to their owner over internal gRPC.
- The owner buffers samples in memory and acknowledges at the configured durability.
- On flush, samples are appended to their series in the current hour bucket as merge operands, alongside any new series dictionary and index entries, in one atomic SlateDB write. The bucket's write generation is bumped.
- SlateDB compaction collapses the merge operands of each series into compressed sample records.
Read path
- Grafana sends PromQL to any reader. Readers open every shard read-only.
- The planner narrows the hour buckets by time, then resolves label matchers against each bucket's Roaring-bitmap postings, smallest first. Resolved matchers are cached per bucket.
- Only the matching series' sample records are loaded, through the block cache.
- Series are joined across shards by label fingerprint and evaluated.
- Range-query steps are cached and reused while the write generation of every bucket they read is unchanged, so dashboards refreshing over old data skip storage entirely.
Storage layout
Each storage shard is one SlateDB database. Keys are scoped by tenant namespace and hour bucket:
Every key starts with this prefix, so each tenant namespace and hour bucket is one contiguous key range. Retention drops whole ranges, and queries scan only the ranges in their time window. The record type comes next:
Indexes are bucket-local. A query over a long range repeats lookups per bucket, and old buckets stay isolated from series churn in the current one. Retention drops whole buckets.
Compared with Mimir
| Concern | Mimir | Plural Metrics |
|---|---|---|
| Recent data | TSDB head in ingester memory, 3× replicated, cut into 2-hour blocks | Writer buffer, then SlateDB on object storage |
| Write fan-out | 3× to ingesters via the hash ring | 1× to the shard owner |
| Historical reads | Store-gateways shard blocks via their own ring | Readers open every shard; one block cache |
| Compaction | Separate compactor ring, block splitting | SlateDB's built-in compactor, per shard |
| Query tiers | Frontend, scheduler, queriers, store-gateways, caches | Readers |
Compatibility
PromQL is checked against the upstream Prometheus promqltest corpus. Unsupported constructs return an error rather than a wrong answer. Supported: native histograms, out-of-order samples, staleness markers, offset and @, subqueries, every stable aggregation and function, and /federate. Not supported:
- Exemplars are dropped, and remote read (
/api/v1/read) is not served. - The lookback delta is fixed at 5 minutes;
timeout,limitandstatsare ignored. - There is no scraper, rule evaluator or alerting. Run Prometheus or Alloy in agent mode, and evaluate rules elsewhere.
- Experimental upstream functions such as
limitk,mad_over_timeandinfo.