Architecture
A Loki-compatible log database on SlateDB. It speaks LogQL and Loki push, and also takes OTLP and Elasticsearch _bulk writes, from two StatefulSets and a bucket.
- 2
- component kinds
- 7
- pods
- 0
- replicated disks
- 0
- hash rings
More writers means more shards from the next hour on. Nothing is copied or rebalanced, and readers scale on their own.
- 9
- component kinds
- 57
- pods
- 18 PVCs
- replicated disks
- 4
- hash rings
New ingesters join the hash ring and take over token ranges; every write fans out to three of them. Scaling down means flushing and handing off in-memory chunks and WALs one ingester at a time.
Write path
- An agent sends a batch to any writer: Loki push, OTLP, or Elasticsearch
_bulk. - Each entry's stream is hashed to a storage shard with the routing epoch in effect at its timestamp. Entries for other shards are forwarded to their owner over internal gRPC.
- The owner buffers entries in memory and acknowledges at the configured durability (
appliedby default). - Every 10 seconds, or at 64 MiB, the buffer is flushed. The flush packs every stream of a time segment into ~1 MiB multi-stream objects, updates label and term postings, and commits it all to SlateDB in one atomic write.
- Writer-side compaction later merges small objects of a segment, so long-lived streams end up in few objects.
Read path
- Grafana sends LogQL to any reader. Readers open every shard read-only and share one block cache.
- For each time segment in range, label matchers intersect Roaring-bitmap postings to find stream IDs.
- Each stream's runs carry their min and max timestamps, so out-of-range objects are rejected before any block is read.
- The surviving blocks are range-read through the RAM + NVMe cache and decoded, then the LogQL pipeline runs over the rows.
- Results from every shard are merged by timestamp and cut at the query's
limit.
| match "…", a Logs extension, uses the full-text term index to nominate rows before decoding. Loki-style line filters (|=, |~) still scan decoded rows, since they must match substrings.
Storage layout
Each storage shard is one SlateDB database. Keys are scoped by tenant namespace and time segment:
Every key starts with this prefix, so each tenant namespace and segment is one contiguous key range. Retention drops whole ranges, and queries scan only the ranges in their time window. The record type comes next:
| matchRows of one stream are stored as independently decoded blocks of 256 rows. Replaced objects stay readable for 10 minutes after compaction, so in-flight queries never lose data underneath them.
Compared with Loki
| Concern | Loki | Plural Logs |
|---|---|---|
| Recent data | Replicated ingester memory and WAL | Writer buffer, then SlateDB on object storage |
| Write fan-out | 3× to ingesters via the hash ring | 1× to the shard owner |
| Index | TSDB index, index gateways, compactor | Per-segment postings in the same SlateDB keyspace |
| Query tiers | Frontend, scheduler, queriers, three caches | Readers with one block cache |
| Full-text search | Line filters only | Line filters, plus indexed | match |
The overview architecture guide walks through each Loki component and the failure modes on both sides.
Compatibility
Against Loki's own LogQL conformance scripts, 1373 of 1385 evaluations match. Stream selectors, line and label filters, json, logfmt, pattern, regexp, line_format, label_format, and all range and vector aggregations are supported. The main differences:
- Empty stream labels are dropped at ingest.
jsonkeeps the last value of a repeated key; Loki keeps the first.approx_count_distinctis exact rather than a HyperLogLog estimate.line_formatandlabel_formatalso accept Jinja templates.