skaidb benchmarks
The latest measured results for the standard comparison scenarios. Numbers are only comparable within a section: each section is one measurement round — all systems in it measured together, in the same environment, driven by the same client harness. Never compare figures across sections. Superseded results live in git history, never alongside current numbers.
Sections: the CRUD comparison suite, skaidb 0.185.0 re-measurement, search workloads vs Elasticsearch, and the MQTT broker vs Mosquitto.
Numbers are for relative comparison on small nodes, not absolute peak throughput. Run-to-run spread on this shared host is ±10–15% when the host is busy; measured on a quiet host with warm runs it tightens to ±3% (five consecutive
write 1cruns: 973/972/935/963/986). Treat differences inside the band as noise, and never compare figures from different sections or environments.
Environment
Host. One Proxmox host — Intel Core i7-8550U (4 cores / 8 threads, 1.8 GHz), 8 GB RAM. Every node is an unprivileged LXC container on this host, bridged on one VLAN. The host also carries unrelated background load; the methodology section below describes how interference is kept out of conclusions.
Nodes. Identical containers: 1 vCPU / 512 MB RAM / 4 GB disk, Debian 12. A 3-node configuration is three such containers.
Client. A dedicated VLAN-local client container for multi-node scenarios; colocated on the server container for the single-node C0 scenario (true loopback).
Versions measured.
| skaidb | PostgreSQL | MongoDB | MariaDB | Elasticsearch | Qdrant |
|---|---|---|---|---|---|
| 0.146.1 | 17.10 | 7.0.37 / 8.0.26 | 11.4 | 8.19.18 | 1.19.1 |
Durability is matched across systems. In each config a write is acknowledged only after the same number of nodes have made it durable:
| Config | Nodes | A write is acked after… | skaidb | MongoDB | PostgreSQL | MariaDB |
|---|---|---|---|---|---|---|
| C0 | 1 | local WAL/journal fsync | RF=1, ONE |
single-member rs | no standbys | binlog only |
| C1 | 2 | both nodes | QUORUM |
w:2 |
sync standby (FIRST 1) |
semi-sync |
| C2 | 2 | the primary only (async replica) | ONE |
w:1 |
async ('') |
semi-sync off |
| C3 | 3 | all 3 nodes | ALL |
w:3 |
FIRST 2 sync standbys |
— ¹ |
| C4 | 3 | any 2 of 3 (quorum) | QUORUM |
w:majority |
ANY 1 standby |
semi-sync ¹ |
¹ MariaDB semi-sync acknowledges after the first replica responds and
has no "wait for N replicas" knob, so true all-3 durability isn't
expressible; its C3 row is the same semi-sync mode as C4 (≈ 2-of-3), a
single measurement marked *.
Workloads (throughput in ops/sec, higher is better):
write 1c— single connection inserting unique keys (durable-write latency floor)write 16c— 16 connections inserting (concurrent write throughput)read 16c— 16 connections, point read by primary key over a 1,000-row tablemixed 16c— 16 connections, 50/50 read/write
Every client connects and authenticates before the timed window opens.
C0 — 1 node, no replication
One node, no peers, client colocated (loopback). skaidb:
replication_factor = 1, consistency ONE. PostgreSQL:
synchronous_standby_names cleared, standbys stopped. MongoDB:
reconfigured to a genuine single-member replica set. MariaDB: semi-sync
off. Elasticsearch runs the same 1 vCPU / 512 MB spec with a 256 MB heap
(its default 1 GB does not fit this node class — the number reflects a
system at its memory floor).
| Workload | skaidb | PostgreSQL | MongoDB 7 | MariaDB | Elasticsearch |
|---|---|---|---|---|---|
| write 1c | 1,877 | 1,692 | 703 | 319 | 79 |
| write 16c | 5,522 | 3,004 | 1,582 | 1,523 | 206 |
| read 16c | 13,660 | 4,143 | 1,673 | 3,209 | 973 |
| mixed 16c | 8,211 | 4,244 | 1,601 | 2,748 | 480 |
skaidb leads every C0 cell: 1.1× PostgreSQL on the single-connection durable-write floor, 1.8× on concurrent writes, 3.3× on reads, 1.9× on mixed.
C1 — 2 nodes, writes wait for both
| Workload | skaidb | PostgreSQL | MongoDB 7 | MongoDB 8 | MariaDB |
|---|---|---|---|---|---|
| write 1c | 863 | 698 | 268 | 197 | 316 |
| write 16c | 3,855 | 2,814 | 1,134 | 607 | 1,666 |
| read 16c | 11,464 | 4,441 | 1,730 | 2,297 | 3,092 |
| mixed 16c | 5,991 | 3,582 | 1,393 | 1,000 | 2,861 |
C2 — 2 nodes, writes wait for the primary only (async replica)
| Workload | skaidb | PostgreSQL | MongoDB 7 | MongoDB 8 | MariaDB |
|---|---|---|---|---|---|
| write 1c | 924 | 981 | 761 | 778 | 348 |
| write 16c | 3,530 | 3,779 | 1,567 | 1,782 | 1,756 |
| read 16c | 11,967 | 4,609 | 1,457 | 2,221 | 3,022 |
| mixed 16c | 6,175 | 4,008 | 1,541 | 1,739 | 3,138 |
C3 — 3 nodes, writes wait for all 3
| Workload | skaidb | PostgreSQL | MongoDB 7 | MongoDB 8 | MariaDB* |
|---|---|---|---|---|---|
| write 1c | 790 | 523 | 171 | 174 | 267* |
| write 16c | 3,158 | 2,433 | 856 | 535 | 1,345* |
| read 16c | 11,721 | 4,573 | 1,382 | 2,139 | 3,029* |
| mixed 16c | 5,358 | 3,167 | 1,239 | 889 | 2,410* |
C4 — 3 nodes, writes wait for 2 of 3 (quorum)
| Workload | skaidb | PostgreSQL | MongoDB 7 | MongoDB 8 | MariaDB |
|---|---|---|---|---|---|
| write 1c | 786 | 629 | 221 | 210 | 263 |
| write 16c | 3,231 | 2,539 | 874 | 1,057 | 1,262 |
| read 16c | 11,577 | 4,528 | 1,573 | 2,273 | 3,047 |
| mixed 16c | 5,260 | 3,318 | 1,163 | 1,578 | 2,350 |
Reading the standings
- skaidb leads every replicated-config cell except C2 writes, where PostgreSQL's async replication edges it inside the noise band (981 vs 924 at 1c, 3,779 vs 3,530 at 16c). Largest margins are reads (2.5–2.6× over PostgreSQL at 3 nodes; 5–8× over MongoDB) and strict-durability writes (C3 write 1c: 1.5× PostgreSQL, 4.5× MongoDB).
- skaidb's reads scale with members (every node coordinates against its full local copy); the other systems read from a single primary regardless of cluster size.
- Durability level barely moves skaidb throughput (C1 ≈ C2 ≈ C3 ≈ C4 within noise at 16 connections): the fsync is group-committed and the peer round-trip pipelined, so stricter acks cost latency headroom, not throughput. MongoDB pays heavily for w:3; PostgreSQL is comparatively flat like skaidb.
- These runs include skaidb's per-row consensus machinery for clustered
UPDATEs (see CLUSTERING.md) — it costs nothing on these workloads because INSERT/SELECT take the plain paths by design.
skaidb 0.185.0 re-measurement (skaidb only)
Its own measurement round — 2026-08-18. Only skaidb was re-run; the comparison tables above are the 0.146.1 round and their PostgreSQL, MongoDB and MariaDB figures were measured in July. Per this document's own rule, do not read a skaidb number here against a competitor number above. The environment also moved: the bench host now also carries a production skaidb node (colocated 2026-08-09), which was not present for the earlier round.
Same host, same node specs, same harness, same four workloads. Warm runs — the first run of each workload is discarded (see Methodology).
| Config | write 1c | write 16c | read 16c | mixed 16c |
|---|---|---|---|---|
C0 — 1 node, ONE |
1,531 | 4,899 | 17,291 | 8,446 |
| C1 — 2 nodes, both | 989 | 4,077 | 14,516 | 5,768 |
| C2 — 2 nodes, primary only | 975 | 4,131 | 14,530 | 6,267 |
| C3 — 3 nodes, all 3 | 742 | 3,299 | 14,484 | 5,585 |
| C4 — 3 nodes, quorum | 765 | 3,253 | 13,612 | 5,825 |
Against the 0.146.1 round — same system, same host, same harness, so this is the most controlled cross-round comparison available, but the colocated production node makes it indicative rather than clean:
| Workload | mean change |
|---|---|
| read 16c | +22.4% |
| mixed 16c | +7.7% |
| write 16c | +3.7% |
| write 1c | −4.2% |
- Reads are up on every single config (+17.6% to +26.6%) — but the A/B below shows this cross-round gain is the client's, not the server's. The bench client is colocated on the 1-vCPU node in C0, so a cheaper client buys the server CPU; the client used here is 39 releases newer than the July round's.
- Replicated writes are flat to up; the one soft spot is C0 (−18% at 1c, −11% at 16c), the only non-replicated config — resolved by the A/B below: 1c was noise, 16c is real but half that size.
write 1cdifferences are the least trustworthy line here: it is one connection deep, so it measures durable-write latency and is the most exposed to the new colocated neighbour.
The C0 A/B (same day, interleaved, one client)
The cross-round C0 anomalies above were re-measured the way this document requires: same-day interleaved legs alternating the 0.146.1 and 0.185.0 server binaries on the same node, fresh data dir per leg, one fixed (current) client, first run of each workload discarded, legs repeated across separated time windows (3 rounds, then a 6-version × 3-round write-only sweep). Findings, all agreeing across legs:
| Workload | 0.146.1 server | 0.185.0 server | change |
|---|---|---|---|
| write 1c | ~1,720 | ~1,750 | flat (the −18% was noise) |
| write 16c | ~5,650–5,920 | ~5,360–5,390 | −5 to −9% |
| read 16c | ~19,500 (p50 0.83 ms) | ~16,900 (p50 0.96 ms) | −14% |
| mixed 16c | ~8,700 | ~8,300 | −5% |
- The +22% cross-round read “gain” decomposes into a client-side improvement masking a server-side regression. With the client held constant, the 0.146.1 server answers point reads 14% faster than 0.185.0. The same pair pinned to single cores on a 32-core host reproduces it (103k vs 92k ops/s server-side): the modern server spends ~1.2 µs more CPU per point read.
- At whole-server granularity the regression looks accretive. Per-version legs (each interleaved against the 0.146.1 reference): 0.156.3 −3%, 0.161.0 −3%, 0.162.6 −5%, 0.163.0 −8%, 0.163.5 −8%, 0.166.0 −11%, 0.176.1 −14%, 0.185.0 −14%. The steps track the 0.162–0.166 read-path work (ValueRef/borrowed views, parallel range scans, the keyset lane, the unified select dispatcher) plus a further step by 0.176 — scan throughput was bought with small per-statement costs that a 1,000-row point-read loop pays on every operation. The scan-heavy shapes those releases targeted (see the read-path RFC measurements) are far faster; this line measures the opposite extreme.
- The engine's share of it DOES localize — to two commits. Measured
with
read_path_breakdown(in-processSession::executeminus parse, 200k ops pinned to one core, 3–4 runs per point, every version built from its own worktree), the bind+dispatch+engine cost moves in two steps and is otherwise flat across 33 releases:
| version | bind+dispatch+engine | step |
|---|---|---|
| 0.161.0 | 0.75 µs | — |
| 0.162.0 | 0.74 µs | flat |
cc7227f streaming projection (borrowed row view) |
0.88 µs | +0.15 |
| 0.163.0 → 0.167.1 | 0.87–0.90 µs | flat |
5ee92b1 plain SELECTs onto ParFold |
1.12 µs | +0.23 |
| 0.168.0 → 0.195.x | 1.07–1.12 µs | flat |
Both are scan optimizations whose fixed per-statement setup a one-row
point read pays without ever using: 5ee92b1 routed order-free,
limit-free projections through the partition-fold machinery, and
parallel_scan_shape rejects a point key outright, so the fold ran on
its serial arm with the plumbing as pure overhead. Point lookups now
skip that route (Cluster::is_pk_point_lookup), which returned the
engine path to ~0.94 µs.
- The cc7227f step was per-statement allocation, and it is gone.
Instruction counts (callgrind over run_simple_select, an absent-key
point read) showed the remaining leg dominated not by the lookup but by
what the planner built and threw away around it: the WHERE-clause
constraint list materialised twice per statement with a clone of every
literal, the primary-key column list cloned per probe, the projection
set built unsized and re-hashed as it grew, and each projected column's
name cloned three times. Reading those by reference brings the leg from
8.8k to 6.3k instructions per statement (−28%) and the harness's
dispatch+lookup stage to 0.75 µs — the 0.161.0 figure. What remains
is the storage layer itself (memtable range probe, read-cache hash and
negative-cache insert), which has not moved.
- Above the engine, it stays accretive. On one instrument (server
process CPU per point read, utime+stime over a fixed single-connection
loopback run) 0.161.0 costs 11.70 µs/op and 0.195.x costs 12.29 µs/op.
The engine accounts for ~0.33 µs of that ~0.59 µs; the rest is spread
above it — request handling, admission gates, per-statement metrics,
protocol encode — and does not localize. It is also where ~90% of the
absolute per-op CPU lives, so it is the larger target even though it
regressed less.
- The write 16c regression is real but small and does not localize.
A 6-version × 15-warm-run interleaved sweep (0.146.1 / 0.178.0 /
0.181.0 / 0.183.3 / 0.184.0 / 0.185.0) puts 0.185.0 at −5% with
every intermediate version inside the ±7% noise band — no single
release explains it, consistent with the same per-statement accretion
as the read side.
- The single-connection lines (write 1c) are unchanged across all of
this — the added cost shows only under concurrency on 1 vCPU, where
every spare microsecond of server CPU is throughput.
skaidb 0.255.0 re-measurement (skaidb only)
Its own measurement round — 2026-09-06, same fleet, same node specs, same four workloads, warm runs. As above: these are skaidb-only figures and must not be read against the competitor columns measured in July.
| Config | write 1c | write 16c | read 16c | mixed 16c |
|---|---|---|---|---|
C0 — 1 node, ONE |
1,858 | 5,520 | 16,645 | 8,505 |
| C1 — 2 nodes, both | 1,008 | 4,328 | 15,164 | 6,920 |
| C2 — 2 nodes, primary only | 992 | 4,158 | 15,597 | 7,057 |
| C3 — 3 nodes, all 3 | 737 | 3,996 | 21,012 | 6,750 |
| C4 — 3 nodes, quorum | 710 | 4,006 | 20,148 | 6,589 |
Three-node reads are up sharply on the 0.185.0 round (C3 +45%, C4 +48%) and 16-connection writes up 1–23%. An interleaved A/B against 0.210.0 — one fixed client, five legs each at C0 — came out flat (write 1c 1,862 vs 1,917 against a 1,621–2,117 spread), so no single release carries the cross-round read gain and the write path has not regressed.
That A/B also produced a methodology rule now recorded in
bench/README.md: the first binary measured after a fleet wipe reads
low. Running the new build first made it look 10% slower at C0 write
1c; reversing the order inverted the gap, and a disk-backed workstation
A/B showed the path flat. Discarding the first run of each workload
does not cover this — alternate which binary goes first.
Search workloads vs Elasticsearch 9.4.4
Workstation round (32 cores / 91 GB, loopback): skaidb
(release build, defaults, auth off) vs Elasticsearch 9.4.4
(Docker, single node, 1 shard / 0 replicas, security off, 4 GB heap).
One 200k-document corpus generated deterministically (zipf vocabulary of
5,000 terms; text title/body, keyword category (50 values), numeric
price/stock) and ingested into both through the same _bulk driver;
ES explicitly mapped, skaidb auto-mapped by its ES gateway. Semantics
verified before timing: hit counts and aggregation buckets agree exactly
on every workload (ES caps total.value at 10k by default; skaidb
reports exact totals). Queries: 16 threads × 10 s, identical bodies via
POST /bench/_search; two interleaved legs (A→B then B→A) agreed within
noise.
| Workload | skaidb | Elasticsearch |
|---|---|---|
_bulk ingest, 8 threads × 1k-doc batches (docs/s) |
73,500 | 104,600 |
match (single term, mixed frequency), size 10 (q/s) |
52,300 | 18,900 |
match_phrase (real bigrams), size 10 (q/s) |
6,700 | 11,300 |
bool: high-freq match must + term filter (q/s) |
41,700 | 17,100 |
terms agg + avg sub-agg over high-freq match (q/s) |
12,300 | 34,800 |
What the rows reflect:
matchleads on the REST gateway's persistent connections (p50 0.21 ms).boolexecutes as a postings intersection: the gateway's auto-mapping declares a.keywordtwin for every string field (mirroring ES dynamic mapping) and translates stringtermfilters to exact matches on it, rather than materializing 10–25k candidate rows per query — 2.4× Elasticsearch on the identical corpus, with identical hit counts.- Keyword-grouped metric aggregations push down as a fold over
fast-field columns (matching doc set → per-segment ord-indexed
accumulators) with exact buckets. Elasticsearch leads this row — its
shard request cache serves repeated
size:0aggregations from cache (on by default), while skaidb recomputes the ~24k-doc fold every query (p50 0.98 ms, proportional to match count). - Phrase and ingest are the two rows where a dedicated search engine leads, at the same order of magnitude.
Vector search vs Qdrant 1.19.1
2026-09-07, one node each on identical containers (2 vCPU / 2 GB —
larger than the C0–C4 nodes above, since an HNSW graph of this size does
not fit 512 MB), 100,000 × 768-dimension vectors, cosine,
HNSW m = 16 / ef_construct = 200 on both sides, and a payload index
on the filter field on both sides (Qdrant a keyword index, skaidb a
secondary index — without one, either system must read rows to evaluate a
filter). Both driven over REST by the same client
(bench/clients/vector_bench.py), 250 queries per point, k = 10,
recall@10 against exact brute force. Filtered queries select one category
holding 10% of the rows.
| ef | skaidb recall | Qdrant recall | skaidb p50 | Qdrant p50 |
|---|---|---|---|---|
| 32 | 0.8899 | 0.9008 | 3.68 ms | 2.95 ms |
| 64 | 0.9698 | 0.9769 | 4.25 ms | 3.39 ms |
| 128 | 0.9962 | 0.9976 | 5.34 ms | 4.46 ms |
| 256 | 0.9999 | 1.0000 | 7.30 ms | 6.46 ms |
Unfiltered, the two are within ~25% on latency at matched recall — the
July round measured 1.4–2.2×, and the difference is dot/l2 kernels
that now vectorize.
Filtered (the same queries plus WHERE cat = …):
| ef | skaidb recall | Qdrant recall | skaidb p50 | Qdrant p50 |
|---|---|---|---|---|
| 32 | 0.9974 | 0.9552 | 11.56 ms | 2.92 ms |
| 64 | 1.0000 | 0.9951 | 15.65 ms | 3.24 ms |
| 128 | 1.0000 | 0.9999 | 22.36 ms | 3.97 ms |
| 256 | 1.0000 | 1.0000 | 33.06 ms | 5.00 ms |
skaidb returns more of the true neighbours here (the filter gates
admission inside the traversal, and a filter the graph cannot satisfy is
answered exactly), but pays 3–4× the latency at matched recall: walking
until k rows PASS the filter visits several times more nodes than
walking until k rows are found. The July round's filtered figure —
0.846 recall from a post-filtered over-fetch — is what this replaced.
Index build, the same corpus, time until the index answers queries:
| skaidb | Qdrant | |
|---|---|---|
| Exact (f32) index | 742 s | 506 s |
| Binary-quantized index | 222 s | 506 s |
QUANTIZED BINARY builds fastest because a Hamming distance over twelve
machine words costs far less than 768 float multiplies, and the exact
rescore recovers the ranking: 0.9949 plain recall at ef = 32 (better
than our own exact index at that ef) and 1.0000 filtered, at a flat
13–15 ms plain / 25–31 ms filtered. Building an index no longer blocks
the node either — see VECTOR.md.
Read the standings the way this document says elsewhere: these are two systems doing different jobs. Qdrant is a specialist that stores vectors and payloads; skaidb is a durable, tunably-consistent database that also indexes vectors, so the same rows carry SQL, full text, geo and time series, and a vector query can join them.
MQTT broker vs Mosquitto 2.0.22
Workstation round, re-measured 2026-09-06 (skaidb 0.255.0, same host,
loopback): skaidb's native MQTT broker (standalone,
state_consistency = one, $SYS off) vs Mosquitto 2.0.22
(persistence false, max_inflight_messages 64, max_queued_messages
1000 — matching skaidb's outbox). One client drove both —
cargo run -p skaidb-mqtt --example mqtt_bench, which speaks the
protocol through this workspace's codec (see bench/README.md);
64-byte payloads, clean sessions, repeated rounds within ~5%.
| Scenario | skaidb | Mosquitto | ratio |
|---|---|---|---|
| QoS 1 acked publish throughput, 64 in flight (msg/s) | 200,000 | 130,000 | 1.54× |
| Fan-out: 50 subscribers, aggregate deliveries (del/s) | 882,000 | 411,000 | 2.15× |
| QoS 0 delivered throughput, 1 pub → 1 sub (msg/s) | 224,000 | 302,000 | 0.74× |
| Delivery latency p50 / p99 (µs) | 19 / 27 | 20 / 31 | — |
Mosquitto wins QoS 0 to a single subscriber, and the flow-control philosophies are why. Offered ~240k msg/s by a paced publisher, Mosquitto delivers every message (it can accept ~480k/s and applies TCP backpressure when it cannot keep up); skaidb reads the publisher flat out and sheds ~6% at the per-session outbox under its QoS 0 drop policy. The comparable figure is what each broker delivers at a publish rate neither client-limits — which is the correction this table needed: an earlier round reported skaidb 229k against Mosquitto 195k by comparing our delivered rate with Mosquitto's throttled accepted rate under a flat-out publisher, two different quantities. Fan-out and acked QoS 1 publishing, where the work is the broker's own routing and acknowledgement paths, remain skaidb's by a wide margin.
Broker RSS after the runs: ~31 MB for skaidb (a full database node) vs
~8 MB (a dedicated broker). Idle connections cost skaidb a registered
descriptor rather than a thread pair (mqtt.idle_park_secs), and parking
is free under load: on/off is within noise on every scenario above.
Mosquitto-client interop (QoS 0/1/2, retained, MQTT 5, persistent
sessions, state surviving a broker restart) is exercised separately —
see MQTT.md.
Methodology
- One system benchmarked at a time; suites run one after another within a single day.
- Clients pre-connect before the timed window; connection setup is never measured.
- Config transitions are verified, not assumed: MariaDB semi-sync via
Rpl_semi_sync_master_clients(the master-side flag alone can lie), MongoDB topology viars.status()member counts, PostgreSQL viapg_stat_replication. - The bench host is shared; interference shows up as temporally-clustered swings. Perf claims about code changes therefore require interleaved same-day A/B runs with alternating binaries, agreeing across legs separated in time — a single leg (even an interleaved one inside an interference window) is not evidence.
- The client must be VLAN-local. Driving the suite from a developer machine measures the network, not the database: RTT from a routed workstation to the bench VLAN is ~16 ms against ~0.05 ms from a container on the bench host — at one connection that is the entire measurement (69 ops/s vs 760 for the identical C4 write). Every figure here comes from a client container on the bench host.
- Discard the first run of each workload. A cold engine reads
consistently low and recovers within one run: C4
write 1cmeasured 677 cold, then 735/736/741/728. Steady-state spread on that workload is ±3%, so a single cold run reads as a ~10% regression that does not exist. - To reproduce: the client harness (
run_suite.sh+ per-system clients) lives on the bench client container; each system's config-switching procedure is described in its scenario section above.