Skip to content

Benchmarks

Every number here is reproducible - run benchmarks/run.py yourself:

pip install "swarmstate[langgraph]" langgraph-checkpoint-sqlite matplotlib
# or: uv add "swarmstate[langgraph]" langgraph-checkpoint-sqlite matplotlib
python benchmarks/run.py --iters 5000 --seed 7

Read the setup before comparing

  • Release build (published wheels / maturin develop --release). Debug builds are several times slower.
  • SwarmStateSaver and InMemorySaver are in-memory; SqliteSaver is file-backed (persists to disk). The put comparison reflects the cost of SQLite-backed persistence as commonly configured; a persistent swarmstate backend (redis/disk) is on the roadmap.
  • Warm cache, single process. Hardware/versions below.

Checkpoint latency (LangGraph interface)

Checkpointer latency

Per-operation latency through LangGraph's BaseCheckpointSaver interface:

Checkpointer put p50 put p99 put throughput get_tuple p50
SwarmStateSaver 5.8 µs 13.9 µs ~158k ops/s 7.5 µs
InMemorySaver 4.0 µs 8.5 µs ~203k ops/s 81.8 µs
SqliteSaver (file) 74.0 µs 323 µs ~9.5k ops/s 14.5 µs
  • put is ~12.8× faster than SqliteSaver - no per-step disk commit.
  • get_tuple (latest) is ~11× faster than InMemorySaver: swarmstate keeps an O(1) "latest checkpoint" pointer, where the reference saver scans keys.

Snapshot cost vs state size

Snapshot scaling

Store.snapshot() is O(1) (persistent data structure + a running byte counter), so it stays flat as state grows, while copy.deepcopy is O(n):

entries Store.snapshot() dict deepcopy speedup
100 0.0003 ms 0.086 ms ~300×
1,000 0.0002 ms 0.86 ms ~4,000×
10,000 0.0002 ms 9.3 ms ~44,000×
50,000 0.0002 ms 52 ms ~236,000×

This is what makes time-travel over the whole checkpoint DB practically free.

Concurrency (free-threaded vs GIL)

The store shards its write locks and releases the GIL on the hot paths. On a free-threaded (no-GIL) CPython build the same code does not serialize behind the interpreter lock. Median of 5 runs, set+get, distinct namespaces per thread:

Build 1 thread 8 threads 8-thread vs GIL
CPython (GIL) ~1.8M ops/s ~130k ops/s baseline
Free-threaded (cp313t) ~2.2M ops/s ~1.8M ops/s ~14x

It's about not collapsing, not linear scaling

This workload is allocation-bound (msgpack buffers, Python result objects), so it does not speed up with more cores on either build. The difference is the shape: under the GIL, adding threads makes it dramatically slower (it collapses); the free-threaded build holds throughput roughly flat, so at 8 threads it does about 10x the GIL build. Single-threaded, free-threaded is ~2x (no per-call GIL machinery). The free-threaded win needs a cp313t interpreter, for which swarmstate ships version-specific wheels. See Architecture.

Batching helps independently: on the free-threaded build at 8 threads, set_many (batches of 50) reaches roughly 3x the throughput of individual sets.

Environment

Machine Apple Silicon (macOS, arm64)
Python 3.14
swarmstate 0.4.0 (release wheel)
langgraph 1.2.x · langgraph-checkpoint 4.x
iters / warmup 4,000 / 400 · seed 7

Numbers vary by hardware - regenerate on yours with the command above.