Glider

Open-source vector database

Vector search with S3 as the source of truth.

Every write is durable in object storage before it is acknowledged. A restart recovers your data; RAM and SSD are disposable caches. Built to go far on less.

from glider_client import Client

client = Client("http://localhost:8080")
client.create_collection("demo", dimensions=3)
docs = client.collection("demo")
docs.upsert([{"id": 1, "vector": [0, 0, 0],
              "metadata": {"color": "red"}}])

hits = docs.query([1, 1, 0.9], k=10)

Measured results

vectors on AWS S3
1M
recall@10
0.998
p95 warm query
29.8ms
p95 durable write
86.9ms

SIFT1M · 128 dimensions · k=10 · S3 Standard · c7g.2xlarge, eu-central-1. Recall is the static mean; latency is from concurrent readers and writers. Methodology & raw results →

10,000 collections · 10M vectors · one server: warm queries over 64 active collections reached 2,457 queries/s at 18.5 ms p95. All 10,000 collections were verified after kill -9 during ingest. Multi-tenant results →

Storage playground

A million vectors. Three places to live.

Pick a spot, follow the blocks, then pull the plug.

1,000,000 vectorsPROCESS ONLINE
Click anywhere to query ↗ 2D projection / fixed seed 0x61de2026
Vectors Cluster centres Selected clusters Nearest neighbours Recent writes
Query a neighbourhood
Query trace · simulated
S3SSDRAM

Route and select in RAM → fetch from RAM, SSD or S3 → rerank in RAM on the CPU.

Fetch: RAM hot cache → SSD cache → S3 range read. Then CPU rerank in RAM: fetched blocks + unsealed tail.

Click the map to route a query through centroids and cached blocks.

Read from S3
—
Served from cache
—
Simulated latency
—

Acknowledged writes lost: 0

Details
Clusters probed
—
Blocks read / candidates
—
RAM / SSD / S3 hits
—
Up to 12 candidate blocks · 8 range GETs · 1 MiB remote

Top-10 neighbours appear as connected points after a query.

Latency = 0.25 ms RAM routing + 0.25 ms RAM selection + 0.5 ms RAM CPU reranking + 0.1 ms per RAM hit + 0.5 ms per SSD hit + 20 ms per batch of four parallel S3 reads. Illustrative, not measured.

S3

Every acknowledged write is durable here.

Details

Immutable logs acknowledge writes; a published root selects sealed packs, indexes and manifests.

SSD

Fetched blocks stay here for the next query.

Details

Slots show capacity, not physical blocks. Least recently used blocks are evicted when full; losing this disk loses no acknowledged writes.

RAM

Routing and recent data live in process memory.

Details

Full vectors enter RAM only as fetched blocks or unsealed writes. The hot block cache is bounded to 4 MiB; a crash clears all process memory.

How the illustration is sized and timed

30,000 sampled dots stand for the selected dataset, already sealed and clustered (including the 100K example). Seeded Gaussian-mixture islands with Poisson-disc centres have varied populations averaging about 4,000 vectors per centroid. Projected islands can overlap; cluster membership represents the original embedding. Routing uses the M31 profile's 32 probes, scaled to 4 for 100K. Reranking is exact within selected blocks and the tail; the overall search is approximate. The projection and five-bit scores use two dimensions, while sizes assume 128-dimensional float32 vectors: 24 B per directory slot, 80 B of five-bit codes + 4 B ID offset + 8 B posting sequence per sealed row, 512 B per centroid, and 640 B per tail/log row. Blocks are sized proportionally up to 120 KiB; packs hold up to eight blocks (960 KiB) plus sketches. Compression, allocator overhead, routing cache files and retired objects are omitted. RAM's hot cache is 4 MiB; SSD defaults to 256 MiB. Latency = 0.25 ms RAM routing + 0.25 ms RAM selection + 0.5 ms RAM CPU reranking + 0.1 ms per RAM hit + 0.5 ms per SSD hit + 20 ms per batch of four parallel S3 range GETs. Trace widths follow these simulated phase times with a 14% minimum per phase; the animation stretches the whole query to 1.2 seconds. Each batch adds 100 new IDs; 32 log objects trigger sealing. Restart compresses the real lease wait and recovery into five visual stages.

Try it on your own data Quickstart

How it works

The bucket holds the truth. The SSD makes it fast.

A write is acknowledged once it is on S3; queries read clustered packs through a local SSD cache.

Architecture guide · Detailed diagram

Why Glider

Small engine. Serious guarantees.

Crash-safe

A restart fences the old writer at the store and replays the log. No operator step.

Retry-safe writes

Reuse a write’s request ID within the retry window to get its original outcome without applying it twice.

Fast warm reads

A local SSD cache serves warm queries. Lose it and you lose no data.

Clusters itself

The server builds its clustered index on its own and rebuilds it as data grows.

Collections

Create and delete collections over HTTP, each with its own dimension and metric.

Metadata filters

Equality, sets, numeric ranges and nested logic, with an exact mode for every match.

Filters

Filter fast. Go exact when it counts.

Filtered search is approximate by default. Add "exact": true and you get every match, up to k.

POST /v1/collections/demo/query
{
  "vector": [1, 1, 0.9],
  "k": 10,
  "filter": {
    "color": "red",
    "price": {"$gt": 1.5, "$lte": 10},
    "tag": {"$in": ["a", "b"]}
  },
  "exact": true
}

Agent memory · MCP

Durable memory for your agents.

glider-mcp gives MCP clients three tools. Memories survive restarts, and request IDs make writes safe to retry within the retained window.

  • remember
  • recall
  • forget

01Start Glider

docker run -d --name glider -p 8080:8080 \
  -v glider-data:/var/lib/glider \
  -e GLIDER_DATA_DIR=/var/lib/glider/data \
  ghcr.io/omerfeyzioglu/glider:latest

02Install the MCP client

pip install \
  "glider-client[mcp] @ git+https://github.com/omerfeyzioglu/glider#subdirectory=clients/python"

03Connect Claude Code

claude mcp add glider \
  -e GLIDER_COLLECTION=memory -- glider-mcp

Quickstart

Three commands to your first query.

Only Docker and curl. This demo uses a local Docker volume that survives container removal. For object storage, run it on S3.

Once it is running, open localhost:8080/console to browse points and query your server.

Keep the server running. Use another terminal for the next steps.

01Start the server

docker run --rm -p 8080:8080 -e GLIDER_DIMENSIONS=3 \
  -v glider-quickstart-data:/var/lib/glider \
  -e GLIDER_DATA_DIR=/var/lib/glider/data \
  ghcr.io/omerfeyzioglu/glider:latest

02Write two points

curl -sS localhost:8080/v1/write \
  -H 'content-type: application/json' \
  -d '{"upsert":[
    {"id":1,"vector":[0,0,0],"metadata":{"color":"red"}},
    {"id":2,"vector":[1,1,1]}]}'

03Find their neighbors

curl -sS localhost:8080/v1/query \
  -H 'content-type: application/json' \
  -d '{"vector":[1,1,0.9],"k":2,"include_metadata":true}'

Good to know

Honest about the edges.

  • One server, with a single committer per collection. No replication or read replicas.
  • The measured 1M-vector collection opens on S3 in a few seconds; larger collections can take longer.
  • Search is approximate by default. Use exact mode for exhaustive results; a cold cache can lower ANN recall.
  • Plain HTTP: put TLS in front with a reverse proxy.