Skip to content

Tile Archives (.gphx)

A tile set can be packed into a single immutable archive instead of one .gph file per tile. The archive can be read from local disk or straight from object storage over HTTP range requests, which is how a large tile set is served from S3/CloudFront without ever landing a copy on the routing host.

A planet build is roughly 1.5M small files: awkward to copy, slow to deploy, and impossible to serve from object storage without one request per tile. One archive replaces all of it.


Format

┌──────────────────────────────┐
│ magic "CALCGPH1"             │  8 bytes
│ u16 version, u16 compression │
│ u64 directory_offset, len    │
│ u64 data_offset              │
│ u32 entry_count              │
│ u32 meta_len + meta bytes    │  valhalla.json + build stamp
├──────────────────────────────┤
│ tile blobs                   │  sorted by (level, tile id), 64-byte aligned
├──────────────────────────────┤
│ directory                    │  20 bytes per entry, zstd-compressed
└──────────────────────────────┘

A directory entry is u32 tile_value (the level+tile-id key already used as the cache key), u64 offset, u32 comp_len, u32 raw_len.

Three properties matter:

  • Sorted by (level, tile id). Tiles that are near each other on the ground are near each other in the file, so a request's tiles can be fetched as a few contiguous byte ranges instead of one request each.
  • 64-byte aligned blobs. Tile parsing casts raw bytes to packed structs, which requires natural alignment, because reading a tile in place from an unaligned offset is undefined behaviour. Padding costs ~32 bytes per tile on average. (Valhalla's tar extract aligns to 512 bytes for the same reason.)
  • Immutable. The directory is read once into memory and trusted from then on. Publish new builds under new keys; never rewrite one in place.

Sidecar index

--sidecar-index writes the same directory as a separate <archive>.idx object. A remote reader fetches that with one small GET rather than ranging into a multi-gigabyte body, and a CDN can cache it independently. France: 21 KB for 2051 tiles. The directory also stays embedded in the archive, so a local file remains self-contained.


calculon-gphx

cargo build -p calculon-tiles --release --bin calculon-gphx

# Pack a tile directory. Every tile is parsed as it is packed, so a corrupt
# tile fails the build instead of a routing request in production.
calculon-gphx build /data/tiles -o france.gphx --sidecar-index

# Or pack straight from a valhalla_build_extract tarball, without unpacking it.
calculon-gphx build planet.tar -o planet.gphx --sidecar-index

calculon-gphx info france.gphx                       # header, per-level stats, bounds
calculon-gphx verify france.gphx --against /data/tiles
calculon-gphx extract france.gphx 2/771149 -o t.gph  # one tile back out
calculon-gphx index france.gphx -o france.gphx.idx   # re-emit the sidecar

# Populate a disk cache ahead of serving (see Prewarming a cache).
calculon-gphx warm france.gphx --cache-dir /var/cache/calculon/tiles --levels 0,1
build option Meaning
--compress zstd\|none Blob compression (default zstd)
--level N zstd level (default 9)
--jobs N Compression threads (default: cores)
--sidecar-index Also write <out>.idx
--meta PATH Metadata JSON to embed (default: valhalla.json beside the tiles, or the member inside a tarball)
--from-tar Force reading the source as a tarball (a file source is detected automatically)

Building from a tarball

valhalla_build_extract produces a plain uncompressed tar whose members are 512-byte aligned, which is what lets Valhalla mmap tiles straight out of it. build exploits the same property: it maps the tarball and reads each tile in place, so packing a 100 GB planet extract needs no unpack step and no ~1.5M files on disk. The tarball's own valhalla.json member is embedded in the archive metadata, exactly as the sibling file would be for a directory build.

That makes the in-region conversion a single step:

# On an instance next to the bucket, not over a laptop link.
aws s3 cp s3://bucket/valhalla/3.6.2/planet.tar - \
  | ... # or download once, then:
calculon-gphx build planet.tar -o planet.gphx --sidecar-index
aws s3 cp planet.gphx    s3://bucket/valhalla/3.6.2/planet.gphx
aws s3 cp planet.gphx.idx s3://bucket/valhalla/3.6.2/planet.gphx.idx

After that nobody needs the tar: servers range-read the archive, and anyone who does want it locally pulls ~2.65x fewer bytes.


Choosing a source

Every source returns byte-identical routes and matrices. They differ in memory and cold-start latency, not in warm throughput.

Measured on France (2051 tiles, 3.8 GB raw), median of 15 warm requests:

Source On disk Startup Idle RSS Cold short/long Warm short/long Settled RSS
Tile directory (mmap) 3.8 GB 481 ms 8 MB 95 / 388 ms 29.8 / 116.2 ms 172 MB
Archive, uncompressed 3.8 GB 33 ms 8 MB 76 / 274 ms 29.6 / 114.9 ms 171 MB
Archive, zstd, no cache 1.45 GB 34 ms 9 MB 112 / 351 ms 29.4 / 114.5 ms 343 MB
Archive, zstd + cache (cold) 1.45 GB 33 ms 9 MB 75 / 224 ms 29.3 / 113.6 ms 284 MB
Archive, zstd + cache (warm) 1.45 GB 33 ms 9 MB 33 / 115 ms 29.3 / 112.6 ms 376 MB
HTTP archive, no cache remote 247 ms 11 MB 165 / 181 ms 61.8 / 136.8 ms 937 MB
HTTP archive + cache (warm) remote 246 ms 11 MB 34 / 116 ms 29.3 / 111.6 ms 378 MB

What that says:

  • Warm latency is a wash. Every source that keeps its tiles locally lands within noise of every other, so the choice between them is a memory and operations decision, not a speed one. A mapped tile reads as fast as a heap one.
  • A warm disk cache is the fastest configuration cold, on long routes by more than 3x against a plain tile directory (115 ms vs 388 ms). The second run of a process, and every run after a restart, reads the working set from local disk.
  • Archives start faster than a tile directory (33 ms vs 481 ms) because they read coverage from the index instead of walking 2051 files.
  • Nothing is loaded before the first request. Idle RSS is 8-11 MB for every configuration; tiles arrive on demand and the first request that needs a cold tile pays for it.

Throughput under concurrent load is flat across sources too. 24 mixed routes on an 8-core machine:

Source 1 worker 8 workers p50 @ 8 p90 @ 8
Tile directory (mmap) 14.3 req/s 45.5 req/s 117.3 ms 299.1 ms
Archive zstd (heap) 14.2 req/s 45.4 req/s 116.8 ms 295.1 ms
Archive zstd + disk cache 14.7 req/s 43.5 req/s 120.9 ms 311.4 ms

Mapped tiles do not lose to heap tiles under parallelism: page faults on a shared mapping are cheaper than the contention they replace.

Loopback hides the thing that matters

The HTTP numbers carry no real network latency. Same-region S3 adds roughly 15-25 ms TTFB per request (p99 nearer 100 ms), so request count is the number to watch, not the byte count.


Memory: mapped tiles vs heap tiles

This is the axis that actually separates the sources.

A tile served from a file mapping costs page cache, which the kernel reclaims under pressure and shares between processes. A tile decompressed onto the heap is anonymous memory the OS cannot reclaim, held until the cache evicts it. Same tiles, very different resident footprint.

RSS on France, after a short and a long route plus 30 warm repeats:

Configuration Idle Settled under load
Tile directory (mmap) 8 MB 172 MB
Archive uncompressed 8 MB 171 MB
Archive zstd, no disk cache 9 MB 343 MB
HTTP archive, no disk cache 11 MB 937 MB
Archive zstd + disk cache 9 MB 376 MB
HTTP archive + disk cache 11 MB 378 MB

A compressed or remote source without a disk cache keeps every tile it has decompressed on the heap, bounded only by --cache-size, and the HTTP case adds the byte ranges it pulled to get there. Those are anonymous pages the kernel cannot reclaim.

Settled figures for the disk-cached rows sit above a tile directory partly because the archive mapping and the transient range buffers are counted too, but the bytes that matter are clean, file-backed pages the OS can drop under pressure. Repeated load plateaus rather than growing.

The disk cache fixes it

--tile-cache-dir writes each decompressed tile to local disk once and maps it back from there. The bytes move from heap to page cache, the cold long route drops from 351 ms to 115 ms once the cache is populated, and the archive still distributes as a single 1.45 GB file.

calculon --tile-archive ./france.gphx --tile-cache-dir /var/cache/calculon/tiles
calculon --tile-archive https://tiles.example.com/france.gphx \
         --tile-cache-dir /var/cache/calculon/tiles
  • Writes are atomic (temp file + rename), so a concurrent reader can never map a half-written tile.
  • The cache uses the Valhalla directory layout, so a populated cache directory is itself a valid --tile-dir.
  • A corrupt entry is dropped and re-fetched rather than served.
  • Sizing: the cache holds decompressed tiles, so a fully warmed France cache is ~3.8 GB of disk. There is no automatic eviction, so point it at a volume you are happy to fill, or pre-warm and treat it as read-mostly.
  • Without a writable cache dir (read-only container, no volume), tiles stay on the heap. The server warns at startup when that is the case.

Choosing

Deployment Recommendation
Local disk, memory matters Uncompressed archive: mapped in place, one file, dir-equivalent RSS
Local disk, size matters zstd archive + --tile-cache-dir
Tiles in S3 zstd archive over HTTP + --tile-cache-dir
Read-only/ephemeral filesystem zstd archive, no cache: budget for the heap via --cache-size

Serving from object storage

calculon --tile-archive https://tiles.example.com/france-2026-08.gphx \
         --tile-cache-dir /var/cache/calculon/tiles

The host must honour Range. A 200 where a 206 was expected is rejected at startup rather than silently treated as tile data, since serving the whole object for a range request would corrupt every tile.

python3 -m http.server does not support Range

It answers 200 with the whole body and Calculon refuses it. Use tools/range_http_server.py from this repo for local testing.

Operational notes:

  • Publish under immutable, versioned keys (france-2026-08.gphx), set Cache-Control: immutable, and front with a CDN.
  • --tile-cache-dir writes fetched tiles to local disk, so restarts start warm and repeat reads never re-cross the network.
  • Requests outside the archive's coverage are rejected from the in-memory directory, before any network IO.

Prewarming a cache

calculon-gphx warm populates a disk cache before a server needs it, so an instance starts against a warm cache instead of paying for its first requests. It is a deployment step, not something the server does: the server must stay useful whether or not warming ran.

# Corridor levels only: small, and every long route crosses them.
calculon-gphx warm https://tiles.example.com/france.gphx \
  --cache-dir /var/cache/calculon/tiles --levels 0,1

# Or just your service area, at every level.
calculon-gphx warm https://tiles.example.com/france.gphx \
  --cache-dir /var/cache/calculon/tiles --bbox 3.5,43.4,5.1,46.0

Measured on France over HTTP, warming levels 0 and 1 (117 tiles, 699 MB):

Startup Idle RSS Cold short/long
Empty cache 282 ms 11 MB 277 / 319 ms
Prewarmed levels 0+1 61 ms 11 MB 106 / 284 ms
Prewarmed bbox corridor (112 tiles, 620 MB) 68 ms 11 MB 107 / 250 ms

Warming took 0.8 s for 699 MB over loopback; from real S3 it is bound by the ~219 MB compressed it has to pull. A re-run over a populated cache is a directory scan: 12 ms, nothing written.

How much it buys depends on how well the warmed set matches real traffic. Levels 0 and 1 cover the corridor every long route uses; a bbox covers a service area. Warming everything is just a slow download of the whole archive, and if you intend to hold it all locally an uncompressed archive mapped in place is the better answer.

Never let warming block a boot

A slow or throttled source must degrade to a partly warmed cache, not a failed instance. Use the caps, and let the unit ignore failure:

[Service]
ExecStartPre=-/usr/bin/calculon-gphx warm https://tiles.example.com/france.gphx \
    --cache-dir /var/cache/calculon/tiles --levels 0,1 --deadline 60 --max-mb 2048
ExecStart=/usr/bin/calculon --tile-archive https://tiles.example.com/france.gphx \
    --tile-cache-dir /var/cache/calculon/tiles

The leading - on ExecStartPre makes failure non-fatal, --deadline and --max-mb bound the work, and both caps stop cleanly so a re-run resumes where it left off.

Other notes for that deployment:

  • Warm onto instance store, not EBS. The cache is disposable and rebuildable by definition, which is what ephemeral NVMe is for.
  • If the tile version is pinned to the image, warm at build time instead. The bandwidth is then paid once rather than per scale-out.
  • Autoscaling is the real trade. An instance that joins the load balancer while still cold serves the slow path. Either gate the health check on warming finishing, or accept a slower first minute; which is right depends on whether your p99 covers scale-out.
  • Warming is idempotent and resumable, so re-running after a half-finished boot costs one directory scan.

Prefetching

Before a search starts, Calculon works out which tiles it will likely need: the neighbourhood of each location at every level, plus the corridor between them at the coarse levels. It then asks the source for them in one go. The source groups those tiles into contiguous byte ranges, merging any gap under 1 MB, and issues the resulting handful of requests in parallel.

A Montpellier→Lyon route needs 21 tiles, fetched as 6 range requests. Without prefetching those would be 21 round trips discovered one at a time inside the A* loop, the difference between ~100 ms and several seconds on real S3.

Tiles already in the disk cache are mapped from there instead of being re-requested, so a restarted process reads its warmed working set from local disk rather than pulling it back over the network.

Prefetching is a pure optimisation: a tile the prediction misses is simply fetched when the search asks for it. Long corridors are capped (MAX_CORRIDOR_TILES) so a Lille→Nice request cannot try to pull a country's worth of level-2 tiles.


Caching

GraphReader layers four levels in front of the source:

Level Storage Notes
Thread-local HashMap of Weak handles ~20 ns, no synchronization
Pinned HashMap of Arc Never evicted
Shared Unbounded map, or byte-weighted LRU Depends on the source (below)
Source dir / archive / HTTP

The shared cache strategy follows what the source hands out:

  • Mapped tiles (tile directory, uncompressed archive, or any source backed by --tile-cache-dir) go in an unbounded map. Their bytes are page cache, so the OS already owns eviction and an application-level bound would only force re-reads of pages the kernel was happy to keep.
  • Heap tiles (compressed archive or HTTP without a disk cache) go in a byte-weighted, sharded LRU bounded by --cache-size.

Two details worth knowing:

  • The bound is bytes, not entries. Tiles range from 15 KB to 85 MB, so a tile count would be meaningless. --cache-size is in MB.
  • Thread-local entries are Weak. A strong reference there would pin every tile a worker thread ever touched, and the byte budget would be fiction, since evicting from the shared cache would free nothing.

Eviction is advisory: a tile evicted while an in-flight search still holds it is freed when that search drops it, so resident memory can briefly exceed the budget. Size it with headroom.


Verifying a build

# Every tile decompresses, parses, and matches the source directory
calculon-gphx verify france.gphx --against /data/tiles

# All three sources must return identical route and matrix JSON
tools/compare_tile_sources.sh test_data/monaco_tiles