Tile Archives (.gphx)
A tile set can be packed into a single immutable archive instead of one .gph
file per tile. The archive can be read from local disk or straight from object
storage over HTTP range requests, which is how a large tile set is served from
S3/CloudFront without ever landing a copy on the routing host.
A planet build is roughly 1.5M small files: awkward to copy, slow to deploy, and impossible to serve from object storage without one request per tile. One archive replaces all of it.
Format
┌──────────────────────────────┐
│ magic "CALCGPH1" │ 8 bytes
│ u16 version, u16 compression │
│ u64 directory_offset, len │
│ u64 data_offset │
│ u32 entry_count │
│ u32 meta_len + meta bytes │ valhalla.json + build stamp
├──────────────────────────────┤
│ tile blobs │ sorted by (level, tile id), 64-byte aligned
├──────────────────────────────┤
│ directory │ 20 bytes per entry, zstd-compressed
└──────────────────────────────┘
A directory entry is u32 tile_value (the level+tile-id key already used as the
cache key), u64 offset, u32 comp_len, u32 raw_len.
Three properties matter:
- Sorted by
(level, tile id). Tiles that are near each other on the ground are near each other in the file, so a request's tiles can be fetched as a few contiguous byte ranges instead of one request each. - 64-byte aligned blobs. Tile parsing casts raw bytes to packed structs, which requires natural alignment, because reading a tile in place from an unaligned offset is undefined behaviour. Padding costs ~32 bytes per tile on average. (Valhalla's tar extract aligns to 512 bytes for the same reason.)
- Immutable. The directory is read once into memory and trusted from then on. Publish new builds under new keys; never rewrite one in place.
Sidecar index
--sidecar-index writes the same directory as a separate <archive>.idx
object. A remote reader fetches that with one small GET rather than ranging into
a multi-gigabyte body, and a CDN can cache it independently. France: 21 KB
for 2051 tiles. The directory also stays embedded in the archive, so a local
file remains self-contained.
calculon-gphx
cargo build -p calculon-tiles --release --bin calculon-gphx
# Pack a tile directory. Every tile is parsed as it is packed, so a corrupt
# tile fails the build instead of a routing request in production.
calculon-gphx build /data/tiles -o france.gphx --sidecar-index
# Or pack straight from a valhalla_build_extract tarball, without unpacking it.
calculon-gphx build planet.tar -o planet.gphx --sidecar-index
calculon-gphx info france.gphx # header, per-level stats, bounds
calculon-gphx verify france.gphx --against /data/tiles
calculon-gphx extract france.gphx 2/771149 -o t.gph # one tile back out
calculon-gphx index france.gphx -o france.gphx.idx # re-emit the sidecar
# Populate a disk cache ahead of serving (see Prewarming a cache).
calculon-gphx warm france.gphx --cache-dir /var/cache/calculon/tiles --levels 0,1
build option |
Meaning |
|---|---|
--compress zstd\|none |
Blob compression (default zstd) |
--level N |
zstd level (default 9) |
--jobs N |
Compression threads (default: cores) |
--sidecar-index |
Also write <out>.idx |
--meta PATH |
Metadata JSON to embed (default: valhalla.json beside the tiles, or the member inside a tarball) |
--from-tar |
Force reading the source as a tarball (a file source is detected automatically) |
Building from a tarball
valhalla_build_extract produces a plain uncompressed tar whose members are
512-byte aligned, which is what lets Valhalla mmap tiles straight out of it.
build exploits the same property: it maps the tarball and reads each tile in
place, so packing a 100 GB planet extract needs no unpack step and no ~1.5M
files on disk. The tarball's own valhalla.json member is embedded in the
archive metadata, exactly as the sibling file would be for a directory build.
That makes the in-region conversion a single step:
# On an instance next to the bucket, not over a laptop link.
aws s3 cp s3://bucket/valhalla/3.6.2/planet.tar - \
| ... # or download once, then:
calculon-gphx build planet.tar -o planet.gphx --sidecar-index
aws s3 cp planet.gphx s3://bucket/valhalla/3.6.2/planet.gphx
aws s3 cp planet.gphx.idx s3://bucket/valhalla/3.6.2/planet.gphx.idx
After that nobody needs the tar: servers range-read the archive, and anyone who does want it locally pulls ~2.65x fewer bytes.
Choosing a source
Every source returns byte-identical routes and matrices. They differ in memory and cold-start latency, not in warm throughput.
Measured on France (2051 tiles, 3.8 GB raw), median of 15 warm requests:
| Source | On disk | Startup | Idle RSS | Cold short/long | Warm short/long | Settled RSS |
|---|---|---|---|---|---|---|
| Tile directory (mmap) | 3.8 GB | 481 ms | 8 MB | 95 / 388 ms | 29.8 / 116.2 ms | 172 MB |
| Archive, uncompressed | 3.8 GB | 33 ms | 8 MB | 76 / 274 ms | 29.6 / 114.9 ms | 171 MB |
| Archive, zstd, no cache | 1.45 GB | 34 ms | 9 MB | 112 / 351 ms | 29.4 / 114.5 ms | 343 MB |
| Archive, zstd + cache (cold) | 1.45 GB | 33 ms | 9 MB | 75 / 224 ms | 29.3 / 113.6 ms | 284 MB |
| Archive, zstd + cache (warm) | 1.45 GB | 33 ms | 9 MB | 33 / 115 ms | 29.3 / 112.6 ms | 376 MB |
| HTTP archive, no cache | remote | 247 ms | 11 MB | 165 / 181 ms | 61.8 / 136.8 ms | 937 MB |
| HTTP archive + cache (warm) | remote | 246 ms | 11 MB | 34 / 116 ms | 29.3 / 111.6 ms | 378 MB |
What that says:
- Warm latency is a wash. Every source that keeps its tiles locally lands within noise of every other, so the choice between them is a memory and operations decision, not a speed one. A mapped tile reads as fast as a heap one.
- A warm disk cache is the fastest configuration cold, on long routes by more than 3x against a plain tile directory (115 ms vs 388 ms). The second run of a process, and every run after a restart, reads the working set from local disk.
- Archives start faster than a tile directory (33 ms vs 481 ms) because they read coverage from the index instead of walking 2051 files.
- Nothing is loaded before the first request. Idle RSS is 8-11 MB for every configuration; tiles arrive on demand and the first request that needs a cold tile pays for it.
Throughput under concurrent load is flat across sources too. 24 mixed routes on an 8-core machine:
| Source | 1 worker | 8 workers | p50 @ 8 | p90 @ 8 |
|---|---|---|---|---|
| Tile directory (mmap) | 14.3 req/s | 45.5 req/s | 117.3 ms | 299.1 ms |
| Archive zstd (heap) | 14.2 req/s | 45.4 req/s | 116.8 ms | 295.1 ms |
| Archive zstd + disk cache | 14.7 req/s | 43.5 req/s | 120.9 ms | 311.4 ms |
Mapped tiles do not lose to heap tiles under parallelism: page faults on a shared mapping are cheaper than the contention they replace.
Loopback hides the thing that matters
The HTTP numbers carry no real network latency. Same-region S3 adds roughly 15-25 ms TTFB per request (p99 nearer 100 ms), so request count is the number to watch, not the byte count.
Memory: mapped tiles vs heap tiles
This is the axis that actually separates the sources.
A tile served from a file mapping costs page cache, which the kernel reclaims under pressure and shares between processes. A tile decompressed onto the heap is anonymous memory the OS cannot reclaim, held until the cache evicts it. Same tiles, very different resident footprint.
RSS on France, after a short and a long route plus 30 warm repeats:
| Configuration | Idle | Settled under load |
|---|---|---|
| Tile directory (mmap) | 8 MB | 172 MB |
| Archive uncompressed | 8 MB | 171 MB |
| Archive zstd, no disk cache | 9 MB | 343 MB |
| HTTP archive, no disk cache | 11 MB | 937 MB |
| Archive zstd + disk cache | 9 MB | 376 MB |
| HTTP archive + disk cache | 11 MB | 378 MB |
A compressed or remote source without a disk cache keeps every tile it has
decompressed on the heap, bounded only by --cache-size, and the HTTP case adds
the byte ranges it pulled to get there. Those are anonymous pages the kernel
cannot reclaim.
Settled figures for the disk-cached rows sit above a tile directory partly because the archive mapping and the transient range buffers are counted too, but the bytes that matter are clean, file-backed pages the OS can drop under pressure. Repeated load plateaus rather than growing.
The disk cache fixes it
--tile-cache-dir writes each decompressed tile to local disk once and maps it
back from there. The bytes move from heap to page cache, the cold long route
drops from 351 ms to 115 ms once the cache is populated, and the archive still
distributes as a single 1.45 GB file.
calculon --tile-archive ./france.gphx --tile-cache-dir /var/cache/calculon/tiles
calculon --tile-archive https://tiles.example.com/france.gphx \
--tile-cache-dir /var/cache/calculon/tiles
- Writes are atomic (temp file + rename), so a concurrent reader can never map a half-written tile.
- The cache uses the Valhalla directory layout, so a populated cache
directory is itself a valid
--tile-dir. - A corrupt entry is dropped and re-fetched rather than served.
- Sizing: the cache holds decompressed tiles, so a fully warmed France cache is ~3.8 GB of disk. There is no automatic eviction, so point it at a volume you are happy to fill, or pre-warm and treat it as read-mostly.
- Without a writable cache dir (read-only container, no volume), tiles stay on the heap. The server warns at startup when that is the case.
Choosing
| Deployment | Recommendation |
|---|---|
| Local disk, memory matters | Uncompressed archive: mapped in place, one file, dir-equivalent RSS |
| Local disk, size matters | zstd archive + --tile-cache-dir |
| Tiles in S3 | zstd archive over HTTP + --tile-cache-dir |
| Read-only/ephemeral filesystem | zstd archive, no cache: budget for the heap via --cache-size |
Serving from object storage
calculon --tile-archive https://tiles.example.com/france-2026-08.gphx \
--tile-cache-dir /var/cache/calculon/tiles
The host must honour Range. A 200 where a 206 was expected is rejected
at startup rather than silently treated as tile data, since serving the whole object
for a range request would corrupt every tile.
python3 -m http.server does not support Range
It answers 200 with the whole body and Calculon refuses it. Use
tools/range_http_server.py from this repo for local testing.
Operational notes:
- Publish under immutable, versioned keys (
france-2026-08.gphx), setCache-Control: immutable, and front with a CDN. --tile-cache-dirwrites fetched tiles to local disk, so restarts start warm and repeat reads never re-cross the network.- Requests outside the archive's coverage are rejected from the in-memory directory, before any network IO.
Prewarming a cache
calculon-gphx warm populates a disk cache before a server needs it, so an
instance starts against a warm cache instead of paying for its first requests.
It is a deployment step, not something the server does: the server must stay
useful whether or not warming ran.
# Corridor levels only: small, and every long route crosses them.
calculon-gphx warm https://tiles.example.com/france.gphx \
--cache-dir /var/cache/calculon/tiles --levels 0,1
# Or just your service area, at every level.
calculon-gphx warm https://tiles.example.com/france.gphx \
--cache-dir /var/cache/calculon/tiles --bbox 3.5,43.4,5.1,46.0
Measured on France over HTTP, warming levels 0 and 1 (117 tiles, 699 MB):
| Startup | Idle RSS | Cold short/long | |
|---|---|---|---|
| Empty cache | 282 ms | 11 MB | 277 / 319 ms |
| Prewarmed levels 0+1 | 61 ms | 11 MB | 106 / 284 ms |
| Prewarmed bbox corridor (112 tiles, 620 MB) | 68 ms | 11 MB | 107 / 250 ms |
Warming took 0.8 s for 699 MB over loopback; from real S3 it is bound by the ~219 MB compressed it has to pull. A re-run over a populated cache is a directory scan: 12 ms, nothing written.
How much it buys depends on how well the warmed set matches real traffic. Levels 0 and 1 cover the corridor every long route uses; a bbox covers a service area. Warming everything is just a slow download of the whole archive, and if you intend to hold it all locally an uncompressed archive mapped in place is the better answer.
Never let warming block a boot
A slow or throttled source must degrade to a partly warmed cache, not a failed instance. Use the caps, and let the unit ignore failure:
[Service]
ExecStartPre=-/usr/bin/calculon-gphx warm https://tiles.example.com/france.gphx \
--cache-dir /var/cache/calculon/tiles --levels 0,1 --deadline 60 --max-mb 2048
ExecStart=/usr/bin/calculon --tile-archive https://tiles.example.com/france.gphx \
--tile-cache-dir /var/cache/calculon/tiles
The leading - on ExecStartPre makes failure non-fatal, --deadline and
--max-mb bound the work, and both caps stop cleanly so a re-run resumes
where it left off.
Other notes for that deployment:
- Warm onto instance store, not EBS. The cache is disposable and rebuildable by definition, which is what ephemeral NVMe is for.
- If the tile version is pinned to the image, warm at build time instead. The bandwidth is then paid once rather than per scale-out.
- Autoscaling is the real trade. An instance that joins the load balancer while still cold serves the slow path. Either gate the health check on warming finishing, or accept a slower first minute; which is right depends on whether your p99 covers scale-out.
- Warming is idempotent and resumable, so re-running after a half-finished boot costs one directory scan.
Prefetching
Before a search starts, Calculon works out which tiles it will likely need: the neighbourhood of each location at every level, plus the corridor between them at the coarse levels. It then asks the source for them in one go. The source groups those tiles into contiguous byte ranges, merging any gap under 1 MB, and issues the resulting handful of requests in parallel.
A Montpellier→Lyon route needs 21 tiles, fetched as 6 range requests. Without prefetching those would be 21 round trips discovered one at a time inside the A* loop, the difference between ~100 ms and several seconds on real S3.
Tiles already in the disk cache are mapped from there instead of being re-requested, so a restarted process reads its warmed working set from local disk rather than pulling it back over the network.
Prefetching is a pure optimisation: a tile the prediction misses is simply
fetched when the search asks for it. Long corridors are capped
(MAX_CORRIDOR_TILES) so a Lille→Nice request cannot try to pull a country's
worth of level-2 tiles.
Caching
GraphReader layers four levels in front of the source:
| Level | Storage | Notes |
|---|---|---|
| Thread-local | HashMap of Weak handles |
~20 ns, no synchronization |
| Pinned | HashMap of Arc |
Never evicted |
| Shared | Unbounded map, or byte-weighted LRU | Depends on the source (below) |
| Source | dir / archive / HTTP |
The shared cache strategy follows what the source hands out:
- Mapped tiles (tile directory, uncompressed archive, or any source backed
by
--tile-cache-dir) go in an unbounded map. Their bytes are page cache, so the OS already owns eviction and an application-level bound would only force re-reads of pages the kernel was happy to keep. - Heap tiles (compressed archive or HTTP without a disk cache) go in a
byte-weighted, sharded LRU bounded by
--cache-size.
Two details worth knowing:
- The bound is bytes, not entries. Tiles range from 15 KB to 85 MB, so a
tile count would be meaningless.
--cache-sizeis in MB. - Thread-local entries are
Weak. A strong reference there would pin every tile a worker thread ever touched, and the byte budget would be fiction, since evicting from the shared cache would free nothing.
Eviction is advisory: a tile evicted while an in-flight search still holds it is freed when that search drops it, so resident memory can briefly exceed the budget. Size it with headroom.
Verifying a build
# Every tile decompresses, parses, and matches the source directory
calculon-gphx verify france.gphx --against /data/tiles
# All three sources must return identical route and matrix JSON
tools/compare_tile_sources.sh test_data/monaco_tiles