Forge ingest throughput

Sustained ingest rate of the Forge storage stack, measured on a dedicated AWS box each time a service image, smelt or the harness changes on main, and nightly.

Loading run data…

History

p5 (held by 95% of windows) Median ● valid ▲ availability errors ○ invalid

Runs

How this is measured

What runs. A single AWS m9gd instance runs every Forge service in Docker using the published main images, pinned by digest for the run. After each run the box deletes all containers, volumes, stored objects and its NVMe data, so every run starts cold.

The workload. The storage-qualification drill's import profile writes objects of about 128 MiB through the S3 gateway, reads each back within a minute and restores ranges of them. On the tier 2 box each per-trigger run writes 500 GB and each nightly run 1,200 GB; on tier 1 they wrote 100 GB and 350 GB.

What the number means. The drill measures ingest in 30-second windows. p5 is the highest rate that at least 95% of those windows held. The median is shown beside it. A run with fewer than 20 steady windows is marked, since its p5 is then its slowest window. Read-back and restore get the same p5 and median, over the same steady windows as ingest, and restore also its ranged GETs per second. Both read streams are served from ingot's local spool on the box's NVMe, which keeps every blob it accepted, so their rates measure ingot's read path from that spool.

The gates. Each gate is the lower of the named instance type's measured S3 upload throughput and NVMe sequential write speed. Each measures one limit in isolation, so the stack's reachable rate sits below it. A gate lights when a valid run's p5 reaches the ceiling measured at the time of that run. A gate without a measurement is drawn dashed.

What is modeled. Storage node services and central services run on one host with a fixed 25 ms round trip added between the two groups, standing in for a storage provider's site about 25 ms from central services in us-east-2. Piri stores blobs in Amazon S3 in the same region, in place of the provider's HDD-backed S3 service; AWS S3 is probably the faster of the two. The added delay has no jitter or loss. The drill client runs on the same host with no added delay, shares the host's CPU with the central services, and hashes every body with SHA-256. Connections are plain HTTP/1.1 where production uses TLS and HTTP/2, so the box opens more TCP connections than production and skips TLS and the proxy hop.

Outcomes. Runs with availability errors keep their numbers and are marked; they do not move the thermometer or light a gate. Invalid measurements (the added latency drifted, a container restarted, an image changed mid-run, or ingest stopped before a steady window) and failed runs are shown and excluded. A run that reaches its time limit before its size keeps its numbers and is marked. Calibration runs, made while the setup is tuned, appear in the table only, and so do experiment runs, which measure one pull request's image of a service against the current main images, back to back on the same box.

Traces. A traced run records a share of its requests as traces. Its details, and the headline when it shows that run, link to the traces in Grafana; the link needs a Grafana login.