
What’s measured
A 4-node synchronous webhook flow that holds the HTTP connection open until the flow returns:1
Webhook trigger
Catches the request on a
/sync URL and holds the connection until the flow finishes.2
Math Helper
Adds
2 + 3.3
Code step
Runs
return inputs.sum + 1 inside an isolated-vm context.4
Webhook response
Returns the result, closing the held connection.
AP_REUSE_SANDBOX=true) — the engine process is reused between jobs.
Results
686 req/s
Peak warm throughput — 16 apps, 160 workers.
~4.5 req/s
Per worker, held flat from 40 to 160 workers — throughput scales linearly with the fleet.
What each tier ran — and what it was actually doing
Only the app and worker counts scale (1:10). Postgres and Redis are a single fixed-size pod each — the same for every row below. CPU is the average across three warm load tests; the singletons’ figures are the whole pod, app/worker are per pod.
Postgres never crosses ~0.65 of a core and Redis never crosses ~0.17, both far below their caps and flat as the fleet quadruples — they are not absorbing a growing share of anything. Workers sit at ≤0.1 of their 0.5-core cap. No tier approaches saturation, which is exactly why each added worker keeps adding throughput. (The singletons are sized this large on purpose — see Test environment — so they provably stay off the critical path; the default Postgres
max_connections=100 would cap the fleet at ~10 apps, which is the artifact behind the earlier “120 cliff”.)
How throughput scales
- Warm scales linearly with the fleet. Per-worker throughput stays flat at ~4.5 req/s from 40 to 160 workers, so total throughput tracks the worker count (185 → 410 → 553 → 686; 3.7× for 4× the fleet). The shared Postgres and Redis singletons are not the wall — they sit near-idle at every fleet size (Postgres under 0.6 of a core, Redis under 0.2, both far below their caps), and raising their resources several-fold does not move the curve. The ceiling is the concurrency-1 worker model: each worker is busy for the whole per-flow time — engine run plus the end-of-run run-log persistence it finishes before taking the next job — so fleet throughput is
workers ÷ per-flow-time, which is linear in the fleet. (The synchronous response reaches the client sooner than that — it is sent at the response step, before the worker wraps up the log write — so client-perceived latency is lower than the worker-busy time that sets throughput.) Per-flow time carries run-to-run variance (the object-store log-write tail), which is why a single run’s curve looks bumpy; the invariant that the per-worker rate holds constant is what shows the scaling is linear.
Why Production Setup recommends 1:10. Apps at 1 vCPU are cheap relative to the worker fleet, and 1:10 is the warm-headroom margin that keeps the app tier from becoming the wall during bursts. See Production Setup.
Latency anatomy
Where the worker’s milliseconds go — warm at peak (16 app · 160 w):
This is the time the worker is occupied per job — and at concurrency 1 it is what sets throughput (
workers ÷ worker-busy-time). The synchronous client sees less: the response is published at the flow’s response step, before the worker finishes persisting the run log, so client-perceived latency runs below the worker-busy figure.
Test environment
- Cluster: GKE
n2-standard-16× 10 nodes,europe-west1-b - Worker: 0.5 vCPU / 1 GB, concurrency 1,
SANDBOX_CODE_ONLY(Node fork +isolated-vm) - App: 1 vCPU / 1 GB
- Object store: same-region GCS bucket (
europe-west1) over the S3-interop endpoint, path-style SigV4 presigned URLs (AP_S3_USE_SIGNED_URLS=true) - Piece bundles: official tarballs served from the Activepieces CDN (
AP_USE_CDN_FOR_BUNDLES=true) - Postgres + Redis: in-cluster singletons, deliberately over-provisioned so they stay off the critical path — Postgres at 3 vCPU / 3 GB with
max_connections=2000(the default 100 would starve the app pools past ~10 apps), durability off, and its data dir on tmpfs; Redis at 2 vCPU / 2 GB withio-threads. Under load both stay near-idle (Postgres<0.6core, Redis<0.2), confirming the worker tier, not the singletons, is the ceiling. - Load:
hey, concurrency matched to worker count (40/80/120/160) so requests don’t queue behind the concurrency-1 workers — latency reflects real service time, not backlog
How to reproduce
benchmark/k8s-sandbox.yaml to the cluster, runs the load test against the app LoadBalancer, and reports warm throughput and the per-run breakdown from worker-pod logs. Set APP_REPLICAS and WORKER_REPLICAS (keeping the 1:10 ratio) to reproduce any row in the results table.
Benchmark your own installation
Load-test and diagnose your own deployment — no cluster scripts required — with the CLI. It publishes a synchronous flow (webhook trigger → data mapper → return response), fires load at its sync webhook endpoint with autocannon, and returns one self-contained diagnostic bundle you can hand to support.AP_WORKER_CONCURRENCY across connected workers) so requests don’t queue and you read real service time, not backlog. Comparing two deployments? Match concurrency to each one’s own slots — never a fixed number, which makes the smaller one queue.
Reference numbers
A real run against the recommended GKE deployment from the results table: 4 workers @ 0.5 vCPU / 1 GB, concurrency 1,SANDBOX_CODE_ONLY, AP_REUSE_SANDBOX=true, same-region GCS with signed URLs, warm, load = concurrency 4 (= slots) × 200 requests.
Run the CLI against your own deployment and compare tier by tier. A number several times larger localizes the problem: RUN ≫ 200 ms means a heavier flow or a CPU-starved worker; storage ≫ 240 ms means a mis-regioned or throttled object store; a large QUEUE with climbing queue depth means you drove more concurrency than you have slots.
Here the read is unambiguous: workers pegged at ~100% of their 0.5-core limit and RUN ≈ 200 ms dominate, while QUEUE (~100 ms at concurrency = slots) and the infra round-trips are small — the deployment is service-bound on worker CPU, so the lever is more/bigger workers, not a code change.
- Exits non-zero if any request fails — usable as a CI gate.
- The throwaway project (and its benchmark flow) is deleted automatically when the run finishes.
If your hardware or sandbox mode differs from the recommended shape, the absolute numbers shift — but the shape holds: match concurrency to your own slot count so nothing queues, then read whether you are queue-bound or service-bound and which tier dominates.