Skip to main content
This benchmark answers one question for the recommended production shape: at the 1:10 app-to-worker ratio, how does throughput scale as you grow the fleet from 40 to 160 workers? It runs the worker-is-the-sandbox model on a real GKE cluster, against a same-region object store with signed URLs and official piece tarballs served from the CDN. Activepieces production architecture: client, app, Redis job queue, Postgres, S3, and one-flow-per-worker execution tier The shape under test: app tier, Redis job queue, Postgres, S3, and a one-flow-per-worker execution tier.

What’s measured

A 4-node synchronous webhook flow that holds the HTTP connection open until the flow returns:
1

Webhook trigger

Catches the request on a /sync URL and holds the connection until the flow finishes.
2

Math Helper

Adds 2 + 3.
3

Code step

Runs return inputs.sum + 1 inside an isolated-vm context.
4

Webhook response

Returns the result, closing the held connection.
The compute is sub-millisecond by design — everything measured below is orchestration (queueing, callbacks, sandbox boot), which is what actually shapes production latency. Each fleet size is held at the recommended 1:10 ratio (1 app per 10 workers) and run warm (AP_REUSE_SANDBOX=true) — the engine process is reused between jobs.

Results

686 req/s

Peak warm throughput — 16 apps, 160 workers.

~4.5 req/s

Per worker, held flat from 40 to 160 workers — throughput scales linearly with the fleet.
Each worker is one sandbox at concurrency 1, hard-capped at 0.5 vCPU / 1 GB. Apps are 1 vCPU / 1 GB. Load concurrency is matched to the worker count so requests don’t queue behind the concurrency-1 workers.

What each tier ran — and what it was actually doing

Only the app and worker counts scale (1:10). Postgres and Redis are a single fixed-size pod each — the same for every row below. CPU is the average across three warm load tests; the singletons’ figures are the whole pod, app/worker are per pod. Postgres never crosses ~0.65 of a core and Redis never crosses ~0.17, both far below their caps and flat as the fleet quadruples — they are not absorbing a growing share of anything. Workers sit at ≤0.1 of their 0.5-core cap. No tier approaches saturation, which is exactly why each added worker keeps adding throughput. (The singletons are sized this large on purpose — see Test environment — so they provably stay off the critical path; the default Postgres max_connections=100 would cap the fleet at ~10 apps, which is the artifact behind the earlier “120 cliff”.)

How throughput scales

  • Warm scales linearly with the fleet. Per-worker throughput stays flat at ~4.5 req/s from 40 to 160 workers, so total throughput tracks the worker count (185 → 410 → 553 → 686; 3.7× for 4× the fleet). The shared Postgres and Redis singletons are not the wall — they sit near-idle at every fleet size (Postgres under 0.6 of a core, Redis under 0.2, both far below their caps), and raising their resources several-fold does not move the curve. The ceiling is the concurrency-1 worker model: each worker is busy for the whole per-flow time — engine run plus the end-of-run run-log persistence it finishes before taking the next job — so fleet throughput is workers ÷ per-flow-time, which is linear in the fleet. (The synchronous response reaches the client sooner than that — it is sent at the response step, before the worker wraps up the log write — so client-perceived latency is lower than the worker-busy time that sets throughput.) Per-flow time carries run-to-run variance (the object-store log-write tail), which is why a single run’s curve looks bumpy; the invariant that the per-worker rate holds constant is what shows the scaling is linear.
Why Production Setup recommends 1:10. Apps at 1 vCPU are cheap relative to the worker fleet, and 1:10 is the warm-headroom margin that keeps the app tier from becoming the wall during bursts. See Production Setup.

Latency anatomy

Where the worker’s milliseconds go — warm at peak (16 app · 160 w): This is the time the worker is occupied per job — and at concurrency 1 it is what sets throughput (workers ÷ worker-busy-time). The synchronous client sees less: the response is published at the flow’s response step, before the worker finishes persisting the run log, so client-perceived latency runs below the worker-busy figure.

Test environment

  • Cluster: GKE n2-standard-16 × 10 nodes, europe-west1-b
  • Worker: 0.5 vCPU / 1 GB, concurrency 1, SANDBOX_CODE_ONLY (Node fork + isolated-vm)
  • App: 1 vCPU / 1 GB
  • Object store: same-region GCS bucket (europe-west1) over the S3-interop endpoint, path-style SigV4 presigned URLs (AP_S3_USE_SIGNED_URLS=true)
  • Piece bundles: official tarballs served from the Activepieces CDN (AP_USE_CDN_FOR_BUNDLES=true)
  • Postgres + Redis: in-cluster singletons, deliberately over-provisioned so they stay off the critical path — Postgres at 3 vCPU / 3 GB with max_connections=2000 (the default 100 would starve the app pools past ~10 apps), durability off, and its data dir on tmpfs; Redis at 2 vCPU / 2 GB with io-threads. Under load both stay near-idle (Postgres <0.6 core, Redis <0.2), confirming the worker tier, not the singletons, is the ceiling.
  • Load: hey, concurrency matched to worker count (40/80/120/160) so requests don’t queue behind the concurrency-1 workers — latency reflects real service time, not backlog

How to reproduce

The script mints a worker token, deploys benchmark/k8s-sandbox.yaml to the cluster, runs the load test against the app LoadBalancer, and reports warm throughput and the per-run breakdown from worker-pod logs. Set APP_REPLICAS and WORKER_REPLICAS (keeping the 1:10 ratio) to reproduce any row in the results table.
This benchmark runs in SANDBOX_CODE_ONLY mode. It does not represent the performance of Activepieces Cloud, which uses a different sandboxing mechanism for multi-tenancy. See Sandboxing.

Benchmark your own installation

Load-test and diagnose your own deployment — no cluster scripts required — with the CLI. It publishes a synchronous flow (webhook trigger → data mapper → return response), fires load at its sync webhook endpoint with autocannon, and returns one self-contained diagnostic bundle you can hand to support.
The CLI provisions a throwaway project for the run (with a high concurrency cap so a project rate limiter can’t queue-throttle the numbers) and deletes it when finished — nothing is left behind, and it never touches your real projects. The API key must be a platform-admin key. Why the numbers are trustworthy. The CLI runs from a different region than your servers, so its client-side latency is polluted by network distance. The numbers that matter are measured server-side instead: the QUEUE/PROVISION/BOOT/RUN split is timed inside the worker, and DB/Redis/S3 round-trips are measured in-region by an admin diagnostics endpoint. Client numbers are shown but marked observational. Concurrency defaults to your execution slotsAP_WORKER_CONCURRENCY across connected workers) so requests don’t queue and you read real service time, not backlog. Comparing two deployments? Match concurrency to each one’s own slots — never a fixed number, which makes the smaller one queue.

Reference numbers

A real run against the recommended GKE deployment from the results table: 4 workers @ 0.5 vCPU / 1 GB, concurrency 1, SANDBOX_CODE_ONLY, AP_REUSE_SANDBOX=true, same-region GCS with signed URLs, warm, load = concurrency 4 (= slots) × 200 requests. Run the CLI against your own deployment and compare tier by tier. A number several times larger localizes the problem: RUN ≫ 200 ms means a heavier flow or a CPU-starved worker; storage ≫ 240 ms means a mis-regioned or throttled object store; a large QUEUE with climbing queue depth means you drove more concurrency than you have slots.
Each section answers one question: Here the read is unambiguous: workers pegged at ~100% of their 0.5-core limit and RUN ≈ 200 ms dominate, while QUEUE (~100 ms at concurrency = slots) and the infra round-trips are small — the deployment is service-bound on worker CPU, so the lever is more/bigger workers, not a code change.
  • Exits non-zero if any request fails — usable as a CI gate.
  • The throwaway project (and its benchmark flow) is deleted automatically when the run finishes.
If your hardware or sandbox mode differs from the recommended shape, the absolute numbers shift — but the shape holds: match concurrency to your own slot count so nothing queues, then read whether you are queue-bound or service-bound and which tier dominates.