Skip to content

Configuration

Scaling and resources

Set how many instances an app runs, how much CPU and memory each one gets, and which port it listens on.

Each service runs on Cloud Run, which starts and stops instances based on traffic. You control the bounds: minimum and maximum instances, CPU and memory per instance, requests per instance, and the port.

Defaults

SettingDefaultAllowed values
Min instances0 (scale to zero)0 to 100
Max instances151 to 100
Concurrency (requests per instance)801 to 1000
CPU11, 2, 4, 8
Memory512Mi128Mi, 256Mi, 512Mi, 1Gi, 2Gi, 4Gi, 8Gi, 16Gi, 32Gi
CPU always allocatedOffOn or off
Port8080Any port your app listens on

Min instances can't be greater than max instances.

Change scaling with the CLI

elula scale                                    # show current settings
elula scale --min 1 --max 5
elula scale --concurrency 40
elula scale --cpu-always                       # also sets min instances to 1 if it was 0
elula scale --memory 1Gi --cpu 1
elula scale --port 3000

elula scale updates the running service right away, without a rebuild. It only works once the app has been deployed. Running it with no options prints the current settings and what they mean.

Change scaling in the dashboard

Open the app's Scaling tab and pick how the app should run:

ChoiceMin–max instancesWhen to use it
Lowest cost0–3Nothing runs when idle; the first visit after a quiet spell takes a few extra seconds
Balanced (default)0–15Free when idle, room to grow when lots of people show up
Always fast1–15One instance always on, so nobody waits for a start; billed around the clock
High traffic2–50Two always on and room for big spikes

For exact numbers open Advanced: a two-handle slider for min and max instances (left: always on, right: the most at once), the same as number fields, and Concurrency. It explains what your numbers mean as you change them. CPU always allocated is under Resources. Click Save. Once the app is deployed, Save applies the change to the running app straight away, without a rebuild. Before the first deploy, it's saved and used by that deploy.

Scale to zero or stay warm

  • Min instances 0: when nobody uses the app, it runs no instances and costs nothing for compute. The first request after an idle period waits for an instance to start (a cold start).
  • Min instances 1 or more: instances stay running, so there are no cold starts. You pay for them while idle.
elula scale --min 1     # always warm
elula scale --min 0     # scale to zero

How it scales out

Concurrency is how many requests one instance handles at the same time (default 80). When every running instance is that busy, Cloud Run starts another one, up to max instances. So the most the app handles at once is about max instances × concurrency: with the defaults, 15 × 80 = 1,200 requests. Past that, new requests wait for a free slot.

Lower concurrency if your app can only handle a few requests at a time (for example heavy CPU work per request): more instances start sooner. The Scaling tab spells out what your numbers mean as you change them.

CPU and memory rules

Cloud Run requires some combinations. Elula checks them before saving:

If you chooseYou also need
4 CPUsAt least 2Gi of memory
8 CPUsAt least 4Gi of memory
More than 4Gi of memoryAt least 2 CPUs
More than 8Gi of memoryAt least 4 CPUs
More than 16Gi of memory8 CPUs
CPU always allocatedAt least 512Mi of memory and at least 1 min instance

CPU always allocated

By default, an instance only gets CPU while it is handling a request. Turn on CPU always allocated in the Scaling tab if your app does work after sending a response, such as background tasks, queues, timers or websockets. It needs at least 1 min instance (there has to be an instance running to keep working) and 512Mi of memory, and it costs more. Turning it on sets min instances to 1 if it was 0.

Port

Cloud Run sends requests to one port in your container and sets PORT to it. If your app reads PORT, leave the default. If it listens on a fixed port, set that port:

elula scale --port 3000

See Deploy a web app or API.

Jobs

For jobs, only CPU and memory apply. Set them in the Scaling tab; they take effect on the next elula deploy. elula scale is for services. See Jobs and schedules.

Cost

Your apps run in your own Google Cloud project, so Google bills compute directly to you. Higher minimum instances, more CPU or memory, and CPU always allocated all increase that bill.