Configuration
Scaling and resources
Set how many instances an app runs, how much CPU and memory each one gets, and which port it listens on.
Each service runs on Cloud Run, which starts and stops instances based on traffic. You control the bounds: minimum and maximum instances, CPU and memory per instance, requests per instance, and the port.
Defaults
| Setting | Default | Allowed values |
|---|---|---|
| Min instances | 0 (scale to zero) | 0 to 100 |
| Max instances | 15 | 1 to 100 |
| Concurrency (requests per instance) | 80 | 1 to 1000 |
| CPU | 1 | 1, 2, 4, 8 |
| Memory | 512Mi | 128Mi, 256Mi, 512Mi, 1Gi, 2Gi, 4Gi, 8Gi, 16Gi, 32Gi |
| CPU always allocated | Off | On or off |
| Port | 8080 | Any port your app listens on |
Min instances can't be greater than max instances.
Change scaling with the CLI
elula scale # show current settings elula scale --min 1 --max 5 elula scale --concurrency 40 elula scale --cpu-always # also sets min instances to 1 if it was 0 elula scale --memory 1Gi --cpu 1 elula scale --port 3000
elula scale updates the running service right away, without a rebuild. It only works once the app has been deployed. Running it with no options prints the current settings and what they mean.
Change scaling in the dashboard
Open the app's Scaling tab and pick how the app should run:
| Choice | Min–max instances | When to use it |
|---|---|---|
| Lowest cost | 0–3 | Nothing runs when idle; the first visit after a quiet spell takes a few extra seconds |
| Balanced (default) | 0–15 | Free when idle, room to grow when lots of people show up |
| Always fast | 1–15 | One instance always on, so nobody waits for a start; billed around the clock |
| High traffic | 2–50 | Two always on and room for big spikes |
For exact numbers open Advanced: a two-handle slider for min and max instances (left: always on, right: the most at once), the same as number fields, and Concurrency. It explains what your numbers mean as you change them. CPU always allocated is under Resources. Click Save. Once the app is deployed, Save applies the change to the running app straight away, without a rebuild. Before the first deploy, it's saved and used by that deploy.
Scale to zero or stay warm
- Min instances 0: when nobody uses the app, it runs no instances and costs nothing for compute. The first request after an idle period waits for an instance to start (a cold start).
- Min instances 1 or more: instances stay running, so there are no cold starts. You pay for them while idle.
elula scale --min 1 # always warm elula scale --min 0 # scale to zero
How it scales out
Concurrency is how many requests one instance handles at the same time (default 80). When every running instance is that busy, Cloud Run starts another one, up to max instances. So the most the app handles at once is about max instances × concurrency: with the defaults, 15 × 80 = 1,200 requests. Past that, new requests wait for a free slot.
Lower concurrency if your app can only handle a few requests at a time (for example heavy CPU work per request): more instances start sooner. The Scaling tab spells out what your numbers mean as you change them.
CPU and memory rules
Cloud Run requires some combinations. Elula checks them before saving:
| If you choose | You also need |
|---|---|
| 4 CPUs | At least 2Gi of memory |
| 8 CPUs | At least 4Gi of memory |
| More than 4Gi of memory | At least 2 CPUs |
| More than 8Gi of memory | At least 4 CPUs |
| More than 16Gi of memory | 8 CPUs |
| CPU always allocated | At least 512Mi of memory and at least 1 min instance |
CPU always allocated
By default, an instance only gets CPU while it is handling a request. Turn on CPU always allocated in the Scaling tab if your app does work after sending a response, such as background tasks, queues, timers or websockets. It needs at least 1 min instance (there has to be an instance running to keep working) and 512Mi of memory, and it costs more. Turning it on sets min instances to 1 if it was 0.
Port
Cloud Run sends requests to one port in your container and sets PORT to it. If your app reads PORT, leave the default. If it listens on a fixed port, set that port:
elula scale --port 3000
Jobs
For jobs, only CPU and memory apply. Set them in the Scaling tab; they take effect on the next elula deploy. elula scale is for services. See Jobs and schedules.
Cost
Your apps run in your own Google Cloud project, so Google bills compute directly to you. Higher minimum instances, more CPU or memory, and CPU always allocated all increase that bill.