Configuring fleet scaling
When you create a fleet, use scaling parameters to control how many workers your Code Engine fleet deploys and how it responds to changes in task load.
All scaling parameters are optional and can be set when you create a fleet.
Scaling parameters in the Code Engine CLI
To create a fleet, you use the ibmcloud ce fleet run or ibmcloud ce fleet create command. It takes the following options for fleet scaling. For a complete list of all available command options, see the CLI reference.
--scale-max-
The upper bound on the total number of concurrent task slots across all workers. Workers scale up to support at most this many concurrent tasks. Defaults to
10. --scale-min-
The minimum number of concurrent task slots that are always kept provisioned, even when there are no pending or running tasks. Workers providing this capacity remain active and are never scaled down automatically.
A fleet with active workers due to this setting, but without pending or running tasks has the Standby status.
Must be less than or equal to
--scale-max. Defaults to0. --scale-spare-
The number of free concurrent task slots maintained above the current number of running tasks. When all workers are busy, the fleet provisions additional workers to restore the spare buffer — up to
--scale-max. New tasks arriving while spare capacity exists start immediately without waiting for a new worker.Must be less than or equal to
--scale-max. Defaults to0. --scale-down-delay-
Specifies how long a worker continues running after it completes its last task before Code Engine scales it down. A worker whose last task completed within the delay window picks up new tasks immediately when they arrive.
Specify a positive integer for the number of seconds. Defaults to
0(workers scale down as soon as they are no longer needed).
Scaling parameters in the Code Engine console
- Maximum container slots
-
The upper bound on the total number of concurrent task slots across all workers. Workers scale up to support at most this many concurrent tasks. Defaults to
10. - Minimum container slots
-
The minimum number of concurrent task slots that are always kept provisioned, even when there are no pending or running tasks. Workers providing this capacity remain active and are never scaled down automatically.
A fleet with active workers due to this setting, but without pending or running tasks has the Standby status.
Must be less than or equal to Maximum container slots. Defaults to
0. - Spare container slots
-
The number of free concurrent task slots maintained above the current number of running tasks. When all workers are busy, the fleet provisions additional workers to restore the spare buffer — up to Maximum container slots. New tasks arriving while spare capacity exists start immediately without waiting for a new worker.
Must be less than or equal to Maximum container slots. Defaults to
0. - Scale-down delay
-
Specifies how long a worker continues running after it completes its last task before Code Engine scales it down. A worker whose last task completed within the delay window picks up new tasks immediately when they arrive.
Specify the number of minutes. Defaults to
0(workers scale down as soon as they are no longer needed).
Workload patterns and relevant parameters
Which parameters to use depends on how tasks are produced, what your latency requirements are, and how stable the fleet configuration is.
One-shot batch workload in the Code Engine CLI
You know all tasks up front, supply them when you create the fleet, and the fleet finishes when the last task completes. This is the simplest model and requires no special scaling configuration beyond setting --scale-max to control
how many tasks run in parallel. Use this approach when you need a dedicated fleet configuration for a batch of tasks.
Relevant option: --scale-max
One-shot batch workload in the Code Engine console
You know all tasks up front, supply them when you create the fleet, and the fleet finishes when the last task completes. This is the simplest model and requires no special scaling configuration beyond setting Maximum container slots to control how many tasks run in parallel. Use this approach when you need a dedicated fleet configuration for a batch of tasks.
Relevant parameter: Maximum container slots
Scheduled batch workload in the Code Engine CLI
You keep a fleet around and add a new batch of tasks on a schedule. Between batches, the fleet has no pending tasks. The gap between batches is predictable and long enough that keeping workers warm is not worthwhile, so there is no benefit
to using --scale-min, --scale-spare, or --scale-down-delay. The fleet configuration remains stable for all batches.
Relevant option: --scale-max
Scheduled batch workload in the Code Engine console
You keep a fleet around and add a new batch of tasks on a schedule. Between batches, the fleet has no pending tasks. The gap between batches is predictable and long enough that keeping workers warm is not worthwhile, so there is no benefit to using Minimum container slots, Spare container slots, or Scale-down delay. The fleet configuration remains stable for all batches.
Relevant parameter: Maximum container slots
Continuous, latency-sensitive workload in the Code Engine CLI
Tasks arrive unpredictably throughout the day — triggered by user actions, sensor readings, or other real-time events. Waiting for a new worker to boot each time a task arrives would add unacceptable latency. Instead, configure the fleet to keep a pool of warm workers ready so that tasks can start immediately.
Relevant options: --scale-max, --scale-min, --scale-spare, --scale-down-delay
- Use
--scale-maxto define the maximum number of container slots that can be provisioned. Your fleet scales up to this concurrency limit based on workload demand. - Use
--scale-minto ensure a baseline worker pool is always running, even when no tasks are queued. Use this option when traffic can drop to zero for a period, but you still want an instant start when it resumes. - Use
--scale-spareto keep a number of free task slots available above the current number of running tasks. When a task arrives, processing begins immediately on an already-running worker, and the fleet provisions more workers in the background to replenish the spare capacity. - Use
--scale-down-delayto prevent workers from scaling down immediately after their last task completes. This setting helps when tasks arrive intermittently with temporary gaps between them, allowing workers to remain available and avoiding unnecessary scale-down and scale-up cycles. Keeping workers warm reduces task startup latency, minimizes worker initialization overhead, and lowers the cost of repeatedly provisioning new workers.
Continuous, latency-sensitive workload in the Code Engine console
Tasks arrive unpredictably throughout the day — triggered by user actions, sensor readings, or other real-time events. Waiting for a new worker to boot each time a task arrives would add unacceptable latency. Instead, configure the fleet to keep a pool of warm workers ready so that tasks can start immediately.
Relevant parameters: Maximum container slots, Minimum container slots, Spare container slots, Scale-down delay
- Use Maximum container slots to define the maximum number of container slots that can be provisioned. Your fleet scales up to this concurrency limit based on workload demand.
- Use Minimum container slots to ensure a baseline worker pool is always running, even when no tasks are queued. Use this option when traffic can drop to zero for a period, but you still want an instant start when it resumes.
- Use Spare container slots to keep a number of free task slots available above the current number of running tasks. When a task arrives, processing begins immediately on an already-running worker, and the fleet provisions more workers in the background to replenish the spare capacity.
- Use Scale-down delay to prevent workers from scaling down immediately after their last task completes. This setting helps when tasks arrive intermittently with temporary gaps between them, allowing workers to remain available and avoiding unnecessary scale-down and scale-up cycles. Keeping workers warm reduces task startup latency, minimizes worker initialization overhead, and lowers the cost of repeatedly provisioning new workers.