
Fleet Management
Fleet management is the control layer for the machines game servers run on: how many hosts are reserved, where, and of which kind. It adds and retires hosts, keeps them healthy and updated, drains them before maintenance and reports utilization. Orchestration then places game server instances on those hosts, or on on-demand capacity when the fleet is full.
Also called
fleet manager, game server fleet management, server fleet management
,
What Does Fleet Management Manage?
Fleet management manages machines, not matches. A game server instance is the process that hosts a match. The hosts are the bare metal or virtual machines those instances run on. Starting, running and stopping instances is game server orchestration. Deciding how many hosts exist, where, and of which kind, and keeping them healthy, is fleet management.
How much of it a studio does depends on how its capacity is paid for:
A reserved fleet is hosts set aside for one studio, usually billed monthly. A monthly bill doesn't change with use, so a fleet costs the same busy or idle, and it exists only in the locations that were rented. Keeping those hosts at the right number, in the right places, and busy is the whole job.
On-demand capacity runs on the provider's machines. A server starts when a match asks for one and is billed while it runs. When no server is running, the cost is $0. The provider manages the machines, so the studio has no fleet of its own to manage.
Many games use both: a reserved fleet for a cost-effective share of overall traffic, and on-demand capacity for everything above it. The reserved part is what a fleet manager manages. The fleet as a resource is covered under server fleet, and how much reserved capacity to hold under warm pools.
Article's key insights
What Private Fleets are: reserved, private capacity for dedicated game server hosting on hardware set aside for your studio in a specific location, with no egress fees and no reason to spin servers down between matches.
Two sides of one axis: everything below the zero line is your own fleet, everything above it is capacity that spilled over to public cloud.
Every bar is an hourly average, not a peak: a 4 vCPU deployment that ran 15 minutes shows up as 1 vCPU used, which is why bursty workloads look quieter than they felt.
Idle green is not always usable green: the idle number sums every host in a city, so a fleet with 3 vCPU free across three hosts still cannot fit a 1.5 vCPU deployment.
Small deployments can round into invisibility: a 0.25 vCPU server running 6 minutes is 0.03 vCPU over the hour, roughly 0.25% of a 12 vCPU fleet. Hover to see the real value.
The last three hours are intentionally hidden so deployments still resolving into an Error state do not inflate your usage.
How Do Fleet Management and Orchestration Work Together?
The two layers share the same data but act on different things. The orchestrator places each instance. Fleet management decides what there is to place it on, and runs its own loop on a slower clock:
Observe. How full each host is over hours and days, host health, and demand by location.
Decide. Whether to add a host, retire one, shift capacity to another location, or take one out for maintenance.
Act. Reserve or release hosts, patch and replace them, and drain a host before it's taken out.
Report. Utilization over time, so the next commitment rests on real numbers.
Starting and stopping instances follows auto scaling policies, and binding one match to one server is server allocation.
The two layers meet at the edge of the fleet. Adding a host takes time, so when every host is full, the orchestrator decides what happens to the next match: make it wait, fail it, or start it on on-demand capacity. Fleet management decides how often that edge is reached. A fleet sized to overflow only at peaks is the basis of hybrid orchestration.
What Does Managing a Server Fleet Involve Day to Day?
Reserved capacity brings work that on-demand capacity leaves with the provider:
Packing. The idle number can overstate how much room a fleet has. Three 4 vCPU hosts each running 3 vCPU show 3 vCPU idle, yet no single host can fit a new 1.5 vCPU server, so it goes to the cloud. Sizing a fleet means reading what fits on each host, not only the total. Our guide to reading Private Fleet utilization charts (September 2026) walks through the example.
Draining. Before a host is patched or retired, new matches have to go elsewhere and running ones have to finish.
Health. A host that fails, or is hit by a provider outage or a DDoS attack, needs replacing without stranding the matches sent to it.
Images. Each host keeps the current server image cached, so instances start without a download when a patch goes live.
Scheduling. Hosts take time to add, so known peaks such as a launch, a major update, a sale or a streamer event are planned ahead, using early signals such as wishlists.
Reading utilization. Hourly averages flatten short matches: a 4 vCPU server that ran for 15 minutes shows as 1 vCPU used. A utilization chart answers how much of the fleet earned its keep over a week, not how busy a host was at one moment.
A note on everything above: managing a fleet as its own layer, by hand or with separate tools, is the traditional way of running game servers. Modern multi-cloud and hybrid orchestrators automate both loops together. They can keep the cost-effectiveness of a reserved fleet while adding on-demand capacity in far more locations: servers start closer to more players, hosts stay busier, and peaks overflow instead of waiting for a new host.
On Edgegap, fleet management and orchestration are one automated layer, so most of this list runs without the studio's involvement. The Private Fleet documentation describes it: hosts in a priority list, automatic recycling of a failed host under the same ID, maintenance states that stop new placements, scheduled hosts, and cloud alarms for overflow spend. What stays with the studio is the decision that adds value: the fleet-to-cloud ratio, balancing cost-effectiveness, coverage and more, rather than the cost of operating the fleet. The documentation notes that finding that ratio "typically takes several iterations" (read October 2026). Before operating a fleet in-house, it's worth asking whether the team's hours return more there or on the ratio itself.
Should You Build, Self-Host or Buy a Fleet Manager?
Three paths are common:
Build your own. Tooling written for the game, on hosts the studio rents. It gives full control, and both loops, for hosts and for instances, are the studio's code to write and run.
Self-host an open-source orchestrator. Usually on Kubernetes, with a game-specific layer that knows which servers are in a match. The software is free to install. The larger cost is operating it: clusters in every region, upgrades, monitoring and an on-call rotation. Assess that full scope, not only the integration work.
Use a managed platform. The provider runs the hosts, the instances and the overflow. The studio sets the rules: where hosts sit, in what priority, and when matches overflow.
The deciding cost is rarely the software. It's the engineers who keep both loops running at every hour, and what else they could be building. On Edgegap, fleet management and orchestration run as one automated layer, sized and placed to keep costs down: Private Fleet hosts in 58 locations, with Edge Cloud overflow across 615+ locations (platform data, 18 September 2026).
A word from our sponsor (ourselves!)
Private Fleets give you dedicated bare-metal compute in 58 locations for persistent, cost-effective, egress-free hosting. When your player base spikes, Edge Cloud automatically absorbs the overflow on demand, keeping online performance steady wherever your players are.
Edgegap's Take (just our opinion, take it with a grain of salt!)
A Fleet Should Ideally Be Tuned, Not Bought Blindly
Reserved capacity is often bought on a forecast: sized for the highest peak a studio expects, and locked in for a year or more. Real traffic rarely matches the forecast, and the gap is paid every month in empty servers. Before launch, a studio seldom knows its actual usage.
That makes flexibility worth as much as the hourly rate. Commit for the term your data supports: a month while traffic is new, 6 or 12 months once it's proven, with on-demand overflow covering everything above the fleet.
Then keep tuning. Each week, compare two numbers per location: idle fleet capacity during peak hours, and overflow to the cloud. Idle at peak means the fleet is too big there. Overflow every evening means it could grow. Adjust one host at a time.
Our Private Fleets work this way: a one-month minimum, lower rates for 6- and 12-month terms, and automatic overflow to Edge Cloud when the fleet is full (pricing).
,










