Discover how Edgegap works

Discover how Edgegap works

Managing Unplanned Traffic Surge: How Edgegap’s Orchestrator Handles Scaling for your Game

Published

Published

Published

Last reviewed for accuracy

Last reviewed for accuracy

Last reviewed for accuracy

You have added Edgegap game servers and orchestration to your game (congrats!) and now you are ready to take the next step – launching your multiplayer game server hosting in the real world.

This can feel daunting. What if it doesn't scale, like so many major projects before it?

This guide is meant to help you prepare for your game launch by highlighting how Edgegap handles a surge in traffic so your game scales in sync with player demand.

This article should be used in tandem with our "Multiplayer Game Pre-Launch Checklist".

How does Edgegap manage unplanned spike load on its platform?

First and foremost, we handle each deployment as a distinct entity on our platform. Every deployment is a pod within our infrastructure. Requests originate from various sources, including the game clients themselves, the matchmaker, or any other means through which users interact with our API.

Once a request gets to our system, we immediately initiate telemetry and geolocation of the player and assign tasks to our workers to provision a server on our infrastructure. Within a second or two, the container is up and running, ready for incoming players.

How do we ensure all the necessary infrastructure is in place to serve these deployments promptly?

We've developed an abstraction layer overseeing 17 different providers, spanning cloud, bare metal, edge and Container-as-a-Service providers. Together, they give us more than 600 data centers worldwide.

Our list of providers mixes those with worldwide service, such as Amazon, Google, and Azure, with localized services specific to a country or region. This complementary approach ensures you have the best coverage, so your game can tap into the world's largest distributed edge network.

Within our platform, these entities are known as Access Points, each with essential information such as capacity, current workload, available resources, and other relevant metrics for decision-making. Thanks to this abstraction layer, you do not need to worry about where your players will come from, or which provider you need to work with.

This allows you to focus on the game, while Edgegap focuses on orchestrating and managing the underlying infrastructure.

What are Edgegap's platform priorities, and in which order?

Edgegap's platform will prioritize the following elements, in this order:

  1. Provide you with a functioning game server, based on the requirements you specify in your application profile.

  2. Get the game server as close as possible to the players, with the shortest networking path, to provide the best player experience, based on a variety of telemetry measurements.

  3. Get you a functioning game server as quickly as possible. The objective is 1-2 seconds.

When we get a surge (and we mean a HUGE amount of traffic in a very SHORT amount of time; e.g., 40+ deployments per second, which means 1,300 players trying to get into a match in a 32-player game every second, that's 4.6 million players in 60 minutes), the list above becomes harder to provide for the first few minutes of the surge. In those rare cases, item 2 (proximity) is the first to be relaxed: some servers may be placed slightly further from players while the system scales up, so that every deployment request is still met.

How long does it typically take for the platform to stabilize in case of an unplanned surge of traffic?

Edgegap's platform routinely handles far more than 40 deployments per second across all the games it hosts, so everyday spikes never hit this ramp. The 10-15 minute window only applies to an extreme case: a single game suddenly adding 40+ deployments per second of unplanned traffic (roughly 4.6 million players in an hour). Even then, servers keep deploying. Deploy times for this game can exceed the 1-2 second target until new capacity settles, which can take up to 10-15 minutes. This is fleet capacity ramp time in an extreme case, not typical server boot time. Under normal conditions, servers start in a median of 2 seconds from cold start (2026, platform data).

What is Edgegap's approach when faced with a sudden, unplanned surge in traffic?

Our Capacity Manager (nicknamed "Capman") is primed to react to incoming traffic, interfacing with the workers requesting resources.

When it detects an increase in demand, it automatically scales up the specific regions where resources are needed by deploying larger machines. If it detects a geographical area that needs additional support, it provisions more machines there to optimize latency and the player experience. All of this is done automatically.

For larger spikes where a lot of CPU is required (i.e. a game server that requires 2 CPUs for each instance), the first scale-up takes between 2 and 5 minutes. After that, the system scales up much faster since it requests larger servers. It adjusts the size of the servers it requests from each provider based on the influx of traffic, requesting smaller or larger servers as the influx accelerates or slows down.

TLDR: for CPU-heavy surges, the first scale-up takes 2 to 5 minutes. Full stabilization during an extreme, unplanned surge (40+ deployments per second from a single game) can take up to 10-15 minutes, and servers keep deploying throughout. Standard daily traffic is not affected.

Additional notes: this does not happen when we know about an event in advance (launch, patches, etc.), nor for normal scaling up and down (typical 24-hour traffic). It only applies to unplanned spikes, e.g. a streamer starts playing your game and a ton of new players join the fun.

Why does this not happen for standard scale-up traffic over 24 hours, and only for an UNPLANNED surge?

To manage daily traffic, our platform is designed to follow the daily peak as it moves west across time zones, scaling up the next active region ahead of time.

By leveraging predictive analytics and monitoring traffic patterns, we anticipate when and where traffic spikes are likely to occur. Our infrastructure proactively adjusts its capacity ahead of these peak periods.

This proactive scaling ensures the necessary resources are available when and where they are needed, maintaining smooth gameplay even during periods of high demand.

What is the behavior when Edgegap is low on access points and newly allocated access points are in a region far away from a player?

When a deployment is requested, the platform makes the best decision based on the resources available. If the platform receives an unplanned large number of requests for a specific region, resources can become scarce in that region while the system scales up.

In such cases, and only while the system scales up, servers may be placed slightly further away than the initial best location.

The objective is to always have capacity, regardless of the situation. These worst-case scenarios happen rarely and are temporary, as the system scales up and learns the new traffic behavior.

What can I do to get the best performance from Edgegap's platform?

Here are a few things you can do to get the best out of our platform ahead of a planned surge in traffic.

Here is the checklist:

  1. Communicate!

    • If you know of something happening in your game, i.e. launch, patches, tournaments, streamers, let the team know about it.

    • Planned events (launches, patches, tournaments) shared with Edgegap ahead of time are not subject to this ramp, because capacity is pre-heated.

    • This is only recommended for large planned events, as the platform learns your traffic pattern over time (i.e. low and high traffic tides over 24 hours) and scales up and down accordingly.

  2. To avoid timeout errors, you can increase the "max_time_to_deploy" value of your AppVersion through either the dashboard or our API.

    • The timeout is the amount of time our system will wait before flagging a game server (a deployment) as in error. During an extreme surge, individual deployments for the game driving the surge can take longer than usual to become ready. Other games on the platform are not affected.

    • Setting this timeout slightly higher prevents deployments that take a little longer during a traffic surge from being flagged as in error by our platform.

  3. We recommend setting up your system (i.e. lobby, matchmaker) to retry should this happen.

    • The platform will always do everything possible to get you a game server through a simple API call. That said, a very (and we mean VERY!) large amount of traffic in a very short time can force our system to return an error message. Build retry logic into your matchmaker with exponential backoff (wait longer between each attempt, never re-poll every second). This is standard resilience practice for any orchestrator under sudden load.

  4. While talking about a disaster recovery plan may seem counterintuitive in a document about performance, we feel that covering every single aspect of every single scenario is at the heart of what we do and how we operate.

    • Our system is based on a microservices architecture, with highly resilient components.

    • Those components are vendor agnostic: they can run on any provider and can be easily deployed or redeployed to add capacity, or to reinstall a new platform in case of catastrophic failure.

    • On top of that, non-critical components are not required for the main service, so orchestrating and hosting game servers always remain the utmost priority.

  5. Prepare with a Load Test!

    • There's no better way to assess the behavior of your game's infrastructure, from Edgegap to game services, than with a load test.

    • We strongly recommend running one ahead of your launch, as it is part of our pre-launch checklist. Check the article for more details.

How do we maintain competitive pricing?

As traffic stabilizes, our automated systems scale the infrastructure back down while keeping the capacity needed to serve your players. Traffic patterns and usage trends are monitored continuously, and predictive analytics anticipate demand so capacity adjusts ahead of time.

Cost optimization goes beyond scaling down. We continually review how we use each provider, including options such as reserved instances. Keeping the infrastructure lean lets us pass those savings on to our customers without trading away performance or reliability.

Conclusion

There is a lot Edgegap does to make sure that whatever traffic volume your game gets, it can scale rapidly to meet player demand.

Still, there is a lot you can do to plan ahead and prevent errors. We hope this guide helps you improve your backend and, ultimately, gives you the peace of mind that your multiplayer is built to scale with player demand, and that your players get a fun, smooth experience.

Written by

Philip Côté (CTO)

Sources and/or content collaboration with

Gabriel Parent (Director) for formatting.

Get your Game Online Easily & in Minutes

Start Integrating Now!

Get your Game Online Easily
& in Minutes

Get your Game Online Easily & in Minutes