Library Registry
Oct 06, 2026
Rahul Rawat

From Theory to Production: Vertical vs Horizontal Scaling.

scalingconcurrencysystem designgcpinfrastructure
From Theory to Production: Vertical vs Horizontal Scaling

I've known scaling for a long time. Vertical scaling, horizontal scaling, concurrency. I've read about them, drawn the diagrams, and explained them in interviews:

  • Vertical scaling → make the machine bigger.
  • Horizontal scaling → add more machines.
  • Concurrency → handle multiple things at the same time.

But I had never actually applied them to a real production system. Knowing the concepts and using them for something that has to handle real traffic turned out to be two very different experiences.

Today, while working with GCP, I had to put real numbers behind those concepts: instance capacity, requests per second, CPU and memory requirements, and how many instances would actually be needed to handle the expected load.

This post is about what changed when the theory met real configuration.

Putting real numbers on it

Let's say my application receives 1,000 requests per second.

After testing, I find that one instance can reliably handle 100 requests per second.

Ignoring redundancy and autoscaling for now, I need roughly:

1,000 / 100 = 10 instances

That's horizontal scaling. Instead of continuously making one machine more powerful, I distribute the workload across multiple instances.

But there's another option. Suppose I upgrade the instance to a machine that can handle 500 requests per second. Now I need only:

1,000 / 500 = 2 instances

That's vertical scaling: fewer, bigger machines.

The definitions were never the hard part. What was new today was making them drive real infrastructure decisions.

Vertical scaling

Vertical scaling means increasing the capacity of an existing machine.

Before:              After:
4 CPU                16 CPU
16 GB RAM            64 GB RAM
100 RPS              500 RPS

You are essentially saying: "This workload is fine on one machine. I just need a stronger machine."

It's simple and often useful, especially when an application isn't designed to run across multiple instances.

But there are limits:

  • There is a ceiling. A machine can only get so large, and larger instances get disproportionately expensive.
  • It's a single point of failure. If everything runs on one instance, your whole application goes down with it.
  • You pay for peak all the time. A big machine costs the same at 3 AM as it does at peak hour.

Horizontal scaling

Horizontal scaling takes a different approach. Instead of making one machine bigger, we add more machines.

              Load Balancer
                   |
       +-----------+-----------+
       |           |           |
   Instance    Instance    Instance
   100 RPS     100 RPS     100 RPS

If each instance handles 100 RPS:

3 instances  → ~300 RPS
10 instances → ~1,000 RPS

Now capacity grows by adding instances. This is where concepts like load balancers, autoscaling, health checks, service discovery and stateless services start becoming important.

That last one matters more in practice than it looks on a diagram. If a user's session lives in one instance's memory, their next request might land on a different instance and the session is gone. So state has to move somewhere shared, like Redis or the database.

This is where scaling stops being just a "machine size" problem. It becomes a distributed systems problem.

Try both

Here are the two side by side. Switch between vertical and horizontal, push the traffic up, and then try taking a machine down. Watch what happens to the one big machine versus the group of small ones.

Vertical vs horizontal
Load balancer
100
100
100
100
100
100
100
100
100
100
Served 1,000 RPSFailing 0 RPS

10 small instances share the 1,000 RPS behind a load balancer.

The math is never exactly 10

Earlier I said 1,000 / 100 = 10 instances. In practice, running every instance at 100% is asking for trouble. Traffic spikes, instances restart, deploys happen. You saw it above: take one instance down out of exactly 10, and some requests start failing.

So you add headroom, usually something like 20–30%:

10 instances × 1.3 ≈ 13 instances

And you keep at least 2 instances running at all times, so one failure doesn't take the whole service down.

This is the calculation I actually went through today. Change the numbers and see how it moves.

How many instances do I deploy?
1,000 / 100 = 10 instances · the bare minimum
1,000 × 1.30 / 100 = 13 · with 30% headroom
max(2, 13) = 13 to deploy
needed for the traffic extra for spikes and failures (3)

If traffic doubles tomorrow, you would need 26 instances.

Where concurrency fits

Scaling and concurrency are related, but they are not the same thing.

Suppose one server with 4 CPU and 16 GB RAM handles 100 requests per second. That doesn't mean it processes only one request at a time. It handles many concurrently:

Request A → waiting for DB
Request B → processing CPU work
Request C → waiting for an API
Request D → reading from cache
Request E → waiting on the network

While Request A is waiting on the database, the server can work on B, C, D or E. That's concurrency: making progress on multiple tasks during overlapping periods of time.

Here are those same five requests on a single CPU core. Press play, or drag the slider through time. Look at the two numbers at the bottom: how many requests are in flight, and how many are actually on the CPU.

Five requests, one CPU core
A
waiting for DB
B
CPU work
C
waiting for API
D
cache
E
waiting on network
using the CPU waiting on something else
Time
450 ms
In flight
3
concurrency
On the CPU
0 / 1
parallelism

Most of the time, several requests are in flight, but only one (or none) is actually using the CPU. The rest are just waiting. That is the difference between concurrency (how many things are in progress) and parallelism (how many things run at the exact same instant). With one core, parallelism is never more than 1, but concurrency can be 4 or 5.

Scaling is about increasing the system's total capacity to handle more work.

Where latency comes in

This is the part I felt most clearly once real numbers were involved. There's a simple formula for how many requests an instance is juggling at once:

requests in flight = requests per second × time per request

If one instance handles 100 RPS and each request takes 200 ms:

100 × 0.2 = 20 requests in flight

Now suppose the database gets slow and requests start taking 400 ms. If that instance can only juggle about 20 requests before it slows down, its throughput drops:

20 / 0.4 = 50 RPS per instance
1,000 / 50 = 20 instances

Same users. Same traffic. Twice the instances, just because each request got slower.

Try it. The traffic is fixed at 1,000 RPS. Only the time per request changes.

Same traffic, different latency1,000 RPS · fixed
In flight
200
RPS × latency
Per instance
100 RPS
20 at once ÷ 200 ms
Instances
10
10 at 200 ms
200 ms

This is the baseline. Now drag the latency and watch the count.

I knew this formula already. But watching it change a real instance count is what made it stick: concurrency, latency and scaling are all part of the same conversation. A slow dependency can force you to scale just as much as a traffic spike.

A simple mental model

The way I think about it now:

  • Concurrency: How many things can my system work on at the same time?
  • Parallelism: How many things can actually execute at the same instant?
  • Vertical scaling: How much more powerful can I make one machine?
  • Horizontal scaling: How many machines can I use to distribute the workload?

These are easy to blur together when you only discuss them. They get much sharper when you have to calculate actual capacity for a real system.

What production adds

The definitions didn't change. What production added was this chain:

Expected traffic
      ↓
Requests per second
      ↓
Capacity per instance
      ↓
Required instances (+ headroom)
      ↓
CPU / memory requirements
      ↓
Instance type
      ↓
Cost
      ↓
Scaling strategy

Instead of saying "we'll horizontally scale the application," you start asking:

  • How many requests per second are we expecting?
  • How many requests can one instance reliably handle?
  • What happens at 2x the expected traffic?
  • What happens if requests get slower?
  • Do we scale on CPU, memory, RPS, latency or queue depth?
  • What happens if one instance dies?

Those questions are much closer to real engineering.

Theory vs production

Scaling is one of those topics where knowing it and doing it are different experiences.

Reading "vertical scaling means increasing the resources of a machine" is easy.

Actually configuring an instance, estimating traffic, calculating capacity, thinking about RPS and latency, and deciding whether you need one larger instance or several smaller ones is where the concept becomes real.

Today, scaling stopped being a system design diagram and became a calculation for a real system.

That's the difference between knowing a concept and having applied it, and today I got to cross that line.

Updates // Newsletter

Stay in the loop.

Receive technical deep-dives and architectural insights directly in your inbox.

NO_SPAM // NO_TRACKING // 0_COST