You are on call for a web service: a load balancer in front of a fleet of instances, a cache in front of a database, and a 99.9 % availability objective. Every minute the model computes queueing delay, timeouts, cache hits and database load from textbook formulas — and what your decisions cost.
What you will learn
Why latency explodes near full utilization (Erlang C), and why autoscaling with a boot delay always arrives late.
How SLOs, error budgets and burn-rate alerts decide when to roll back.
How a cold cache turns into a database outage, and which levers buy time: shedding, degrading, warming.
Simulator
Time 0 min
▶Instance serving
⚙Instance booting
!Instance on the bad build
·Free slot
•Incoming requests
Controls
Floor for the fleet. Raising it launches instances at once — they still need the boot delay before serving.
Target tracking on utilization; no new scale-out while instances are still booting. Off = exactly the minimum.
Lower = more headroom and more cost. Measured utilization cannot exceed 100 %, so a saturated fleet grows only step by step.
Longer TTL = more hits, but answers can be older (mean age ≈ TTL/2).
Loads hot keys for 6 minutes (+12 % of the cache per minute) at the price of 600 extra database queries/s.
Share of the low-priority 30 % (prefetch, batch, crawlers) rejected at the load balancer with “retry later”.
The heavy feature adds 20 ms of CPU and one database query per request. Off = graceful degradation.
Redeploys the last good build on a fresh fleet (5 min, billed twice), then switches traffic. Pressing again restarts the preparation; with no bad build live it only costs money.
Indicators
Error rate
0.00%
normal
p99 latency
342ms
normal
Error budget left (30 days)
50.0%
normal
Fleet utilization
52%
normal
Burn rate (1 h)
0.0 ×
Requests
1125 req/s
Instances serving
10
Instances booting
0
Cache hit rate
77 %
Database utilization
19 %
Traffic shed
0 %
Fleet cost
4.00 $/h
Cost so far
0.00 $
Mean age of cached answers
30 s
Traffic on the bad build
0 %
Recommendations available
100 %
Trend
Crisis scenarios
Level 1 · Bad release
A new build goes out at 09:10. The month has already been rough: only 20 % of the error budget is left. Minutes after the deploy the burn-rate page fires. Protect the budget.
Error budget left at the end ≥ 18.5 %
Average error rate ≤ 0.65 % after the deploy
Fleet cost ≤ $9.50
Level 2 · Flash crowd
A link to the service is spreading fast and a traffic surge is expected some time this morning — nobody knows when, or how big. New instances need 8 minutes to boot today. When it comes, keep errors and latency down without burning money on idle capacity.
Average error rate ≤ 0.2 %
Average p99 latency ≤ 400 ms
Average traffic shed ≤ 5 %
Total cost ≤ $21
Recommendations available ≥ 85 % of the time
Level 3 · Cold cache
At the midday peak a maintenance script flushes the whole cache. Every request now goes to the database, which was sized for the usual 85 % hit rate. Get the service back without overloading the database.
Average error rate ≤ 1.5 %
Database utilization never above 90 % after the first minute
Mean age of cached answers ≤ 90 s on average
Total cost ≤ $9
Recommendations available ≥ 80 % of the time
Average traffic shed ≤ 5 %
Basis — the model behind the numbers
Every relation the simulator uses, with its source. Constants marked as assumptions are illustrative calibrations.
Requests follow a daily curve plus disturbances; the per-minute count is random (Poisson, normal approximation) with a little burstiness.
λ(t) = base × (1 + 0.25·sin(2π(t + clock − 6 h)/24 h)) × surge(t), clock = time of day at the start (peak at 12:00); count/min ≈ N(60λ, √(60λ)) × (1 + N(0, 0.02))[6]Assumption: worker counts, service times, database capacity, cache size, refill speed, prices and the bad build’s error rate are illustrative values for a mid-size web service.
Erlang C: the probability that a request has to wait for a free worker in an M/M/N system.
N = instances × 16 workers, a = λ·S; C(a, N) = B / (1 − (a/N)(1 − B)), B = Erlang B[3][6]
Waiting time tail: the chance of waiting longer than t falls exponentially; requests still waiting at the 2-s timeout fail. Beyond capacity, the excess fails.
p99 latency from the service-time and waiting-time quantiles.
p99 ≈ S·ln 100 + ln(C/0.01)/(N/S − λ) (service + waiting quantile, an approximation)[3][6]Approximation: adding the service and waiting quantiles is not the exact p99 of their sum (it can be a little high or low); the queue is treated as steady within each minute because requests take milliseconds.
Little’s law: busy workers = arrival rate × time in service.
TTL cache: with random requests, each miss starts a TTL period during which requests hit.
hit = warm × rT/(1 + rT), r = λ / 20,000 objects; mean age of a cached answer ≈ T/2[5]
Cache misses load the database; its queueing delay slows every request that touches it, which also fills the app workers.
DB load = λ·q·(1 − hit); query time = 5 ms/(1 − ρ_db) (≤ 250 ms); S = S_app + (1 − hit)·q·query time[6]Assumption: worker counts, service times, database capacity, cache size, refill speed, prices and the bad build’s error rate are illustrative values for a mid-size web service.
SLO and error budget: the burn rate says how many times faster than allowed the budget is being spent.
budget = 1 − SLO = 0.1 %; burn = error rate / 0.1 %; Δbudget per min = burn / 43,200; burn (1 h) = mean error rate over the last 60 min / 0.1 % (window pre-filled with the opening minute)[1][2]
Target-tracking autoscaler with a boot delay and a cooldown.
desired = ⌈serving × utilization / target⌉ (utilization saturates at 100 %); new instances serve after the boot delay[6]Assumption: worker counts, service times, database capacity, cache size, refill speed, prices and the bad build’s error rate are illustrative values for a mid-size web service.
Other operating constants used by the model.
16 workers/instance · app time 50 ms (+20 ms and +1 query with the feature on) · 2 queries/request · DB 4,000 queries/s · 20,000 hot objects · cache refill τ = 30 min (slower while the DB is saturated) · warm-up job +12 %/min for 6 min, +600 queries/s · timeout 2 s · 30 % low-priority traffic · bad build +5 % errors, ×1.25 CPU · rollback 5 min · $0.40 per instance-hour · scale-in by ≤ 20 % of the fleet after 10 quiet minutes · up to 40 instancesAssumption: worker counts, service times, database capacity, cache size, refill speed, prices and the bad build’s error rate are illustrative values for a mid-size web service.
Randomness: a seeded mulberry32 generator; distributions used — uniform, exponential (inverse CDF), normal (Box–Muller), Poisson (Knuth). The seed is shown and shareable.
M. Harchol-Balter — Performance Modeling and Design of Computer Systems: Queueing Theory in Action (M/M/k, server farms, capacity provisioning) — Cambridge University Press, 2013