You are on the operations desk of a payment switch that sits between merchants and card issuers. Every minute about 400 transactions per second arrive; each is scored for fraud, waits for a free connection slot, and is forwarded to its issuer for an answer. The model computes queueing, timeouts, client retries, the fraud-screen trade-off and stand-in exposure from textbook formulas. Amounts are in generic currency units (u). Educational simulation only — not financial, legal or investment advice.
What you will learn
Why retries can lock an overloaded switch in a retry storm even after capacity is added — and how load shedding breaks the loop.
How a fraud score threshold trades fraud losses against good customers declined, and why the right threshold depends on the fraud base rate.
How one slow issuer fills every connection slot (Little’s law), and what timeouts and stand-in processing cost.
Simulator
Time 0 min
▶Server in service
⚙Server starting
·Free server rack position
✓Issuer answering normally
⌛Issuer slow or down
⇄Stand-in processing on
•Incoming transactions
Controls
Each server holds 32 connection slots. Added servers take 5 minutes to start; every server is billed, starting or not.
Admits at most 90 % of capacity and answers the excess at once with “try again later”, instead of letting it queue and time out.
How merchants retry a technical failure (up to 3 retries). Immediate = in the next minute; backoff = randomized exponential backoff, mean delays 1, 2 and 4 minutes.
Transactions scoring at or above this are declined. Scores are in standard deviations of genuine traffic: lower means more fraud stopped and more good customers declined.
Scores this far below the threshold get an extra customer challenge instead of a decision: 85 % of genuine customers complete it, 5 % of fraudsters pass. 0 = off.
How long a connection slot waits for the issuer’s answer. After that the switch sends a reversal and uses stand-in, or declines “issuer unavailable”.
When the issuer does not answer in time, the switch approves on its behalf up to this amount. Every stand-in approval is exposure the issuer never checked. 0 = off.
Indicators
Good-customer approval rate
99.9%
normal
Authorization time
330ms
normal
Fraud rate (share of approved amount)
10.4bp
normal
Connection slots busy
69%
normal
Approval rate, healthy issuers
99.9 %
Good customers declined by the fraud rule
1.3 ‰
Transactions sent to step-up
0.0 %
Fraud stopped
31 %
New transactions
400 tx/s
Offered load (new + retries)
400 tx/s
Retries
0 tx/s
Shed (answered “try later”)
0 tx/s
Dropped in the queue
0.0 %
Issuer timeouts
0.0 %
Stand-in approvals
0 tx/s
Servers in service
6
Servers starting
0
Server cost rate
36 u/h
Server cost so far
0 u
Stand-in exposure
0.00 M u
Fraud approved so far
0.00 M u
Good transactions lost so far
0.0 k tx
Transactions waiting to retry
0 tx
Issuer group B latency (mean)
250 ms
Trend
Crisis scenarios
Level 1 · Peak sale surge
A big online sale opens at minute 5 and doubles traffic to about 800 transactions per second for 45 minutes. The switch runs 6 servers, sized for a normal day at about 70 % busy. Merchants retry every technical failure at once. Keep good customers approved and answers fast without buying capacity you do not need.
Average good-customer approval ≥ 95 % from the sale on
Average authorization time ≤ 500 ms
Server cost ≤ 75 u
Level 2 · Fraud wave
At minute 5 a batch of stolen card details starts being used: the fraud share of traffic jumps from 0.1 % to 1 %. The decline threshold is set for normal days (3σ) and step-up is off. Bring the fraud rate down without turning away good customers.
Average fraud rate ≤ 25 bp
Average good-customer approval ≥ 98 %
Good customers declined by the rule ≤ 5 ‰
Level 3 · Issuer outage
At minute 5 the issuers of group B — a quarter of all traffic — slow to a mean answer time of 12 seconds. The switch waits up to 8 s for an answer, has no stand-in, and runs 6 servers. Keep the customers of the healthy issuers flowing, serve group B as far as you safely can, and keep stand-in exposure and cost in check.
Average approval for healthy-issuer customers ≥ 90 %
Average good-customer approval ≥ 86 %
Stand-in exposure ≤ 9 M u
Server cost ≤ 65 u
Basis — the model behind the numbers
Every relation the simulator uses, with its source. Constants marked as assumptions are illustrative calibrations.
New transactions arrive at a base rate times the sale surge; the per-minute count is random (Poisson, normal approximation) with a little burstiness.
λ(t) = base × surge(t) (2-min lag); count/min ≈ N(60λ, √(60λ)) × (1 + N(0, 0.03))[2]Assumption: slot counts, service and issuer times, score separation, step-up pass rates, wasted-work share, amounts and prices are illustrative values for a mid-size switch, not figures of any real network.
Binormal fraud score: genuine and fraud scores are two normal curves d′ apart; the threshold picks a point on the ROC curve.
genuine score ~ N(0,1), fraud ~ N(d′,1), d′ = 2.5; FPR(t) = 1 − Φ(t), TPR(t) = 1 − Φ(t − d′); AUC = Φ(d′/√2) ≈ 0.96[6][7]Assumption: slot counts, service and issuer times, score separation, step-up pass rates, wasted-work share, amounts and prices are illustrative values for a mid-size switch, not figures of any real network.
Cost-optimal threshold: decline when the likelihood ratio exceeds the cost ratio weighted by the base rate — ten times more fraud moves it ln 10 / d′ ≈ 0.9σ lower.
Step-up: scores in the band below the threshold are challenged instead of decided.
scores in [t − b, t) are challenged: genuine pass 85 %, fraud pass 5 %; scores ≥ t declined[10]Assumption: slot counts, service and issuer times, score separation, step-up pass rates, wasted-work share, amounts and prices are illustrative values for a mid-size switch, not figures of any real network.
Issuer latency is exponential; a slot is held for the latency or the timeout, whichever comes first.
L ~ Exp(mean m); E[min(L, T)] = m(1 − e^(−T/m)), P(L > T) = e^(−T/m); S = 80 ms + forwarded × Σ share·E[min(L,T)][2][9]
Erlang C: the chance a transaction waits for a free slot, and the chance it waits longer than the 2-s queue timeout.
P(W > 2 s) = C(a, N)·e^(−(N/S − λ)·2 s), a = λS; mean wait = C / (N/S − λ)[1][2]
Beyond capacity the switch also spends work on requests it drops later, so useful throughput falls as load rises; load shedding refuses the excess cheaply.
λ ≥ N/S: goodput = (N/S − ω·λ)/(1 − ω), ω = 0.3; load shedding admits ≤ 0.9·N/S and answers the rest at once[4][5]Assumption: slot counts, service and issuer times, score separation, step-up pass rates, wasted-work share, amounts and prices are illustrative values for a mid-size switch, not figures of any real network.
Client retries: each technical failure is retried up to three times, at once or with randomized exponential backoff.
failed attempt → retry with p = 0.95, ≤ 3 retries; immediate: next minute; backoff: delay ~ Exp(mean 1, 2, 4 min)[4][5]Assumption: slot counts, service and issuer times, score separation, step-up pass rates, wasted-work share, amounts and prices are illustrative values for a mid-size switch, not figures of any real network.
Issuer timeout, reversal and stand-in: amounts are lognormal, so the share under the limit and the volume approved follow from the normal CDF.
issuer timeout → reversal; stand-in approves if amount ≤ limit: P = Φ((ln L − μ)/σ), volume = e^(μ+σ²/2)·Φ((ln L − μ − σ²)/σ)[9][11]
Fraud rate in basis points of approved amount; 13 bp is used as a reference scale for small remote card payments.
32 slots per server, 2–24 servers, +5 min boot, 6 u per server-hour · switch time 80 ms · queue timeout 2 s · issuer latency 250 ms (group B = 25 % of traffic, ±15 % per minute) · amounts lognormal, median 40 u, σ = 1, fraud ×1.5 · ω = 0.3 · shedding at 90 % · reversal = 80 ms of switch timeAssumption: slot counts, service and issuer times, score separation, step-up pass rates, wasted-work share, amounts and prices are illustrative values for a mid-size switch, not figures of any real network.
Randomness: a seeded mulberry32 generator; distributions used — uniform, exponential (inverse CDF), normal (Box–Muller), Poisson (Knuth). The seed is shown and shareable.
M. Harchol-Balter — Performance Modeling and Design of Computer Systems: Queueing Theory in Action (M/M/k, capacity provisioning) — Cambridge University Press, 2013