When Does Liquidity Stress Need a Distribution?
The taxonomy of liquidity uncertainty, the ladder from scenarios to Monte Carlo, and the single assumption that outweighs every stress narrative combined
A bank's ALCO reviews four stress scenarios each quarter. The committee debates idiosyncratic versus systemic, short versus prolonged, picks runoff rates and haircuts, signs off, and moves to the next agenda item. Is that enough? Sometimes yes. Sometimes dangerously not. The question is not whether to stress-test — every regulator on the planet insists on that. The question is how far up the ladder of complexity to climb. And here is the part that rarely gets said aloud: climbing too far can be worse than not climbing at all, if the data cannot support the altitude or the governance cannot breathe the thinner air.
What follows is an attempt to draw the map. Not a sales pitch for Monte Carlo. Not a defense of simple scenarios. A map — showing where each method earns its keep, where it breaks down, and which single assumption, once you reach the probabilistic rungs, dwarfs everything else in value impact.
The numbers come from a single illustrative institution — call it Avelmont — with three major liquidity flows and a stress-testing framework under review.
The four-cell taxonomy: what kind of uncertainty are you facing?
Before choosing a method, you need to classify the problem. Every cash flow on a bank's balance sheet lives in one of four cells, depending on two binary questions: Is the timing known or random? and Is the amount known or random?
This sounds elementary, but the taxonomy has teeth. Misclassifying a flow — treating a timing problem as an amount problem, or vice versa — produces answers that are not merely imprecise but structurally wrong. A prepayment on a mortgage is a timing event: the amount repaid is the outstanding balance, but when the borrower exercises that option is random. Modeling it as an uncertain amount at a fixed date misses the cliff risk entirely. Conversely, a deposit that matures on a known date but whose renewal volume is uncertain is an amount problem. Randomizing the maturity date instead wastes modeling effort on the wrong axis.
The bottom-right cell — time random, amount random — is where most of the modeling difficulty concentrates. Non-maturity deposits live here: you do not know when the depositor will leave, and you do not know how much she will take when she does. Credit-line drawdowns live here too: the borrower may draw any amount, at any time, and tends to draw most aggressively precisely when the bank least wants to fund the outflow.
The key insight is not that the bottom-right cell is "harder" in some abstract sense. It is that flows in the bottom-right cell require more machinery to handle honestly, and applying bottom-right machinery to a top-left flow wastes resources while adding no signal. Conversely, applying top-left methods to a bottom-right flow suppresses the very risk you are trying to measure.
Four methods that need no probability at all
Once the flows are classified, the next question is: what method do you use to stress them? And here the instinct to reach for a probability distribution should be resisted — at least initially. Four perfectly legitimate methods require no distributional assumption whatsoever.
Named scenarios
A committee picks a narrative: "large depositor leaves within a week," or "wholesale market closes for three months." Basel's own framework requires four cells of this type: idiosyncratic versus systemic, crossed with short-lived versus prolonged. The scenario is a story with numbers attached. It does not claim to know how likely the story is — only what happens if it occurs. This is not a weakness. It is a feature. The scenario says: if this happens, here is our gap. The board can debate whether the story is plausible without ever arguing about tail percentiles.
Reverse stress testing
Instead of picking a scenario and measuring the damage, reverse stress testing works backwards: what combination of events would break us? Find the conditions, not the probability. If it turns out that the bank fails only when deposit runoff exceeds 25% and wholesale markets close and the central bank facility is exhausted — all at once — then the board has learned something concrete. Whether that joint event has a 0.1% or a 0.01% probability is, for many purposes, secondary to knowing that it exists.
Historical replay
Take another institution's crisis and apply it to your own balance sheet. Northern Rock's deposit run. SVB's overnight collapse. The 2020 repo-market freeze. The advantage is that these are not hypothetical — they happened. The disadvantage is survivorship bias: history tells you what did happen, not the full space of what could have happened. But as a plausibility check, a replay is hard to beat. If your bank would have survived SVB's deposit trajectory applied to your own liability profile, that is meaningful information.
Sensitivity grids
Which parameter moves the number most? Vary one input at a time — deposit runoff rate, wholesale rollover probability, haircut on collateral — and record the gap. The output is a table: rows of inputs, columns of values, no probabilities attached. The grid does not tell you how likely a given parameter value is. It tells you how much it matters — which is often the more useful question for resource allocation.
Three conditions that justify climbing higher
When should a bank move beyond scenarios and into full distributional territory — Monte Carlo simulation, fitted tails, probabilistic stress? Not when the quant team is bored. Not when the regulator hints. Only when three conditions hold simultaneously. If any one fails, the distribution adds cost without adding insight — and may actively mislead.
Condition 1: the data must earn the tail
A probability distribution is only as honest as the data behind it. If your deposit-runoff history covers five calm years, the tail you fit is fiction. The model will produce a 99th-percentile number, of course — models always produce numbers — but the tail will reflect five years of mild variation extrapolated by a parametric assumption, not five years of evidence that includes a genuinely bad month. At Avelmont, the deposit series begins in 2006. It contains the 2008 crisis and the 2020 freeze in-sample. That is enough to anchor a tail. A series that begins in 2015 does not have that anchor, and any tail it produces is the modeler's opinion dressed as data.
Condition 2: the decision must actually turn on the tail
Suppose you run the full simulation and find that the liquidity buffer required at the 50th percentile is $1.15 billion and at the 99th percentile is $1.22 billion. The difference is $70 million — less than the rounding error in most balance-sheet projections. Does the board's decision change? If the buffer is $1.2 billion either way, the distribution has added modeling cost, validation cost, and governance burden without changing the answer. The right test is simple: does the action differ between the median and the tail? If not, the scenario was sufficient.
Condition 3: governance must own the premises
A Monte Carlo model has inputs. Some of those inputs are estimated from data. Others — the most consequential ones, as we will see — are assumed. If the board cannot explain why the model uses a correlation of 0.75 between deposit runoff and credit-line draws, the number the model produces is an orphan. It has precision without ownership. And a precise number from unowned premises is worse than a range from owned ones, because precision creates false confidence. The board signs the buffer. The board must understand, in plain language, what the model assumes and why. If the most expensive assumption cannot survive a two-minute explanation at the committee table, the model is not ready.
The correlation trap: the most expensive assumption no one tests
Suppose Avelmont passes all three conditions and builds a Monte Carlo model. Three flows dominate the stressed liquidity gap:
| Flow | Stressed Mean (M) | Volatility |
|---|---|---|
| Deposit runoff | 900 | 30% |
| Credit-line drawdowns | 350 | 50% |
| Wholesale rollover failure | 250 | 10% |
Each flow has a mean and a volatility. The question that controls the tail — and therefore the buffer — is not how large each flow is individually. It is how they move together. And this is where the modeling becomes treacherous, because the answer to that question cannot be settled by regression. There is no dataset long enough, granular enough, and crisis-heavy enough to pin down the correlation between deposit runoff and credit-line draws with any precision. The estimate will always be wide. And the buffer swings enormously across that width.
Model the three flows as independent (zero correlation). The combined 97.5% tail of the gap comes out at 892M. That is, under the assumption that a depositor leaving has nothing to do with a borrower drawing a credit line, the worst-case gap at the 97.5th percentile is 892M.
Now introduce correlation. At 50% pairwise correlation, the tail jumps to 1,114M — an increase of 222M from a single parameter change. The flows are now assumed to worsen together: when deposits run, credit lines get drawn, and wholesale fails to roll, with moderate co-movement. The physics is plausible. In a systemic event, all three channels deteriorate simultaneously. But is 50% the right number? Nobody knows with any precision.
At 75% correlation — the number Avelmont has published in its ILAAP — the tail rises to 1,222M. Another 108M, from moving one parameter by 25 percentage points. At full correlation (100%), the flows move in lockstep: 1,338M.
The range from independence to full correlation is 446M — call it 450M. From a single assumption. Not a scenario choice. Not a runoff-rate debate. One correlation parameter that the risk team cannot estimate from data with any useful confidence interval, and that the board probably has never discussed.
And this is the trap. The model produces a number — 1,222M, say — with apparent precision. The board reads it as a finding. But the number rests on a correlation of 0.75 that was chosen because it "seemed conservative," or because a peer bank used it, or because the quant who built the model had to put something in the cell. The premise is unowned. The precision is cosmetic.
The inequality that reconciles plan with stress
Step back from the model for a moment and consider a structural problem that every Treasury and Risk function must solve jointly, whether they use scenarios or Monte Carlo.
The funding plan is a calendar. It is deterministic: Treasury knows when each wholesale tranche matures, when each deposit campaign launches, when each bond issuance is scheduled. It lives in the world of dates and amounts. It is Treasury's instrument.
The stress test is a distribution — or, at minimum, a set of scenarios. It is stochastic: Risk models what could go wrong, how much could run, how deeply capacity could erode. It lives in the world of probabilities and tails. It is Risk's instrument.
These two objects live in different arithmetic. One is a schedule. The other is a random variable. But they must meet. The planned capacity — what Treasury expects to have available, according to the funding plan — must exceed the stressed need — what Risk says might materialize in the tail. The gap between the two is the buffer.
If you use scenarios only, you get one gap number per scenario. The buffer is sized to cover the worst scenario the committee endorses. If you use Monte Carlo, you get a distribution of gaps — and you can read the 97.5th percentile, or the Expected Shortfall, or any other risk measure the governance structure requires. The distribution does not replace the plan; it tells you whether the plan is sufficient for the tail it must withstand.
The buffer sits in the gap. It is sized by the stress test (Risk's domain) and funded by the plan (Treasury's domain). When Risk says the 97.5% tail requires 1,222M of liquidity, and Treasury's plan delivers 1,400M, the 178M surplus is the margin of safety. When the plan delivers only 1,100M, there is a 122M shortfall — and either the plan must be revised (issue more, issue longer, diversify sources) or the stress assumption must be re-examined (is 75% correlation truly the right anchor?).
What moves the number most
If Avelmont's risk committee has a finite amount of governance capital — time at the table, willingness to debate assumptions, appetite for technical nuance — where should it spend that capital? The sensitivity analysis answers this question with arithmetic, not rhetoric.
Three inputs were varied across their plausible range:
| Input | Low Value | High Value | Buffer Impact (M) |
|---|---|---|---|
| Pairwise correlation | 0% (independence) | 100% (full) | 450 |
| Capacity erosion rate (haircuts under stress) | 10% | 40% | 344 |
| Stress scenario switch (idiosyncratic vs. systemic) | Idiosyncratic | Systemic | ~100 |
Look at the ratios. Correlation outranks the scenario switch by roughly 4 to 1 in value impact. The capacity erosion assumption — how severely haircuts bite under stress — is worth 344M, comparable to correlation and far above the scenario choice. Yet in a typical ALCO meeting, the scenario narrative ("are we running the idiosyncratic or the systemic?") absorbs most of the discussion time. The correlation assumption — buried in a model specification document that no one at the table has read — moves the number four times as much.
This is not an argument against discussing scenarios. Scenarios have value: they build institutional intuition, they are communicable, they anchor the stress narrative. But spending governance capital on the correlation assumption is four times more valuable per minute of committee time than debating which scenario to endorse. And the capacity erosion assumption — what happens to the haircuts on your collateral when you most need to monetize it — is worth three times the scenario switch.
Where most banks should stand on the ladder
The ladder exists for a reason. Each rung serves a purpose, and the right rung depends on the institution — its data history, its balance-sheet complexity, its governance maturity.
Most banks belong on the middle rungs. Named scenarios with sensitivity grids. Reverse stress tests that reveal the breaking point. Historical replays that ground the scenarios in events that actually occurred. These methods are not approximations to Monte Carlo. They are complete tools for their purpose, and their purpose — sizing a buffer, testing a plan, informing a board — does not require a probability distribution in most cases.
Monte Carlo earns its keep only when all three conditions hold: the data is long enough to contain genuine stress, the decision turns on the tail, and the governance owns the premises. When those conditions are met, the distributional approach adds real value — it distinguishes the 90th percentile from the 99th, it quantifies the gap distribution rather than a single point, and it forces a conversation about correlation that scenarios let you avoid.
But even then — even with decades of data, a decision that swings on the tail, and a board that understands every premise — the model's most expensive input is not a parameter the quants estimate from regression. It is a correlation that the board must own. The number is 450M at Avelmont. Whatever it is at your institution, it is almost certainly larger than the scenario debate that fills your committee agenda.
The question is not "scenarios or Monte Carlo?" as if one were the answer and the other the wrong turn. The question is: does your institution have the data, the decision sensitivity, and the governance ownership to justify the probabilistic rung? If yes, climb — but audit the correlation first. If no, stay on the scenario rungs and invest your governance capital in the sensitivity grid, which will tell you what matters without pretending to know how likely it is.
And if you take only one number from this entire analysis, let it be the ratio: 4 to 1. Correlation over scenario choice. That is where the money is. That is where the governance capital should flow. Everything else is commentary.
The worked example uses Avelmont, a fictional institution, with illustrative parameters. The taxonomy, ladder, and sensitivity analysis described here reflect one analytical framework — not the only one — for deciding when liquidity stress testing requires a probability distribution.
Continue Learning
FTP and All-In Loan Pricing
Build a bank's all-in transfer price from the ground up.
This course takes you inside the mechanics of Funds Transfer Pricing — from constructing the funding curve and modeling deposit behavioral maturity, to layering in the liquidity term structure, contingent buffer costs, expected credit loss, and capital charges for IRRBB. You'll build each component in hands-on labs on a live balance sheet, learning to price loans incrementally and defend every basis point to ALCO. Designed for ALM practitioners, treasury professionals, and risk managers in both developed and emerging markets.
Intermediate · 9 phases · 56 lessons · 18 labs · 12 deep dives · 10h · Instructors: Andre Camatta & Diogo Gobira