medium · The Black-Scholes PDE and the Greeks
You are short one European option and you delta-hedge it continuously. You mark and hedge the option using the Black–Scholes formula with a fixed implied volatility , so its price satisfies
with the constant risk-free rate. The stock, however, actually evolves under the real-world measure as
where the realized volatility need not equal . Your self-financing portfolio consists of the short option, shares held long, and the remaining cash earning the rate .
Show that the instantaneous mark-to-market P&L of this delta-hedged position is locally deterministic (no term) and equals
For a vanilla call or put, . State the condition on realized versus implied volatility under which this short, delta-hedged position makes money, and reconcile it with the phrase "short gamma."
Take a single ATM call: , strike , , implied , and 30 trading days to expiry (, with days per year). Suppose that over the next trading day () the stock does not move at all, so realized volatility that day is . Compute the day's hedging P&L in dollars.
Over the portfolio value changes by three pieces: the short option , the long shares , and interest on the cash :
Apply Itô to , and here is the point a candidate must not miss: the quadratic variation of is driven by the realized volatility, since we are computing the actual P&L of the actual path, not a fictitious one. So
Substitute, and use so that the term cancels against . Every stochastic (hence every ) is gone:
Now use that the option is marked with implied vol, so it obeys the Black–Scholes PDE with :
Insert this for . The and terms cancel identically, leaving only the two second-order terms:
The drift has vanished entirely, and there is no : the instantaneous P&L is deterministic given the current state. That is the whole content of the result — the randomness in the option is exactly offset by the randomness in the hedge, and what remains is a clean bet on volatility.
For a call or put , so precisely when , i.e. realized volatility comes in below the implied volatility you sold at. Being short a positive-gamma option is being "short gamma": you collect the option premium (theta) and profit as long as the stock fails to move as much as the implied vol you charged. If the stock realizes more vol than implied, the gamma cost of rehedging overwhelms the premium and you lose. The magnitude is scaled by the dollar gamma , which is why traders track "dollar gamma" as the local size of the vol bet.
With the ATM-call gamma is
Compute the inputs:
The realized move is zero, so and the P&L is pure premium capture:
The short position gains about $0.046 on the day.
Closing note. Because , this equals minus the option's one-day theta: gives per day, and a short position earns . That is the gamma–theta identity (the PDE) read off directly. The tempting error is to run Itô with the implied vol in ; then everything cancels to zero and the whole point — that hedging error is a vol arbitrage — disappears.
hard · Variational inference and the asymmetry of KL
Fix a separation and let the target density on be the equal-weight mixture
The variational family is the set of Gaussians . Throughout, .
1. Determine (the forward KL, integrating against the true ). State the general principle from exponential-family theory that identifies the minimizer, and give in closed form.
2. Now consider the reverse KL objective . (a) Show that , and use this to write explicitly up to the single term . (b) Prove that is stationary in for every . Then prove the following, which is the crux: at any point that is also stationary in , the second derivative equals . Conclude that the symmetric, mass-covering configuration is always a local minimum of the reverse KL — there is no destabilizing bifurcation of it as grows.
3. Exhibit a second family of (asymptotic) stationary points near , , and compute . Compare this to the value of at the symmetric local minimum of part 2 as (you may work to leading order in ). Which configuration is the global reverse-KL optimum for large , and how does that contrast with the forward-KL answer of part 1?
Minimizing over is the same as maximizing , since does not involve . For ,
The general principle: the Gaussian family is an exponential family with sufficient statistics , and minimizing over an exponential family forces the expected sufficient statistics of to equal those of (the stationarity conditions read ). Here that means match the mean and variance.
For the mixture, by symmetry and (each component contributes variance about mean ). Hence
Forward KL returns a single broad Gaussian straddling both modes: mass-covering, mean-seeking.
Write the mixture as a single expression:
Multiplying by and taking logs gives the stated identity
and note the mixing constant cancelled. Now, with having entropy and ,
Collecting the log terms, , so
The tools are two Gaussian differentiation identities for : integrating by parts against the Gaussian gives
the second because the Gaussian density solves the heat equation . Apply to , with and :
Stationarity in at . Then is symmetric and is odd, so . Since , we get for every .
Curvature in . Differentiate again: . Evaluated at ,
This looks like it could go negative for large — the tempting route to a pitchfork. The step a candidate misses is to impose the -stationarity condition simultaneously. From , a symmetric stationary point satisfies
Substituting into the curvature,
Also (an odd integrand in ), so the Hessian is diagonal there; combined with at a -minimum, is a genuine local minimum of the reverse KL. Such a symmetric stationary point exists for every : writing , one has as and as , so has a solution. The mass-covering configuration never destabilizes — there is no bifurcation.
Try . Using the closed form of ,
For , . Under , concentrates at , so and
The interpretation is exact: the mode-locked Gaussian captures one component carrying mass , and pays the irreducible for ignoring the other half. Checking stationarity confirms this is (asymptotically) a critical point: since , and at .
The symmetric branch, by contrast, grows. At with large , use , so (from ). Minimizing over gives , i.e. , and
So although the mass-covering point is a real local minimum, its cost blows up quadratically, while the mode-locked minimum stays at .
| configuration | law | reverse KL as |
|---|---|---|
| forward-KL optimum | — (this is ) | |
| reverse-KL symmetric local min | ||
| reverse-KL global min |
Conclusion. For well-separated modes the global reverse-KL optimum is a single mode, — zero-forcing / mode-seeking — whereas forward KL returns the mass-covering . This is the asymmetry of KL made quantitative.
Two new problems every morning at 8am · every day so far