Saturday 29 August 2026

Quant interview

Why a delta-hedged short option lives and dies by gamma

medium · The Black-Scholes PDE and the Greeks

You are short one European option and you delta-hedge it continuously. You mark and hedge the option using the Black–Scholes formula with a fixed implied volatility σ, so its price V(S,t) satisfies

∂tV+12σ2S2∂SSV+rS∂SV−rV=0,

with r the constant risk-free rate. The stock, however, actually evolves under the real-world measure as

dSt=μStdt+σrStdWt,

where the realized volatility σr need not equal σ. Your self-financing portfolio Π consists of the short option, Δt=∂SV(St,t) shares held long, and the remaining cash V−ΔtSt earning the rate r.

  1. Show that the instantaneous mark-to-market P&L of this delta-hedged position is locally deterministic (no dW term) and equals dΠ=12ΓtSt2(σ2−σr2)dt,Γt=∂SSV(St,t).

  2. For a vanilla call or put, Γ>0. State the condition on realized versus implied volatility under which this short, delta-hedged position makes money, and reconcile it with the phrase "short gamma."

  3. Take a single ATM call: S=100, strike K=100, r=0, implied σ=20%, and 30 trading days to expiry (T=30/252, with 252 days per year). Suppose that over the next trading day (dt=1/252) the stock does not move at all, so realized volatility that day is 0. Compute the day's hedging P&L in dollars.

Solution

The one substitution that does the work

Over [t,t+dt] the portfolio value changes by three pieces: the short option −dV, the long shares ΔtdS, and interest on the cash r(V−ΔtS)dt:

dΠ=−dV+ΔtdS+r(V−ΔtS)dt.

Apply Itô to V(St,t), and here is the point a candidate must not miss: the quadratic variation of S is driven by the realized volatility, since we are computing the actual P&L of the actual path, not a fictitious one. So

dV=∂tVdt+∂SVdS+12∂SSVσr2S2dt.

Substitute, and use Δt=∂SV so that the ∂SVdS term cancels against ΔtdS. Every stochastic dS (hence every dW) is gone:

dΠ=−∂tVdt−12∂SSVσr2S2dt+rVdt−rS∂SVdt.

Now use that the option is marked with implied vol, so it obeys the Black–Scholes PDE with σ:

∂tV=−12σ2S2∂SSV−rS∂SV+rV.

Insert this for ∂tV. The rS∂SV and rV terms cancel identically, leaving only the two second-order terms:

dΠ=12∂SSVσ2S2dt−12∂SSVσr2S2dt.

dΠ=12ΓtSt2(σ2−σr2)dt

The drift μ has vanished entirely, and there is no dW: the instantaneous P&L is deterministic given the current state. That is the whole content of the result — the randomness in the option is exactly offset by the randomness in the hedge, and what remains is a clean bet on volatility.

Sign for a vanilla short position

For a call or put Γ>0, so dΠ>0 precisely when σ>σr, i.e. realized volatility comes in below the implied volatility you sold at. Being short a positive-gamma option is being "short gamma": you collect the option premium (theta) and profit as long as the stock fails to move as much as the implied vol you charged. If the stock realizes more vol than implied, the gamma cost of rehedging overwhelms the premium and you lose. The magnitude is scaled by the dollar gamma 12ΓS2, which is why traders track "dollar gamma" as the local size of the vol bet.

The one-day number

With r=0 the ATM-call gamma is

Γ=ϕ(d1)SσT,d1=σT2.

Compute the inputs:

T=30/252=0.34504,σT=0.069008,d1=0.034504.

ϕ(d1)=e−d12/22π=0.39875,Γ=0.39875100×0.069008=0.057785.

The realized move is zero, so σr=0 and the P&L is pure premium capture:

dΠ=12ΓS2σ2dt=12(0.057785)(1002)(0.04)(1252).

dΠ=12(0.057785)(10000)(0.04)(0.0039683)=0.0459.

The short position gains about $0.046 on the day.

Closing note. Because r=0, this equals minus the option's one-day theta: Θ=−Sσϕ(d1)2T gives Θ/252=−0.0459 per day, and a short position earns +0.0459. That is the gamma–theta identity Θ+12σ2S2Γ=0 (the r=0 PDE) read off directly. The tempting error is to run Itô with the implied vol σ in dV; then everything cancels to zero and the whole point — that hedging error is a vol arbitrage — disappears.

Statistics in machine learning

Fitting one Gaussian to two: which way does the KL point?

hard · Variational inference and the asymmetry of KL

Fix a separation μ>0 and let the target density on ℝ be the equal-weight mixture

p(x)=12𝒩(x;−μ,1)+12𝒩(x;μ,1).

The variational family is the set of Gaussians 𝒬={qm,v=𝒩(m,v):m∈ℝ,v>0}. Throughout, KL(a‖b)=𝔼a[log(a/b)].

1. Determine q⋆=\argminq∈𝒬KL(p‖q) (the forward KL, integrating against the true p). State the general principle from exponential-family theory that identifies the minimizer, and give q⋆ in closed form.

2. Now consider the reverse KL objective F(m,v)=KL(qm,v‖p). (a) Show that logp(x)=−12log(2π)−x2+μ22+logcosh(μx), and use this to write F explicitly up to the single term g(m,v):=𝔼x~𝒩(m,v)[logcosh(μx)]. (b) Prove that (m,v)=(0,v) is stationary in m for every v>0. Then prove the following, which is the crux: at any point (0,v⋆) that is also stationary in v, the second derivative ∂m2F equals 1/v⋆>0. Conclude that the symmetric, mass-covering configuration is always a local minimum of the reverse KL — there is no destabilizing bifurcation of it as μ grows.

3. Exhibit a second family of (asymptotic) stationary points near m=±μ, v≈1, and compute limμ→∞KL(𝒩(μ,1)‖p). Compare this to the value of F at the symmetric local minimum of part 2 as μ→∞ (you may work to leading order in μ). Which configuration is the global reverse-KL optimum for large μ, and how does that contrast with the forward-KL answer of part 1?

Solution

1. Forward KL: moment matching

Minimizing KL(p‖q)=𝔼p[logp]−𝔼p[logq] over q is the same as maximizing 𝔼p[logq], since 𝔼p[logp] does not involve q. For q=𝒩(m,v),

𝔼p[logq]=−12log(2πv)−𝔼p[(x−m)2]2v.

The general principle: the Gaussian family is an exponential family with sufficient statistics (x,x2), and minimizing KL(p‖q) over an exponential family forces the expected sufficient statistics of q to equal those of p (the stationarity conditions ∇η[logZ(η)−η⊤𝔼pT]=0 read 𝔼qT=𝔼pT). Here that means match the mean and variance.

For the mixture, 𝔼p[x]=0 by symmetry and 𝔼p[x2]=1+μ2 (each component contributes variance 1 about mean ±μ). Hence

q⋆=𝒩(0,1+μ2).

Forward KL returns a single broad Gaussian straddling both modes: mass-covering, mean-seeking.

2a. A closed form for the reverse KL

Write the mixture as a single expression:

𝒩(x;μ,1)+𝒩(x;−μ,1)=12πe−(x2+μ2)/2(eμx+e−μx)=22πe−(x2+μ2)/2cosh(μx).

Multiplying by 12 and taking logs gives the stated identity

logp(x)=−12log(2π)−x2+μ22+logcosh(μx),

and note the mixing constant log2 cancelled. Now, with q=𝒩(m,v) having entropy 12log(2πev) and 𝔼q[x2]=m2+v,

F(m,v)=−12log(2πev)⏟𝔼q[logq]−(−12log(2π)−m2+v+μ22+g(m,v)).

Collecting the log terms, −12log(2πev)+12log(2π)=−12−12logv, so

F(m,v)=−12−12logv+m2+v+μ22−g(m,v),g(m,v)=𝔼𝒩(m,v)[logcosh(μx)].

2b. The symmetric configuration is always a local minimum

The tools are two Gaussian differentiation identities for G(m,v)=𝔼𝒩(m,v)[f(x)]: integrating by parts against the Gaussian gives

∂mG=𝔼[f′(x)],∂vG=12𝔼[f″(x)],

the second because the Gaussian density solves the heat equation ∂vϕ=12∂x2ϕ. Apply to f(x)=logcosh(μx), with f′=μtanh(μx) and f″=μ2sech2(μx):

∂mg=μ𝔼q[tanh(μx)],∂vg=12μ2𝔼q[sech2(μx)].

Stationarity in m at m=0. Then q=𝒩(0,v) is symmetric and tanh is odd, so 𝔼q[tanh(μx)]=0. Since ∂mF=m−∂mg, we get ∂mF(0,v)=0 for every v.

Curvature in m. Differentiate again: ∂m2F=1−∂m2g=1−μ2𝔼q[sech2(μx)]. Evaluated at m=0,

∂m2F(0,v)=1−μ2𝔼𝒩(0,v)[sech2(μx)].

This looks like it could go negative for large μ — the tempting route to a pitchfork. The step a candidate misses is to impose the v-stationarity condition simultaneously. From ∂vF=−12v+12−∂vg, a symmetric stationary point (0,v⋆) satisfies

−12v⋆+12−12μ2𝔼𝒩(0,v⋆)[sech2(μx)]=0⟹μ2𝔼[sech2(μx)]=1−1v⋆.

Substituting into the curvature,

∂m2F(0,v⋆)=1−(1−1v⋆)=1v⋆>0.

Also ∂m∂vF(0,v)=0 (an odd integrand in x), so the Hessian is diagonal there; combined with ∂v2F>0 at a v-minimum, (0,v⋆) is a genuine local minimum of the reverse KL. Such a symmetric stationary point exists for every μ: writing R(v)=1−μ2𝔼𝒩(0,v)[sech2(μx)], one has 1/v→+∞>R as v→0 and 1/v→0<R→1 as v→∞, so 1/v=R(v) has a solution. The mass-covering configuration never destabilizes — there is no bifurcation.

3. The mode-locked branch, and who wins globally

Try q=𝒩(μ,1). Using the closed form of logp,

logq−logp=−(x−μ)22+x2+μ22−logcosh(μx)=μx−logcosh(μx).

For x>0, μx−logcosh(μx)=log2−log(1+e−2μx). Under q, x concentrates at μ≫0, so 𝔼q[log(1+e−2μx)]→0 and

limμ→∞KL(𝒩(μ,1)‖p)=log2.

The interpretation is exact: the mode-locked Gaussian captures one component carrying mass 12, and pays the irreducible −log12=log2 for ignoring the other half. Checking stationarity confirms this is (asymptotically) a critical point: ∂mF=μ−μ𝔼q[tanh(μx)]→0 since tanh(μx)→1, and ∂vF→−12v+12=0 at v=1.

The symmetric branch, by contrast, grows. At (0,v⋆) with large μ, use logcosh(μx)≈|μx|−log2, so g≈μ2v/π−log2 (from 𝔼|x|=2v/π). Minimizing F≈−12logv+v+μ22−μ2v/π over v gives 2πv=2μ, i.e. v⋆≈2μ2/π, and

F(0,v⋆)≈(12−1π)μ2≈0.18μ2 ⟶ ∞.

So although the mass-covering point is a real local minimum, its cost blows up quadratically, while the mode-locked minimum stays at log2.

configuration law reverse KL as μ→∞
forward-KL optimum 𝒩(0,1+μ2) — (this is \argminKL(p‖q))
reverse-KL symmetric local min 𝒩(0,≈2μ2/π) ~(12−1π)μ2
reverse-KL global min 𝒩(±μ,1) →log2≈0.693

Conclusion. For well-separated modes the global reverse-KL optimum is a single mode, 𝒩(±μ,1) — zero-forcing / mode-seeking — whereas forward KL returns the mass-covering 𝒩(0,1+μ2). This is the asymmetry of KL made quantitative.

Closing notes


Two new problems every morning at 8am · every day so far