Tuesday 25 August 2026

Quant interview

Powers of a geometric Brownian motion

medium · Ito's lemma

Let St follow a geometric Brownian motion under a fixed measure,

dSt=μStdt+σStdWt,

with S0>0, constants μ∈ℝ, σ>0, and Wt a standard Brownian motion generating the filtration ℱt. For a real exponent n define Yt=Stn.

  1. Derive the stochastic differential equation satisfied by Yt.
  2. Compute 𝔼[Stn] in closed form.
  3. Determine every value of n for which the process Stn is a martingale, and evaluate the nonzero solution numerically for μ=0.05, σ=0.20.
Solution

Step 1: Ito on the power map

Take f(x)=xn, so f′(x)=nxn−1 and f″(x)=n(n−1)xn−2. Ito's lemma applies because f is C2 on (0,∞) and St stays strictly positive (GBM never hits zero). With dSt=μStdt+σStdWt the quadratic variation term is 12f″(St)(dSt)2=12n(n−1)Stn−2σ2St2dt. Hence

dYt=nStn−1dSt+12n(n−1)Stn−2σ2St2dt.

Substituting dSt and collecting terms, each of which carries a factor Stn=Yt,

dYt=Yt[(nμ+12n(n−1)σ2)dt+nσdWt].

The step a candidate most often botches is dropping the 12n(n−1)σ2 contribution: Yt is itself a geometric Brownian motion, but with an inflated (or deflated) drift, not simply drift nμ.

Step 2: The expectation

Write the drift coefficient as a:=nμ+12n(n−1)σ2 and the volatility as b:=nσ. Then Yt solves a linear SDE with constant coefficients, whose solution is

Yt=Y0exp((a−12b2)t+bWt).

Taking expectations and using 𝔼[ebWt]=e12b2t (moment generating function of Wt~𝒩(0,t)), the −12b2 and +12b2 cancel and only the drift survives:

𝔼[Stn]=𝔼[Yt]=S0neat.

Explicitly,

𝔼[Stn]=S0nexp((nμ+12n(n−1)σ2)t).

Equivalently, taking 𝔼[dYt] directly kills the dWt term (it is a true stochastic integral of a square-integrable integrand, hence a martingale with zero mean) and leaves the ODE m′(t)=am(t) with m(0)=S0n, giving the same answer.

Step 3: When is Stn a martingale?

A nonnegative Ito process is a martingale exactly when its drift vanishes identically (the dWt part is already a local martingale, and here it is a genuine martingale since 𝔼∫0tYs2ds<∞). So we need

nμ+12n(n−1)σ2=0.

Factor out n:

n[μ+12(n−1)σ2]=0.

The trivial root n=0 gives the constant process Yt≡1. The nonzero root is

n=1−2μσ2.

For μ=0.05, σ=0.20:

n=1−2(0.05)0.04=1−2.5=−1.5.

Answers

  1. dYt=Yt[(nμ+12n(n−1)σ2)dt+nσdWt].
  2. 𝔼[Stn]=S0nexp((nμ+12n(n−1)σ2)t).
  3. n=0 or n=1−2μ/σ2; the nontrivial value is n=−1.5.

Closing note. The tempting wrong answer to part 2 is S0nenμt, which ignores the convexity correction and is only correct when n∈{0,1}. The martingale exponent n=1−2μ/σ2 is the same object that appears in the general solution of the stationary form of barrier/perpetual-option ODEs; its emergence here from a one-line drift condition is worth remembering.

Statistics in machine learning

Why EM climbs, where it stops, and how slowly

medium · The EM algorithm: monotonicity and the ELBO

Let x be observed data and z a latent variable, with a joint density p(x,z;θ) for a scalar parameter θ ranging over an open interval. Write the observed-data log-likelihood ℓ(θ)=logp(x;θ)=log∫p(x,z;θ)dz, the posterior p(z∣x;θ)=p(x,z;θ)/p(x;θ), and the EM auxiliary function Q(θ′∣θ)=𝔼z~p(·∣x;θ)[logp(x,z;θ′)]. One EM step maps θ↦M(θ)=\argmaxθ′Q(θ′∣θ). Assume throughout that all densities are strictly positive and smooth in θ, that differentiation under the integral sign is valid, and that the maximizers involved are interior with nonsingular curvature.

1. Prove that for every pair θ,θ′, ℓ(θ′)−ℓ(θ)=[Q(θ′∣θ)−Q(θ∣θ)]+KL(p(·∣x;θ)\|p(·∣x;θ′)), and deduce ℓ(M(θ))≥ℓ(θ). State precisely which term the M-step controls and which term is nonnegative for free.

2. Show that any fixed point θ\*=M(θ\*) satisfies ℓ′(θ\*)=0. Conclude what monotonicity alone does, and does not, guarantee about the point EM converges to.

3. Define the expected complete-data information Icom=−∂θ′2Q(θ′∣θ\*)|θ′=θ\*, the observed information Iobs=−ℓ″(θ\*), and the missing information Imis=Icom−Iobs. Prove that M′(θ\*)=ImisIcom, so EM converges locally at linear rate equal to the fraction of missing information. Then work the following instance exactly: z1,…,zn~iidN(μ,σ2) with σ2 known and μ unknown, but only z1,…,zm (1≤m<n) are observed; the remaining n−m are missing completely at random and treated as latent. Derive M, give the exact convergence rate, and confirm it equals Imis/Icom.

Solution

1. The ELBO decomposition and monotonicity

Start from the pointwise identity, valid for any z, logp(x;θ′)=logp(x,z;θ′)−logp(z∣x;θ′). The left side does not depend on z, so we may take 𝔼z~p(·∣x;θ) of both sides without changing it: ℓ(θ′)=𝔼z∣x;θ[logp(x,z;θ′)]⏟=Q(θ′∣θ)−𝔼z∣x;θ[logp(z∣x;θ′)]. Doing the same with θ′=θ gives ℓ(θ)=Q(θ∣θ)−𝔼z∣x;θ[logp(z∣x;θ)]. Subtract: ℓ(θ′)−ℓ(θ)=[Q(θ′∣θ)−Q(θ∣θ)]+𝔼z∣x;θ[logp(z∣x;θ)p(z∣x;θ′)]. The last expectation is exactly KL(p(·∣x;θ)‖p(·∣x;θ′)), which proves the identity.

The key facts:

Both brackets nonnegative gives ℓ(M(θ))−ℓ(θ)≥0. (Note only Q(θ′∣θ)≥Q(θ∣θ) is needed, not full maximization — this is why generalized EM is also monotone.)

2. Fixed points are stationary — and that is all

The step a candidate misses is Fisher's identity: for every θ, ∂θ′Q(θ′∣θ)|θ′=θ=ℓ′(θ). Indeed ∂θ′Q(θ′∣θ)=𝔼z∣x;θ[∂θ′logp(x,z;θ′)], and at θ′=θ this is ∫∂θp(x,z;θ)p(x,z;θ)p(x,z;θ)p(x;θ)dz=∂θ∫p(x,z;θ)dzp(x;θ)=∂θp(x;θ)p(x;θ)=ℓ′(θ), using the assumed interchange of ∂θ and ∫.

Now if θ\*=M(θ\*) is an interior maximizer of Q(·∣θ\*), its first-order condition is ∂θ′Q(θ′∣θ\*)|θ′=θ\*=0, which by Fisher's identity equals ℓ′(θ\*). Hence ℓ′(θ\*)=0.

Conclusion. Monotonicity plus this shows EM's limit set consists of stationary points of ℓ. It buys nothing more: a stationary point may be a local maximum, a saddle, or (with a flat direction) a plateau. Monotone ascent does not distinguish these, and it says nothing about reaching the global maximum. That ℓ never decreases is a statement about the sequence of values, not about where they land.

3. The convergence rate is the fraction of missing information

Define M implicitly by the M-step stationarity ∂θ′Q(M(θ)∣θ)=0. Differentiate in θ: Q20M′(θ)+Q11=0,M′(θ)=−Q11Q20, where Q20=∂θ′2Q and Q11=∂θ∂θ′Q, evaluated at (M(θ),θ). At the fixed point −Q20(θ\*∣θ\*)=Icom.

To identify Q11, differentiate Fisher's identity ℓ′(θ)=∂θ′Q(θ′∣θ)|θ′=θ totally in θ (both the first and second argument move): ℓ″(θ)=Q20(θ∣θ)+Q11(θ∣θ). At θ\* this reads −Iobs=−Icom+Q11, so Q11(θ\*∣θ\*)=Icom−Iobs=Imis. Therefore M′(θ\*)=−Q11Q20=ImisIcom=1−IobsIcom. Since 0≤Imis≤Icom, this lies in [0,1): EM contracts, but at a rate that degrades to 1 as the latent structure hides more of the information.

The Gaussian instance. With σ2 known, the complete-data log-likelihood is −12σ2∑i=1n(zi−μ)2+const. Given the current μ, the posterior over each missing zi is N(μ,σ2), so Q(μ′∣μ)=−12σ2[∑i=1m(zi−μ′)2+∑i=m+1n𝔼[(zi−μ′)2]]+const, and maximizing in μ′ replaces each missing zi by its conditional mean μ: M(μ)=1n(∑i=1mzi⏟S+(n−m)μ). This map is affine with slope (n−m)/n everywhere, so from any start EM converges linearly (not just locally) to the fixed point μ\*=S/m=z¯obs, at exact rate ρ=n−mn Check against the information formula: Icom=−∂μ′2Q=n/σ2, Iobs=m/σ2, Imis=(n−m)/σ2, so Imis/Icom=(n−m)/n, matching. If half the data are missing the rate is 1/2; as m→1 it approaches 1−1/n and EM crawls.

Closing note

The instance is deliberately absurd — the MLE is just the observed-sample mean z¯obs=S/m, obtained in closed form. That is the point: even here EM can be made arbitrarily slow purely by increasing the fraction of missing data, and monotonicity is powerless to prevent it. Two tempting errors worth flagging: (i) believing monotone ascent implies convergence to the global maximum (Part 2: only stationarity is guaranteed); and (ii) believing a guaranteed increase at every step implies fast convergence (Part 3: the rate is the fraction of missing information and can sit just below 1).


Two new problems every morning at 8am · every day so far