medium · Ito's lemma
Let follow a geometric Brownian motion under a fixed measure,
with , constants , , and a standard Brownian motion generating the filtration . For a real exponent define .
Take , so and . Ito's lemma applies because is on and stays strictly positive (GBM never hits zero). With the quadratic variation term is . Hence
Substituting and collecting terms, each of which carries a factor ,
The step a candidate most often botches is dropping the contribution: is itself a geometric Brownian motion, but with an inflated (or deflated) drift, not simply drift .
Write the drift coefficient as and the volatility as . Then solves a linear SDE with constant coefficients, whose solution is
Taking expectations and using (moment generating function of ), the and cancel and only the drift survives:
Explicitly,
Equivalently, taking directly kills the term (it is a true stochastic integral of a square-integrable integrand, hence a martingale with zero mean) and leaves the ODE with , giving the same answer.
A nonnegative Ito process is a martingale exactly when its drift vanishes identically (the part is already a local martingale, and here it is a genuine martingale since ). So we need
Factor out :
The trivial root gives the constant process . The nonzero root is
For , :
Closing note. The tempting wrong answer to part 2 is , which ignores the convexity correction and is only correct when . The martingale exponent is the same object that appears in the general solution of the stationary form of barrier/perpetual-option ODEs; its emergence here from a one-line drift condition is worth remembering.
medium · The EM algorithm: monotonicity and the ELBO
Let be observed data and a latent variable, with a joint density for a scalar parameter ranging over an open interval. Write the observed-data log-likelihood the posterior , and the EM auxiliary function One EM step maps . Assume throughout that all densities are strictly positive and smooth in , that differentiation under the integral sign is valid, and that the maximizers involved are interior with nonsingular curvature.
1. Prove that for every pair , and deduce . State precisely which term the M-step controls and which term is nonnegative for free.
2. Show that any fixed point satisfies . Conclude what monotonicity alone does, and does not, guarantee about the point EM converges to.
3. Define the expected complete-data information , the observed information , and the missing information . Prove that so EM converges locally at linear rate equal to the fraction of missing information. Then work the following instance exactly: with known and unknown, but only () are observed; the remaining are missing completely at random and treated as latent. Derive , give the exact convergence rate, and confirm it equals .
Start from the pointwise identity, valid for any , The left side does not depend on , so we may take of both sides without changing it: Doing the same with gives . Subtract: The last expectation is exactly , which proves the identity.
The key facts:
Both brackets nonnegative gives . (Note only is needed, not full maximization — this is why generalized EM is also monotone.)
The step a candidate misses is Fisher's identity: for every , Indeed , and at this is using the assumed interchange of and .
Now if is an interior maximizer of , its first-order condition is , which by Fisher's identity equals . Hence .
Conclusion. Monotonicity plus this shows EM's limit set consists of stationary points of . It buys nothing more: a stationary point may be a local maximum, a saddle, or (with a flat direction) a plateau. Monotone ascent does not distinguish these, and it says nothing about reaching the global maximum. That never decreases is a statement about the sequence of values, not about where they land.
Define implicitly by the M-step stationarity . Differentiate in : where and , evaluated at . At the fixed point .
To identify , differentiate Fisher's identity totally in (both the first and second argument move): At this reads , so . Therefore Since , this lies in : EM contracts, but at a rate that degrades to as the latent structure hides more of the information.
The Gaussian instance. With known, the complete-data log-likelihood is . Given the current , the posterior over each missing is , so and maximizing in replaces each missing by its conditional mean : This map is affine with slope everywhere, so from any start EM converges linearly (not just locally) to the fixed point , at exact rate Check against the information formula: , , , so , matching. If half the data are missing the rate is ; as it approaches and EM crawls.
The instance is deliberately absurd — the MLE is just the observed-sample mean , obtained in closed form. That is the point: even here EM can be made arbitrarily slow purely by increasing the fraction of missing data, and monotonicity is powerless to prevent it. Two tempting errors worth flagging: (i) believing monotone ascent implies convergence to the global maximum (Part 2: only stationarity is guaranteed); and (ii) believing a guaranteed increase at every step implies fast convergence (Part 3: the rate is the fraction of missing information and can sit just below ).
Two new problems every morning at 8am · every day so far