hard · Martingales and optional stopping for Brownian motion
Let be standard Brownian motion, let , and let
First, almost surely: Brownian motion has , so it cannot stay in a bounded interval forever.
Now the justification that matters. is a martingale, but optional stopping does not apply to an arbitrary unbounded stopping time. What rescues the argument is that the stopped process never leaves , so it is uniformly bounded, hence uniformly integrable, and almost surely. Bounded convergence gives
Writing and using that ,
is a martingale, so for every . The left side is bounded by , so the right side is too, and letting with monotone convergence on the right and bounded convergence on the left gives . Then
Note how the finiteness of came out of the argument rather than being assumed; assuming it is the usual gap in a hurried solution.
itself is not a martingale, so exponentiate. For any , is a martingale; take , whose equals , and observe
so is a martingale with . On the stopped interval stays in , so lies in — bounded again, which is what licenses optional stopping. With ,
As both numerator and denominator vanish; to first order they are and , so , matching part 1. As , , as it must.
Note. The tempting move is to apply optional stopping to and stop there, but that only reproduces part 1 in disguise: it gives one equation relating and , two unknowns. The exponential martingale is the right tool precisely because it is bounded on the interval and eliminates from the equation entirely.
medium · MAP estimation and the geometry of sparsity
Observe a single with known, and put a Laplace prior on the parameter, with density .
Up to terms free of ,
so multiplying by (which changes the objective but not its minimizer) the MAP estimator minimizes
This is the lasso objective in one dimension. The penalty weight is a ratio of scales: a tighter prior (small ) or noisier data (large ) both penalize more.
The objective is convex but not differentiable at , so use the subdifferential. Away from zero, stationarity reads , giving when (consistent only if ) and when (only if ). At zero, optimality requires , that is . Collecting the cases,
With the objective is , smooth everywhere, and setting the derivative to zero gives pure linear shrinkage:
which is zero only when .
The distinction is not about how heavy the prior's tails are; it is about the derivative at the origin. The Laplace penalty has a kink, so its subdifferential at is the whole interval , and any data gradient smaller in magnitude than can be absorbed — zero is a genuine minimizer for a whole range of . The Gaussian penalty is differentiable with derivative exactly at the origin, so it exerts no force there: the stationarity condition always balances at a nonzero whenever . Sparsity comes from non-differentiability at zero, and this is why penalties with threshold while does not.
Note. This is a statement about the posterior mode only. Under the Laplace prior the posterior mean is a smooth, strictly monotone function of and is never exactly zero, so the sparsity is an artifact of choosing the mode as the summary — worth remembering before describing the lasso as "the Bayesian estimator" for a Laplace prior.
Two new problems every morning at 8am · every day so far