Thursday 13 August 2026

Quant interview

Two-sided exit, with and without drift

hard · Martingales and optional stopping for Brownian motion

Let (Bt)t≥0 be standard Brownian motion, let a,b>0, and let

τ=inf{t≥0:Bt∉(−b,a)}.

  1. Show that ℙ(Bτ=a)=ba+b. The stopping time is unbounded, so say explicitly what justifies optional stopping here.
  2. Show that 𝔼[τ]=ab.
  3. Let Xt=μt+Bt with μ>0, and let τμ be the exit time of X from (−b,a). Compute ℙ(Xτμ=a) in closed form, and verify that it recovers part 1 as μ→0.
Solution

1. The exit distribution

First, τ<∞ almost surely: Brownian motion has lim suptBt=+∞, so it cannot stay in a bounded interval forever.

Now the justification that matters. B is a martingale, but optional stopping does not apply to an arbitrary unbounded stopping time. What rescues the argument is that the stopped process Bt∧τ never leaves [−b,a], so it is uniformly bounded, hence uniformly integrable, and Bt∧τ→Bτ almost surely. Bounded convergence gives

𝔼[Bτ]=limt→∞𝔼[Bt∧τ]=B0=0.

Writing p=ℙ(Bτ=a) and using that Bτ∈{a,−b},

pa−(1−p)b=0⟹p=ba+b.

2. The expected exit time

Mt=Bt2−t is a martingale, so 𝔼[Bt∧τ2]=𝔼[t∧τ] for every t. The left side is bounded by max(a,b)2, so the right side is too, and letting t→∞ with monotone convergence on the right and bounded convergence on the left gives 𝔼[τ]=𝔼[Bτ2]<∞. Then

𝔼[Bτ2]=pa2+(1−p)b2=ba2+ab2a+b=ab.

𝔼[τ]=ab

Note how the finiteness of 𝔼[τ] came out of the argument rather than being assumed; assuming it is the usual gap in a hurried solution.

3. Adding drift

X itself is not a martingale, so exponentiate. For any θ, eθBt−θ2t/2 is a martingale; take θ=−2μ, whose θ2/2 equals 2μ2, and observe

e−2μXt=e−2μBt−2μ2t,

so Mt=e−2μXt is a martingale with M0=1. On the stopped interval X stays in [−b,a], so Mt∧τμ lies in [e−2μa,e2μb] — bounded again, which is what licenses optional stopping. With pμ=ℙ(Xτμ=a),

1=pμe−2μa+(1−pμ)e2μb

pμ=e2μb−1e2μb−e−2μa=1−e−2μb1−e−2μ(a+b)

As μ→0 both numerator and denominator vanish; to first order they are 2μb and 2μ(a+b), so pμ→b/(a+b), matching part 1. As μ→∞, pμ→1, as it must.

Note. The tempting move is to apply optional stopping to Xt−μt and stop there, but that only reproduces part 1 in disguise: it gives one equation relating 𝔼[τμ] and pμ, two unknowns. The exponential martingale is the right tool precisely because it is bounded on the interval and eliminates τμ from the equation entirely.

Statistics in machine learning

Where a kink in the prior comes from

medium · MAP estimation and the geometry of sparsity

Observe a single y~N(θ,σ2) with σ2 known, and put a Laplace prior on the parameter, θ~Laplace(0,b) with density 12be−|θ|/b.

  1. Show that the MAP estimator solves a penalized least-squares problem, and identify the penalty weight in terms of σ2 and b.
  2. Solve that problem in closed form.
  3. Repeat with a Gaussian prior θ~N(0,τ2), and explain precisely why one of the two estimators can return exactly zero and the other cannot.
Solution

1. From posterior to penalty

Up to terms free of θ,

−logp(θ∣y)=(y−θ)22σ2+|θ|b+const,

so multiplying by σ2 (which changes the objective but not its minimizer) the MAP estimator minimizes

12(y−θ)2+λ|θ|,λ=σ2b.

This is the lasso objective in one dimension. The penalty weight is a ratio of scales: a tighter prior (small b) or noisier data (large σ2) both penalize more.

2. Soft thresholding

The objective is convex but not differentiable at 0, so use the subdifferential. Away from zero, stationarity reads θ−y+λsign(θ)=0, giving θ=y−λ when θ>0 (consistent only if y>λ) and θ=y+λ when θ<0 (only if y<−λ). At zero, optimality requires 0∈−y+λ[−1,1], that is |y|≤λ. Collecting the cases,

θ^MAP=sign(y)(|y|−λ)+.

3. The Gaussian prior, and the source of the difference

With θ~N(0,τ2) the objective is (y−θ)22σ2+θ22τ2, smooth everywhere, and setting the derivative to zero gives pure linear shrinkage:

θ^ridge=τ2τ2+σ2y,

which is zero only when y=0.

The distinction is not about how heavy the prior's tails are; it is about the derivative at the origin. The Laplace penalty has a kink, so its subdifferential at 0 is the whole interval [−λ,λ], and any data gradient smaller in magnitude than λ can be absorbed — zero is a genuine minimizer for a whole range of y. The Gaussian penalty is differentiable with derivative exactly 0 at the origin, so it exerts no force there: the stationarity condition always balances at a nonzero θ whenever y≠0. Sparsity comes from non-differentiability at zero, and this is why ℓq penalties with q≤1 threshold while q>1 does not.

Note. This is a statement about the posterior mode only. Under the Laplace prior the posterior mean is a smooth, strictly monotone function of y and is never exactly zero, so the sparsity is an artifact of choosing the mode as the summary — worth remembering before describing the lasso as "the Bayesian estimator" for a Laplace prior.


Two new problems every morning at 8am · every day so far