Wednesday 26 August 2026

Quant interview

The maximum of Brownian motion, with and without drift

hard · Brownian hitting times and the reflection principle

Let (Wt)t≥0 be a standard Brownian motion on a filtered probability space, W0=0. Write MT=max0≤t≤TWt and let Φ denote the standard normal CDF.

1. Using the reflection principle, prove that for a≥0 and b≤a, ℙ(MT≥a, WT≤b)=ℙ(WT≥2a−b), and deduce the joint density of (MT,WT) on the region {m≥0, w≤m}.

2. From part 1, identify the marginal law of MT and compute 𝔼[MT] in closed form.

3. Now let Xt=μt+Wt be Brownian motion with constant drift μ∈ℝ, and fix a level a>0. Let τa=inf{t≥0:Xt=a}. Compute ℙ(τa≤T)=ℙ(max0≤t≤TXt≥a) in closed form. Then evaluate it numerically for μ=1, a=2, T=1.

Solution

Part 1: reflection principle and the joint density

Fix a≥0 and b≤a. On the event {MT≥a} the stopping time τa=inf{t:Wt=a} satisfies τa≤T. By the strong Markov property, the post-τa process W~s=Wτa+s−a is a standard Brownian motion independent of Fτa, and by symmetry −W~ has the same law. Reflecting the path after τa about the level a therefore produces an equally likely path, and it sends the endpoint WT=w to 2a−w.

Under this reflection the event {MT≥a, WT≤b} (endpoint at least a−b below the barrier) maps bijectively onto {MT≥a, WT≥2a−b}. But since b≤a we have 2a−b≥a, so {WT≥2a−b}⊆{MT≥a} automatically. Hence ℙ(MT≥a, WT≤b)=ℙ(MT≥a, WT≥2a−b)=ℙ(WT≥2a−b).

The key step a candidate misses is checking 2a−b≥a, which is exactly what lets us drop the constraint MT≥a on the right.

Now extract the density. Write ϕT(x)=12πTe−x2/2T. Differentiating ℙ(MT≥a, WT≤b)=∫2a−b∞ϕT(x)dx in b gives ∫a∞g(m,b)dm=ϕT(2a−b) (chain rule contributes +1 from ∂b(−(2a−b))). Differentiating again in a: −g(a,b)=∂aϕT(2a−b)=−2(2a−b)TϕT(2a−b). Therefore the joint density is g(m,w)=2(2m−w)T2πTexp(−(2m−w)22T),m≥0, w≤m.

Part 2: law of MT and its mean

Setting b=a in part 1 gives ℙ(MT≥a, WT≤a)=ℙ(WT≥a). Adding ℙ(MT≥a, WT>a)=ℙ(WT>a) (since WT>a forces MT≥a), ℙ(MT≥a)=2ℙ(WT≥a)=ℙ(|WT|≥a). So MT=d|WT|, the half-normal law. Since WT~N(0,T), 𝔼[MT]=𝔼[|WT|]=T𝔼[|Z|]=T·2π=2Tπ.

Part 3: adding drift via Girsanov

The reflection principle relies on the symmetry of Brownian increments and fails once there is drift; one cannot simply write 2ℙ(XT≥a). The clean route is a change of measure.

Define ℚ on FT by dℙdℚ|FT=exp(μXT−12μ2T). By Girsanov's theorem, under ℚ the process Xt is a standard Brownian motion (the density removes the drift). Writing MTX=maxt≤TXt, ℙ(MTX≥a)=𝔼ℚ[1{MTX≥a}eμXT−12μ2T], and under ℚ the pair (MTX,XT) has the density g from part 1. Since a>0, the constraint m≥a subsumes m≥0: ℙ(MTX≥a)=e−μ2T/2∫−∞∞eμw∫max(a,w)∞2(2m−w)T2πTe−(2m−w)2/2Tdmdw. The inner integral, with u=2m−w, is ϕT(max(2a−w,w)). Splitting at w=a (where 2a−w=w): ℙ(MTX≥a)=e−μ2T/2[∫a∞eμwϕT(w)dw+∫−∞aeμwϕT(2a−w)dw]. Completing the square in the first integral, μw−w22T−μ2T2=−(w−μT)22T, gives Φ(μT−aT). In the second, substitute v=2a−w; the exponent becomes −(v+μT)22T times e2μa, giving e2μaΦ(−μT−aT). Hence  ℙ(max0≤t≤TXt≥a)=Φ(μT−aT)+e2μaΦ(−μT−aT). 

Sanity check. At μ=0 this collapses to 2Φ(−a/T)=2ℙ(WT≥a), matching part 2.

Numerical value for μ=1, a=2, T=1: Φ(−1)+e4Φ(−3)=0.15866+54.598×0.0013499≈0.15866+0.07370=0.2324.

So the drifted Brownian motion reaches level 2 within one time unit with probability ≈0.232.

Closing note. The tempting wrong answer in part 3 is 2ℙ(XT≥a)=2Φ(μT−aT), obtained by blindly reusing the driftless reflection identity. The e2μa factor is precisely the correction the change of measure supplies; forgetting it (or getting the sign of μ wrong inside the second Φ) is the standard way to lose the problem.

Statistics in machine learning

Gamma-Poisson updating: shrinkage weights and the predictive law

easy · Bayesian conjugacy and posterior updating

Let λ>0 carry the prior λ~Gamma(α,β) in the shape-rate parametrization, i.e. with density

π(λ)=βαΓ(α)λα−1e−βλ,α,β>0.

Conditional on λ, observe X1,…,Xn i.i.d. with Xi∣λ~Poisson(λ). Write S=∑i=1nXi.

  1. Show that the posterior λ∣X1:n is again Gamma, and give its two parameters.

  2. Write the posterior mean 𝔼[λ∣X1:n] as a convex combination of the prior mean and the maximum-likelihood estimate X¯=S/n. Identify the weight on the MLE and its limit as n→∞, and state the operational meaning of β that this reveals.

  3. Derive the posterior predictive law of a fresh draw Xn+1∣X1:n (with Xn+1∣λ~Poisson(λ), independent of the past given λ). Give the pmf in closed form and name the distribution.

Solution

1. Conjugacy

The Poisson likelihood is p(x1:n∣λ)=∏iλxie−λxi!∝λSe−nλ (as a function of λ; the xi! are constants). Multiplying by the prior kernel,

π(λ∣x1:n)∝λα−1e−βλ·λSe−nλ=λ(α+S)−1e−(β+n)λ.

This is the kernel of a Gamma density, and since the posterior is a genuine probability density the normalizing constant is forced. Hence

λ∣X1:n~Gamma(α+S, β+n).

The key structural fact is that the Poisson likelihood, as an exponential family in λ, has sufficient statistic (S,n); the Gamma prior is the conjugate family whose hyperparameters live in the same coordinates, so updating is just addition of (S,n) to (α,β).

2. Posterior mean as shrinkage

For Gamma(a,b) the mean is a/b, so

𝔼[λ∣X1:n]=α+Sβ+n.

Split the numerator to expose the two estimators:

α+Sβ+n=ββ+n·αβ+nβ+n·Sn.

So the posterior mean is a convex combination

𝔼[λ∣X1:n]=(1−w)αβ⏟prior mean+wX¯⏟MLE,w=nβ+n.

The weight on the data is w=n/(β+n)→1 as n→∞: the prior washes out at rate O(1/n). The step a candidate skips is reading off what β is: it enters exactly where n does, so β is a prior sample size — a count of pseudo-observations (equivalently, prior units of exposure), with α the pseudo-total of events. The prior Gamma(α,β) is worth β observations carrying α events.

3. Posterior predictive

Write α′=α+S, β′=β+n for the posterior parameters. The predictive is the Poisson likelihood averaged against the posterior:

P(Xn+1=k∣X1:n)=∫0∞λke−λk!·β′α′Γ(α′)λα′−1e−β′λdλ.

Collect the powers and use ∫0∞λc−1e−dλdλ=Γ(c)/dc with c=k+α′, d=β′+1:

P(Xn+1=k∣X1:n)=β′α′k!Γ(α′)·Γ(k+α′)(β′+1)k+α′.

Rearrange:

P(Xn+1=k∣X1:n)=Γ(k+α′)k!Γ(α′)(β′β′+1)α′(1β′+1)k,k=0,1,2,…

This is a negative binomial law with (real-valued) size r=α′=α+S and success probability p=β′β′+1=β+nβ+n+1:

Xn+1∣X1:n ~ NegBin(r=α+S, p=β+nβ+n+1).

Its mean is r(1−p)/p=α′/β′=𝔼[λ∣X1:n], and its variance r(1−p)/p2=α′β′·β′+1β′>α′β′ exceeds the mean: the predictive is overdispersed relative to a plug-in Poisson, because it carries the posterior uncertainty in λ that a plug-in throws away.

Closing note

A quick coherence check that also gives a slicker route: because updating just adds (S,n) to (α,β), processing the data one point at a time and processing all n at once yield the same posterior — Bayesian updating is associative here precisely because the sufficient statistic is additive. The tempting error in part 3 is to "plug in" λ^=X¯ and report Poisson(X¯); that discards the integration over λ and understates predictive variance.

Answers. (1) Gamma(α+S, β+n). (2) 𝔼[λ∣X1:n]=ββ+nαβ+nβ+nX¯, weight w=n/(β+n)→1; β is a prior sample size. (3) NegBin(r=α+S, p=β+nβ+n+1) with the pmf displayed above.


Two new problems every morning at 8am · every day so far