This is a port of the following article.

数学書「Brownian Motion, Martingales, and Stochastic Calculus (Le Gall)」を読んで|moni掲題の本のレビューです。これに感化されてこの本を手に取る人間が存在するのかという話は一旦さておき、備忘録の意味も込めて書き残したいと思います。あと数学書に限らず、専門書は必ず複数回読むべき(オチを知った状態で読み返すことで目的意識がクリアになる)という信条に基づき、このレビューが 2 周目の役割を果たすことを期待しています。 Brownian Motion, Martingales, and Stochastic Calculus This book offers a rigorous and self-contained presentation o link.sprnote(ノート)

This is my review of the book named in the title. Whether anyone will actually be inspired by this review to pick up the book is a separate question, but I want to leave these notes here for my own reference as well. I also believe that technical books—not just mathematics books—should always be read more than once, because rereading them when you already know where they are going makes your purpose much clearer. I hope writing this review will serve as my second pass through the book.

Brownian Motion, Martingales, and Stochastic CalculusThis book offers a rigorous and self-contained presentation of stochastic integration and stochastic calculus within the general framework of continuous semimartingales. The main tools of stochastic calculus, including Itô’s formula, the optional stopping theorem and Girsanov’s theorem, are treated in detail alongside many illustrative examples. The book also contains an introduction to Markov processes, with applications to solutions of stochastic differential equations and to connections between Brownian motion and partial differential equations. The theory of local times of semimartingales is discussed in the last chapter. Since its invention by Itô, stochastic calculus has proven to be one of the most important techniques of modern probability theory, and has been used in the most recent theoretical advances as well as in applications to other fields such as mathematical finance. Brownian Motion, Martingales, and Stochastic Calculus provides astrong theoretical background to the reader interested in such developments. Beginning graduate or advanced undergraduate students will benefit from this detailed approach to an essential area of probability theory. The emphasis is on concise and efficient presentation, without any concession to mathematical rigor. The material has been taught by the author for several years in graduate courses at two of the most prestigious French universities. The fact that proofs are given with full details makes the book particularly suitable for self-study. The numerous exercises help the reader to get acquainted with the tools of stochastic calculus.SpringerLink

Preface

This was the first specialist book in pure mathematics that I had read in quite some time since finishing graduate school. It belongs to Springer's Graduate Texts in Mathematics series and is therefore aimed at graduate students. My master's specialization was partial differential equations, so I expected my background in another area of analysis to let me move through it smoothly. Instead, the material was quite difficult, and I struggled a great deal. Some of my knowledge had faded, and I was also paying the price for studying carelessly as an undergraduate, so I worked through the gaps by constantly bouncing questions off LLMs (Claude and Gemini). I started on May 1 and, judging by my slow progress, expected it to take the entire month. In the end, I finished my first pass on May 25, although I skipped almost all the exercises. I will use the rest of the month to review the material and write this article.

My reason for reading the book came from investing. Financial assets such as stock prices are modeled as geometric Brownian motion, the solution of the following stochastic differential equation:

dXt=σXtdBt+rXtdtdX_t = \sigma X_t \, dB_t + r X_t \, dt

Before reading this book, I could not even state a precise definition of the Brownian motion (Bt)t0(B_t)_{t \ge 0} that appears here. The equation above is also merely convenient notation, not a mathematical definition, although it is extremely useful in practice. A stochastic process X=(Xt)t0X = (X_t)_{t \ge 0} is defined to satisfy the stochastic differential equation above when, for every t0t \ge 0,

Xt=X0+0tσXsdBs+0trXsds()X_t = X_0 + \int_0^t \sigma X_s \, dB_s + \int_0^t r X_s \, ds \tag{$\ast$}

holds, along with a few additional conditions. The second term on the right, 0tσXsdBs\displaystyle \int_0^t \sigma X_s \, dB_s, is called a stochastic integral. Almost surely (a.s.ωΩ\text{a.s.}\omega \in \Omega), however, the sample path tBt(ω)t \mapsto B_t(\omega) of Brownian motion is not of bounded variation, so this term cannot be defined as a Stieltjes integral. Thus, even after being told that this is the definition of a solution to a stochastic differential equation, we are still left with a mysterious term.

For geometric Brownian motion (Xt)t0(X_t)_{t \ge 0} satisfying ()(\ast), applying Itô's formula to logXt\log X_t gives the explicit solution

Xt=X0exp(σBt+(rσ22)t).X_t = X_0 \exp \left( \sigma B_t + \left(r - \dfrac{\sigma^2}{2} \right)t \right).

In one dimension, Itô's formula states that for a continuous semimartingale XX and a twice continuously differentiable function fC2(R)f \in C^2(\mathbb R), f(X)f(X) is again a continuous semimartingale and

f(Xt)=f(X0)+0tf(Xs)dXs+120tf(Xs)dX,Xs.f(X_t) = f(X_0) + \int_0^t f'(X_s) \, dX_s + \dfrac{1}{2} \int_0^t f''(X_s) \, d \langle X, X \rangle_s.

To understand the geometric Brownian motion ()(\ast) mathematically, we must therefore define Brownian motion, interpret what (semi)martingales represent, and give the stochastic integral a mathematical meaning. With that motivation, the title Brownian Motion, Martingales, and Stochastic Calculus was a perfect fit and drew me to the book. Geometric Brownian motion does appear in Chapter 8, but astonishingly, it receives less than a page of coverage (lol). The chapter is mainly concerned with the existence and uniqueness of solutions under (local) Lipschitz continuity conditions. Its lack of interest in equations that can be solved explicitly also reminded me that I had returned to the world of pure mathematics.

Overall Assessment

Difficulty

As I said at the beginning, I think this is a difficult book. In particular, it leaves large gaps for the reader to fill. I repeatedly encountered that familiar experience of a mathematics book calling something “obvious,” “elementary,” “straightforward,” or “easy to verify” when it looked anything but obvious. As one Amazon review says, “It is a good book, but not one for beginners in probability theory.” I think readers need to be thoroughly familiar with Lebesgue integration and probability theory—or, more precisely, arguments involving finite measures. For example, assuming Brownian motion (Bt)t0(B_t)_{t \ge 0} is defined on a particular probability space (Ω,F,P)(\Omega, \mathscr F, P), the book describes Wiener measure WW as follows:

Consider C(R+,R)C(\mathbb R_+, \mathbb R), the set of continuous functions from [0,)[0, \infty) to R\mathbb R. Define the σ\sigma-algebra C\mathscr C as the smallest σ\sigma-algebra for which the coordinate map ww(t)w \mapsto w(t) is measurable for every t0t \ge 0. This agrees with the Borel σ\sigma-algebra when C(R+,R)C(\mathbb R_+, \mathbb R) is equipped with the topology of locally uniform convergence. For Brownian motion B=(Bt)t0B = (B_t)_{t \ge 0}, the map ΩωB(ω)C(R+,R)\Omega \ni \omega \mapsto B_\cdot(\omega) \in C(\mathbb R_+, \mathbb R) is measurable. Wiener measure—or the law of Brownian motion—W(dw)W(dw) is the image of the probability measure P(dω)P(d \omega) under this map. Wiener measure is characterized by its values on cylinder sets. Given 0=t0<t1<<tn0 = t_0 < t_1 < \cdots < t_n and A0,AnB(R)A_0, \cdots A_n \in \mathscr B(\mathbb R) , the Wiener measure of A={wC(R+,R):i,w(ti)Ai}A = \{ w \in C(\mathbb R_+, \mathbb R) : \forall i, \, w(t_i) \in A_i \} is W(A)=P(i,BtiAi)=1A0(0)A1××Andx1dxni=1n(2π(titi1))1/2exp(i=1n(xixi1)22(titi1)) W(A) = P(\forall i, \, B_{t_i} \in A_i) = \mathbf{1}_{A_0}(0) \displaystyle \int_{A_1 \times \cdots \times A_n} \dfrac{dx_1 \dots dx_n}{\prod_{i=1}^n (2 \pi (t_i - t_{i-1}))^{1/2}} \exp \left(- \sum_{i=1}^n \dfrac{(x_i - x_{i-1})^2}{2(t_i - t_{i-1})} \right), where x0=0x_0 = 0. This uniquely defines the law of Brownian motion.

If you can accept this reasonably smoothly, you probably have what it takes to continue reading. If, like me, your head fills with question marks, reviewing measure theory first would be wise. Here is one example of the background that the book leaves unstated:

Let πt:C(R+,R)ww(t)R\pi_t: C(\mathbb R_+, \mathbb R) \ni w \mapsto w(t) \in \mathbb R be the coordinate map. Then C\mathscr C is the smallest σ\sigma-algebra containing every inverse image πt1(A),AB(R)\pi_t^{-1}(A), \, A \in \mathscr B(\mathbb R). On the other hand, C(R+,R)C(\mathbb R_+, \mathbb R) is a Fréchet space under the countable family of seminorms wk=sup0tkw(t)\displaystyle \|w \|_k = \sup_{0 \le t \le k} |w(t)|, and its Borel σ\sigma-algebra agrees with C\mathscr C. (Ask an AI for the proof.)

A map F:ΩC(R+,R)F: \Omega \to C(\mathbb R_+, \mathbb R) is measurable if and only if πtF\pi_t \circ F is measurable for every t0t \ge 0. (Proof omitted.) This makes Brownian motion measurable, so Wiener measure can be defined as the pushforward measure W=PB1W = P \circ B^{-1}.

Let

D={i=0nπti1(Ai):0=t0<t1<<tn,AiB(R)}\mathscr D = \left\{ \displaystyle \bigcap_{i=0}^n \pi_{t_i}^{-1}(A_i) : 0 = t_0 < t_1 < \cdots < t_n, \, A_i \in \mathscr B(\mathbb R) \right\}

be the collection of all cylinder sets. It is a π\pi-system and satisfies σ(D)=C\sigma(\mathscr D) = \mathscr C. By Dynkin's π\pi-λ\lambda theorem, a measure is uniquely determined once it is defined on D\mathscr D. (The rest is omitted.) Consequently, defining Wiener measure using a different Brownian motion BB' on a different probability space produces the same measure.

Prerequisites

As explained above, measure theory and probability theory are essential. It goes without saying that calculus, linear algebra, set theory, and topology are also required. The book frequently uses arguments involving Hilbert spaces, so functional analysis is essential as well. A little complex analysis appears, but knowing the Cauchy–Riemann equations seems sufficient; like me, you can probably have forgotten everything else without much trouble.

One disappointment was that the book assumes prior knowledge of discrete-time martingales. For continuous-time martingales, particularly estimates such as Doob's inequalities, it reduces the arguments to discrete time by using the separability of R+\mathbb R_+ and the (right) continuity of sample paths. The underlying discrete-time estimates appear without proof in an appendix, but I do not regard reading propositions without proofs as studying. An enthusiastic reader might want to read the following book by the same author either before or after this one.

Evaluation

I think it is a good book. Its proofs and arguments are mathematically rigorous, and it contains relatively few typographical errors—not none—so it should help readers build a solid theoretical framework in their minds. For someone with a strong foundation in analysis, it may be a useful introduction to stochastic analysis. After reading it once, you can return to it as a reference book when your memory begins to fade.

Its main shortcoming is, again, the use of propositions about discrete martingales without proof. The book also invokes results such as Kolmogorov's extension theorem without proving them, but those lie somewhat outside the main path of stochastic analysis and did not bother me much. Martingales, even in discrete time, form part of the theory's core, so I hesitate to call the book “self-contained.” Still, it is already almost 300 pages long, so I understand the decision to leave some material to another book. A coherent preparation for research in stochastic analysis might be to take measure theory and functional analysis in the third year of university, then read Measure Theory, Probability, and Stochastic Processes, followed by Brownian Motion, Martingales, and Stochastic Calculus, in the fourth year.

Chapter Summaries

Flowchart

Flowchart
Flowchart
Flowchart

The theory develops according to the flowchart above. The author also included a flowchart in the preface, but I redrew it as a Mermaid diagram. In the following sections, I will look back and forth between the notes I kept and the book and write down my own understanding of each chapter.

Chapter 1: Gaussian Distributions

(I assigned the chapter titles myself; here and below, they do not follow the book.)

This chapter reviews the material on Gaussian distributions needed for the rest of the book. If you know probability theory and functional analysis, this part should be straightforward. For me, it was the only chapter I could move through quickly. The Gaussian white noise defined here gives the Wiener integral 0tf(s)dBs\displaystyle \int_0^t f(s) \, dB_s introduced in Chapter 2. This is a special case of a stochastic integral—an integral with respect to Brownian motion whose integrand is deterministic—and the resulting random variable is again Gaussian.

The book defines white noise for a more general σ\sigma-finite measure space (E,E,μ)(E, \mathscr E, \mu), without even assuming that EE is separable. To read only this book, however, it seems sufficient to consider E=R+E = \mathbb R_+ with μ\mu equal to Lebesgue measure. Put simply, without introducing the definition of a Gaussian space, a Gaussian white noise G:L2(R+)L2(Ω)G: L^2(\mathbb R_+) \to L^2(\Omega) is an isometric linear map, and therefore preserves inner products, such that G(f)G(f) follows a mean-zero Gaussian distribution; that is, G(f)N(0,0f(t)2dt)\displaystyle G(f) \sim \mathscr N \left(0, \int_0^\infty f(t)^2 \, dt \right). The book shows that such a GG always exists on a suitable (Ω,F,P)(\Omega, \mathscr F, P). Its isometry also gives, for 0st0 \le s \le t,

E[G(1[0,s])G(1[0,t])]=E[G(1[0,s])2]+E[G(1[0,s])G(1(s,t])]=01[0,s](r)dr+01[0,s](r)1(s,t](r)dr=s.\begin{aligned}E[G(\mathbf 1_{[0, s]}) G(\mathbf 1_{[0, t]})] & = E[G(\mathbb 1_{[0, s]})^2] + E[G(\mathbb 1_{[0, s]})G(1_{(s, t]})] \\&= \int_0^\infty \mathbf 1_{[0, s]}(r) \, dr + \int_0^\infty \mathbf 1_{[0, s]}(r) \mathbf 1_{(s, t]} (r) \, dr \\&= s.\end{aligned}

In particular, Cov(G(1[0,s]),G(1[0,t]))=st:=min(s,t)\operatorname{Cov}(G(\mathbf 1_{[0, s]}), G(\mathbf 1_{[0, t]})) = s \wedge t := \min(s, t).

Chapter 2: Brownian Motion

Brownian motion (Bt)t0(B_t)_{t \ge 0} is defined as a stochastic process of the form Bt=G(1[0,t])B_t = G(\mathbf 1_{[0, t]}) for a Gaussian white noise GG, with continuous sample paths. In fact, the statement that (Bt)t0(B_t)_{t \ge 0} is Brownian motion without assuming continuity is equivalent to saying that it is a mean-zero Gaussian process satisfying Cov(Bs,Bt)=st\operatorname{Cov}(B_s, B_t) = s \wedge t. It could therefore be defined without introducing the more elaborate L2L^2-isometry GG. I think Gaussian white noise is introduced first both to make the existence of Brownian motion obvious and to define the Wiener integral naturally by

stf(r)dBr:=G(f1(s,t]).\int_s^t f(r) \, dB_r := G(f \mathbf 1_{(s, t]}).

Even if Brownian motion were defined first, the Wiener integral could still be obtained by noting that step functions are dense in L2(R+)L^2(\mathbb R_+), that 1(s,t]BtBs\mathbf 1_{(s, t]} \mapsto B_t - B_s is an L2L^2-isometry, and that linearity and the closedness of Gaussian spaces allow an extension to L2(R+)L^2(\mathbb R_+). Thus, which one comes first is a minor question.

We have already established the existence of GG, but Brownian motion requires continuous sample paths. Another important concept from this point onward is a modification (X~t)tT(\tilde X_t)_ {t \in T} of a stochastic process (Xt)tT(X_t)_{t \in T}, meaning that Xt=X~tX_t = \tilde X_t almost surely for every tTt \in T. Kolmogorov's continuity theorem (Kolmogorov's lemma) gives the following sufficient condition for a process to have a modification with continuous sample paths:

q,ε,C>0,E[d(Xs,Xt)q]Cts1+ε.\exists q, \varepsilon, C > 0, \, E[d(X_s, X_t)^q] \le C|t - s|^{1 + \varepsilon}.

In fact, it implies local Hölder continuity. For the Brownian motion defined by GG without any prior guarantee of continuity, BtBsN(0,ts)B_t - B_s \sim \mathscr N(0, t - s) implies that, for every q>0q > 0,

E[BtBsq]=Cq(ts)q/2.E[|B_t - B_s|^q] = C_q (t-s)^{q/2}.

This establishes the existence of Brownian motion, meaning a continuous modification.

Chapter 3: Martingales and Stopping Times

This chapter introduces a filtration (Ft)t[0,](\mathscr F_t)_{t \in [0, \infty]} and stopping times TT. At first, the definition of a stopping time, t0,{Tt}Ft\forall t \ge 0, \, \{ T \le t \} \in \mathscr F_t, meant little to me. It turns out to be an extremely important concept used repeatedly in later proofs. I think the following proposition is particularly important:

Let (Xt)t0(X_t)_{t \ge 0} be a stochastic process with continuous sample paths taking values in a metric space (E,d)(E,d). If FEF \subset E is closed, then TF:=inf{t0:XtF}T_F := \inf \{ t \ge 0 : X_t \in F \} is a stopping time.

A frequently used proof technique is to define Tn:=inf{t0:Xtn}T_n := \inf \{t \ge 0 : |X_t| \ge n \} for a stochastic process (Xt)t0(X_t)_{t \ge 0} with continuous sample paths, and then define the stopped process XTnX^{T_n} by XtTn=XtTnX^{T_n}_t = X_{t \wedge T_n}. In other words, continuity and the intermediate value theorem mean that each sample path stops forever once Xt=n|X_t| = n. By definition, XtTnn|X^{T_n}_t| \le n. Since TnT_n \nearrow \infty as nn \to \infty, we have XtTnXtX_{t \wedge T_n} \to X_t, reducing arguments about the stochastic process to the bounded case. To see that limnTn=\displaystyle \lim_{n \to \infty} T_n = \infty, monotonicity gives a limit M=limnTn[0,]\displaystyle M = \lim_{n \to \infty} T_n \in [0, \infty]. If M<M < \infty, then nN,TnM\forall n \in \mathbb N, T_n \le M, which implies nN,tn[0,M],Xtnn\forall n \in \mathbb N, \exists t_n \in [0, M], |X_{t_n}| \ge n. On the other hand, continuity of the sample paths and the extreme value theorem give supt[0,M]Xt<\displaystyle \sup_{t \in [0, M]} |X_t| < \infty, a contradiction.

This was also my first encounter with martingales. Even after seeing the definition E[MtFs]=MsE[M_t \mid \mathscr F_s] = M_s and hearing an AI explain it as a “fair game,” my reaction was little more than “well, of course.” Yet this property is extremely strong and leads to a wide range of results. One of the first is the martingale convergence theorem:

Let (Mt)t0(M_t)_{t \ge 0} be a martingale with right-continuous sample paths. The following are equivalent: (1) ZL1(Ω),t0,Mt=E[ZFt]\exists Z \in L^1(\Omega), \, \forall t \ge 0, \, M_t = E[Z \mid \mathscr F_t] ; (2) {Mt}t0\{M_t\}_{t \ge 0} is uniformly integrable; and (3) ZL1(Ω),limtE[ZMt]=0\displaystyle \exists Z \in L^1(\Omega), \, \lim_{t \to \infty} E[|Z - M_t|] = 0. When these conditions hold, writing the limit as MM_\infty gives Mt=E[MFt]M_t = E[M_\infty \mid \mathscr F_t].

Another important result is the following corollary of the optional stopping theorem:

Let (Mt)t0(M_t)_{t \ge 0} be a martingale with right-continuous sample paths, and let TT be a stopping time. Then MT=(MtT)t0{M^T} = (M_{t \wedge T})_{t \ge 0} is also a martingale. If (Mt)t0(M_t)_{t \ge 0} is uniformly integrable, then (MtT)t0(M_{t \wedge T})_{t \ge 0} is uniformly integrable as well, and MtT=E[MTFt]M_{t \wedge T} = E[M_T \mid \mathscr F_t].

Brownian motion is itself a martingale; in fact, it has the still stronger property of independent increments. As discussed in connection with stopping times, even if a martingale (Mt)t0(M_t)_{t \ge 0} such as Brownian motion is not uniformly integrable, MtTnM_{t \wedge T_n} is bounded—a much stronger condition—so we can apply convergence theorems and the L2L^2 properties discussed below.

I may sound authoritative here, but uniform integrability is a concept useful in finite-measure settings such as probability theory. Having specialized in partial differential equations, where I used little beyond Lebesgue measure, I could not even state its definition. I peppered the AI with questions and reviewed results such as Vitali's convergence theorem. I regard my university and graduate-school years as the period when I studied hardest, but this drove home that no matter how much you study, a day will come when you wish you had studied more.

Chapter 4: Processes of Bounded Variation and Semimartingales

Here comes part two of “I wish I had studied more.” The chapter introduces functions of bounded variation and their associated signed measures. It is easier to follow if you understand how the Hahn decomposition of a signed measure μ\mu and the associated Jordan decomposition μ=μ+μ\mu = \mu^+ - \mu^- give the representation μ=μ++μ|\mu| = \mu^+ + \mu^- of the total variation measure, and how a function of bounded variation ff can be expressed, using its variation

V0t(f):=sup{i=1nf(ti)f(ti1):0=t0<t1<<tn=t},V_0^t(f) := \sup \left\{ \sum_{i=1}^n |f(t_i) - f(t_{i-1})| : 0 = t_0 < t_1 < \cdots < t_n = t \right\},

as the difference of two increasing functions:

f(t)=12(V0t(f)+f(t))12(V0t(f)f(t))=:f+(t)f(t).f(t) = \dfrac{1}{2} (V_0^t(f) + f(t)) - \dfrac{1}{2}(V_0^t(f) - f(t)) =: f^+(t) - f^-(t).

Their Stieltjes measures df±df^{\pm} agree with the Jordan components μ±\mu^{\pm}. Here again, I paid the price for my unserious attitude of assuming that a rough understanding was enough because the topic was unrelated to partial differential equations.

A continuous semimartingale itself is not a particularly difficult concept: XX is a continuous semimartingale if it can be written X=M+VX = M + V, where MM is a continuous local martingale and VV is a continuous process of bounded variation. A continuous local martingale is defined by the existence of stopping times (Tn)nN(T_n)_{n \in \mathbb N} such that TnT_n \nearrow \infty and MTnM^{T_n} is uniformly integrable for every nNn \in \mathbb N. Moreover, for a continuous local martingale with M0L1M_0 \in L^1, the stopping times Tn=inf{t0:Mtn}T_n = \inf\{ t \ge 0 : |M_t| \ge n \} defined earlier satisfy these conditions, so in practice we can always use them. Surprisingly, a process (Mt)t0(M_t)_{t \ge 0} is both a local martingale and of bounded variation only when Mt=M0M_t = M_0. Thus, if—as in this book—the definition of a process of bounded variation includes V0=0V_0 = 0, the semimartingale decomposition X=M+VX = M + V is unique.

It also follows that nontrivial (local) martingales, including Brownian motion, are not of bounded variation. Nevertheless, the sum of squared variations called the quadratic variation M,Mt\langle M, M \rangle_t, which appeared in Itô's formula at the beginning, does converge. More precisely:

Let (Mt)t0(M_t)_{t \ge 0} be a continuous local martingale. For every t>0t > 0, let 0=t0n<<tpnn=t0 = t_0^n < \cdots < t_{p_n}^n = t be an increasing sequence of partitions of [0,t][0, t]—that is, n,i,j,tin=tjn+1\forall n, \, \forall i, \, \exists j, \, t_i^n = t_j^{n+1}—such that limnmax1ipn(tinti1n)=0\displaystyle \lim_{n \to \infty} \max_{1 \le i \le p_n} (t_i^n - t_{i-1}^n) = 0 . Then i=1pn(MtinMti1n)2\displaystyle \sum_{i=1}^{p_n} (M_{t_i^n} - M_{t_{i-1}^n})^2 converges in probability as nn \to \infty. Denoting the limit by M,Mt\langle M, M \rangle_t, the process M,Mt\langle M, M \rangle_t is characterized as the unique continuous increasing process for which Mt2M,MtM_t^2 - \langle M, M \rangle_t is a continuous local martingale.

In particular, Brownian motion (Bt)t0(B_t)_{t \ge 0} satisfies B,Bt=t\langle B, B \rangle_t = t. This follows from its independent increments, since

E[Bt2Fs]=E[{(BtBs)+Bs}2Fs]=E[(BtBs)2]+2E[BtBs]Bs+Bs2=ts+Bs2,\begin{aligned}E[B_t^2 \mid \mathscr F_s] &= E[\{ (B_t - B_s) + B_s \}^2 \mid \mathscr F_s] \\&= E[(B_t - B_s)^2] + 2E[B_t - B_s]B_s + B_s^2 \\&= t - s + B_s^2,\end{aligned}

so Bt2tB_t^2 - t is a martingale.

Adding a process of bounded variation does not change the limit of these sums of squared variations. Expanding gives

i=1pn{(Mtin+Vtin)(Mti1n+Vti1n)}2=i=1pn(MtinMti1n)2+2i=1pn(MtinMti1n)(VtinVti1n)+i=1pn(VtinVti1n)2.\begin{aligned}& \sum_{i=1}^{p_n} \{ (M_{t_i^n} + V_{t_i^n}) - (M_{t_{i-1}^n} + V_{t_{i-1}^n}) \}^2 \\= & \sum_{i=1}^{p_n} (M_{t_i^n} - M_{t_{i-1}^n})^2 + 2\sum_{i=1}^{p_n}(M_{t_i^n} - M_{t_{i-1}^n})(V_{t_i^n} - V_{t^n_{i-1}}) + \sum_{i=1}^{p_n}(V_{t^n_i} - V_{t^n_{i-1}})^2.\end{aligned}

Then, by uniform continuity of the sample paths on [0,t][0, t], for sufficiently large nn,

i=1pn(MtinMti1n)(VtinVti1n)max1ipnMtinMti1ni=1pnVtinVti1nεV0t(f),\begin{aligned}& \left| \sum_{i=1}^{p_n}(M_{t^n_i} - M_{t^n_{i-1}})(V_{t^n_i} - V_{t^n_{i-1}}) \right| \\ \le & \max_{1 \le i \le p_n} | M_{t^n_i} - M_{t^n_{i-1}} | \sum_{i=1}^{p_n} |V_{t^n_i} - V_{t^n_{i-1}}| \\ \le & \varepsilon V_0^t(f),\end{aligned}

and the other term is handled similarly. The quadratic variation of a continuous semimartingale X=M+VX = M + V is therefore naturally defined as X,X=M,M\langle X, X \rangle = \langle M, M \rangle. Quadratic variation also admits the bilinear extension X,Y=12(X+Y,X+YX,XY,Y)\langle X, Y \rangle = \dfrac{1}{2} (\langle X+Y, X+Y \rangle - \langle X, X \rangle - \langle Y, Y \rangle) . Since M,Mt\langle M, M \rangle_t is a continuous increasing process, X,Yt\langle X, Y \rangle_t is a continuous process of bounded variation.

Chapter 5: Stochastic Integrals

This is the heart of the book, occupying more than 50 pages by itself. One of the chapter's goals is to define the stochastic integral

0tHsdXs=0tHsdMs+0tHsdVs\int_0^t H_s \, dX_s = \int_0^t H_s \, dM_s + \int_0^t H_s \, dV_s

for a semimartingale—I will omit “continuous” because it is cumbersome—X=M+VX = M + V and a suitable stochastic process (Ht)t0(H_t)_{t \ge 0} serving as the integrand. The integral with respect to dVsdV_s is a Stieltjes integral, so if HtH_t is locally bounded, meaning that sup0stHs<\displaystyle \sup_{0 \le s \le t }|H_s| < \infty almost surely for every t0t \ge 0, it can be defined because 0tHsdVs0tHsdVs<\displaystyle \left| \int_0^t H_s \, dV_s \right| \le \int_0^t |H_s| \, |dV_s| < \infty. The remaining goal is therefore to define stochastic integration with respect to a local martingale.

By analogy with the Riemann–Stieltjes integral, we expect a property such as

limni=1pnf(cin)(g(tin)g(ti1n))=0tf(s)dg(s),cin[ti1n,tin].\displaystyle \lim_{n \to \infty} \sum_{i=1}^{p_n} f(c^n_i) (g(t^n_i) - g(t^n_{i-1})) = \int_0^t f(s) \, dg(s), \,\, \forall c^n_i \in [t^n_{i-1}, t^n_i].

This is indeed true to a point. When (Xt)t0(X_t)_{t \ge 0} is a continuous semimartingale and (Ht)t0(H_t)_{t \ge 0} is continuous, one important convergence theorem for stochastic integrals states that

plimni=1pnHti1n(XtinXti1n)=0tHsdXs.\operatorname*{plim}_{n \to \infty} \sum_{i=1}^{p_n} H_{t^n_{i-1}} (X_{t^n_i} - X_{t^n_{i-1}}) = \int_0^t H_s \, dX_s.

When HtH_t is itself a semimartingale, the following observation shows why the integrand must be evaluated at the left endpoint ti1nt^n_{i-1} of each subinterval: choosing a different point changes the limit in probability.

i=1pnHtin(XtinXti1n)=i=1pnHti1n(XtinXti1n)+i=1pn(HtinHti1n)(XtinXti1n)p0tHsdXs+H,Xt.\begin{aligned}& \sum_{i=1}^{p_n} H_{t^n_{i}} (X_{t^n_i} - X_{t^n_{i-1}}) \\= & \sum_{i=1}^{p_n} H_{t^n_{i-1}} (X_{t^n_i} - X_{t^n_{i-1}}) + \sum_{i=1}^{p_n} (H_{t^n_i} - H_{t^n_{i-1}}) (X_{t^n_i} - X_{t^n_{i-1}}) \\\xrightarrow{p} & \int_0^t H_s \, dX_s + \langle H, X \rangle_t.\end{aligned}

Chapter 4 showed that the quadratic covariation H,Xt\langle H, X \rangle_t is identically zero when either argument is of bounded variation. Conversely, when neither is of bounded variation, it is not zero; the effect of integrating with respect to a process that is not of bounded variation appears as the quadratic covariation.

We want to define stochastic integration with respect to a local martingale (Mt)t0(M_t)_{t \ge 0}. Very loosely speaking, it is enough to construct the theory under reasonably well-behaved assumptions because we can reduce to stopped processes. Indeed, 0tHsdMs\displaystyle \int_0^t H_s \, dM_s can be viewed as stopping at the constant stopping time T=tT=t, and more generally 0THsdMs=0HsdMsT\displaystyle \int_0^T H_s \, dM_s = \int_0^\infty H_s \, dM^T_s. Also speaking loosely and without justification, the analogy with the Riemann–Stieltjes sums above means it suffices to consider M0=0M_0 = 0. The definition uses the Hilbert-space structure of L2L^2, and the following powerful property promotes a local martingale to a martingale:

Let (Mt)t0(M_t)_{t \ge 0} be a continuous local martingale with M0=0M_0 = 0. The following are equivalent: (1) MM is an L2L^2-bounded martingale, meaning supt0E[Mt2]<\displaystyle \sup_{t \ge 0} E[M_t^2] < \infty; and (2) E[M,M2]<E[\langle M, M \rangle_\infty^2] < \infty. When these hold, Mt2M,MtM_t^2 - \langle M, M \rangle_t is a uniformly integrable martingale, and in particular E[M2]=E[M,M]E[M_\infty^2] = E[\langle M, M \rangle_\infty].

Chapter 2 defined the Wiener integral using Gaussian white noise. Yet even there, explicit calculations can only be performed for step functions using stdBs=G(1(s,t])=BtBs\displaystyle \int_s^t dB_s = G(\mathbf 1_{(s, t]}) = B_t - B_s. For a general fL2(R+)f \in L^2(\mathbb R_+), the only definition provided uses the density of step functions to approximate it. Since the Wiener integral is a special case of the stochastic integral, we may as well resign ourselves to constructing stochastic integrals in the same way.

Let H2\mathbb H^2 be the set of all L2L^2-bounded martingales (Mt)t0(M_t)_{t \ge 0} with M0=0M_0 = 0. It is a Hilbert space with inner product (M,N)H2=E[M,N]=E[MN](M, N)_{\mathbb H^2} = E[\langle M, N \rangle_\infty] = E[M_\infty N_\infty]. Fix MHM \in \mathbb H, and define the Hilbert space L2(M)L^2(M) to consist of stochastic processes satisfying HL2(M):=E[0Hs2M,Ms]1/2<\displaystyle \|H\|_{L^2(M)} := E \left[\int_0^\infty H_s^2 \, \langle M, M \rangle_s \right]^{1/2} < \infty, where the integral is a Stieltjes integral. A dense subspace consists of all stochastic processes of the form

Hs(ω)=i=1pH(i1)(ω)1(ti1,ti](s),H_s(\omega) = \sum_{i=1}^p H_{(i - 1)} (\omega) \mathbf 1_{(t_{i-1}, t_i]}(s),

where each H(i1)H_{(i-1)} is bounded and Fti1\mathscr F_{t_{i-1}}-measurable. Defining the stochastic integral for such processes by

0tHsdMs=(HM)t=H(i1)(MtitMti1t),\int_0^t H_s \, dM_s = (H \cdot M)_t = H_{(i-1)} (M_{t_i \wedge t} - M_{t_{i-1} \wedge t}),

we obtain

HM,HMt=i,j=1pH(i1)(MtitMti1t),H(j1)(MtjtMtj1t)=i=1pH(i1)(MtitMti1t),H(i1)(MtitMti1t)=i=1pH(i1)2(M,MtitM,Mti1t)=0tHs2dM,Ms.\begin{aligned}\langle H \cdot M, H \cdot M \rangle_t &= \sum_{i, j = 1}^p \langle H_{(i-1)} (M_{t_i \wedge t} - M_{t_{i-1} \wedge t}), H_{(j-1)}(M_{t_j \wedge t} - M_{t_{j-1} \wedge t}) \rangle \\&= \sum_{i=1}^p \langle H_{(i-1)} (M_{t_i \wedge t} - M_{t_{i-1} \wedge t}), H_{(i-1)}(M_{t_i \wedge t} - M_{t_{i-1} \wedge t}) \rangle \\&= \sum_{i=1}^p H_{(i-1)}^2(\langle M, M \rangle_{t_i \wedge t} - \langle M, M \rangle_{t_{i-1} \wedge t}) \\&= \int_0^t H_s^2 \, d \langle M, M \rangle_s.\end{aligned}

The intermediate steps can be established using approximations of the quadratic covariation M,Nt\langle M, N \rangle_t. Furthermore, H(i1)(MtitMti1t)H_{(i-1)} (M_{t_i \wedge t} - M_{t_{i-1} \wedge t}) is a martingale for each ii, so HMH \cdot M belongs to H2\mathbb H^2, and the correspondence HHMH \mapsto H \cdot M is an isometry. Extending this map defines the stochastic integral HMH2H \cdot M \in \mathbb H^2 for every HL2(M)H \in L^2(M) and MH2M \in \mathbb H^2.

Extending this construction to local martingales shows that the stochastic integral is a local martingale, but of course does not imply that it belongs to H2\mathbb H^2. If, however, for some t[0,]t \in [0, \infty],

E[0tHs2dM,Ms]<,E\left[ \int_0^t H_s^2 \, d\langle M, M \rangle_s \right] < \infty,

then viewing it as an integral with respect to the stopped process MtM^t for T=tT=t gives (HM)tH2(H \cdot M)^t \in \mathbb H^2. In particular, it is a martingale and satisfies the isometry

E[(0tHsdMs)2]=E[0tHs2dM,Ms].E \left[ \left( \int_0^t H_s \, dM_s \right)^2 \right] = E \left[ \int_0^t H_s^2 \, d \langle M, M \rangle_s \right].

This chapter contains much more important material, but I will finish by discussing Itô's formula, which appeared at the beginning, and the explicit solution of geometric Brownian motion, although the existence of that solution has not yet been guaranteed at this point. The formal statement of Itô's formula is as given earlier, but the following form is useful for actual calculations:

df(Xt)=f(Xt)dXt+12f(Xt)dX,Xt.df(X_t) = f'(X_t) \, dX_t + \dfrac{1}{2} f''(X_t) \, d\langle X, X \rangle_t.

Suppose XtX_t satisfies the stochastic differential equation for geometric Brownian motion, dXt=σXtdBt+rXtdtdX_t = \sigma X_t \, dB_t + rX_t \, dt. Applying Itô's formula to logXt\log X_t gives

dlogXt=dXtXt12dX,XtXt2=σXtdBt+rXtdtXt12σ2Xt2dB,BtXt2=σdBt+rdt12σ2dB,Bt=σdBt+(r12σ2)dt,\begin{aligned}d \log X_t &= \dfrac{dX_t}{X_t} - \dfrac{1}{2} \dfrac{d \langle X, X \rangle_t}{X_t^2} \\&= \dfrac{\sigma X_t \, dB_t + r X_t \, dt}{X_t} - \dfrac{1}{2} \dfrac{\sigma^2 X_t^2 \, d \langle B, B \rangle_t}{X_t^2} \\&= \sigma \, dB_t + r \, dt - \dfrac{1}{2} \sigma^2 \, d \langle B, B \rangle_t \\&= \sigma \, dB_t + \left(r - \dfrac{1}{2} \sigma^2 \right) dt,\end{aligned}

which yields the solution. These formal calculations are justified by the following two associative laws:

  • K(HX)=(HK)XK \cdot (H \cdot X) = (HK) \cdot X
  • HX,KY=(HK)X,Y\langle H \cdot X, K \cdot Y \rangle = (HK) \cdot \langle X, Y \rangle

Apply them respectively to 1XX\dfrac{1}{X} \cdot X and 1X2X,X\dfrac{1}{X^2} \cdot \langle X, X \rangle, then substitute Xt=X0+((σX)B)t+((rX)1Ωid)tX_t = X_0 + ((\sigma X) \cdot B)_t + ((rX) \cdot \mathbf 1_\Omega\text{id})_t.

Chapter 6: Markov Processes

This chapter has a rather different flavor from what came before. Chapter 2 already established the Markov property of Brownian motion, but I think the purpose here is to examine it for a broader class of stochastic processes. In particular, the chapter considers the strong continuity assumptions of a Feller semigroup. This reminded me nostalgically of the Hille–Yosida theorem from functional analysis, although I had completely forgotten what that theorem actually says.

A Feller semigroup (Qt)t0(Q_t)_{t \ge 0} assigns a probability measure Qt(x,)Q_t(x, \cdot) to each xRdx \in \mathbb R^d, and the operators

Qtf(x)=Rdf(y)Qt(x,dy)Q_tf(x) = \int_{\mathbb R^d} f(y) \, Q_t(x, dy)

form a C0C_0-semigroup on C0(Rd)C_0(\mathbb R^d). An Rd\mathbb R^d-valued stochastic process XtX_t is a Markov process with semigroup (Qt)t0(Q_t)_{t \ge 0} if, for every fL(Rd)f \in L^\infty(\mathbb R^d) and 0st0 \le s \le t,

E[f(Xt)Fs]=Qtsf(Xs)=Rdf(x)Qts(Xs,dx).E[f(X_t) \mid \mathscr F_s] = Q_{t-s}f(X_s) = \int_{\mathbb R^d} f(x) \, Q_{t-s}(X_s, dx).

In other words, even with all information up to Fs\mathscr F_s, the distribution of XtX_t depends only on the current value XsX_s and is Qts(Xs,)Q_{t-s}(X_s, \cdot). We have not established the existence of such a Markov process, although one can always be constructed. If XX is a Markov process, repeatedly taking conditional expectations shows inductively that its finite-dimensional marginal distribution, for 0=t0<t1<<tn0 = t_0 < t_1 < \cdots < t_n and A0,AnB(Rd)A_0, \cdots A_n \in \mathscr B(\mathbb R^d), is

P(i,XtiAi)=A0A1A2AnQtntn1(xn1,dxn)Qt2t1(x1,dx2)Qt1(x0,dx1)γ(dx0),\begin{aligned}& P(\forall i, \, X_{t_i} \in A_i) \\ =& \int_{A_0}\int_{A_1}\int_{A_2} \cdots \int_{A_n} Q_{t_n - t_{n-1}}(x_{n-1}, dx_n) \cdots Q_{t_2 - t_1}(x_1, dx_2)Q_{t_1}(x_0, dx_1) \gamma(dx_0),\end{aligned}

where γ\gamma is the law of X0X_0, γ=PX01\gamma = P \circ X_0^{-1}. Thus, the distribution of a Markov process with Feller semigroup (Qt)t0(Q_t)_{t \ge 0} is uniquely characterized by the distribution of X0X_0. A dd-dimensional Brownian motion is a Markov process with a Feller semigroup when Qt(x,dy)=pt(yx)dyQ_t(x, dy) = p_t(y - x) \, dy, using the Gaussian kernel pt(x)=1(2πt)d/2exp(x22t)p_t(x) = \dfrac{1}{(2 \pi t)^{d/2}} \exp \left(- \dfrac{|x|^2}{2t} \right). Substituting this into the expression for the marginal distributions recovers the values of Wiener measure on cylinder sets described at the beginning.

Very loosely, the ordinary Markov property says that for every s0s \ge 0, conditional on the information in Fs\mathscr F_s, the distribution of the Markov process (Xs+t)t0(X_{s + t})_{t \ge 0} equals that of a Markov process (Xt)t0(X'_t)_{t \ge 0} with X0=XsX'_0 = X_s. Applied to Brownian motion (Bt)t0(B_t)_{t \ge 0}, conditional on Fs\mathscr F_s we have Bs+t=(d)βt+BsB_{s+t} \overset{(d)} = \beta_t + B_s, where βt\beta_t is another Brownian motion and =(d)\overset{(d)}{=} denotes equality in distribution. Independent increments and Bs+t=(Bs+tBs)+Bs=(d)βt+BsB_{s+t} = (B_{s+t} - B_s) + B_s \overset{(d)}{=} \beta_t + B_s show that knowing Fs\mathscr F_s provides no information beyond fixing BsB_s. Consequently, (Bs+tBs)t0(B_{s+t} - B_s)_{t \ge 0} is a Brownian motion independent of Fs\mathscr F_s. The strong Markov property asserts that the same remains true when the time ss is replaced by a stopping time TT, and Markov processes with Feller semigroups satisfy it.

Chapter 7: Partial Differential Equations

This chapter is completely independent, so you can safely skip it if you like. The author even mentions in Chapter 8 that it is independent of Chapter 7. The first half examines properties of harmonic functions satisfying Laplace's equation from the perspective of Brownian motion. These results can also be proved by ordinary partial-differential-equation methods without stochastic processes, so although the viewpoint is interesting, it does not produce especially novel results. I think this part mainly prepares for the second half.

The second half is more interesting and discusses the recurrence—or transience—and asymptotic behavior of dd-dimensional Brownian motion. By analogy with the fact that a discrete-time random walk is recurrent in two or fewer dimensions, always returning to its starting point, and transient in three or more dimensions, consider Brownian motion with B0=xRd{0}B_0 = x \in \mathbb R^d \setminus \{0\} and the stopping time Ur=inf{t0:Bt=r}U_r = \inf \{ t \ge 0 : |B_t| = r \} for 0r<x0 \le r < |x|. When d2d \ge 2, P(U0<)=0P(U_0 < \infty) = 0, while for r>0r > 0,

P(Ur<)={1d=2(rx)d2d3.P(U_r < \infty) = \begin{cases}1 \quad & d = 2 \\\left(\dfrac{r}{|x|}\right)^{d-2} \quad & d \ge 3.\end{cases}

In two dimensions, the probability of reaching the origin—and hence any particular point other than xx—is zero, yet the process eventually visits every arbitrarily small ball. In three or more dimensions, even that is not guaranteed.

Chapter 8: Stochastic Differential Equations

I have already written what I wanted to say in the preface and Chapter 5. As in the proof of existence and uniqueness for ordinary differential equations, the chapter assumes Lipschitz continuity of σ,b\sigma, b in the stochastic differential equation

dXt=σ(t,Xt)dBt+b(t,Xt)dt    defXt=X0+0tσ(s,Xs)dBs+0tb(s,Xs)ds()\begin{aligned}& dX_t = \sigma(t, X_t) \, dB_t + b(t, X_t) \, dt \\\overset{\text{def}}{\iff} & X_t = X_0 + \int_0^t \sigma(s, X_s) \, dB_s + \int_0^t b(s, X_s) \, ds\end{aligned} \tag{$\ast \ast$}

and then grinds through the calculations using Grönwall's inequality and Picard iteration. After proving existence and uniqueness, it examines three stochastic differential equations, including geometric Brownian motion. When an equation can be solved, Itô's formula gives an explicit solution; when it cannot, the book studies its behavior using stopping times.

In relation to the Markov property, under the Lipschitz conditions the solution to ()(\ast \ast) is a Markov process with a Feller semigroup and therefore satisfies the strong Markov property. Its generator can also be written explicitly. The book does not develop the subject much further, however, so perhaps this is intended as an invitation to pursue the area if it interests you.

Chapter 9: Local Time

This is a rather mysterious chapter—or perhaps I should say that considering a mysterious stochastic process called local time somehow produces all sorts of properties, which means I did not understand it very well. At first I thought local time had been introduced to generalize Itô's formula. Then, in the final section on the Kallianpur–Robbins law, local time suddenly appears in the study of the asymptotic behavior of two-dimensional Brownian motion and makes it possible to calculate a limiting distribution. Perhaps it has many applications. Since this is the only example the book presents, I cannot really tell.

For a semimartingale XX and aRa \in \mathbb R, the increasing process La(X)=(Lta(X))t0L^a(X) = (L^a_t(X))_{t \ge 0} characterized by the limit

Lta(X)=limε0+1ε0t1{aXsa+ε}dX,XsL^a_t(X) = \lim_{\varepsilon \to 0^+} \dfrac{1}{\varepsilon} \int_0^t \mathbf 1_{\{a \le X_s \le a + \varepsilon \}} \, d \langle X, X \rangle_s

is called local time. It is continuous in t0t \ge 0 and càdlàg in aRa \in \mathbb R. As a generalization of Itô's formula, if ff is convex on R\mathbb R, then f(X)f(X) is a semimartingale and

f(Xt)=f(X0)+0tf(Xs)dXs+12Lta(X)df(a).f(X_t) = f(X_0) + \int_0^t f'_-(X_s) \, dX_s + \dfrac{1}{2} \int_{-\infty}^\infty L^a_t(X) \, df'_-(a).

Here ff'_- is the left derivative of ff, and dfdf'_- is the Stieltjes measure associated with the increasing function ff'_-. That said, this book gives little detail about what Lta(X)L^a_t(X) actually looks like, beyond facts such as Lt0(B)=(d)BtL^0_t(B) \overset{(d)} = |B_t|, so I do not really understand its practical usefulness. Chapters 7 through 9 are fundamentally intended to show some of the ways in which the material developed through Chapter 6 can be applied. Perhaps the lingering feeling that “this might be useful somehow...?” is meant to nudge the reader toward research in stochastic analysis.