Cheng S. A Short Course on the Lebesgue Integral and Measure Theory (web draft, 2004)(53s)_MCat_
PDF · 53 pages · 430.3 KB
Open PDF file
Downloaded expository text (a web draft dated August 5, 2004) by Steve Cheng, not Phil's own work, kept in a folder of math book downloads. It begins with the limits of the Riemann integral, then covers sigma algebras, measures, measurable functions, convergence theorems, Lp spaces, construction of Lebesgue measure, Fubini, change of variables, density of C0-infinity functions, and Egorov's theorem. Exercises and a bibliography are included.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
A Short Course on the Lebesgue Integral and
Measure Theory
Steve Cheng
August 5, 2004
Contents
1 Motivation for the Lebesgue integral 2
2 Basic measure theory 4
3 Measurable functions 8
4 Definition of the Lebesgue Integral 11
5 Convergence theorems 15
6 Some Results of Integration Theory 18
7 Lpspaces 23
8 Construction of Lebesgue Measure 28
9 Lebesgue Measure in Rn32
10 Riemann integrability implies Lebesgue integrability 34
11 Product measures and Fubini’s Theorem 36
12 Change of variables in Rn39
13 Vector-valued integrals 41
14 C∞
0functions are dense in Lp(Rn) 42
15 Other examples of measures 47
16 Egorov’s Theorem 50
17 Exercises 51
1
18 Bibliography 52
Preface
This article develops the basics of the Lebesgue integral and measure theory.
In terms of content, it adds nothing new to any of the existing textbooks on
the subject. But our approach here will be to avoid unduly abstractness and
absolute generality, instead focusing on producing proofs of useful results as
quickly as possible.
Much of the material here comes from lecture notes from a short real analysis
course I had taken, and the rest are well-known results whose proofs I had
worked out myself with hints from various sources. I typed this up mainly for
my own benefit, but I hope it will be interesting for anyone curious about the
Lebesgue integral (or higher mathematics in general).
I will be providing proofs of every theorem. If you are bored reading them,
you are invited to do your own proofs. The bibliography outlines the background
you need to understand this article.
Copyright matters
Permission is granted to copy, distribute and/or modify this document under
the terms of the GNU Free Documentation License, Version 1.2 or any later
version published by the Free Software Foundation; with no Invariant Sections,
with no Front-Cover Texts, and with no Back-Cover Texts.
1 Motivation for the Lebesgue integral
If you have followed the rigorous definition of the Riemann integral in RorRn,
you may be wondering why do we need to study yet another integral. After all,
why should we even care to integrate nasty functions like:
D(x) =/braceleftbigg1, x∈Q
0, x∈R\Q
Rephrased in another way, D(x) is actually the indicator function1of the
set
S={x∈Q} ⊂R,
and we want to find its “length”. Continuing to rephrase this question, sup-
pose we are taking many real-valued measurements xof a particular physical
phenomenon. What is the probability, say, that xis rational? If we assume
1For any set S, this is the function χSdefined by χS(x) = 1 ifx∈SandχS(x) = 0 if
x /∈S. Math people call this the “characteristic function”, while probability people call it the
“indicator function” instead.
2
the measurements are distributed normally with a mean of µand a standard
deviation of σ, then this is given by:
Pr[X∈Q] =/integraldisplay
S1√
2πσ2e−1
2(x−µ
σ)2
dx
So wild sets like Sare theoretically worth considering, and it does not work to
use Riemann integral to evaluate the above probability.
Another limitation to the Riemann integral is with limits. If a sequence of
functionsfnis uniformly convergent (on a closed interval, or more generally a
compact set A⊆Rn), then we can interchange limits for the Riemann integral:
lim
n→∞/integraldisplay
Afn(x)dx=/integraldisplay
Alim
n→∞fn(x)dx,
but the criterion of uniform convergence is often too restrictive, e.g. when
integrating Fourier series. On the other hand, it can be proven with the Lebesgue
integral that the interchange is valid under weaker conditions (e.g. the functions
fnis bounded above somehow, and they converge pointwise ).
As an added benefit, some sophisticated results concerning the Riemann
integral, such as the Change of Variables Theorem in Rn, are more easily proven
using the Lebesgue integral, with its arsenal of limit theorems.
Finally, the Riemann integral does not deal with integration over “infinite
bounds” very well. For example, the standard way to compute the probability
integral/integraldisplay+∞
−∞e−1
2x2dx
goes like this:
/parenleftBig/integraldisplay+∞
−∞e−1
2x2dx/parenrightBig2
=/integraldisplay+∞
−∞e−1
2x2dx/integraldisplay+∞
−∞e−1
2y2dy
=/integraldisplay+∞
−∞/integraldisplay+∞
−∞e−1
2(x2+y2)dxdy
=/integraldisplay
R2e−1
2(x2+y2)dxdy
=/integraldisplay2π
0/integraldisplay∞
0e−1
2r2rdrdθ (using polar coordinates)
= 2π/bracketleftBig
−e−1
2r2/bracketrightBigr=∞
r=0
= 2π.
So /integraldisplay+∞
−∞e−1
2x2dx=√
2π.
The above computation seems easy, and although it can be justified using the
Riemann integral alone, it is not entirely trivial, but it is with the Lebesgue
3
integral. (For example, why should/integraltext+∞
−∞/integraltext+∞
−∞be the same as/integraltext
R2? Note that
in the Riemann theory, the iterated integral and the area integral are proven to
be equal only for bounded sets of integration.)
You will probably be able to find other sorts of limitations with the Riemann
integral.
2 Basic measure theory
The setting of abstract integration is measure theory, which tells us what the
areas or volumes of various sets are. Essentially we are given some function
µof sets which returns the area or volume — formally called the measure —
of the given set. i.e. We assume at the beginning that such a function µhas
already been defined for us. The abstract approach of the Lebesgue integral has
the obvious advantage that the theory can be applied to many other measures
besides volume in Rn.
We begin with the axioms of measure theory.
Definition 2.1. LetXbe any non-empty set. A sigma algebra2of subsets of
Xis a family Aof subsets of X, with the properties:
1.Ais non-empty.
2.IfE∈ A, thenX\E∈ A.
3.If{En}n∈Nis a sequence of sets in A, then their union is in A. That is,
Ais closed under countable unions.
The pair (X,A) is called a measurable space , and the sets in Aare called
themeasurable sets .
Notice that the axioms always imply that X∈ A. Also, by De Morgan’s
laws,Ais closed under countable intersections as well as countable union.
Needless to say, we cannot insist that Ais closed under arbitrary unions
or intersections, as that would force A= 2XifAcontains all the singleton
sets. That would be uninteresting. On the other hand, we want closure under
countable set operations, rather than just finite ones, as we will want to take
countable limits.
Example 2.1.LetXbe any (non-empty) set. Then A= 2Xis a sigma algebra.
Example 2.2.LetXbe any (non-empty) set. Then A={X,∅}is a sigma
algebra.
To get non-trivial sigma algebras to work with we need the following, a very
unconstructive(!) construction:
If we have a family of sigma algebras on X, then the intersection of all the
sigma algebras from this family is also a sigma algebra on X. If all of the sigma
2I do not know why it has such a ridiculous name, other than the fact that it is often
denoted by the Greek letter.
4
algebras from the family contains some fixed G ⊆ 2X, then the intersection of
all the sigma algebras from the family, of course, is a sigma algebra containing
G.
Now if we are given G, and we take allthe sigma algebras on Xthat contain
G, and intersect all of them, we get the smallest sigma algebra that contains G.
Definition 2.2. The smallest sigma algebra containing any given G ⊆ 2X, as
constructed above, is denoted /angbracketleftG/angbracketright, and is also called the sigma algebra generated
byG.
The following is an often-used sigma algebra.
Definition 2.3. IfXis a topological space, we can construct the sigma algebra
/angbracketleftT /angbracketright, where Tis the set of all open sets. This is called the Borel sigma algebra
and is denoted B(X). When topological spaces are involved, we will always take
the sigma algebra to be the Borel sigma algebra unless stated otherwise.
B(X), being generated by the open sets, then contains all open sets, all
closed sets, and countable unions and intersections of open sets and closed sets.
It seems unlikely, however, that every set in B(X) is expressible as a countable
union and/or intersection of open sets and closed sets, although it is tempting
to think that.
By the way, Theorem 9.3shows the Borel sigma algebra is generally not all
of 2X.
Sigma algebras are the domain on which measures are defined.
Definition 2.4. Let (X,A) be a measurable space. A positive measure on this
space is a function µ:A → [0,∞] such that
1.µ(∅) = 0
2.Countable additivity : For any sequence of mutually disjoint setsEn∈ A,
µ/parenleftbigg∞/uniondisplay
n=1En/parenrightbigg
=∞/summationdisplay
n=1µ(En).
The set (X,A,µ) will be called a measure space . Whenever convenient we
will abbreviate this expression, as in “let Xbe a measure space”, etc. Also, in
this article, when we say “measure”, we will be dealing with positive measures
only. (There are also theories about signed measures and complex measures.)
Example 2.3.LetXbe an arbitrary set, and Abe a sigma algebra on X. Define
µ:A → [0,∞] as
µ(A) =/braceleftbigg|A|,ifAis a finite set
∞,ifAis an infinite set .
This is called the counting measure .
We will be able to model the infinite series/summationtext∞
n=1anin Lebesgue integration
theory by using X=Nand the counting measure, since integrals are essentially
sums of the integrand values weighted by areas or measures.
5
Example 2.4.X=Rn, andA=B(Rn). We can construct the Lebesgue measure
λwhich assigns to the rectangle [ a1,b1]× ··· × [an,bn] inRnits expected n-
dimensional volume ( b1−a1)···(bn−an). Of course this measure should also
assign the correct volumes to the usual geometric figures, as well as for all the
other sets in A.
The existence of such a measure will be demonstrated later.
Intuitively, defining the volume of the rectangle only should suffice to uniquely
also determine the volume of the other sets, since the volume of every set can
be approximated by the volume of many small rectangles. Indeed, we will later
show this intuition to be true. In fact, you will see that most theorems using
Lebesgue measure really depend only on the definition of the volume of the
rectangle.
Example 2.5.Any probability measure (as defined by the usual axioms of prob-
ability) is actually a measure in our sense. For example,
Pr[Z∈B] =µ(B) =/integraldisplay
B1√
2πe−1
2t2dt,3
where the integration, of course, is with respect to the Lebesgue measure on
the real line. Other examples include the uniform distribution, the Poisson
distribution, and so forth.
Before we begin the prove more theorems, I must mention that we will be
operating on the quantity ∞as if it were a number, even though you may have
been told this is “wrong” by some teachers. It is true, of course, that certain
algebraic properties of Rwould fail to hold with ∞included (i.e. R∪{∞,−∞}
is not a field), but the crucial point in real analysis is that ∞obeys the usual
ordering rules when used in inequalities. The rules we adopt are the following:
a≤ ∞
∞+∞=∞
a· ∞=∞(a/negationslash= 0)
0· ∞= 0
The first three rules are self-explanatory. The last rule may need explaining:
when integrating functions, we often want to ignore “isolated” singularities, e.g.
at zero for/integraltext1
0dx/√x. The point 0 is supposed to have “measure zero”, so
even though the function is ∞there, the area contribution at that point should
still be 0 = 0 · ∞. Hence the rule. At this point a warning should be issued:
the additive cancellation rule will not work with∞. The danger should be
sufficiently illustrated in the proofs of the following theorems.
Theorem 2.1. The following are easy facts about measures:
1.It is finitely additive.
3I guess this is my favorite integral. It’s got all the important numbers in it — well, except
fori.
6
2.Monotonicity: If E,F∈ A, andE⊆F, thenµ(E)≤µ(F).
3.IfE⊆Fhas finite measure ( µ(E)<∞), thenµ(F\E) =µ(F)−µ(E).
4.IfAorBhas finite measure, then µ(A∪B) =µ(A) +µ(B)−µ(A∩B).
Proof. The first fact is obvious. For the second fact, we have
µ(F) =µ((F\E)/unionmultiE) =µ(F\E) +µ(E),
andµ(F\E)≥0. For the third fact, just subtract µ(E) from both sides. (The
funny union symbol means that the union is disjoint.) Of course the fact that
Ais a sigma algebra is used throughout to know that the new sets also belong
toA.)
For the fourth fact, we decompose each of A,B, andA∪Binto disjoint
parts, to obtain the following:
µ(A) =µ(A∩Bc) +µ(A∩B).
µ(B) =µ(B∩Ac) +µ(B∩A).
µ(A∪B) =µ(A∩Bc) +µ(Ac∩B) +µ(A∩B).
Adding the first two equations and then substituting in the third one,
µ(A) +µ(B) =µ(A∩Bc) +µ(A∩B) +µ(B∩Ac) +µ(B∩A)
=µ(A∪B) +µ(A∩B).
Since one of AorBhas finite measure, so does A∩B⊆A,B, by the second
fact, so we may subtract µ(A∩B) from both sides. Of course if one of AorB
has infinite measure, the resulting equation says nothing interesting. /square
The preceding theorem, as well as the next ones, are quite intuitive and you
should have no trouble remembering them.
Theorem 2.2. Let(X,A,µ)be measure space, and let E1⊆E2⊆E3⊆ ···
be subsets in Awith union E. (The sets Enare said to increase to E, and
henceforth we will write {En} /arrownortheastEfor this.) Then
µ(E) =µ/parenleftbigg∞/uniondisplay
n=0En/parenrightbigg
= lim
n→∞µ(En).
Proof. The setsEkandEcan be written as the disjoint unions
Ek=E1∪(E2\E1)∪(E3\E2)∪ ··· ∪ (Ek\Ek−1)
E=E1∪(E2\E1)∪(E3\E2)∪ ···,
(and setE0=∅), so that
µ(E) =∞/summationdisplay
k=1µ(Ek\Ek−1) = lim
n→∞n/summationdisplay
k=1µ(Ek\Ek−1) = lim
n→∞µ(En)./square
7
Theorem 2.3. For anyEn∈ A,
µ/parenleftbigg∞/uniondisplay
n=1En/parenrightbigg
≤∞/summationdisplay
n=1µ(En).
Proof.
µ/parenleftbigg∞/uniondisplay
n=1En/parenrightbigg
=µ/parenleftbigg∞/uniondisplay
n=1En\(E1∪E2∪ ··· ∪En−1)/parenrightbigg
=∞/summationdisplay
n=1µ/parenleftbig
En\(E1∪E2∪ ··· ∪En−1)/parenrightbig
≤∞/summationdisplay
n=1µ(En)./square
Theorem 2.4. Let{En} /arrowsoutheastE(that is,Enare decreasing and their intersection
isE), andµ(E1)<∞. Then
lim
n→∞µ(En) =µ(E) =µ/parenleftbigg∞/intersectiondisplay
n=1En/parenrightbigg
.
Proof. We have {E1\En} /arrownortheast(E1\E). So
µ(E1\E) =µ/parenleftbigg∞/uniondisplay
n=1(E1\En)/parenrightbigg
= lim
n→∞µ(E1\En),
µ(E1)−µ(E) = lim
n→∞[µ(E1)−µ(En)] =µ(E1)−lim
n→∞µ(En),
and cancel µ(E1) on both sides. /square
3 Measurable functions
To do integration theory, we of course need functions to integrate. You should
not expect that arbitrary functions can be integrated, but only the “measurable”
ones. The following definition is not difficult to motivate.
Definition 3.1. Let (X,A) and (Y,B) be measurable spaces. A map f:X→Y
ismeasurable if
for allB∈ B,the setf−1(B) = [f∈ B] is in A.
Example 3.1.A constant map is always measurable, for f−1(B) is either ∅or
X.
Theorem 3.1. The composition of two measurable functions is measurable.
Proof. Immediate from the definition. /square
8
Theorem 3.2. Let(X,A)and (Y,B)be measurable spaces, and suppose H
generates the sigma algebra B:/angbracketleftH/angbracketright=B. A function f:X→Yis measurable
if and only if for every V∈ H,f−1(V)is inA.
Proof. The “only if” part is just the definition of measurability. For the “if”
direction, define G={f−1(V) :V∈ H} , and also C={V∈ B:f−1(V)∈ /angbracketleftG/angbracketright} .
It is easily checked that Cis a sigma algebra on Y, and it contains H, and hence it
is actually equal to B. That is, for every V∈ B,f−1(V) is in/angbracketleftG/angbracketright ⊆ /angbracketleftA/angbracketright =A./square
Corollary 3.3. All continuous functions (between topological spaces) are mea-
surable.
A comment about infinities again. There is a natural topology on [ −∞,+∞]
and [0,∞] that make them look like closed intervals. Some denote [ −∞,+∞] by
R(“the extended real numbers”). However, for convenience, I will just denote it
as plain R. So keep in mind that when we prove our theorems, we have to make
sure that they work (or do not work) when infinite quantities are introduced.
Theorem 3.4. Let(X,A)be a measurable space. A map f:X→Ris mea-
surable if and only if [f >c ] =f−1((c,+∞])∈ A for allc∈R.
Proof. LetBbe the set of all open intervals ( a,b), along with {−∞},{+∞}. Let
Hbe the set of all intervals ( c,+∞]. (a,b,c are finite.) Evidently Hgenerates
B:
(a,b) = [−∞,b)∩(a,+∞],
[−∞,b) =∞/uniondisplay
n=1[−∞,b−1
n] =∞/uniondisplay
n=1R\(b−1
n,+∞].
{+∞}=∞/intersectiondisplay
n=1(n,+∞].
{−∞} =R\∞/uniondisplay
n=1(−n,+∞].
In turn, Bgenerates the Borel sigma algebra on R. Applying Theorem 3.2to
the generator Hgives the result. (As B ⊆ /angbracketleftH/angbracketright andH ⊆ /angbracketleftB/angbracketright together mean
/angbracketleftH/angbracketright=/angbracketleftB/angbracketright.) /square
Remark 3.5.We can replace ( c,+∞], in the statement of the theorem, by
[c,+∞], [−∞,c], etc. and there is no essential difference.
The “countable union with 1 /n” trick used in the proof is widely applicable.
It may be of interest to note that the Archimedean property of the real num-
bers is being used here — the same proof will not work with non-Archimedean
ordered fields.
9
Theorem 3.6. Letfnbe a sequence of measurable R-valued functions. Then
the functions
sup
nfn,inf
nfn,max
nfn,min
nfn,lim sup
nfn,lim inf
nfn
(the limits are pointwise) are all measurable.
Proof. Ifg(x) = supnfn(x), then [g>c ] =/uniontext
n[fn>c], and we apply Theorem
3.4. Similarly, if g(x) = inf nfn(x), then [g < c ] =/uniontext
n[fn< c]. The rest can
be expressed as in terms of supremums and infimums (over a countable set), so
they are measurable also. /square
Not surprisingly, we will need to do arithmetic in integration theory, so we
better know that
Theorem 3.7. Iff,g:X→Rare measurable, then so are f+g,fg, andf/g.
Proof. Consider the countable union
[f+g<c ] =/uniondisplay
r∈Q[f <c−r]∩[g<r ].
The set equality is justified as follows: Clearly f(x)<c−randg(x)<rtogether
implyf(x) +g(x)< c. Conversely, if we set g(x) =t, thenf(x)< c−t, and
we can increase tslightly to a rational number rsuch thatf(x)< c−r, and
g(x)<t<r . This shows that f+gis measurable (by Theorem 3.4).
Since [ −f <c ] = [f >−c], we see that −fis measurable.
Therefore the functions
f+(x) = max {+f(x),0}(positive part of f)
f−(x) = max {−f(x),0}(negative part of f)
are measurable (from Theorem 3.6and Example 3.1). Sincef=f+−f−,fis
measurable if f+andf−are measurable separately also.
Since
fg= (f+−f−)(g+−g−) =f+g+−f+g−−f−g++f−g−,
to prove that fgis measurable, it suffices to assume that fandgare both
non-negative. Then just as with the sum,
[fg<c ] =/uniondisplay
r∈Q[f <c/r ]∩[g<r ].
Finally, for 1 /g,
[1/g<c ] =/braceleftbigg
[1/c<g,cg> 0]∪[1/c>g,cg> 0], c/negationslash= 0
[g<0], c = 0./square
10
Remark 3.8.You probably have already noticed there may be difficulty in defin-
ing what the arithmetic operations mean when the operands are infinite (or
when dividing by zero). The usual way to deal with these problems is to simply
redefine the functions whenever they are infinite to be some fixed value. In
particular, if the function gis obtained by changing the original measurable
functionfon a measurable setAto be a constant c, we have:
[g∈B] =/parenleftbig
[g∈B]∩A/parenrightbig
∪/parenleftbig
[g∈B]∩Ac/parenrightbig
[g∈B]∩A=/braceleftbigg
A, c ∈B
∅, c /∈B
[g∈B]∩Ac= [f∈B]∩Ac,
so the resultant function gis also measurable. Very conveniently, any sets like
A= [f= +∞],[f= 0] are automatically measurable. Thus the gaps in the
previous proof with respect to infinite values can be repaired with this device.
As a final note, one intermediate result from the proof is quite useful and
should be formally recognized:
Theorem 3.9. AnR-valued function fis measurable if and only if f+andf−
are measurable. Moreover, if fis measurable, so is |f|=f++f−.
Remark 3.10.Of course the converse to the second statement is not true. You
may construct a counterexample to convince yourself of this fact.
4 Definition of the Lebesgue Integral
The idea behind Riemann integration is to try to measure the sums of area of
the rectangles “below a graph” of a function and then take some sort of limit.
The Lebesgue integral uses a similar approach: we perform integration on the
“simple” functions first:
Definition 4.1. A function is simple if its range is a finite set.
AnR-valued simple function ϕalways has a representation
ϕ=n/summationdisplay
k=1akχEk,
whereakare the distinct values of ϕ, andEk=ϕ−1({ak}). Conversely, any
expression of the above form, where akneed not be distinct, and Ekis not neces-
sarilyϕ−1({ak}), also defines a simple function. For the purposes of integration,
however, we will require that Ekbe measurable, and that they partition X. It
should be mentioned that χSis measurable if and only if Sis.
11
Definition 4.2. Let (X,µ) be a measure space. The Lebesgue integral, over
X, of aR+-valued measurable simple function ϕis defined as
/integraldisplay
Xϕdµ =/integraldisplay
Xn/summationdisplay
k=1akχEkdµ=n/summationdisplay
k=1akµ(Ek).
(We restrict ϕto being non-negative for now, to avoid mixed + ∞,−∞ on
the right-hand side.) Needless to say, the quantity on the right represents the
sum of the areas below the graph of ϕ.
It had better be the case that the value of the integral does not depend
on the representation of ϕ. Ifϕ=/summationtext
iaiχAi=/summationtext
jbjχBj, whereAiandBj
partitionX(soAi∩BjpartitionX), then
/summationdisplay
iaiµ(Ai) =/summationdisplay
j/summationdisplay
iaiµ(Ai∩Bj) =/summationdisplay
j/summationdisplay
ibjµ(Ai∩Bj) =/summationdisplay
jbjµ(Bj).
The second equality follows because the value of ϕisai=bjonAi∩Bj, so
ai=bjwheneverAi∩Bj/negationslash=∅. So fortunately the integral is well-defined.
Using the same algebraic manipulations just now, you can prove that if we
have two simple functions ϕ≤ψ, then/integraltext
Xϕdµ≤/integraltext
Xψdµ (monotonicity of the
integral).
Theorem 4.1. The Lebesgue integral (for non-negative simple functions) is
linear.
Proof. Clearly/integraltext
Xcϕdµ =c/integraltext
Xϕdµ. And ifϕ=/summationtext
iaiχAi,ψ=/summationtext
jbjχBj, we
have
/integraldisplay
Xϕdµ +/integraldisplay
Xψdµ =/summationdisplay
iaiµ(Ai) +/summationdisplay
jbjµ(Bj)
=/summationdisplay
i/summationdisplay
jaiµ(Ai∩Bj) +/summationdisplay
j/summationdisplay
ibjµ(Bj∩Ai)
=/summationdisplay
i/summationdisplay
j(ai+bj)µ(Ai∩Bj)
=/integraldisplay
X(ϕ+ψ)dµ. /square
Next, we integrate non-simple measurable functions like this:
Definition 4.3. Letf:X→[0,+∞] be measurable. Consider te set Sfof all
measurable simple functions 0 ≤ϕ≤f, and define the integral of foverXas
/integraldisplay
Xfdµ = sup
ϕ∈Sf/integraldisplay
Xϕdµ.
12
Intuitively, the simple functions in Sfare supposed to approximate fas
close as we like, and we find the integral of fby computing the integrals of
these approximations. But logically we need to know that these approximations
really do exist. This is the essence of the following theorem.
Theorem 4.2 (Approximation Theorem). Letf:X→[0,∞]be measur-
able. Then there exists a sequence of non-negative functions {ϕn} /arrownortheastf, meaning
ϕnare increasing pointwise and converging pointwise to f. Moreover, if fis
bounded, it becomes possible for the ϕnto converge to funiformly.
Proof. We prove the second statement first. Let Nbe any integer >supf, and
set
ϕn=N2n/summationdisplay
k=1k−1
2nχEn,k, E n,k=f−1/parenleftbigg/bracketleftbiggk−1
2n,k
2n/parenrightbigg/parenrightbigg
.
We have 0 ≤f−ϕn<2−nuniformly. The detailed verification is left to the
reader.
The construction for the first statement is very similar. Set
ϕn=n2n/summationdisplay
k=1k−1
2nχEn,k+χFn, F n= [f≥n].
I leave it to you to check that 0 ≤f(x)−ϕn(x)<2−nwhenevern>f (x), and
ϕn(x) =nwheneverf(x) =∞. /square
In case you were worrying about whether this new definition of the integral
agrees with the old one in the case of the non-negative simple functions, well, it
does. Use monotonicity to prove this.
Definition 4.4. Iffis not necessarily non-negative, we define
/integraldisplay
Xfdµ =/integraldisplay
Xf+dµ−/integraldisplay
Xf−dµ,
provided that the two integrals on the right are not both ∞.
Of course we will want to integrate over subsets of Xalso. This can be
accomplished in two ways. Let Abe a measurable subset of X. Either we
simply consider integrating over the measure space restricted to subsets of A,
or we define /integraldisplay
Afdµ =/integraldisplay
XfχAdµ.
Ifϕis non-negative simple, a simple working out of the two definitions of the
integral over Ashows that they are equivalent. To prove this for the case of
arbitrary measurable functions, we will need the tools of the next section.
Let us note some other basic properties of our integral:
13
Letf,gbe non-negative. Since
Scf=c·Sf={cϕ:ϕ∈Sf},0≤c<∞.
we have (we freely omit the “ dµ” and/or the integration limit “ X” when they
are implied by the context)/integraldisplay
cf=c/integraldisplay
f.
This rule about constant multiplication also holds for fandcnot necessarily
non-negative, as you can easily check, but proving linearity requires the tools of
the next section.
Moreover, if 0 ≤f≤g, thenSf⊆Sg, and therefore/integraltext
f≤/integraltext
g. In particular,
ifA⊆B, thenfχA≤fχB, so
/integraldisplay
Af≤/integraldisplay
Bf.
Unsurprisingly/integraltext
f≤/integraltext
galso holds if f,gare not necessarily non-negative, and
that is proven by considering the positive and negative parts of f,gseparately.
Then since −|f| ≤f≤ |f|, we also obtain
−/integraldisplay
|f| ≤/integraldisplay
f≤/integraldisplay
|f|,i.e./vextendsingle/vextendsingle/vextendsingle/integraldisplay
f/vextendsingle/vextendsingle/vextendsingle≤/integraldisplay
|f|.
(This last inequality is sometimes called the “generalized triangle inequality”,
as integrals can be viewed as an advanced form of summing.)
Next, we make one more definition related to integrals.
Definition 4.5. A (µ-)measurable set is said to have ( µ-)measurable zero if
µ(E) = 0.
Typical examples of a measure-zero set are the singleton points in Rn, and
lines and curves in Rn,n≥2. By countable additivity, any countable set in Rn
has measure zero also.
A particular property is said to hold almost everywhere if the set of points
for which the property fails to hold is a set of measure zero. For example, “a
function vanishes almost everywhere”.
Clearly, if you integrate anything on a set of measure zero, you get zero.
Assuming that linearity of the integral has been proved, we can demonstrate
the following intuitive result.
Theorem 4.3. A measurable function f:X→[0,∞]vanishes almost every-
where if and only if/integraltext
Xf= 0.
Proof. LetA= [f= 0], andµ(Ac) = 0. Then
/integraldisplay
Xf=/integraldisplay
Xf·(χA+χAc) =/integraldisplay
XfχA+/integraldisplay
XfχAc=/integraldisplay
Af+/integraldisplay
Acf= 0 + 0.
14
Conversely, if/integraltext
Xf= 0, consider [ f >0] =/uniontext
n[f >1
n]. We have
µ[f >1
n] =/integraldisplay
[f>1
n]1 =n/integraldisplay
[f>1
n]1
n≤n/integraldisplay
Xf= 0.
Henceµ[f >0] = 0. /square
5 Convergence theorems
The following theorems are another feature of the Lebesgue integral that make
it so much better than the Riemann definition.
Theorem 5.1 (Monotone Convergence Theorem). Let(X,µ)be a mea-
sure space. Let fnbe non-negative measurable functions increasing pointwise to
f. Then /integraldisplay
Xfdµ =/integraldisplay
X/parenleftBig
lim
n→∞fn/parenrightBig
dµ= lim
n→∞/integraldisplay
Xfndµ.
Proof.fis measurable because it is a limit of measurable functions. Since fnis
an increasing sequence of functions bounded by f, their integrals is an increasing
sequence of numbers bounded by/integraltext
Xf; thus the following limit exists:
lim
n→∞/integraldisplay
Xfndµ≤/integraldisplay
Xfdµ.
Next we show the inequality in the other direction.
Take any 0 <t< 1. Given a fixed ϕ∈Sf, letAn= [fn−tϕ≥0]. TheAn
are obviously increasing.
If for a particular x∈X, we have ϕ(x) = 0, then x∈An, for alln.
Otherwise, ϕ(x)>0, sof(x)≥ϕ(x)>tϕ (x), and there is going to be some n
for whichfn(x)≥tϕ(x), i.e.x∈An. HenceX=/uniontext
nAn.
For allµ-measurable sets E, define a new measure on Xby
ν(E) =/integraldisplay
Etϕdµ.
Then/integraldisplay
Xtϕdµ =ν(X) =ν/parenleftBig/uniondisplay
nAn/parenrightBig
= lim
n→∞ν(An) = lim
n→∞/integraldisplay
Antϕdµ
≤lim
n→∞/integraldisplay
Anfndµ, since onAnwe havetϕ≤fn
≤lim
n→∞/integraldisplay
Xfndµ.
t/integraldisplay
Xϕdµ≤lim
n→∞/integraldisplay
Xfndµ, and take limit t→1.
/integraldisplay
Xϕdµ≤lim
n→∞/integraldisplay
Xfndµ, and take sup over ϕ∈Sf. /square
15
Using this theorem, we are now in the position to prove linearity of the
Lebesgue integral for non-simple functions. Given any two non-negative mea-
surable functions f,g, by the approximation theorem (Theorem 4.2), we know
that are non-negative simple functions {ϕn} /arrownortheastf, and {ψn} /arrownortheastg. Then
{ϕn+ψn} /arrownortheastf+g, and so
/integraldisplay
f+g= lim
n→∞/integraldisplay
ϕn+ψn= lim
n→∞/integraldisplay
ϕn+/integraldisplay
ψn=/integraldisplay
f+/integraldisplay
g.
(The second equality follows because we already know the integral is linear
for simple functions. For the first and third equality we apply the Monotone
Convergence Theorem.)
And iff,gnot necessarily non-negative, then
/integraldisplay
f+g=/integraldisplay
(f+−f−) + (g+−g−)
=/integraldisplay
f++g+−(f−+g−)
=/integraldisplay
f++g+−/integraldisplay
f−+g−
=/integraldisplay
f++/integraldisplay
g+−(/integraldisplay
f−+/integraldisplay
g−) =/integraldisplay
f+/integraldisplay
g.
(Only at the fourth equality we apply what we had just proved for non-negative
functions. The third and last equality are just by the definition of the integral.)
Here is another application of the Monotone Convergence Theorem.
Theorem 5.2 (Beppo Levi). Letfn:X→[0,∞]be measurable. Then
/integraldisplay∞/summationdisplay
n=1fn=∞/summationdisplay
n=1/integraldisplay
fn.
Proof. LetgN=/summationtextN
n=1fn, andg=/summationtext∞
n=1fn. The Monotone Convergence
Theorem applies to gN, and:
/integraldisplay
g=/integraldisplay
lim
N→∞gN= lim
N→∞/integraldisplay
gN= lim
N→∞N/summationdisplay
n=1/integraldisplay
fn=∞/summationdisplay
n=1/integraldisplay
fn. /square
There are many more applications like this. We postpone those for now,
since you will probably be even more amazed by the next convergence theorem.
We first need a lemma.
Lemma 5.3 (Fatou’s Lemma). Letfn:X→[0,∞]be measurable. Then
/integraldisplay
lim inf
n→∞fn≤lim inf
n→∞/integraldisplay
fn.
16
Proof. Setgn= inf k≥nfk, so thatgn≤fn, and{gn} /arrownortheastlim inf nfn. Then
/integraldisplay
lim inf
nfn=/integraldisplay
lim
ngn= lim
n/integraldisplay
gn= lim inf
n/integraldisplay
gn≤lim inf
n/integraldisplay
fn./square
Remark 5.4.By adding and subtracting a constant, the hypotheses may be
weakened to allow functions that are bounded below by any fixed number, not
just non-negative functions. (This lower bound condition cannot be dropped.)
The same considerations apply to the Monotone Convergence Theorem.
The following definition is used to formulate a crucial hypothesis of the
theorem that is about to follow.
Definition 5.1. A function f:X→Ris called integrable if it is measurable
and/integraltext
X|f|<∞.
It is immediate that fis integrable if and only if f+andf−are both
integrable. It is also helpful to know, that/integraltext
|f|<∞must imply |f|<∞
almost everywhere.
Theorem 5.5 (Dominated Convergence Theorem). Let(X,µ)be a mea-
sure space. Let fn:X→Rbe a sequence of measurable functions converging
pointwise to f. Moreover, suppose that there is an integrable functiongsuch
that|fn| ≤g, for alln. Thenfnandfare also integrable, and
lim
n→∞/integraldisplay
X|fn−f|dµ= 0.
Proof. Obviouslyfnandfare integrable. Also, 2 g−|fn−f|is measurable and
non-negative. By Fatou’s lemma,
/integraldisplay
lim inf
n(2g− |fn−f|)≤lim inf
n/integraldisplay
(2g− |fn−f|).
Sincefnconverges to f, the left-hand quantity is just/integraltext
2g. The right-hand
quantity is:
lim inf
n/parenleftbigg/integraldisplay
2g−/integraldisplay
|fn−f|/parenrightbigg
=/integraldisplay
2g+ lim inf
n/parenleftbigg
−/integraldisplay
|fn−f|/parenrightbigg
=/integraldisplay
2g−lim sup
n/integraldisplay
|fn−f|.
Since/integraltext
2gis finite, it may be cancelled from both sides. Then we obtain
lim sup
n/integraldisplay
|fn−f| ≤0,i.e. lim
n→∞/integraldisplay
|fn−f|= 0. /square
Remark 5.6.It obviously suffices to only require that fnconverge to fpointwise
almost everywhere, or that |fn|is bounded above by galmost everywhere. (Of
course, iffnonly converges to falmost everywhere, then the theorem would
not automatically say that fis measurable.)
17
Remark 5.7.By the generalized triangle inequality, we conclude from the hy-
potheses that also
lim
n→∞/integraldisplay
Xfndµ=/integraldisplay
Xfdµ,
which is usually how this theorem is applied.
Remark 5.8.The theorem also holds for continuous limits of functions, not just
countable limits. That is, if we have a continuous sequence of functions, say ft,
0≤t<1, we can also say
lim
t→1/integraldisplay
X|ft−f|dµ= 0 ;
for given any sequence {an}convergent to 1, we can apply the theorem to fan.
Since this can be done for any sequence convergent to 1, the above limit is
established.
6 Some Results of Integration Theory
This section contains some nice applications proven using the convergence the-
orems from the last section.
Theorem 6.1 (Generalization of Beppo Levi). LetXbe a measure space,
andfn:X→Rbe measurable functions, with/integraltext/summationtext|fn|=/summationtext/integraltext
|fn|<∞. Then
∞/summationdisplay
n=1/integraldisplay
fn=/integraldisplay∞/summationdisplay
n=1fn.
Proof. LetgN=/summationtextN
n=1fn,g= lim supN→∞gN, andh=/summationtext∞
n=1|fn|. Then
|gN| ≤h=|h|. Since/integraltext
|h|<∞by hypothesis, we have |h|<∞almost
everywhere, so/summationtext∞
n=1fnis absolutely convergent almost everywhere. That is,
gNconverges pointwise to galmost everywhere.
By the Dominated Convergence Theorem, lim
N→∞/integraltext
gN=/integraltext
g, whence
∞/summationdisplay
n=1/integraldisplay
fn= lim
N→∞N/summationdisplay
n=1/integraldisplay
fn= lim
N→∞/integraldisplayN/summationdisplay
n=1fn=/integraldisplay∞/summationdisplay
n=1fn. /square
Example 6.1.Here’s a perhaps unexpected application. Suppose we have a
countable set of real numbers an,m,n,m∈N. Letµbe the counting measure on
N. Then/integraltext
m∈Nan,mdµ=/summationtext∞
m=1an,m. Moreover, Theorem 6.1says that we can
sum either along nfirst ormfirst and get the same results (/summationtext∞
n=1/summationtext∞
m=1an,m=/summationtext∞
m=1/summationtext∞
n=1an,m) if the double sum is absolutely convergent. Of course, this
fact can also be proven in an entirely elementary way.
Theorem 6.2. Letg:X→[0,∞]be measurable in the measure space (X,A,µ).
Let
ν(E) =/integraldisplay
Egdµ, E ∈ A.
18
Thenνis a measure on (X,A), and for any measurable function fonX,
/integraldisplay
Xfdν =/integraldisplay
Xfgdµ,
often written as dν=gdµ.
Proof. We proveνis a measure. ν(∅) = 0 is trivial. For countable additivity,
let{En}be measurable with union E, so thatχE=/summationtext∞
n=1χEn, and
ν(E) =/integraldisplay
Egdµ =/integraldisplay
XgχEdµ
=/integraldisplay
X∞/summationdisplay
n=1gχEndµ=∞/summationdisplay
n=1/integraldisplay
XgχEndµ=∞/summationdisplay
n=1ν(En).
Next, iff=χEfor someE∈ A, then
/integraldisplay
Xfdν =/integraldisplay
XχEdν=ν(E) =/integraldisplay
XχEgdµ =/integraldisplay
Xfgdµ.
By linearity, we see that/integraltext
fdν =/integraltext
fgdµ wheneverfis non-negative simple.
For general non-negative f, we use a sequence of simple approximations {ϕn} /arrownortheast
f, so{ϕng} /arrownortheastfg. Then by the monotone convergence,
/integraldisplay
Xfdν = lim
n→∞/integraldisplay
Xϕndν= lim
n→∞/integraldisplay
Xϕngdµ =/integraldisplay
Xlim
n→∞ϕngdµ =/integraldisplay
Xfgdµ.
Finally, for fnot necessarily non-negative, we apply the above to its positive
and negative parts, and use linearity. /square
The procedure of proving some fact about integrals by first reducing to the
case of simple functions and non-negative functions is used quite often. (It will
get quite monotonous if we had to detail the procedure every time we use it, so
we won’t anymore if the circumstances permit.)
Also, we should note that if fis only measurable but not integrable, then
the integrals of f+orf−might be infinite. If both are infinite, the integral of f
is not defined, although the equation of the theorem might still be interpreted
as saying that the left-hand and right-hand sides are undefined at the same
time. For this reason, and for the sake of the clarity of our exposition, we will
not bother to modify the hypotheses of the theorem to state that fmust be
integrable.
Problem cases like this also occur for some of the other theorems we present,
and there I will also not make too much of a fuss about these problems, trusting
that you understand what happens when certain integrals are undefined.
Theorem 6.3 (Change of variables). LetX,Y be measure spaces, and
g:X→Y,f:Y→Rbe measurable. Then
/integraldisplay
X(f◦g)dµ=/integraldisplay
Yfdν,
whereν(B) =µ(g−1(B))is a measure defined for all measurable B⊆Y.
19
Proof. First suppose f=χB. LetA=g−1(B)⊆X. Thenf◦g=χA, and we
have
/integraldisplay
Yfdν =/integraldisplay
YχBdν=ν(B) =µ(g−1(B)) =µ(A) =/integraldisplay
X(f◦g)dµ.
Since both sides of the equation are linear in f, the equation holds whenever f
is simple. Applying the “standard procedure” mentioned above, the equation is
then proved for all measurable f. /square
Remark 6.4.The change of variables theorem can also be applied “in reverse”.
Suppose we want to compute/integraltext
Yfdν, whereνis already given to us. Further
assume that gis bijective and its inverse is measurable. Then we can define
µ(A) =ν(g(A)), and it follows that/integraltext
Yfdν =/integraltext
X(f◦g)dµ.
Our theorem (especially when stated in the reverse form) is clearly related
to the usual “change of variables” theorem in calculus. If g:X→Yis a bijec-
tion between open subsets of Rn, and both it and its inverse are continuously
differentiable (i.e. g is a diffeomorphism ), andν=λis the Lebesgue measure
inRn, then (as we shall prove rigorously in Lemma 12.1),
µ(A) =λ(g(A)) =/integraldisplay
A|det Dg|dλ.
Appealing to Theorems 6.2and6.3, we obtain:
Theorem 6.5 (Differential change of variables in Rn).Letg:X→Ybe
a diffeomorphism of open sets in Rn. IfA⊆Xis measurable, and f:Y→R
is measurable, then
/integraldisplay
g(A)fdλ =/integraldisplay
A(f◦g)dµ=/integraldisplay
A(f◦g)· |det Dg|dλ.
The next two theorems are the Lebesgue versions of well-known results about
the Riemann integral.
Theorem 6.6 (First Fundamental Theorem of Calculus). LetI⊆Rbe
an interval, and f:I→Rbe integrable (with Lebesgue measure in R). Then
the function
F(x) =/integraldisplayx
af(t)dt
is continuous. Furthermore, if fis continuous at x, thenF/prime(x) =f(x).
Proof. To prove continuity, we compute:
F(x+h)−F(x) =/integraldisplayx+h
xf(t)dt=/integraldisplay
If(t)·χ[x,x+h](t)dt.
20
(Naturally, when h<0,χ[x,x+h]should be interpreted as −χ[x+h,x]here.) Since
|f·χ[x,x+h]| ≤ |f|, by the Dominated Convergence Theorem (along with remark
5.8),
lim
h→0F(x+h)−F(x) = lim
h→0/integraldisplay
If(t)·χ[x,x+h](t)dt
=/integraldisplay
Ilim
h→0f(t)·χ[x,x+h](t)dt
=/integraldisplay
If(t)·χ{x}(t)dt= 0.
The proof of differentiability is the same as for the Riemann integral:
/vextendsingle/vextendsingle/vextendsingle/vextendsingleF(x+h)−F(x)
h−f(x)/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraltextx+h
x(f(t)−f(x))dt
h/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle
≤/integraltext
[x,x+h]|f(t)−f(x)|dt
|h|
≤supt∈[x,x+h]|f(t)−f(x)| · |h|
|h|,
which goes to zero as hdoes. /square
The First Fundamental Theorem was easy, but the Second Fundamental
Theorem (which states that/integraltextb
af/prime=f(b)−f(a)) is not entirely trivial. The
difficulty is that we should not assume as hypotheses that f/primeis continuous, or
even that it is Lebesgue-integrable. It turns out that a theorem without such
strong hypotheses is possible; we will not reproduce its proof here, but just
settle for a weaker version:
Theorem 6.7 (Second Fundamental Theorem of Calculus). Suppose
f: [a,b]→Ris measurable and bounded above and below. If f=g/primefor some
g, then/integraldisplayb
af(x)dx=g(b)−g(a).
Proof. We first note that/vextendsingle/vextendsingle/vextendsingleg(x+h)−g(x)
h/vextendsingle/vextendsingle/vextendsinglecan be bounded by a constant using the
Mean Value Theorem, and a constant is obviously integrable on a finite interval.
21
Then
/integraldisplayb
af(x)dx=/integraldisplayb
alim
h→0g(x+h)−g(x)
hdx
= lim
h→0/integraldisplayb
ag(x+h)−g(x)
hdx
= lim
h→01
h/bracketleftBig/integraldisplayb+h
a+hg(x)dx−/integraldisplayb
ag(x)dx/bracketrightBig
= lim
h→01
h/bracketleftBig/integraldisplayb+h
bg(x)dx−/integraldisplaya+h
ag(x)dx/bracketrightBig
=g(b)−g(a).
(The last equality follows from the First Fundamental Theorem and that gmust
be continuous at aandbif it is differentiable there.) /square
Remark 6.8.Ifg/prime(x) exists, then it can also be computed as the countable limit
limn→∞n(g(x+ 1/n)−g(x)), thus showing that g/primeis measurable. Thus we can
drop the hypothesis that fis measurable in Theorem 6.7.
Remark 6.9.You might have noticed that I cheated a bit in the proof, in as-
suming the integral of a function is invariant under horizontal translations. But
of course, this can be proven readily using the fact that Lebesgue measure is
translation-invariant.
The following theorems are often not found in calculus texts even though
they are quite important for applications.
Theorem 6.10 (Continous dependence on integral parameter). Let
(X,µ)be a measure space, Tbe any metric space (e.g. Rn), andf:X×T→R,
withf(·,t)being measurable for each t∈T. Consider the function
F(t) =/integraldisplay
x∈Xf(x,t).
Then we have Fcontinuous at t0∈Tif the following conditions are met:
1.For eachx∈X,f(x,·)is continuous at t0∈I.
2.There is an integrable function gsuch that |f(x,t)| ≤g(x)for allt∈T.
Proof.
lim
t→t0/integraldisplay
x∈Xf(x,t) =/integraldisplay
x∈Xlim
t→t0f(x,t) =/integraldisplay
x∈Xf(x,t0). /square
Theorem 6.11 (Differentiation under the integral sign). Using the same
notation as Theorem 6.10, withTbeing an open real interval, we have
F/prime(t) =d
dt/integraldisplay
x∈Xf(x,t) =/integraldisplay
x∈X∂
∂tf(x,t)
if the following conditions are satisfied:
22
1.For eachx∈X,∂
∂tf(x,t)exists.
2.There is an integrable function gsuch that/vextendsingle/vextendsingle∂
∂tf(x,t)/vextendsingle/vextendsingle≤g(x)for allt∈T.
Proof. This theorem is often proven by using iterated integrals and switching
the order of integration, but that method is theoretically troublesome because it
requires more stringent hypotheses. It is easier, and better, to prove it directly
from the definition of the derivative.
The straightforward computation yields:
lim
h→0F(t+h)−F(t)
h= lim
h→0/integraldisplay
x∈Xf(x,t+h)−f(x,t)
h
=/integraldisplay
x∈Xlim
h→0f(x,t+h)−f(x,t)
h=/integraldisplay
x∈X∂
∂tf(x,t)
(noting that/vextendsingle/vextendsingle/vextendsinglef(x,t+h)−f(x,t)
h/vextendsingle/vextendsingle/vextendsingleis bounded by g(x)). /square
Remark 6.12.It is easy to see that we may generalize Theorem 6.11 toTbeing
any open set in Rn, taking partial derivatives. I won’t write it out in full because
the notation is somewhat complicated.
Example 6.2.Check that the function Γ( x) =/integraltext∞
0e−ttx−1dt,x> 0 is continu-
ous, and differentiable with the obvious formula for the derivative.
7 Lpspaces
The contents in this section do not have applications in this article, but they
are so well known that it would not do justice to omit them.
Definition 7.1. LetXbe a measure space, and let p∈[1,∞). The space Lp
consists of all measurable functions f:X→Rsuch that
/integraldisplay
|f|p<∞.
Definition 7.2. For eachf∈Lp, define
/bardblf/bardblp=/parenleftBig/integraldisplay
|f|p/parenrightBig1/p
.
(Iff /∈Lp, this quantity is of course defined as ∞.)
As suggested by the notation, /bardbl·/bardblpis a real norm on the vector space Lp,
provided that we declare two functions to be equivalent if they differ only on
a set of measure zero (so that /bardblf/bardblp= 0 if and only if f= 0 as equivalence
classes). Only the verification of the triangle inequality presents any difficulties
— this will be solved by the theorems below.
Definition 7.3. Two numbers p,q∈(1,∞) are called conjugate exponents
when1
p+1
q= 1.
23
Theorem 7.1 (H¨ older’s inequality). ForR-valued measurable functions f
andg,/vextendsingle/vextendsingle/vextendsingle/integraldisplay
fg/vextendsingle/vextendsingle/vextendsingle≤/integraldisplay
|f||g| ≤ /bardblf/bardblp/bardblg/bardblq.
Proof. The first inequality is trivial. For the second inequality, since it only
involves absolute values, for the rest of the proof we may assume that f,gare
non-negative.
If/bardblf/bardblp= 0, then |f|p= 0 almost everywhere, and so f= 0 andfg= 0
almost everywhere too. Thus the inequality is valid in this case. (Similarly
when /bardblg/bardblq= 0.)
If/bardblf/bardblpor/bardblg/bardblqis infinite, the inequality is trivial.
So we now assume these two quantities are both finite and non-zero. Define
F=f//bardblf/bardblp,G=g//bardblg/bardblq, so that /bardblF/bardblp=/bardblG/bardblq= 1. We must then show that/integraltext
FG≤1.
To do this, we employ the fact that log is concave:
1
plogs+1
qlogt≤log/parenleftbiggs
p+t
q/parenrightbigg
,0≤s,t≤ ∞
or,
s1
pt1
q≤s
p+t
q.
Substitutes=Fp,t=Gq, and integrate both sides:
/integraldisplay
FG≤1
p/integraldisplay
Fp+1
q/integraldisplay
Gq=1
p/bardblF/bardblp
p+1
q/bardblG/bardblq
q=1
p+1
q= 1. /square
You may have seen a special case of this theorem, for p=q= 2, as the
Cauchy-Schwarz inequality.
Theorem 7.2 (Minkowski’s inequality). ForR-valued measurable functions
fandg,
/bardblf+g/bardblp≤ /bardblf/bardblp+/bardblg/bardblp.
Proof. The inequality is trivial when p= 1 or when /bardblf+g/bardblp= 0. Also, since
/bardblf+g/bardblp≤ /bardbl|f|+|g|/bardblp, it again suffices to consider only the case when f,gare
non-negative.
Next, we employ the convexity of t/mapsto→tp,p>1,
/parenleftbiggs+t
2/parenrightbiggp
≤sp+tp
2,0≤s,t≤ ∞.
When we substitute s=f,t=g, we get (f+g)p≤2p−1(fp+gp). This
inequality shows that if /bardblf+g/bardblpis infinite, then one of /bardblf/bardblpor/bardblg/bardblpis also
infinite, so Minkowski’s inequality holds true in that case.
We may now assume /bardblf+g/bardblpis finite. We write:
/integraldisplay
(f+g)p=/integraldisplay
f(f+g)p−1+/integraldisplay
g(f+g)p−1.
24
By H¨ older’s inequality, and noting that ( p−1)q=pfor conjugate exponents,
/integraldisplay
f(f+g)p−1≤ /bardblf/bardblp/vextenddouble/vextenddouble(f+g)p−1/vextenddouble/vextenddouble
q=/bardblf/bardblp/parenleftBig/integraldisplay
|f+g|p/parenrightBig1/q
=/bardblf/bardblp/bardblf+g/bardblp/q
p.
A similar inequality holds for/integraltext
g(f+g)p−1. Putting these together:
/bardblf+g/bardblp
p≤(/bardblf/bardblp+/bardblg/bardblp)/bardblf+g/bardblp/q
p.
Dividing by /bardblf+g/bardblp/q
pyields the desired result. /square
Definition 7.4. A measure space ( X,µ) has finite measure ifµ(X) is finite.
Theorem 7.3. Let(X,µ)have finite measure. Then whenever 1≤r<p< ∞,
Lp⊆Lr. Moreover, the inclusion map from LptoLris continuous.
Proof. Iff∈Lp, apply the H¨ older inequality with conjugate exponentsp
rand
s=p
p−r:
/bardblf/bardblr
r=/integraldisplay
|f|r≤/parenleftBig/integraldisplay
|f|r·p
r/parenrightBigr/p/parenleftBig/integraldisplay
1s/parenrightBig1/s
=/bardblf/bardblr
pµ(X)1/s,
and so
/bardblf/bardblr≤ /bardblf/bardblpµ(X)1/rs=/bardblf/bardblpµ(X)1
r−1
p<∞.
To show continuity of the inclusion map, replace fwithf−gabove where
/bardblf−g/bardblp<ε. /square
Example 7.1./integraltext1
0x−1
2dx= 2<∞, so automatically/integraltext1
0x−1
4dx < ∞. On the
other hand, the condition that µ(X)<∞is indeed necessary:/integraltext∞
1x−2dx<∞,
but/integraltext∞
1x−1dx=∞.
Theorem 7.4. Letfn:X→Rbe measurable functions converging (almost
everywhere) pointwise to f, and |fn| ≤gfor someg∈Lp. Thenf,fn∈Lp,
andfnconverges to fin the Lpnorm, meaning:
lim
n→∞/parenleftBig/integraldisplay
|fn−f|p/parenrightBig1/p
= lim
n→∞/bardblfn−f/bardblp= 0.
Proof. |fn−f|pconverges to 0 and |fn−f|p≤(2g)p∈L1. Apply the Dominated
Convergence Theorem on these functions. /square
One wonders whether the converse is true: if fnconverges to fin the Lp
norm, do the functions fnthemselves converge pointwise to f? The answer is
no (the counterexamples are not difficult), but we do have the following.
Theorem 7.5. Let{fn}be a Cauchy sequence in Lp. Then it has a subsequence
converging pointwise almost everywhere.
25
Corollary 7.6. Iflim
n→∞/bardblfn−f/bardblp= 0then there is a subsequence {fn(k)}con-
verging tofpointwise almost everywhere.
Corollary 7.6is used in the following result.
Theorem 7.7. Lpis a complete metric space. (This means every Cauchy
sequence in Lpconverges.)
The proofs of these theorems are collected in the next section.
It is also possible to define “ L∞”:
Definition 7.5. LetXbe a measure space, and f:X→Rbe measurable. A
numberM∈[0,∞] is an almost-everywhere upper bound for |f|if|f| ≤M
almost everywhere. The infimum of all almost-everywhere upper bounds for |f|
is denoted by /bardblf/bardbl∞.
Definition 7.6. L∞is the set of all measurable functions fwith/bardblf/bardbl∞<∞.
Its norm is given by /bardblf/bardbl∞.
The use of the subscript “ ∞” is explained by the following theorem, whose
proof is left to the reader:
Theorem 7.8. Ifµ(X)<∞, then lim
p→∞/bardblf/bardblp=/bardblf/bardbl∞.
Remark 7.9.Observe that there is a similar thing for vectors /vector a= (a1,...,a n)∈
Rn:
lim
p→∞/bardbl/vector a/bardblp= lim
p→∞/parenleftbig
|a1|p+···+|an|p/parenrightbig1/p= max( |a1|,...,|an|) =/bardbl/vector a/bardbl∞.
This remark may serve as a hint.
If we (naturally) define the conjugate exponent of p= 1 to be q=∞,
the H¨ older inequality remains valid, since |fg| ≤ |f|/bardblg/bardbl∞almost everywhere.
Integrating both sides gives /bardblfg/bardbl1≤ /bardblf/bardbl1/bardblg/bardbl∞.
Before closing, we mention two more results. From Theorem 7.4, it is clear
that
Theorem 7.10. Letf:X→R∈Lp(X),1≤p <∞. Then for any ε >0,
there exist simple functions ϕ:X→Rsuch that
/bardblϕ−f/bardblp=/parenleftBig/integraldisplay
Rn|ϕ−f|pdλ/parenrightBig1/p
<ε.
(This is also true for p=∞.) This may be summarized by saying that the
set of all simple functions is dense in Lp(in the topological sense).
WhenX=Rnwith Lebesgue measure λ, we have another result. The set of
infinitely differentiable functions on Rnwith compact support4, denoted C∞
0,
4The support of a function ψ:X→Ris the closure of the set {x∈X:ψ(x)/negationslash= 0}. “ψhas
compact support” means that the support of ψis compact (when X=Rn, same as closed
and bounded).
26
is also dense in Lp(Rn,λ), 1≤p<∞. (This fact will be fully proven in Section
14.)
These two facts are typically used as tools to prove other theorems about
functionsf∈Lp. One first proves that a certain theorem holds for all simple
ϕ(orϕ∈C∞
0), and then prove that the same result holds for arbitrary f∈Lp
by approximating such fbyϕ. An example of this procedure follows.
Theorem 7.11 (Riemann-Lebesgue Lemma). For allf∈L1(R),
lim
|ω|→∞/integraldisplay
Rf(x) sin(ωx)dx= 0.
Proof. For convenience, assume ω>0.
Suppose first that f=ψ∈C∞
0. Sinceψhas compact support, integrating
ψ(x) sin(ωx) over Ris the same as integrating over some compact interval [ a,b]
containing the support of ψ. And since every function involved is infinitely
differentiable, we may use integration by parts ( u=ψ(x),dv= sin(ωx)dx):
/integraldisplayb
aψ(x) sin(ωx)dx=−ψ(x)cos(ωx)
ω/vextendsingle/vextendsingle/vextendsingle/vextendsingleb
a+/integraldisplayb
acos(ωx)
ωψ/prime(x)dx.
As|ψ|and|ψ/prime|are continuous on [ a,b], they are bounded by constants Mand
M/primerespectively. We have,
/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplayb
aψ(x) sin(ωx)dx/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤1
ω|ψ(b) cos(ωb)−ψ(a) cos(ωa)|+1
ω/integraldisplayb
a|cos(ωx)ψ/prime(x)|dx
≤2M
ω+(b−a)M/prime
ω→0, ω→ ∞.
Thus we have proven the result when f=ψ∈C∞
0. Now suppose fis
arbitrary. By the denseness of C∞
0, for every ε >0 we can find ψ∈C∞
0such
that/bardblf−ψ/bardbl1<ε. Then
/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplay
Rf(x) sin(ωx)dx−/integraldisplay
Rψ(x) sin(ωx)dx/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤/integraldisplay
R/vextendsingle/vextendsingle/vextendsingle/parenleftbig
f(x)−ψ(x)/parenrightbig
sin(ωx)/vextendsingle/vextendsingle/vextendsingledx
≤/integraldisplay
R|f(x)−ψ(x)|dx
<ε,
or, /vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplay
Rf(x) sin(ωx)dx/vextendsingle/vextendsingle/vextendsingle/vextendsingle<ε+/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplay
Rψ(x) sin(ωx)dx/vextendsingle/vextendsingle/vextendsingle/vextendsingle.
We take lim supω→∞of both sides:
lim sup
ω→∞/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplay
Rf(x) sin(ωx)dx/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤ε+ 0.
Butε>0 is arbitrary, so we must have:
lim
ω→∞/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplay
Rf(x) sin(ωx)dx/vextendsingle/vextendsingle/vextendsingle/vextendsingle= lim sup
ω→∞/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplay
Rf(x) sin(ωx)dx/vextendsingle/vextendsingle/vextendsingle/vextendsingle= 0. /square
27
8 Construction of Lebesgue Measure
We now come to actually construct Lebesgue measure, as promised. The idea
is to extend an existing measure µwhich has been only partially defined, to an
“outer measure” µ∗. The extension is remarkably simple and intuitive:
µ∗(E) = inf
A1,A2,...∈A
E⊆S
nAn/summationdisplay
nµ(An),for allE⊆X.
(One interesting point to note: the definition of µ∗“works” with pretty much
any non-negative function µ, provided that µis defined on some set A ⊆ 2X
with∅,X∈ A, andµ(∅) = 0. In fact, for the first few proofs, these are the
only formal properties of µthat we need. Of course later we will need stronger
conditions on µ, such as additivity.)
Lemma 8.1. µ∗has the following properties:
1.µ∗(∅) = 0 .
2.It is monotone: µ∗(E)≤µ∗(F)whenE⊆F⊆X.
3.µ∗(A)≤µ(A)for allA∈ A.
4.µ∗is countably subadditive: if E1,E2,...⊆X, then
µ∗(/uniontext
nEn)≤/summationtext
nµ∗(En).
Proof. The first three properties are obvious. For the fourth, first observe
that if any of the µ∗(En) is infinite, there is nothing to prove. Otherwise,
letε >0. For each En, by the definition of µ∗, there are sets {An,m}m∈ A
coveringEn, with/summationtext
mµ(An,m)≤µ∗(En)+ε/2n. All of the An,mtogether cover/uniontext
nEn, and/summationtext
n,mµ∗(An,m)≤/summationtext
nµ∗(En) +ε. Sinceεwas arbitrary, we have/summationtext
n,mµ∗(An,m)≤/summationtext
nµ∗(En). /square
There is of course no guarantee that µ∗satisfies all the properties of a proper
measure on 2X. (It often does not.) Instead we will claim that µ∗is a proper
measure on the following subcollection of 2X:
M={B∈2X|µ∗(B∩E) +µ∗(Bc∩E) =µ∗(E) for allE⊆X}
={B∈2X|µ∗(B∩E) +µ∗(Bc∩E)≤µ∗(E) for allE⊆X}.
(The two subcollections are the same since by subadditivity we always have
µ∗(B∩E) +µ∗(Bc∩E)≥µ∗(E).)
Lemma 8.2. Mis a sigma algebra.
Proof. It is immediate from the definition that Mis closed under taking com-
plements, and that ∅ ∈ M . We first show Mis closed under finite intersection
28
(and hence under finite union). Let B,C∈ M .
µ∗/parenleftbig
(B∩C)∩E/parenrightbig
+µ∗/parenleftbig
(B∩C)c∩E/parenrightbig
=µ∗(B∩C∩E) +µ∗/parenleftbig
(Bc∩C∩E)∪(B∩Cc∩E)∪(Bc∩Cc∩E)/parenrightbig
≤µ∗(B∩C∩E) +µ∗(Bc∩C∩E) +µ∗(B∩Cc∩E) +µ∗(Bc∩Cc∩E)
=µ∗(C∩E) +µ∗(Cc∩E),by definition of B∈ M
=µ∗(E),by definition of C∈ M.
ThusB∩C∈ M .
We now have to show that if B1,B2,...∈ M ,/uniontext
nBn∈ M . We may assume
thatBnare disjoint, for otherwise we just consider B/prime
n=Bn\(B1∪···∪Bn−1),
which are in Mby the previous paragraph.
We shall need to know that, for all N≥1, and allE⊆X,
µ∗/parenleftbiggN/uniondisplay
n=1Bn∩E/parenrightbigg
=N/summationdisplay
n=1µ∗(Bn∩E).
The proof will be by induction on N. The statement is trivial for N= 1. For
the induction step, let DN=/uniontextN
n=1Bnwhich are increasing and are all in M.
Then
µ∗(DN+1∩E) =µ∗/parenleftbig
DN∩(DN+1∩E)/parenrightbig
+µ∗/parenleftbig
Dc
N∩(DN+1∩E)/parenrightbig
=µ∗(DN∩E) +µ∗(BN+1∩E)
=N+1/summationdisplay
n=1µ∗(Bn∩E) (induction hypothesis).
Using the fact just proven, we now have:
µ∗(E) =µ∗(DN∩E) +µ∗(Dc
N∩E)
=N/summationdisplay
n=1µ∗(Bn∩E) +µ∗(Dc
N∩E)
≥N/summationdisplay
n=1µ∗(Bn∩E) +µ∗/parenleftbigg/parenleftBig∞/uniondisplay
n=1Bn/parenrightBigc
∩E/parenrightbigg
(monotonicity).
TakingN→ ∞ , we obtain:
µ∗(E)≥∞/summationdisplay
n=1µ∗(Bn∩E) +µ∗/parenleftbigg/parenleftBig∞/uniondisplay
n=1Bn/parenrightBigc
∩E/parenrightbigg
≥µ∗/parenleftbigg/parenleftBig∞/uniondisplay
n=1Bn/parenrightBig
∩E/parenrightbigg
+µ∗/parenleftbigg/parenleftBig∞/uniondisplay
n=1Bn/parenrightBigc
∩E/parenrightbigg
(subadditivity).
But this shows/uniontext∞
n=1Bn∈ M . /square
29
Lemma 8.3. IfB1,B2,...∈ M are disjoint, then µ∗(/uniontext
nBn) =/summationtext
nµ∗(Bn).
Proof. The finite case µ∗(/uniontextN
n=1Bn) =/summationtextN
n=1µ∗(Bn) was proven in the previous
lemma (set E=X). By monotonicity,/summationtextN
n=1µ∗(Bn)≤µ∗(/uniontext∞
n=1Bn). Taking
N→ ∞ gives/summationtext∞
n=1µ∗(Bn)≤µ∗(/uniontext∞
n=1Bn). Inequality in the other direction
is implied by subadditivity of µ∗. /square
We now know µ∗satisfies all the properties of a measure on M. In order for
µ∗to be a sane extension of µtoM, we need to impose some conditions on µ.
The ones that work from experience are:
1.Ashould be an algebra , meaning that it is non-empty, and closed under
intersection and finite union (and intersection).
2.IfA1,...,A n∈ A are disjoint, then µ(/uniontext
iAi) =/summationtext
iµ(Ai). It follows that
µis monotone and finitely subadditive.
3.Also, ifA1,A2,...∈ A are disjoint, and/uniontext
iAihappens to be in A, the
previous equation must also hold. (This is equivalent to requiring that
µ(/uniontext
iAi)≤/summationtext
iµ(Ai).)
Then we have the following important result. (By the way, the name of
“Carath´ eodory Extension Process” is often used to refer to the constructions in
this section.)
Theorem 8.4. A ⊆ M , andµ∗(A) =µ(A)for allA∈ A.
Thusµ∗is a measure extending of µonto the sigma algebra Mcontaining
the algebra A. (Mmust also then contain the sigma algebra Bgenerated by A,
although Mmay be larger than B.)
Proof. FixA∈ A. For any E⊆Xandε > 0, by definition we can find
A1,A2,...∈ A withE⊆/uniontext
nAnand/summationtext
nµ(An)≤µ∗(E) +ε. Then we have
µ∗(A∩E) +µ∗(Ac∩E)≤µ∗/parenleftBig
A∩/uniondisplay
nAn/parenrightBig
+µ∗/parenleftBig
Ac∩/uniondisplay
nAn/parenrightBig
≤/summationdisplay
nµ∗(A∩An) +/summationdisplay
nµ∗(Ac∩An)
≤/summationdisplay
nµ(A∩An) +/summationdisplay
nµ(Ac∩An)
=/summationdisplay
nµ(An) (finite additivity of µ)
≤µ∗(E) +ε.
εbeing arbitrary, µ∗(A∩E) +µ∗(Ac∩E)≤µ∗(E), showing that A∈ M .
30
Now we show µ∗(A) =µ(A). Noteµ(A)≥µ∗(A) is always true. Con-
sider anyA1,A2,...∈ A withA⊆/uniontext
nAn. By countable subadditivity and
monotonicity of µ,
µ(A) =µ/parenleftBig/uniondisplay
nA∩An/parenrightBig
≤/summationdisplay
nµ(A∩An)≤/summationdisplay
nµ(An).
This implies µ(A)≤µ∗(A), directly from the definition of µ∗. /square
Our final results for this section concern the uniqueness of this extension.
Theorem 8.5. Assumeµ(X)<∞. LetBbe the sigma algebra generated from
the algebra A. Ifνis another measure on B, which agrees with µ∗onA, then
µ∗andνagree on Bas well.
Proof. LetB∈ B.
µ∗(B) = inf
A1,A2,...∈A
B⊆S
nAn/summationdisplay
nµ(An) = inf/summationdisplay
nν(An)≥infν/parenleftBig/uniondisplay
nAn/parenrightBig
≥infν(B).
Soµ∗(B)≥ν(B). Similarly, we have µ∗(Bc)≥ν(Bc), so thatµ∗(X)−µ∗(B)≥
ν(X)−ν(B) =µ∗(X)−ν(B), orµ∗(B)≤ν(B). (In fact, this actually proves
µ∗andνagree on Malso, ifνis defined on M.) /square
But the hypothesis that µ(X)<∞is clearly too restrictive. The fix is easy:
Definition 8.1. A measure space ( X,B,µ) issigma-finite , if there are mea-
surable sets X1,X2,...⊆X, such that/uniontext
nXn=Xandµ(Xn)<∞for all
n.
(Clearly, we may as well assume that the Xnare increasing in this definition.)
Theorem 8.6. Theorem 8.5holds also in the case that (X,B,µ)is sigma-finite.
Proof. Let{Xn} /arrownortheastX,µ(Xn)<∞as in the definition of sigma-finiteness. (Of
course we also need to assume that the Xncan be chosen from A.) For each
B∈ B, Theorem 8.5says thatµ∗(B∩Xn) =ν(B∩Xn). Taking limits as
n→ ∞ givesµ∗(B) =ν(B). /square
The restriction that Xbe sigma-finite is not too severe, since the usual spaces
such as Rnaresigma-finite. Sigma-finiteness also comes back in the theorems
of Section 11.
Finally, the corollary below is just Theorem 8.6restated without reference
to the outer measure:
Corollary 8.7 (Uniqueness of measures). LetBbe the sigma algebra gen-
erated by the algebra A. Then if two measures µandνagree on A, andXis
sigma-finite (under either µorν), thenµ=νonB.
31
9 Lebesgue Measure in Rn
In this section we rigorously construct the n-dimensional volume measure on Rn.
As hinted before, the idea is to define the measure for rectangles and then use
the extension process in Section 8. Unfortunately, there is some grunt work to
do in order to verify that the hypotheses of those theorems are indeed satisfied.
Our setting will be the collection Rof rectangles I1× ··· ×IninRn, where
Ikis any open, half-open or closed, bounded or unbounded, interval in R. The
following definition is merely an abstracted version of the formal facts we need
about these rectangles.
Definition 9.1. LetXbe any set. A semi-algebra is any R ⊆ 2Xwith the
following properties:
1.The empty set is in R.
2.The intersection of any two sets in Ris also in R.
3.For any set in R, its complement is expressible as a finite disjoint union
of other elements of R.
It is easy to see, although tiresome to write down formally, that the collection
of all rectangles is indeed a semi-algebra. But immediately from this, we can
automatically construct an algebra A:
Theorem 9.1. The set Aof all finite disjoint unions of elements of a semi-
algebra Ris an algebra on X.
Proof. We check the properties for an algebra:
1.The empty set is trivially in A.
2.IfA=/unionmultitext
iRi, andB=/unionmultitext
jSj, whereRiandSidenote a finite number of
sets chosen from R, thenA∩B=/unionmultitext
iRi∩/unionmultitext
jSj=/unionmultitext
i,jRi∩Sj∈ A.
3.IfA=/unionmultitext
iRi, thenAc=/intersectiontext
iRc
i=/intersectiontext
i/unionmultitext
jSi,jfor someSi,j∈ R. The finite
intersection belongs to Aby the previous step. So Ac∈ A.
4.Finally, given Ai=/unionmultitext
jRi,j, for a finite number of i, letD0=∅, andDi=
Di−1/unionmulti(Ai\Di−1)∈ A. Then/uniontext
iAi=/uniontext
iDi=/unionmultitext
i(Di\Di−1)∈ A./square
By the way, the sigma algebra generated by Awill contain the Borel sigma
algebra: every open set U∈Rnobviously can be written as a union of open
rectangles, and in fact we can use a countable union of open rectangles. For,
given an arbitrary collection of rectangles covering U⊆Rn, there always exists
acountable subcover5.
The volume of A=I1×···×Inis naturally defined as λ(A) =λ(I1)·····λ(In),
λ(Ik) being the length of the interval Ik, with the usual rules about multiplying
zeroes and infinities together in force.
5If you are not aware of this theorem, you are invited to prove it yourself.
32
Now let’s suppose that the rectangle Ahas been partitioned into a disjoint
smaller rectangles. Then the sum of the volumes of the smaller rectangles, as we
have defined it, should equal the volume of A. This is true, of course, although
it is again tedious to write down formally. Essentially, one draws a rectangular
grid onAusing the boundaries of the smaller rectangles, and show that the sum
of the volumes of each cell in the grid equals the volume of A, by applying the
distributive property of multiplication over addition.
In the even more general case, suppose A∈ Ais a disjoint union of rectangles,
butAis not necessarily in the semi-algebra of rectangles. The volume of Ais
defined as the sum of the volumes of the component rectangles. This is obvious
— we mention it only to note that, although Amay certainly have different
decompositions ( A=/uniontext
iRi=/uniontext
jSj), the volume sum is always the same.
To see this, simply take the common refinement Ri∩Sj. Then/summationtext
iλ(Ri) =/summationtext
i/summationtext
jλ(Ri∩Sj) by applying the result of the previous paragraph on each
rectangleRi. But/summationtext
jλ(Rj) equals this double sum also.
It follows easily then, that λis finitely additive. Thus there is only one
final thing left to show: if the disjoint union of A1,A2,...∈ A isC∈ A,
thenλ(C) =/summationtext∞
i=1λ(Ai). From monotonicity and taking limits we always have
λ(C)≥/summationtext∞
i=1λ(Ai).
We showλ(C)≤/summationtext∞
i=1λ(Ai). Observe that this also ought to be true if C
is contained in, but not necessarily equal to, the union of the Ai, andAineed
not be disjoint at all. Henceforth these are our new hypotheses.
Suppose first that Chappens to be compact, and Aiare all open. In other
words, {Ai}form an open cover of the compact set C. So there is a finite
subcoverA1,...,A n. By finite subadditivity, we have λ(C)≤/summationtextn
i=1λ(Ai)≤/summationtext∞
i=1λ(Ai).
Now continue to assume that Cis compact, but Aiare not open. But
it is easy to make the Aiopen and still cover C, by slightly expanding each
Ai. In particular, stipulate that the volume of each new Aigrows by at most
ε/2i. (We can assume the Aiare plain rectangles, rather than disjoint unions of
them, and that they are bounded, since Cis bounded.) Then we have λ(C)≤/summationtext∞
i=1λ(Ai) +ε, andε>0 is arbitrary.
All that remains is the case that Cis not compact. If Cis bounded, so is its
closure, and hence by the Heine-Borel theorem, Cis compact. Similarly take
the closure of the Ai. But taking closures of elements of Adoes not change
their volumes.
Finally consider Cunbounded. But C∩[−N,N ]nis bounded for each N,
so from the previous case we have λ(C∩[−N,N ]n)≤/summationtext∞
i=1λ(Ai). It is easily
checked that the limit as N→ ∞ of the left side is exactly λ(C).
Thus, using the theorems of Section 8, we can conclude:
Theorem 9.2. Lebesgue measure in Rnexists, and it is uniquely determined,
given our hypotheses.
It is obvious from our constructions that Lebesgue measure is invariant under
translations of sets. It is also invariant under other rigid motions (rotations,
reflections); this will be a consequence of Lemma 12.1 and some linear algebra.
33
An interesting question to ask is whether there are any sets that are not
Borel, or that cannot be assigned any volume. The following theorem gives a
classic example (and should also serve to convince you why our strenuous efforts
are necessary).
Theorem 9.3 (Vitali). There exists a non-measurable set in [0,1]using Lebesgue
measure. In other words, Lebesgue measure cannot be defined consistently for
allsubsets of [0,1].
Proof. The key fact in this proof is translation-invariance. In particular, given
any measurable H⊆[0,1], define its “shift with wrap-around”:
H⊕x={h+x:h∈H, h +x≤1} ∪ {h+x−1 :h∈H, h +x>1}.
Thenλ(H⊕x) =λ(H).
Define two real numbers to be equivalent if their difference is rational. The
interval [0,1] is partitioned by this equivalence relation. Compose a set H⊂
[0,1] consisting of exactly one element from each equivalence class, and also say
0/∈H. Then (0,1] equals the disjoint union of all H⊕r, forr∈[0,1)∩Q.
Consequently, by countable additivity,
1 =λ((0,1]) =/summationdisplay
r∈[0,1)∩Qλ(H⊕r) =/summationdisplay
r∈[0,1)∩Qλ(H),
a contradiction, because the sum on the right can only be 0 or ∞. HenceH
cannot be measurable. /square
A fact related to these matters is that Lebesgue measure is complete , mean-
ing ifµ(A) = 0, then every B⊆Ais Lebesgue-measurable and µ(B) = 0. (This
follows directly from the construction of the outer measure in the previous sec-
tion.) On the other hand, one can show that the Lebesgue measure restricted
to the Borel sets in Rnisnotcomplete. This means a slight complication in
the theorems we prove about Lebesgue measure, but fortunately the extension
process allows us to complete any (sigma-finite) measure if necessary.
10 Riemann integrability implies Lebesgue in-
tegrability
You have probably already suspected that any function that any Riemann-
integrable function is also Lebesgue-integrable, and certainly with the same
values for the two integrals. We shall prove this fact here. Let us first review
the definition of the Riemann integral.
LetA⊂Rmbe a (bounded) rectangle. Usually a bounded function f:A→
Ris said to be (proper-) Riemann-integrable if the supremum of its lower sums
and the infimum of its upper sums are equal:
sup
P/braceleftBig/summationdisplay
R∈Pµ(R)·inf
x∈Rf(x)/bracerightBig
= inf
P/braceleftBig/summationdisplay
R∈Pµ(R)·sup
x∈Rf(x)/bracerightBig
34
(where Pdenotes a rectangular partition of A).
We can rephrase the definition by considering not just lower sums and upper
sums forf, but the integral of any simple function s≤fors≥f(simple with
respect to a rectangular partition). Such simple functions are obviously both
Riemann- and Lebesgue- integrable with the same values for the integral. It is
also easily seen that for every such s≤fthere exists some lower sum for f(in
the usual sense) such that the integral of sis less than or equal to that lower
sum. Similarly for the upper simple functions and the upper sums. Therefore
sup
all simple s≤f/braceleftBig/integraldisplay
As/bracerightBig
= sup
P/braceleftBig/summationdisplay
R∈Pµ(R)·inf
x∈Rf(x)/bracerightBig
,
inf
all simple s≥f/braceleftBig/integraldisplay
As/bracerightBig
= inf
P/braceleftBig/summationdisplay
R∈Pµ(R)·sup
x∈Rf(x)/bracerightBig
,
and we may equivalently define fto be Riemann-integrable if the supremum
of the integrals of the lower simple functions is equal to the infimum of the
integrals of the upper simple functions.
It follows from the usual arguments, that if s1ands2are simple with s1≤
f≤s2, then/integraltext
As1≤/integraltext
As2, and thatfis Riemann-integrable if and only if there
exists a sequence of lower simple functions ln≤f, and upper simple functions
un≥fsuch that
lim
n→∞/integraldisplay
Aln=/integraldisplay
Af= lim
n→∞/integraldisplay
Aun.
This more relaxed definition of Riemann integrability is easier to work with
in the proof of the following theorem.
Theorem 10.1. LetA⊂Rmbe a rectangle. If f:A→Ris proper Riemann-
integrable, then it is also Lebesgue-integrable (with respect to Lebesgue measure)
with the same value for the integral.
Proof.fis Riemann-integrable, so choose a sequence of simple functions ln≤
f≤unwith lim n/integraltext
Aln=/integraltext
Af= lim n/integraltext
Aun. Letgn(x) = max k≤nlk(x), so that
gn(x) increase to g(x) = supngn(x). By our construction, we have
ln≤gn≤g≤f≤un.
Since the integrals of lnandunconverge onto each other, we know that g
is Riemann-integrable. Riemann-integrating and applying limits to the above
inequality,
lim
n→∞/integraldisplay
Aln≤lim
n→∞/integraldisplay
Agn≤/integraldisplay
Ag≤/integraldisplay
Af≤lim
n→∞/integraldisplay
Aun.
Thus the non-negative function f−ghas a Riemann integral of zero, and so
f−g= 0 almost everywhere with respect to Lebesgue measure6. In turn,f−g
must be measurable. (This follows from Lebesgue measure being complete.)
6This follows from a very famous theorem about Riemann integrability, whose proof you
can find in [ Spivak2 ] or [ Munkres ].
35
On the other hand, gis measurable, because gnandlnare, sofis measur-
able. Since |f|is bounded and Ahas finite measure, fmust also be Lebesgue-
integrable. Then the same inequality above with the Riemann integrals changed
to Lebesgue integrals shows that the Lebesgue and Riemann integrals of fare
the same. /square
The following theorem concerns the absolutely convergent improper Rie-
mann integral as defined in [ Munkres ].
Theorem 10.2. LetA⊆Rmbe open, and f:A→Rbe locally bounded on A
and continuous almost everywhere on A. Iffis improper-Riemann-integrable
(i.e./integraltext
A|f|<∞), then its Lebesgue integral exists with the same value for the
integral. Also/integraltext
A|f|diverges simultaneously for the improper Riemann integral
and the Lebesgue integral.
Proof. Suppose first that f≥0. LetCnbe a sequence of compact Jordan-
measurable subsets of Awhose union is AandCn⊂interiorCn+1. Then
/integraldisplay
Af= lim
n→∞/integraldisplay
Cnf.
Note that/integraltext
Cnfis valid as both a Riemann and Lebesgue integral, by Theo-
rem10.1, and it can also be written as the Lebesgue integral/integraltext
Af·χCnwhich
converges monotonically, as n→ ∞ , to the Lebesgue integral/integraltext
Af. This must
of course be equal to the left side of the equation above, which is the improper
Riemann integral.
For general f, repeating the same reasoning for the non-negative functions
|f|,f+,f−, in turn proves the theorem. /square
11 Product measures and Fubini’s Theorem
Fubini’s theorem concerns integrals in “multiple dimensions” and their eval-
uation using iterated integrals. The concept should be familiar from multi-
dimensional calculus, so I won’t launch myself into an extended discussion here.
But before we start writing down integral signs, we need to discuss the
measurability of the sets involved in multiple integration.
Definition 11.1. Let (X,A) and (Y,B) be two measurable spaces. A measur-
able rectangle inX×Yis a set of the form A×B, whereA∈ A andB∈ B.
The sigma algebra generated by all the measurable rectangles is denoted by
A ⊗ B , and this will be the sigma algebra we use for X×Y.
Theorem 11.1. LetEbe a measurable set from (X×Y,A ⊗ B ). Let
Ex={y∈Y: (x,y)∈E}, x∈X.
Ey={x∈X: (x,y)∈E}, y∈Y .
ThenEy∈ A,Ex∈ B.
36
Proof. We prove the theorem for Ey; the proof for Exis the same.
We consider the collection D={E∈ A ⊗ B :Ey∈ A} .
SupposeE=A×Bis a measurable rectangle. Then Ey=Awheny∈ B,
otherwiseEy=∅. In both cases Ey∈ A, soE∈ D.
Dis a sigma algebra, because:
1.∅y={x∈X: (x,y)∈ ∅} =∅ ∈ A , so∅ ∈ D .
2.IfEn∈ D, then (/uniontextEn)y=/uniontextEn
y∈ A. So/uniontextEn∈ D.
3.IfE∈ D, andF=Ec, thenFy={x∈X: (x,y)∈F}={x∈X:
(x,y)/∈E}=X\Ey∈ A. SoEc∈ D.
ThusDis a sigma algebra containing the measurable rectangles, i.e. Dis all of
A ⊗ B . /square
Theorem 11.2. Let(X×Y,A ⊗ B )and(Z,C)be measurable spaces.
Iff:X×Y→Zis measurable, then the functions fy:X→Z,fx:Y→Z
obtained by holding one variable fixed are also measurable.
Proof. Again we consider only fy. LetIy:X→X×Ybe defined by Iy(x) =
(x,y). Then given E∈ A ⊗ B , by Theorem 11.1.I−1
y(E) =Ey∈ A , soIyis a
measurable function. But fy=f◦Iy. /square
The next theorem on multiple integration requires the following technical
tool, which comes equipped with a definition.
Definition 11.2. A family Aof subsets of Xis amonotone class if it is closed
under increasing unions and decreasing intersections.
The intersection of any set of monotone classes is a monotone class. The
smallest monotone class containing a given set Gis the intersection of all mono-
tone classes containing G. This construction is analogous to the one for sigma
algebras, and the result is also said to be the monotone class generated by G.
Theorem 11.3 (Monotone Class Theorem). IfAbe an algebra on X, then
the monotone class generated by Ais the same as the sigma algebra generated
byA.
Proof. Since a sigma algebra is a monotone class, the generated sigma algebra
contains the generated monotone class M. So we only need to show Mis a
sigma algebra.
We first claim that Mis actually closed under complementation. Let M/prime=
{S∈ M :X\S∈ M} ⊆ M . This is a monotone class, and it contains the
algebra A. SoM=M/primeas desired.
To prove that Mis closed under countable unions, we only need to prove
that it is closed under finite unions, for it is already closed under countable
increasing unions.
First letA∈ A, andN(A) ={B∈ M :A∪B∈ M} ⊆ M . Again this is a
monotone class containing the algebra A; thus N(A) =M.
37
Finally, let S∈ M , with the same definition of N(S). The last paragraph,
rephrased, says that A ⊆ N (S). And N(S) is a monotone class containing A
by the same arguments as the last paragraph. Thus N(S) =Mas desired. /square
Theorem 11.4. Let(X,A,µ),(Y,B,ν)be sigma-finite measure spaces.
IfE∈ A ⊗ B , then
1.ν(Ex)is a measurable function of x∈X.
2.µ(Ey)is a measurable function of y∈Y.
Proof. We concentrate on µ(Ey). Let {Xm} /arrownortheastX, withµ(Xm)<∞. Fixm
for now and let D={E∈ A ⊗ B :µ(Ey∩Xm) is a measurable function of y}.
Dis equal to A ⊗ B , because:
1.IfE=A×Bis a measurable rectangle, then µ(Ey∩Xm) =µ(A∩Xm)·
χB(y) which is a measurable function of y.
IfEis a finite disjoint union of measurable rectangles En, thenµ(Ey∩
Xm) =/summationtext
nµ(En
y∩Xm) which is also measurable.
The measurable rectangles form a semi-algebra (just like the rectangles in
Rn). Therefore, applying Theorem 9.1,Dcontains the algebra of finite
disjoint unions of measurable rectangles.
2.IfEnare increasing sets in D(not necessarily measurable rectangles), then
µ((/uniontextEn)y∩Xm) =µ(/uniontextEn
y∩Xm) = lim n→∞µ(En
y∩Xm) is measurable,
so/uniontextEn∈ D.
Similarly, if Enare decreasing sets in D, then using limits we see that/intersectiontextEn∈ D. (Here it is crucial that En
y∩Xmhave finite measure, for the
limiting process to be valid.)
3.These arguments show that Dis a monotone class, and it contains the
monotone class generated by the algebra of finite unions of measurable
rectangles. By the Monotone Class Theorem, Dmust therefore be the
same as the sigma algebra A ⊗ B .
We now know that for each E∈ A⊗B ,µ(Ey∩Xm) is measurable, for every
m. Taking limits as m→ ∞ , we conclude that µ(Ey) is also measurable. /square
One thing has been deliberately left out of our discussion so far: the con-
struction of the product measure µ⊗ν, which, as in the case of Rn, should
assign a measure µ(A)ν(B) to the measurable rectangle A×B. The problem is
that we cannot prove countable additivity of µ⊗νdefined this way. We take
an indirect route instead, defining it by iterated integrals:
Theorem 11.5. Let(X,A,µ),(Y,B,ν)be sigma-finite measure spaces. There
exists a unique product measure µ⊗ν:A ⊗ B → [0,∞], with
(µ⊗ν)(E) =/integraldisplay
x∈X/integraldisplay
y∈YχE(x,y)dν
/bracehtipupleft /bracehtipdownright/bracehtipdownleft /bracehtipupright
ν(Ex)dµ=/integraldisplay
y∈Y/integraldisplay
x∈XχE(x,y)dµ
/bracehtipupleft /bracehtipdownright/bracehtipdownleft /bracehtipupright
µ(Ey)dν.
38
Proof. Letλ1(E) denote the double integral on the left, and λ2(E) denote the
one on the right. (These integrals exist by Theorem 11.4.) It is obvious that
bothλ1andλ2are countably additive, so they are both measures on A ⊗ B .
Moreover, if E=A×B, then just expanding the two integrals shows λ1(E) =
µ(A)ν(B) =λ2(E). NowX×Yis sigma-finite if XandYare, so by uniqueness
of measures7(Corollary 8.7),λ1=λ2on all of A ⊗ B . /square
There’s not much work left for our final theorems:
Theorem 11.6 (Fubini). Let(X,A,µ)and(Y,B,ν)be sigma-finite measure
spaces. Iff:X×Y→Risµ⊗ν-integrable, then
/integraldisplay
X×Yfd(µ⊗ν) =/integraldisplay
x∈X/bracketleftBig/integraldisplay
y∈Yf(x,y)dν/bracketrightBig
dµ=/integraldisplay
y∈Y/bracketleftBig/integraldisplay
x∈Xf(x,y)dµ/bracketrightBig
dν.
This equation also holds when f≥0(if it is merely measurable, not integrable).
Proof. The casef=χEis just Theorem 11.5. Since all three integrals are
additive, they are equal for non-negative simple f, and hence also for all other
non-negative f, by approximation and monotone convergence. From linearity
onf=f+−f−, we see that they are equal for any integrable f. (We will allow
∞ − ∞ to occur on a set of measure zero.) /square
Theorem 11.7 (Tonelli). Let(X,A,µ)and(Y,B,ν)be sigma-finite measure
spaces, and f:X×Y→Rbeµ⊗ν-measurable. Then fisµ⊗ν-integrable if
and only if /integraldisplay
x∈X/bracketleftBig/integraldisplay
y∈Y|f(x,y)|dν/bracketrightBig
dµ<∞
(or withXandYreversed).
Consequently, if any one of these conditions hold, then it is valid to switch
the order of integration when integrating f.
Proof. Immediate from Fubini’s theorem applied to the function |f|. /square
Remark 11.8.Note that it is possible that/integraltext
y∈Y|f(x,y)|dν=∞on a set of
measure zero in X, and still have integrability, and conversely.
12 Change of variables in Rn
This section will be devoted to completing the proof of the differential change
of variables formula, Theorem 6.5. As we noted in the remarks preceding that
theorem, it suffices to prove the following.
7If you prefer, you can also prove this theorem using the Monotone Class Theorem instead.
39
Lemma 12.1. Letg:X→Ybe a diffeomorphism between open sets in Rn.
Then for all measurable sets A⊆X,
λ(g(A)) =/integraldisplay
g(A)1 =/integraldisplay
A|det Dg|.
Proof. We first begin with two simple reductions.
(I)It suffices to prove the lemma locally.
That is, suppose there exists an open cover of X,{Uα}, so that the
equation of the lemma holds for measurable Acontained inside one of the
Uα. Then the equation actually holds for all measurable A⊆X.
Proof. By taking a countable subcover, we may assume there are only
countably many Ui. Define the disjoint measurable sets Ei=Ui\(U1∪
··· ∪Ui−1), which cover X. Also define the two measures:
µ(A) =λ(g(A)), ν(A) =/integraldisplay
A|det Dg|.
Now letA⊆Xbe any measurable set. We have A∩Ei⊆Ui, soµ(A∩
Ei) =ν(A∩Ei) by hypothesis. Therefore,
µ(A) =µ/parenleftBig/uniondisplay
iA∩Ei/parenrightBig
=/summationdisplay
iµ(A∩Ei) =/summationdisplay
iν(A∩Ei) =ν(A).
(II)Suppose the lemma holds for two diffeomorphisms gandh, and all mea-
surable sets. Then it holds for the composition diffeomorphism g◦h(and
all measurable sets).
Proof. For any measurable A,
/integraldisplay
g(h(A))1 =/integraldisplay
h(A)|det Dg|=/integraldisplay
A|(det Dg)◦h|·|det Dh|=/integraldisplay
A|det D(g◦h)|.
The second equality follows from Theorem 6.5applied to the diffeomor-
phismh, which is valid once we know λ(h(B)) =/integraltext
B|det Dh|for all mea-
surableB.
We proceed to prove the lemma by induction, on the dimension n.
Base case n= 1.CoverXby a countable set of bounded intervals IkinR. By
Reduction I, it suffices to prove the lemma for measurable sets contained
in each of the Ikindividually. By the uniqueness of measures (Corollary
8.7), it also suffices to show µ=νonly for the intervals [ a,b], (a,b), etc.
But this is just the Fundamental Theorem of Calculus:
/integraldisplay
g([a,b])1 =|g(b)−g(a)|=|/integraldisplayb
ag/prime|=/integraldisplayb
a|g/prime|.
40
(For the last equality, remember that the g/primemust be either positive on all
of [a,b] or negative on all of [ a,b]. If the interval is open or half-open, we
may not be able to apply the Fundamental Theorem, but the preceding
equation can still be obtained via a limiting procedure.)
Induction step. Locally (i.e. on a sufficiently small open set around each
pointx∈X),gcan always be factored8asg=hk◦ ··· ◦h2◦h1, where
eachhiis a diffeomorphism and fixes one coordinate of Rn. By Reduction
I, it suffices to consider this local case only. By Reduction II, it suffices to
prove the lemma for each of the diffeomorphisms hi.
So suppose gfixes one coordinate. For convenience in notation, assume g
fixes the last coordinate: g(u,v) = (hv(u),v), foru∈Rn−1,v∈R, andhv
are functions on (open subsets of) Rn−1, in fact diffeomorphisms. Clearly
hvare one-to-one, and most importantly, det D hv(u) = det Dg(u,v)/negationslash= 0.
Next, let a measurable set Abe given, and consider its projection V=
{v∈R: (u,v)∈A}, and its cross-section Uv={u∈Rn−1: (u,v)∈A}.
We now apply Fubini’s theorem and the induction hypothesis on hv:
/integraldisplay
g(A)1 =/integraldisplay
v∈V/integraldisplay
hv(Uv)1
=/integraldisplay
v∈V/integraldisplay
u∈Uv|det Dhv(u)|
=/integraldisplay
v∈V/integraldisplay
u∈Uv|det Dg(u,v)|=/integraldisplay
A|det Dg|. /square
13 Vector-valued integrals
This section, in short, is a remark that everything we have done so far generalizes
to vector-valued functions, the vectors being from real finite-dimensional spaces
(orC). These often occur in applications.
Givenf:X→Rnmeasurable, let {ek}denote the standard basis vectors
inRn, and {fk}the components of fwith respect to this basis. Of course we
define/integraldisplay
f=n/summationdisplay
k=1/parenleftBig/integraldisplay
fk/parenrightBig
ek,
provided the integrals on the right exist. Generally, to say that/integraltext
Xfexists, we
do not allow any one of the components to be infinite, for this is usually not
useful when n≥2.
It is not hard to see that fis measurable if and only if each fkis.
It follows immediately that this integral is linear, and hence, the definition
is independent of the basis, as it ought to be. That is, if {ek}isanybasis of
Rn, the sum in the definition does not change.
8The proof of this fact can be found in [ Munkres ], but it is not difficult to prove it yourself.
Hint: Inverse Function Theorem.
41
Since by our convention that the components of the integral are not allowed
to be infinite, it seems we do not need to define separately what it means for f
to be integrable. But we define it, because we want to take note of some facts
(admittedly they are not very interesting): fbeing integrable means/integraltext
/bardblf/bardbl<∞.
The norm can be arbitrary, for in Rn, every norm is equivalent: if /bardbl·/bardbl1and
/bardbl·/bardbl2are any two given norms, then there always exist constants α,β > 0 such
thatα/bardbl·/bardbl2≤ /bardbl·/bardbl 1≤β/bardbl·/bardbl2. Thus being integrable in one norm implies integra-
bility in another norm. In particular, by using the norm /bardblx/bardblΣ=/summationtextn
k=1|xk|, we
see thatfis integrable if and only if its components fkare integrable.
We want to show that /bardbl/integraltext
f/bardbl ≤/integraltext
/bardblf/bardbl; this basic inequality will enable us
to make estimates without having to separate components. As a start, this is
clearly true if fis a simple function, i.e. f=/summationtextm
j=1ajχEjforaj∈Rn. It is
also trivial if/integraltext
/bardblf/bardbl=∞. To prove the inequality for the other f, we use the
following easy lemma.
Lemma 13.1. Letf:X→Rnbe integrable. There exists a sequence of simple
functionsϕj:X→Rnconverging pointwise to f, with
lim
j→∞/integraldisplay
/bardblϕj−f/bardbl= 0.
Proof. By equivalence of norms, it suffices to prove this only for the norm /bardbl·/bardblΣ
as defined above.
For eachfk+, by the approximation theorem in R, there exists measurable
simpleϕk+
jincreasing to fk+. Similarly for fk−. Letϕk
j=ϕk+
j−ϕk−
j. Using
the Dominated Convergence Theorem applied to each component,
lim
j→∞/integraldisplay
/bardblϕj−f/bardblΣ= lim
j→∞/integraldisplayn/summationdisplay
k=1|ϕk
j−fk|= 0.
We also note that by the Monotone Convergence Theorem, the ϕjsatisfy
lim
j→∞/integraltext
ϕj=/integraltext
f. /square
We finish our demonstration of the generalized triangle inequality. Let ϕj
as in the lemma. Then
/vextenddouble/vextenddouble/vextenddouble/integraldisplay
ϕj/vextenddouble/vextenddouble/vextenddouble≤/integraldisplay
/bardblϕj/bardbl ≤/integraldisplay
/bardblf/bardbl+/integraldisplay
/bardblϕj−f/bardbl,
so that
lim
j→∞/vextenddouble/vextenddouble/vextenddouble/integraldisplay
ϕj/vextenddouble/vextenddouble/vextenddouble=/vextenddouble/vextenddouble/vextenddoublelim
j→∞/integraldisplay
ϕj/vextenddouble/vextenddouble/vextenddouble=/vextenddouble/vextenddouble/vextenddouble/integraldisplay
f/vextenddouble/vextenddouble/vextenddouble≤/integraldisplay
/bardblf/bardbl+ lim
j→∞/integraldisplay
/bardblϕj−f/bardbl=/integraldisplay
/bardblf/bardbl.
14 C∞
0functions are dense in Lp(Rn)
This section is devoted to the result that the space of C∞
0functions is dense in
Lp(Rn), which was discussed at the end of Section 7.
42
Theorem 14.1. Letf:Rn→R∈Lp(Rn),1≤p <∞. Then for any ε >0,
there exists ψ∈C∞
0such that
/bardblψ−f/bardblp=/parenleftBig/integraldisplay
Rn|ψ−f|pdλ/parenrightBig1/p
<ε.
Our strategy for proving this theorem is straightforward. Since we already
know that the simple functions ϕ=/summationtext
iaiχEiare dense in Lp, we should try
approximating χEibyC∞
0functions. Since C∞
0functions are non-zero on com-
pact sets, it stands to reason that we should approximate the sets Eiby compact
setsKi. If this can be done, then it suffices to construct the C∞
0functions on
the setsKi.
Our constructions start with this last step. You might even have seen some
of these constructions before.
Lemma 14.2. LetAbe a compact rectangle in Rn. Then there exists φ∈C∞
0
which is positive on the interior of Aand zero elsewhere.
Proof. Consider the infinitely differentiable function
f(x) =/braceleftbigg
e−1/x2, x> 0
0, x ≤0.
IfA= [0,1], thenφ(x) =f(x)·f(1−x) is the desired function of C∞
0. (Draw
pictures!)
IfA= [a1,b1]× ··· × [an,bn], then we let
φA(x) =φ/parenleftbiggx1−a1
b1−a1/parenrightbigg
···φ/parenleftbiggxn−an
bn−an/parenrightbigg
. /square
Lemma 14.3. For anyδ >0, there exists an infinitely differentiable function
h:R→[0,1]such thath(x) = 0 forx≤0andh(x) = 1 forx≥δ.
Proof. Take the function φfrom Lemma 14.2 for the rectangle [0 ,δ], and let
h(x) =/integraltextx
−∞φ(t)dt/integraltext∞
−∞φ(t)dt. /square
Theorem 14.4. LetUbe open, and K⊂Ucompact. Then there exists ψ∈
C∞
0which is positive on Kand vanishes outside some other compact set L,
K⊂L⊂U.
Proof. For eachx∈U, letAx⊂Ube a bounded open rectangle containing x,
whose closure Axlies inU. The {Ax}together form an open cover of K. Take
a finite subcover {Axi}. Then the compact rectangles {Axi}also coverK.
From Lemma 14.2, obtain functions ψi∈C∞
0that are positive on Axiand
vanish outside Axi. Letψ=/summationtext
iψi∈C∞
0.ψvanishes outside L=/uniontext
iAxi,
which is compact. /square
43
Corollary 14.5. In Theorem 14.4, it is even possible to require in addition that
0≤ψ(x)≤1for allx∈Rnandψ(x) = 1 forx∈K.
Proof. Letψbe from Theorem 14.4. Sinceψis positive on the compact set K,
it has a positive minimum δthere. Take the function hof Lemma 14.3 for this
δ. The new candidate function is h◦ψ. /square
As we have said, we must now approximate arbitrary Borel sets B∈ B(Rn)
by compact sets. (We will also need approximation by open sets.) It turns out
that this part of the proof is purely topological, and generalizes to other metric
spacesXbesides Rn. Henceforth we consider the more general case.
Letddenote the metric for the metric space X.
Theorem 14.6. Let(X,B(X),µ)be a finite measure space, and let B∈ B(X).
For everyε > 0, there exists a closed set Vand an open set Usuch that
V⊆B⊆Uandµ(U\V)<ε.
Proof. LetMbe the set of all B∈ B(X) for which the statement is true. We
show that Mis a sigma algebra containing all the open sets in X.
1.LetB∈ M withVandUas above. Then Vcopen⊇Bc⊇Ucclosed,
andµ(Vc)−µ(Uc)<ε. This shows Bc∈ M .
2.LetBn∈ M . ChooseVnandUnfor eachBnsuch thatµ(Un\Vn)<ε/2n.
LetU=/uniontext
nUnwhich is open, and V=/uniontext
nVn, so thatV⊆/uniontext
nBn⊆U.
Of courseVis not necessarily closed, but WN=/uniontextN
n=1Vnare, and these
WNincrease to V. Henceµ(V\WN)→0 asN→ ∞ , meaning that for
large enough N,µ(V\WN)<ε.
Next, we have
U\WN= (U\V)/unionmulti(V\WN)
⊆/uniondisplay
n(Un\Vn)∪(V\WN),
µ(U\WN) =µ(U\V) +µ(V\WN)
≤/summationdisplay
n(Un\Vn) +µ(V\WN)<ε+ε.
This shows that/uniontext
nBn∈ M .
3.LetBbe open, and A=Bc. Also let d(x,A) = inf y∈Ad(x,y) be the
distance from x∈XtoA. SetDn={x∈X:d(x,A)≥1/n}.Dnis
closed, because d(·,A) is a continuous function, and [1 /n,∞] is closed.
Clearlyd(x,A)≥1/n> 0 impliesx∈Ac=B, but since Ais closed, the
converse is also true: for every x∈Ac=B,d(x,A)>0. Obviously the
Dnare increasing, so we have just shown that they in fact increase to B.
Henceµ(B\Dn)<εfor large enough n. ThusB∈ M . /square
44
The case that µis not a finite measure is taken care of, as you would expect,
by taking limits like we did for sigma-finite measures in Section 8. But since
compact and open sets are involved, we need stronger hypotheses:
1.There exists {Kn} /arrownortheastX, withKncompact and µ(Kn)<∞.
2.There exists {Xn} /arrownortheastX, withXnopen andµ(Xn)<∞.
It is easily seen that these properties are satisfied by X=Rnand the
Lebesgue measure λ, as well as many other “reasonable” measures µonB(Rn).
We will discuss this more later.
We assume henceforth that Xandµhave the properties just listed.
Theorem 14.7. LetB∈ B(X)withµ(B)<∞. For every ε>0, there exists
a compact set Vand an open set Usuch thatK⊆B⊆Uandµ(U\K)<ε.
Proof. It suffices to show that µ(U\B)<εandµ(B\K)<εseparately.
Existence of K.Since {B∩Kn} /arrownortheastB, there exists some nsuch thatµ(B)−
µ(B∩Kn)<ε/2.
For thisn, define the finite measure µKn(E) =µ(E∩Kn), forE∈ B(X).
By Theorem 14.6, there are sets V⊆B⊆U,Vclosed, and µKn(B\V)≤
µKn(U\V)<ε/2. SinceXnis compact, it is closed. Then K=V∩Kn
is also closed, and hence compact, because it is contained in the compact
setKn. We have,
µ(B\K) =µ(B)−µ(B∩Kn) +µ(B∩Kn)−µ(K)
=µ(B)−µ(B∩Kn) +µKn(B)−µKn(V)<ε
2+ε
2.
Existence of U.For everyn, define the finite measure µXn(E) =µ(E∩Xn),
forE∈ B(X). By Theorem 14.6, there are sets Vn⊆B⊆Un,Unopen,
andµXn(Un\B)≤µXn(Un\Vn)<ε/2n.
LetU=/uniontext
nUn∩Xn⊇B. We have,
µ(U\B)≤/summationdisplay
nµ(Un∩Xn\B) =/summationdisplay
nµXn(Un\B)<ε. /square
We return to the case of X=Rn.
Theorem 14.8. LetB∈ B(Rn)withµ(B)<∞. For every ε>0, there exists
ψ∈C∞
0such that
/bardblψ−χB/bardblp=/parenleftBig/integraldisplay
Rn|ψ−χB|pdµ/parenrightBig1/p
<ε.
45
Proof. By Theorem 14.7, there is compact Kand openU,K⊆B⊆U,µ(U\
K)< ε. From Corollary 14.5, there isψ∈C∞
0such thatψ= 1 onK,ψ= 0
outsideU, and 0 ≤ψ≤1. Then
/integraldisplay
Rn|ψ−χB|p=/integraldisplay
Rn\U0 +/integraldisplay
U\Bψp+/integraldisplay
B\K(1−ψ)p+/integraldisplay
K0
≤µ(U\B) +µ(B\K)
=µ(U)−µ(B) +µ(B)−µ(K)
<ε. /square
Proof of Theorem 14.1.Letϕ=/summationtext
iaiχEi,ai/negationslash= 0 be a simple function such
that/bardblϕ−f/bardblp< ε/ 2. Letψi∈C∞
0such that /bardblψi−χEi/bardblp< ε/ 2|ai|. (Note
thatEimust have finite measure; otherwise ϕwould not be integrable.) Let
ψ=/summationtext
iaiψi. Then (Minkowski’s inequality),
/bardblf−ψ/bardblp≤ /bardblf−ϕ/bardblp+/bardblϕ−ψ/bardblp
≤ /bardblf−ϕ/bardblp+/summationdisplay
i|ai| · /bardblχEi−ψi/bardblp<ε. /square
Actually, even the last part of theorem can be generalized to spaces other
thanRn: instead of infinitely differentiable functions with compact support, we
consider continuous functions, defined on the metric space X, with compact
support. In this case, a topological argument must be found to replace Lemma
14.2. This is easy:
Lemma 14.9. LetAbe any compact set in X. Then there exists a continuous
functionφ:X→Rwhich is positive on the interior of Aand zero elsewhere.
Proof. LetC=X\interiorA, soCis closed. Then φ(x) =d(x,C) works.
(d(x,C) was defined in the proof of Theorem 14.6.) /square
The proof of Theorem 14.4 goes through verbatim for metric spaces X,
provided that Xislocally compact . This means: given any x∈Xand an open
neighborhood Uofx, there exists another open neighborhood Vofx, such that
Vis compact and V⊆U.
Finally, we need to consider when properties (1) and (2) (in the remarks
preceding Theorem 14.7) are satisfied. These properties are somewhat awkward
to state, so we will introduce some new conditions instead.
Definition 14.1. A measure µon a topological space Xislocally finite if for
eachx∈X, there is an open neighborhood Uofxsuch thatµ(U)<∞.
It is easily seen that when µis locally finite, then µ(K)<∞forevery
compact set K.
Definition 14.2. A topological space Xisstrongly sigma-compact if there
exists a sequence of open sets Xnwith compact closure, and {Xn} /arrownortheastX.
46
IfXis strongly sigma-compact, and µis locally finite, then properties (1) and
(2) are automatically satisfied. It is even true that strong sigma-compactness
implies local compactness in a metric space. (The proof requires some topology
and is left as an exercise.) Then we have the following theorem:
Theorem 14.10. LetXbe a strongly sigma-compact metric space, and µbe
any locally finite measure on B(X). Then the space of continuous functions with
compact support is dense in Lp(X,B(X),µ),1≤p<∞.
15 Other examples of measures
Since so far we have chiefly worked only in Rnwith Lebesgue measure, it should
be of interest to give a few more useful examples of measures.
k-dimensional volume of a k-dimensional manifold
A manifold is a generalization of curves and surfaces to higher dimensions, and
sometimes even to spaces other than Rn. But here we shall concentrate on differ-
entiable manifolds inside Rn; the theory is elucidated in [ Spivak2 ] or [Munkres ].
Here we give a definition of the k-dimensional volume for k-dimensional mani-
folds which does not require those dreaded “partitions of unity”.
Suppose ak-dimensional manifold M⊆Rnis covered by a single coordinate
chartα:U→M,U⊆Rkopen. Let D αdenote the n-by-kmatrix
Dα=/bracketleftbiggdα
dt1dα
dt2...dα
dtk/bracketrightbigg
.
(More precisely, each vectordα
dtiis represented by a column vector in the standard
basis of Rn. Actually our definition works using any orthonormal basis also.)
Define, for any vectors v1,...,v k∈Rn(again represented in an orthnormal
basis):
V(v1,...,v k) =/radicalBig
det/bracketleftbigv1v2... v k/bracketrightbigtr/bracketleftbigv1v2... v k/bracketrightbig
=/radicalBig
det [vi·vj]i,j=1,...,k,
This is the k-dimensional volume of a k-dimensional parallelopiped spanned by
the vectors v1,...,v kinRn. One easily shows that this volume is invariant under
orthogonal transformations, and that it agrees with the usual k-dimensional
volume (as defined by the Lebesgue measure) when the parallelopiped lies in
the subspace Rk×0⊆Rn.
Thek-dimensional volume of any E∈ B(M) is defined as:
ν(E) =/integraldisplay
α−1(E)V(Dα)dλ.
(Sinceαis continuous, α−1(E)∈ B(Rk).)
47
The integrand, of course, is supposed to represent “infinitesimal” elements of
surface area ( k-dimensional volume), or approximations of the surface area of E
by polygons that are “close” to E. As indicated by the quotation marks, these
assertions about “surface area” are completely non-rigorous, and we won’t be-
labour to prove them, since the equation above isour definition of k-dimensional
volume. But it should be pointed out that there are better theories of k-
dimensional volume available, which are intrinsic to the sets being measured,
instead of our computational theory. (I don’t know these other theories well
enough though.)
Back to our definitions. If Mis not covered by a single coordinate chart,
but more than one, say αi:Ui→M,i= 1,2,..., then partition Mwith
V1=α1(U1),Vi=αi(Ui)\Vi−1, and define
ν(E) =/summationdisplay
i/integraldisplay
α−1
i(E∩Vi)V(Dα)dλ.
It is left as an exercise to show that ν(E) is well-defined: it is independent of
the coordinate charts αiused forM.
Finally, the scalar integral of f:M→RoverMis simply
/integraldisplay
Mfdν.
And the integral of a differential form ωon an oriented manifold Mis
/integraldisplay
p∈Mω/parenleftbig
p;T(p)/parenrightbig
dν,
whereT(p) is an orthonormal frame of the tangent space of Matp, oriented
according to the given orientation of M. (If you don’t know what I’m talking
about, just ignore this definition — essentially it generalizes the line and surface
integrals in calculus.)
Again it is not hard to show that the formulae I have given are exactly
equivalent to the classical ones for evaluating scalar integrals and integrals of
differential forms, which are of course needed for actual computations. But there
are several advantages to our new definitions. First is that they are elegant: they
are mostly coordinate-free, and all the different integrals studied in calculus have
been unified to the Lebesgue integral by employing different measures. In turn,
this means that the nice properties and convergence theorems we have proven
all carry over to integrals on manifolds.
For example, everybody “knows” that on a sphere, any circular arc Chas
“measure zero”, and so may be ignored when integrating over the sphere. To
prove this rigorously using our definitions, we only have to remark that ν(C) =
0, sinceλ(α−1(C)) = 0 for a coordinate chart αfor the sphere.
Stieltjes measure
The definition of the Stieltjes measure is best motivated by probability theory.
Suppose we have a random variable Zwith distribution µ, and the cumulative
48
distribution function F:R→[0,1] — by definition, they satisfy F(z) = Pr[Z≤
z] =µ/parenleftbig
[−∞,z]/parenrightbig
. It follows that Fis (non-strict) increasing, and µ((a,b]) =
F(b)−F(a).
The idea here is to try to reverse this procedure: given any increasing func-
tionF, can we construct a measure µonB(R) that assigns, to any interval
(a,b], a “length” of F(b)−F(a)?
Actually we will need to impose some conditions on Ffirst. Since Fis in-
creasing, it always has only a countable number of discontinuities, and these
discontinuities must all be jump discontinuities. At these jumps, we will insist
thatFis right-continuous, i.e. lim x/arrowsoutheastaF(x) =F(a). Otherwise, taking count-
able limits may fail: for example, if at the point a,Fis left-continuous instead
of right-continuous, then
µ/parenleftbig
(a,b]/parenrightbig
=F(b)−F(a)
/negationslash= lim
x/arrowsoutheasta(F(b)−F(x))
= lim
x/arrowsoutheastaµ/parenleftbig
(a,x]/parenrightbig
.
(Of course, the preference of “right” over “left” comes from our convention that
we used intervals ( a,b] that are open on the left and closed on the right.)
We will also insist that F(x)<∞forx/negationslash=−∞,+∞, so that the subtraction
F(b)−F(a) makes sense. However, it should be allowed that, say, F(−∞) =−∞
(andF(x) does not have to be in [0 ,1] either). This allows sigma-finite measures
to be constructed.
With the necessary conditions now stated, we can begin the construction of
µ, which is not much different from the construction of the Lebesgue measure
onR— not surprising, since F(x) =xis exactly the Lebesgue measure on R.
First, it is easily checked, by drawing pictures of the intervals ( a,b], that
they actually form a semi-algebra on R. (Just ignore the point −∞ for now.)
The measure µon the generated algebra Ais defined in the obvious way, and
it follows from the same arguments as in Section 9thatµis finitely additive.
Countable additivity requires the typical approximation arguments. Suppose
that we have Jn∈ A with infinite disjoint union I∈ A. By monotonicity we
automatically have/summationtext∞
n=1µ(Jn)≤µ(I), so we only have to prove the other
inequality. We can assume that Jnare simple intervals ( an,bn], instead of finite
disjoint unions of intervals. Also assume, for now, that Iis the finite interval
(a,b].
SinceFis right-continuous, for every ε>0, there exists δ>0 such that
0≤F(b)−F(a)<F(b)−F(a+δ) +ε.
Also there exists δn>0 such that
0≤F(bn+δn)−F(an)<F(bn)−F(an) +ε
2n.
There exists a finite set n1,...,n ksuch that (a+δ,b]⊆/uniontextk
i=1(ani,bni+δni],
49
since the open sets ( an,bn+δ) cover the compact set [ a+δ,b]. Therefore,
F(b)−F(a+δ)≤k/summationdisplay
i=1F(bni+δni)−F(ani)
≤∞/summationdisplay
n=1F(bn+δn)−F(an),
F(b)−F(a)≤2ε+∞/summationdisplay
n=1F(bn)−F(an).
and we take ε→0.
Finally, the proof for general I=I1/unionmulti ··· /unionmultiIk∈ A just follows from finite
additivity and that the finite sum of limits equals the limit of the finite sum.
Infinite intervals are handled in the same way as in Section 9.
Thus using the theorems of Section 8,µcan thereby be extended to a measure
onB(R). This is called the Stieltjes measure onR, and the integral
/integraldisplay
Rgdµ =/integraldisplay
RgdF
is the Stieltjes integral, and it generalizes the Riemann-Stieltjes integral that
is sometimes studied in real analysis courses. (The Riemann-Stieltjes integral
is defined by taking limits of Riemann-like sums/summationtext
ig(ξi)·(F(xi)−F(xi−1)).
Showing this limit exists requires some effort, however.)
Of course, if Fis differentiable, the Stieltjes integral just reduces to
/integraldisplay
RgdF =/integraldisplay
Rg·F/primedλ.
Lastly, we should mention that if we admit signed measures , which are dif-
ferences of two (positive) measures, then the condition that Fbe increasing can
even be relaxed. We will not pursue that theory here though.
16 Egorov’s Theorem
The following theorem does not really belong in a first course, but it is quite a
surprising and interesting result, and I want to record its proof.
Theorem 16.1 (Egorov). Let(X,µ)be a measure space of finite measure, and
fn:X→Rbe a sequence of measurable functions convergent almost everywhere
tof. Then given any ε>0, there exists a measurable subset A⊆Xsuch that
µ(X\A)<εand the sequence fnconverges uniformly to fonA.
Proof. First define
Bn,m=∞/intersectiondisplay
k=n/bracketleftBig
|f−fk|<1
m/bracketrightBig
.
50
Fixm. For mostx∈X,fn(x) converges to f(x), so there exists nsuch that
|fk(x)−f(x)|<1/mfor allk≥n, sox∈Bn,m. Thus we see {Bn,m}n/arrownortheastX\C
(Cis some set of measure zero).
We construct the set Ainductively as follows. Set A0=X\C. For each
m > 0, since {Am−1∩Bn,m}n/arrownortheastAm−1, we haveµ(Am−1\Bn,m)→0, so we
can choose n(m) such that
µ(Am−1\Bn(m),m)<ε
2m.
Furthermore set
Am=Am−1∩Bn(m),m.
SinceAm/unionmulti(Am−1\Bn(m),m) =Am−1, we have
µ(Am)>µ(Am−1)−ε
2m
>µ(X)−ε
2−ε
4− ··· −ε
2m≥µ(X)−ε.
The setsAmare decreasing, so letting
A=∞/intersectiondisplay
m=1Am=∞/intersectiondisplay
m=1Bn(m),m,
we haveµ(A)≥µ(X)−ε, orµ(X\A)≤ε. Finally, for x∈A,x∈Bn(m),mfor
allm, showing that |f(x)−fk(x)|<1/mwheneverk≥n(m). This condition
is uniform for all x∈A. /square
17 Exercises
I have been suggested to provide some more exercises to this text. Here they
are.
1.InRnwith Lebesgue measure, find an uncountable set of measure zero.
2.Show that if f:Rn→Ris continuous and equal to zero almost everywhere,
thenfis in fact equal to zero everywhere.
3.Find a sequence of integrable functions fnsuch thatfn(x)→0 for every
xbut/integraltext
fn→ ∞ .
4.Letf∈L1(Rn), 1≤p<∞. Compute the limits
lim
h→0/integraldisplay
|f(x+h)−f(x)|pdx, lim
/bardblh/bardbl→∞/integraldisplay
|f(x+h)−f(x)|pdx.
5.Letf∈L1(Rn). Show that
lim
/bardbly/bardbl→∞/integraldisplay
Rnf(x)ei/angbracketleftx,y/angbracketrightdx= 0.
51
6.Letfn∈Lp,1<p< ∞be a sequence of functions converging to f∈Lp
almost everywhere, and suppose there is a constant Msuch that /bardblfn/bardblp≤
Mfor alln. Then for each g∈Lq,
/integraldisplay
fg= lim
n→∞/integraldisplay
fng.
Hint: Use Egovov’s Theorem and a density argument. Is this also true for
p= 1?
7.Letfn∈Lp,1≤p <∞be a sequence of functions converging f∈Lep
almost everywhere. Prove that fnconverges to fin theLepnorm if and
only if /bardblfn/bardblp→ /bardblf/bardblp.
8.Letf∈Lp(R),g∈Lq(R), with 1 ≤p,q≤ ∞ . Show that the function
F(x) =/integraltextx
0f(t)dtis defined and continuous for all x∈R, and that the
functionh(x) = (|x|+ 1)−aF(x)g(x) is in L1(R), the constant abeing
larger than 2 −1
p−1
q.
9.Letf∈L1(R). Show that the series
∞/summationdisplay
n=11√nf(x√n)
is convergent for almost all x∈R.
10.Letfbe in L1(R), andgbe a continuous periodic function with period 1.
Show that
lim
n→∞/integraldisplay∞
−∞f(x)g(nx)dx=/integraldisplay∞
−∞f(x)dx/integraldisplay1
0g(y)dy.
18 Bibliography
The following outlines the prerequesites for this article. (Although I’m not
suggesting that you must first know everything here before you read this article;
you could be learning these as you go along.)
First, you need a respectable first-year calculus course, dealing with limits
rigorously. The course I took used [ Spivak1 ], possibly the best math book ever.
You probably should be at least somewhat familiar with multi-dimensional
calculus, if only to have a motivation for the theorems we prove (e.g. Fu-
bini’s Theorem, Change of Variables). I learned multi-dimensional calculus
from [ Spivak2 ] and [ Munkres ]. As you’d expect, these are theoretical books,
and not very practical, but we will need a few elementary results that these
books prove.
Point-set topology is also introduced in the study of multi-dimensional cal-
culus. We will not need a deep understanding of that subject here, but just the
52
basic definitions and facts about open sets, closed sets, compact sets, continous
maps between topological spaces, and metric spaces. I don’t have particular
references for these, as it has become popular to learn topology with Moore’s
method (as I have done), where you are given lists of theorems that you are
supposed to prove alone.
The last book, [ Rosenthal ], (not a prerequesite) is what I mostly referred
to while writing up Section 8. It contains applications to probability of the
abstract measure stuff we do here, and it is not overly abstract. I recommend
it, and it’s cheap too.
I don’t mention any of the standard real analysis or measure theory books
here, since I don’t have them handy, and this text is supposed to supplant a fair
portion of these books anyway. But surely you can find references elsewhere.
References
[Spivak1] Michael Spivak, Calculus (3rd ed.). Publish or Perish, 1994; ISBN
0-914098-89-6.
[Spivak2] Michael Spivak, Calculus on Manifolds . Perseus, 1965; ISBN 0-
8053-9021-9.
[Munkres] James R. Munkres, Analysis on Manifolds . Westview Press, 1991;
ISBN 0-201-51035-9.
[Rosenthal] Jeffrey S. Rosenthal, A First Look at Rigorous Probability Theory .
World Scientific, 2000; ISBN 981-02-4303-0.
53