A_Simple_Introduction_to_Particle_Physic
PDF · 140 pages · 942.9 KB
Open PDF file
An arXiv paper (0810.3328, hep-th) by four Baylor University authors, not by Phil, kept in the particle physics folder. It reviews classical physics (Hamilton's principle, Noether's theorem, Lorentz transformations, electrodynamics), then group theory and Lie groups (SU(2), SU(3)), quantum field theory and quantization, gauge theory, and symmetry breaking. It ends with the Standard Model, its particles and forces.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
BU-HEPP-08-20
A Simple Introduction to Particle Physics
Part I - Foundations and the Standard Model
Matthew B. Robinson,1Karen R. Bland,2
Gerald B. Cleaver,3and Jay R. Dittmann4
Department of Physics, One Bear Place # 97316
Baylor University
Waco, TX 76798-7316
Abstract
This is the first of a series of papers in which we present a brief introduction to the rele-
vant mathematical and physical ideas that form the foundation of Particle Physics, in-
cluding Group Theory, Relativistic Quantum Mechanics, Quantum Field Theory and
Interactions, Abelian and Non-Abelian Gauge Theory, and the SU(3)
SU(2)
U(1)
Gauge Theory that describes our universe apart from gravity. Our approach, at first, is
an algebraic exposition of Gauge Theory and how the physics of our universe comes
out of Gauge Theory.
With an algebraic understanding of Gauge Theory and the relevant physics of the
Standard Model from this paper, in a subsequent paper we will “back up” and refor-
mulate Gauge Theory from a geometric foundation, showing how it connects to the
algebraic picture initially built in these notes.
Finally, we will introduce the basic ideas of String Theory, showing both the geometric
and algebraic correspondence with Gauge Theory as outlined in the first two parts.
These notes are not intended to be a comprehensive introduction to any of the ideas
contained in them. Their purpose is to introduce the “forest” rather than the “trees”.
The primary emphasis is on the algebraic/geometric/mathematical underpinnings
rather than the calculational/phenomenological details. Among the glaring omis-
sions are CPT theorems, evaluations of Feynman Diagrams, Renormalization, and
Anomalies. The topics were chosen according to the authors’ preferences and agenda.
These notes are intended for a student who has completed the standard undergrad-
uate physics and mathematics courses. The material in the first part is intended as a
review and is therefore cursory. Furthermore, these notes should not and will not in
any way take the place of the related courses, but rather provide a primer for detailed
courses in QFT, Gauge Theory, String Theory, etc., which will fill in the many gaps left
by this paper.
[email protected]
2karen [email protected]
3gerald [email protected]
[email protected]:0810.3328v1 [hep-th] 18 Oct 2008
Contents
1 Part I — Preliminary Concepts 5
1.1 Review of Classical Physics . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
1.1.1 Hamilton’s Principle . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
1.1.2 Noether’s Theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
1.1.3 Conservation of Energy . . . . . . . . . . . . . . . . . . . . . . . . . . 7
1.1.4 Lorentz Transformations . . . . . . . . . . . . . . . . . . . . . . . . . . 7
1.1.5 A More Detailed Look at Lorentz Transformations . . . . . . . . . . . 9
1.1.6 Classical Fields . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
1.1.7 Classical Electrodynamics . . . . . . . . . . . . . . . . . . . . . . . . . 11
1.1.8 Classical Electrodynamics Lagrangian . . . . . . . . . . . . . . . . . . 12
1.1.9 Gauge Transformations . . . . . . . . . . . . . . . . . . . . . . . . . . 13
1.2 References and Further Reading . . . . . . . . . . . . . . . . . . . . . . . . . 14
2 Part II — Algebraic Foundations 15
2.1 Introduction to Group Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
2.1.1 What is a Group? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
2.1.2 Finite Discrete Groups and Their Organization . . . . . . . . . . . . . 17
2.1.3 Group Actions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18
2.1.4 Representations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
2.1.5 Reducibility and Irreducibility — A Preview . . . . . . . . . . . . . . 23
2.1.6 Algebraic Definitions . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24
2.1.7 Reducibility Revisited . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
2.2 Introduction to Lie Groups . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33
2.2.1 Classification of Lie Groups . . . . . . . . . . . . . . . . . . . . . . . . 34
1
2.2.2 Generators . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38
2.2.3 Lie Algebras . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 40
2.2.4 The Adjoint Representation . . . . . . . . . . . . . . . . . . . . . . . . 42
2.2.5SO(2). . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43
2.2.6SO(3). . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43
2.2.7SU(2). . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 45
2.2.8SU(2)and Physical States . . . . . . . . . . . . . . . . . . . . . . . . . 45
2.2.9SU(2)forj=1
2. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 48
2.2.10SU(2)forj= 1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50
2.2.11SU(2)for Arbitrary j. . . . . . . . . . . . . . . . . . . . . . . . . . . . 50
2.2.12 Root Space . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51
2.2.13 Adjoint Representation of SU(2) . . . . . . . . . . . . . . . . . . . . . 57
2.2.14SU(2)for Arbitrary j::: Again . . . . . . . . . . . . . . . . . . . . . . 59
2.2.15SU(3). . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61
2.2.16 What is the Point of All of This? . . . . . . . . . . . . . . . . . . . . . 64
2.3 References and Further Reading . . . . . . . . . . . . . . . . . . . . . . . . . 65
3 Part III — Quantum Field Theory 66
3.1 A Primer to Quantization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 66
3.1.1 Quantum Fields . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 66
3.1.2 Spin-0 Fields . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 67
3.1.3 Why SU(2)for Spin? . . . . . . . . . . . . . . . . . . . . . . . . . . . . 68
3.1.4 Spin1
2Particles . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69
3.1.5 The Lorentz Group . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71
3.1.6 The Dirac Sea Interpretation of Antiparticles . . . . . . . . . . . . . . 73
3.1.7 The QFT Interpretation of Antiparticles . . . . . . . . . . . . . . . . . 74
2
3.1.8 Lagrangians for Scalars and Dirac Particles . . . . . . . . . . . . . . . 75
3.1.9 Conserved Currents . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76
3.1.10 The Dirac Equation with an Electromagnetic Field . . . . . . . . . . . 76
3.1.11 Gauging the Symmetry . . . . . . . . . . . . . . . . . . . . . . . . . . 77
3.2 Quantization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 81
3.2.1 Review of What Quantization Means . . . . . . . . . . . . . . . . . . 81
3.2.2 Canonical Quantization of Scalar Fields . . . . . . . . . . . . . . . . . 82
3.2.3 The Spin-Statistics Theorem . . . . . . . . . . . . . . . . . . . . . . . . 86
3.2.4 Left-Handed and Right-Handed Fields . . . . . . . . . . . . . . . . . 87
3.2.5 Canonical Quantization of Fermions . . . . . . . . . . . . . . . . . . . 89
3.2.6 Insufficiencies of Canonical Quantization . . . . . . . . . . . . . . . . 90
3.2.7 Path Integrals and Path Integral Quantization . . . . . . . . . . . . . 91
3.2.8 Interpretation of the Path Integral . . . . . . . . . . . . . . . . . . . . 93
3.2.9 Expectation Values . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 94
3.2.10 Path Integrals with Fields . . . . . . . . . . . . . . . . . . . . . . . . . 95
3.2.11 Interacting Scalar Fields and Feynman Diagrams . . . . . . . . . . . 98
3.2.12 Interacting Fermion Fields . . . . . . . . . . . . . . . . . . . . . . . . . 102
3.3 Final Ingredients . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103
3.3.1 Spontaneous Symmetry Breaking . . . . . . . . . . . . . . . . . . . . 104
3.3.2 Breaking Local Symmetries . . . . . . . . . . . . . . . . . . . . . . . . 106
3.3.3 Non-Abelian Gauge Theory . . . . . . . . . . . . . . . . . . . . . . . . 107
3.3.4 Representations of Gauge Groups . . . . . . . . . . . . . . . . . . . . 109
3.3.5 Symmetry Breaking Revisited . . . . . . . . . . . . . . . . . . . . . . . 110
3.3.6 Simple Examples of Symmetry Breaking . . . . . . . . . . . . . . . . 112
3.3.7 A More Complicated Example of Symmetry Breaking . . . . . . . . . 114
3
3.4 Particle Physics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 115
3.4.1 Introduction to the Standard Model . . . . . . . . . . . . . . . . . . . 115
3.4.2 The Gauge and Higgs Sector . . . . . . . . . . . . . . . . . . . . . . . 116
3.4.3 The Lepton Sector . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 120
3.4.4 The Quark Sector . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 124
3.5 References and Further Reading . . . . . . . . . . . . . . . . . . . . . . . . . 126
4 The Standard Model — A Summary 127
4.1 How Does All of This Relate to Real Life? . . . . . . . . . . . . . . . . . . . . 127
4.2 The Fundamental Forces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 128
4.3 Categorizing Particles . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 129
4.4 Elementary Particles . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 130
4.4.1 Elementary Fermions . . . . . . . . . . . . . . . . . . . . . . . . . . . 130
4.4.2 Elementary Bosons . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 132
4.5 Composite Particles . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 133
4.6 Visualizing It All . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 134
5 A Look Ahead 135
4
1 Part I — Preliminary Concepts
1.1 Review of Classical Physics
1.1.1 Hamilton’s Principle
Nearly all physics begins with what is called a Lagrangian for a particle, which is initially
defined as the kinetic energy minus the potential energy,
LT V
whereT=T(q;_q)andV=V(q). Then, the Action is defined as the integral of the
Lagrangian from an initial time to a final time,
SZtf
tidtL(q;_q)
It is important to realize that Sis a “functional” of the particle’s world-line in (q;_q)space,
not a function. This means that it depends on the entire path (q;_q), rather than a given
point on the path. The only fixed points on the path are q(ti),q(tf),_q(ti), and _q(tf). The
rest of the path is generally unconstrained, and the value of Sdepends on the entire path.
Hamilton’s Principle says that nature extremizes the path a particle will take in going
fromq(ti)at timetito position q(tf)at timetf. In other words, the path that extremizes
the action will be the path the particle will travel.
But, because Sis a functional, depending on the entire path in (q;_q)space rather than a
point, it cannot be extremized in the “Calculus I” sense of merely setting the derivative
equal to 0. Instead, we must find the path for which the action is “stationary”. This means
that the first-order term in the Taylor Expansion around that path will vanish, or S= 0
at that path.
To find this, consider some arbitrary path (q;_q). If it is a path that minimizes the action,
then we will have
0 =S=Ztf
tidtL(q;_q) =Ztf
tidtL(q+q;_q+_q) S
=Ztf
tidtL(q;_q) +Ztf
tidt
q@L
@q+_q@L
@_q
S
=Ztf
tidt
q@L
@q+@L
@_qd
dtq
Integrating the second term by parts, and taking the variation of qto be at 0 at tiandtf,
S=Ztf
tidt
q@L
@q qd
dt@L
@_q
=Ztf
tidtq@L
@q d
dt@L
@_q
= 0
5
The only way to guarantee this for an arbitrary variation qfrom the path (q;_q)is to
demand
d
dt@L
@_q @L
@q= 0
This equation is called the Euler-Lagrange equation, and it produces the equations of
motion of the particle.
The generalization to multiple coordinates qi(i= 1;:::;n ) is straightforward:
d
dt@L
@_qi @L
@qi= 0 (1.1)
1.1.2 Noether’s Theorem
Given a Lagrangian L=L(q;_q), consider making an infinitesimal transformation
q!q+q
whereis some infinitesimal constant. This transformation will give
L(q;_q)!L(q+q;_q+_q) =L(q;_q) +q@L
@q+_q@L
@_q
If the Euler-Lagrange equations of motion are satisfied, so that@L
@q=d
dt@L
@_q, then under
q!q+q,
L!L+q@L
@q+_q@L
@_q=L+qd
dt@L
@_q+@L
@_qd
dtq=L+d
dt@L
@_qq
So, underq!q+q, we haveL=d
dt @L
@_qq
. We define the Noether Current ,j, as
j@L
@_qq
Now, if we can find some transformation qthat leaves the action invariant, or in other
words such that S= 0, thendj
dt= 0, and therefore the current jis a constant in time. In
other words, jisconserved .
As a familiar example, consider a projectile, described by the Lagrangian
L=1
2m( _x2+ _y2) mgy (1.2)
This will be unchanged under the transformation x!x+, whereis any constant
(here,q= 1 in the above notation), because x!x+)_x!_x. So,j=@L
@_qq=m_xis
conserved. We recognize m_xas the momentum in the x-direction, which we expect to be
conserved by conservation of momentum.
So in summary, Noether’s Theorem merely says that whenever there is a continuous
symmetry in the action, there is a corresponding conserved quantity.
6
1.1.3 Conservation of Energy
Consider the quantity
dL
dt=d
dtL(q;_q) =@L
@qdq
dt+@L
@_qd_q
dt+@L
@t
BecauseLdoes not depend explicitly on time,@L
@t= 0, and therefore
dL
dt=@L
@q_q+@L
@_qq=d
dt@L
@_q
_q+@L
@_qq=d
dt@L
@_q_q
where we have used the Euler-Lagrange equation to get the second equality. So, we have
dL
dt=d
dt @L
@_q_q
, or
d
dt@L
@_q_q L
= 0 (1.3)
For a general non-relativistic system, L=T V, so@L
@_q=@T
@_qbecauseVis a function of q
only, and normally
T/_q2)@L
@_q_q= 2T
So,@L
@_q_q L= 2T (T V) =T+V=E, the total energy of the system, which is conserved
according to (1.3). We identify T+VHas the Hamiltonian , or total energy function,
of the system.
Furthermore, we define@L
@_qpto be the momentum of the system. Then, the relationship
between the Lagrangian and the Hamiltonian is the Legendre transformation
p_q L=H
1.1.4 Lorentz Transformations
Consider some event that occurs at spatial position (x;y;z )T, at timet. (The superscript
Tdenotes the transpose, so this is a column vector.) We arrange this event in a column
4-vector as (ct;x;y;z )T, wherecis the speed of light (the units of cgive each element the
same units). A more useful notation is to refer to this vector as a= (ct;x;y;z )T, where
= 0;1;2;3. This 4-vector, with the index raised, is called a “vector”, or a “contravariant
vector”. Then, we define the row vector a= ( ct;x;y;z ). This is called a “covector”, or
a “covariant vector”. In general, the sign of the 0thcomponent (the component in the first
position) changes when going from vector to covector.
7
There is something very deep going on here regarding the geometrical picture between
vectors and covectors, but we will not discuss it until the next paper in this series.
The dot product between two such vectors (a covector and vector) is then defined as the
product with one index raised and the other lowered. Whenever indices are contracted
in such a way, it is understood that they are to be summed over.1
ab=ab=a0b0+a1b1+a2b2+a3b3= a0b0+a1b1+a2b2+a3b3
Or, plugging in the spacetime notation from above, where
a= (ct1;x1;y1;z1)Tand b= (ct2;x2;y2;z2)T
we have
ab=ab= c2t1t2+x1x2+y1y2+z1z2
We can also discuss the differential version of this. If s= (ct;x;y;z ), thends2= c2dt2+
dx2+dy2+dz2.
In his theory of Special Relativity, Einstein postulated that all inertial reference frames
are equivalent, and that the speed of light is the same in all frames. To put this in more
mathematical terms, if observers in different inertial frames 1and2each see an event,
they will see, respectively,
ds2
1= c2dt2
1+dx2
1+dy2
1+dz2
1
ds2
2= c2dt2
2+dx2
2+dy2
2+dz2
2
We then demand that ds2
1=ds2
2. To do this, we must find a modification of the standard
Galilean transformations that will leave ds2unchanged. The derivation for the correct
transformations can be found in any introductory or modern physics text, so we merely
quote the result. If we assume that frame 2is moving only in the z-direction with respect
to frame 1(and that their x,y, andzaxes are aligned), then we find that the transforma-
tions are
t2=
(ct1 z1)
z2=
(z1 ct1) (1.4)
where=v
cand
=1p
1 2. These transformations, which preserve ds2when transform-
ing one frame to another, are called Lorentz Transformations .
Discussions of the implications of these transformations, including time dilation, length
contraction, and the relationship between energy and mass can be found in most intro-
ductory texts. You are encouraged to review the material if you are not familiar with
it.
1because we are summing over components, we can write aborab— they mean the same thing
8
1.1.5 A More Detailed Look at Lorentz Transformations
As we have seen, we have a quantity ds2= c2dt2+dx2+dy2+dz2, which does not
change under transformations (1.4). Thinking of physical ideas this way, in terms of
“what doesn’t change when something else changes”, will prove to be an extraordinarily
powerful approach. In order to understand Special Relativity in such a way, we begin
with a simpler example.
Consider a spatial rotation around, say, the z-axis (or, equivalently, mixing the xandy
coordinates). Such a transformation is called an Euler Transformation , and takes the
form
t0=t
x0=xcos+ysin
y0= xsin+ycos
z0=z (1.5)
whereis the angle of rotation, called the Euler Angle . We can simultaneously express a
Lorentz transformation as a sort of “rotation” that mixes a spatial dimension and a time
dimension, as follows (these transformations are equivalent to (1.4):
t0=tcosh xsinh
x0= tsinh+xcosh
y0=y
z0=z (1.6)
whereis defined by the relationship = tan.
We denote a transformation mixing two spatial dimensions simply a Rotation , whereas
a transformation mixing a spatial dimension and a time dimension is a Boost . Any two
frames whose origins coincide at t=t0= 0 can be transformed into each other through
some combination of rotations and boosts.
To rephrase this in more precise language, given a 4-vector x, it will be related to the
equivalent 4-vector in another frame, x0, by some matrix L, according to x0=L
x
(where the summation convention discussed earlier is in effect for the repeated index).
We also introduce what is called the Metric matrix,
==0
BB@ 1 0 0 0
0 1 0 0
0 0 1 0
0 0 0 11
CCA
In general, () 1.
9
Using the metric, the dot product of any 4-vector x= (ct;x;y;z )Tcan be easily written
asx2=xx=xx= c2t2+x2+y2+z2. In general, a Lorentz transformation can be
defined as a matrix L
(including boosts and rotations) that leaves xxunchanged.
For example, a scalar, or an object with no uncontracted indices, like orxx, is simply
invariant under Lorentz transformations ( !,xx!xx).
A vector, or an object with only one uncontracted index, like xorab
, transforms ac-
cording to x0=L
x, or(ab
)0=L
(ab
).
Now, consider the dot product x2=xx=xx. Ifx2is invariant, then x02=x2)
x0x0=L
L
xxdemands that L
L
=. So, the constraint for Lorentz
transformations is that they are the set of all matrices such that
L
L
=
We take this to be the defining constraint for a Lorentz transformation.
1.1.6 Classical Fields
When deriving the Euler-Lagrange equations, we started with an action Swhich was an
integral over time only ( SR
dtL). If we are eventually interested in a relativistically
acceptable theory, this is obviously no good because it treats time and space differently
(the action is an integral over time but not over space).
So, let’s consider an action defined not in terms of the Lagrangian, but of the “Lagrangian
per unit volume”, or the Lagrangian Density L. The Lagrangian will naturally be the
integral ofLover all space, L=R
dnxL. The integral is in n-dimensions, so dnxmeans
dx1dx2dx2dxn.
Now, the action will be S=R
dtL=R
dtdnxL. In the normal 1+3 dimensional Minkowski
spacetime we live in, this will be S=R
dtd3xL=R
d4xL.
Before,Ldepended not on t, but on the path q(t),_q(t). In a similar sense, Lwill not
depend on xandt, but on what we will refer to as Fields ,(x;t) =(x), which exist in
spacetime.
Following a nearly identical argument as the one leading to (1.1), we get the relativistic
field generalization
@@L
@(@i)
@L
@i= 0
for multiple fields i(i= 1;:::;n ).
10
Noether’s Theorem says that, for !+, we have a current j@L
@(@), and if!
+leavesL= 0, then@j= 0) @j0
@t+rj= 0, wherej0is the Charge Density ,
andjis the Current Density . The total charge will naturally be QR
allspaced3xj0.
Finally, we also have a Hamiltonian Density and momentum
H @L
@__ L (1.7)
@L
@_(1.8)
One final comment for this section. For the remainder of these notes, we will ultimately
be seeking a relativistic field theory, and therefore we will never make use of Lagrangians.
We will always use Lagrangian densities. We will always use the notation Linstead ofL,
but we will refer to the Lagrangian densities simply as Lagrangians. We drop the word
“densities” for brevity, and because there will never be ambiguity.
1.1.7 Classical Electrodynamics
We choose our units so that c=0=0= 1. So, the magnitude of the force between two
chargesq1andq2isF=q1q2
4r2. In these units, Maxwell’s equations are
rE= (1.9)
r B @E
@t=J (1.10)
rB= 0 (1.11)
r E+@B
@t= 0 (1.12)
If we define the Potential 4-vectorA= (;A), then we can define B=r Aand
E= r @A
@t. Writing Band Ethis way will automatically solve the homogenous
Maxwell equations, (1.11) and (1.12).
Then, we define the totally antisymmetric Electromagnetic Field Strength Tensor Fas
F@A @A=0
BB@0 Ex Ey Ez
Ex0 BzBy
EyBz 0 Bx
Ez ByBx 01
CCA
We define the 4-vector current as J= (;J). It is straightforward, though tedious, to
11
show that
@F+@F+@F= 0)rB= 0 and r E+@B
@t= 0
@F=J)rE=and r B @E
@t=J
1.1.8 Classical Electrodynamics Lagrangian
Bringing together the ideas of the previous sections, we now want to construct a La-
grangian density Lwhich will, via Hamilton’s Principle, produce Maxwell’s equations.
First, we know that Lmust be a scalar (no uncontracted indices). From our intuition with
“Physics I” type Lagrangians, we know that kinetic terms are quadratic in the derivatives
of the fundamental coordinates (i.e.1
2m_x2=1
2m(dx
dt)(dx
dt)). The natural choice is to take
Aas the fundamental field. It turns out that the correct choice is
LEM= 1
4FF JA (1.13)
(note that the F2term is quadratic in @A). So,
S=Z
d4x
1
4FF JA
(1.14)
Taking the variation of (1.14) with respect to A,
S=Z
d4x
1
4FF 1
4FF JA
=Z
d4x
1
2FF JA
=Z
d4x
1
2F(@A @A) JA
=Z
d4x
F@A JA
Integrating the first term by parts, and choosing boundary conditions so that Avanishes
at the boundaries,
=Z
d4x
@FA JA
=Z
d4x
@F J
A
So, to have S= 0, we must have @F=J, and if this is written out one component
at a time, it will give exactly the inhomogenous Maxwell equations (1.9) and (1.10). And
12
as we already pointed out, the homogenous Maxwell equations become identities when
written in terms of A.
As a brief note, the way we have chosen to write equation (1.13), in terms of a “potential”
A, and the somewhat mysterious antisymmetric “field strength” F, is indicative of an
extremely deep and very general mathematical structure that goes well beyond classical
electrodynamics. We will see this structure unfold as we proceed through these notes.
We just want to mention now that this is not merely a clever way of writing electric and
magnetic fields, but a specific example of a general theory.
1.1.9 Gauge Transformations
Gauge Transformations are usually discussed toward the end of an undergraduate
course on E&M. Students are typically told that they are extremely important, but the
reason why is not obvious. We will briefly introduce them here, and while their signifi-
cance may still not be transparent, we will return to them several times throughout these
notes.
Given some specific potential A, we can find the field strength action as in (1.14). How-
ever,Adoes not uniquely specify the action. We can take any arbitrary function (x),
and the action will be invariant under the transformation
A!A0=A+@ (1.15)
or
A!A0= ( @
@t;A+r)
Under this transformation, we have
F0=@A0 @A0=@(A+@) @(A+@)
=@A @A+@@ @@
=F(1.16)
So,F0=F.
Furthermore, JA!JA+J@. Integrating the second term by parts with the usual
boundary conditions,
Z
d4xJ@= Z
d4x(@J)
But, according to Maxwell’s equations, @J=@@F0becauseFis totally anti-
symmetric. So, both FandJ@are invariant under (1.15), and therefore the action of
Sis invariant under (1.15).
13
While the importance of gauge transformations may not be obvious at this point, it will
become perhaps the most important idea in particle physics. As a note before moving on,
recall previously when we mentioned the idea of “what doesn’t change when something
else changes” when talking about Lorentz transformations. A gauge transformation is
exactly this (in a different context): the fundamental fields are changed by , but the
equations which govern the physics are unchanged.
In the next section, we provide the mathematical tools to understand why this idea is so
important.
1.2 References and Further Reading
The material in this section can be found in nearly any introductory text on Classical
Mechanics, Classical Electrodynamics, and Relativity. The primary sources for these notes
are [3], [12], and [13].
For further reading, we recommend [6], [18], [19], [22], [33], and [34].
14
2 Part II — Algebraic Foundations
2.1 Introduction to Group Theory
There are several symbols in this section which may not be familiar. We therefore provide
a summary of them for reference.
N=f0;1;2;3;:::g
Z=f0;1;2;3;:::g
Q= Rational Numbers
R= Real Numbers
C= Complex Numbers
Zn=Zmod n)is read \implies"
i is read \if and only if"
8is read \for every"
9is read \there exists"
2is read \in"
3is read \such that"_ = is \represented by"
is \subset of"
is \dened as"
Now that we have reviewed the primary ideas of classical physics, we are almost ready to
start talking about particle physics. However, there is a bit of mathematical “machinery”
we will need first. Namely, Group Theory .
Group theory is, in short, the mathematics of symmetry. We are going to begin talking
about what will seem to be extremely abstract ideas, but eventually we will explain how
those ideas relate to physics. As a preface of what is to come, the most foundational idea
here is, as we said before, “what doesn’t change when something else changes”. A group
is a precise and well-defined way of specifying the thing or things that change.
2.1.1 What is a Group?
To begin with, we define the notion of a Group . This definition may seem cryptic, but it
will be explained in the paragraphs that follow.
A group, denoted (G;?), is a set of objects, denoted G, and some operation on those objects, de-
noted?, subject to the following:
1.8g1; g22G,g1?g22Galso. (closure)
2.8g1;g2;g32G, it must be true that (g1?g2)?g3=g1?(g2?g3). (associativity)
3.9g2G, denotede,38gi2G; e?gi=gi?e=gi. (identity)
4.8g2G;9h2G3h?g =g?h =e, (soh=g 1). (inverse)
Now we explain what this means. By “objects” we literally mean anything. We could be
talking about ZorR, or we could be talking about a set of Easter eggs all painted different
colors.
15
The meaning of “some operation”, which we are calling ?, can literally be anything you
can do to those objects. A formal definition of what ?means could be given, but it will be
easier to understand with examples.
Note: The definition of a group doesn’t demand that gi?gj=gj?gi. This is a very
important point, but we will discuss it in more detail later. We mention it now so it is not
new later.
Example 1: (G;?) = (Z;+)
Consider the set Gto beZ, and the operation to be ?= +, or simply addition.
We first check closure. If you take any two elements of Zand add them to-
gether, is the result in Z? In other words, if a;b2Z, isa+b2Z? Obviously
the answer is yes; the sum of two integers is an integer, so closure is met.
Now we check associativity. If a;b;c2Z, it is trivially true that a+ (b+c) =
(a+b) +c. So, associativity is met.
Now we check identity. Is there an element e2Zsuch that when you add e
to any other integer, you get that same integer? Clearly the integer 0 satisfies
this. So, identity is met.
Finally, is there an inverse? For any integer a2Z, will there be another integer
b2Zsuch thata+b=e= 0? Again, this is obvious, a 1= ain this case. So,
inverse is met.
So,(G;?) = (Z;+)is a group.
Example 2: (G;?) = (R;+)
Obviously, any two real numbers added together is also a real number.
Associativity will hold (of course).
The identity is again 0.
And finally, once again, awill be the inverse of any a2R.
Example 3: (G;?) = (R;)(multiplication)
Closer is met; two real numbers multiplied together give a real number.
Associativity obviously holds.
Identity also holds. Any real number a2R, when multipled by 1 is a.
Inverse, on the other hand, is trickier. For any real number, is there another
real number you can multiply by it to get 1? The instinctive choice is a 1=1
a.
But, this doesn’t quite work because of a= 0. This is the only exception, but
because there’s an exception, (R;)is not a group.
Note: If we take the set R f0ginstead of R, then (R f0g;)is a group.
16
Example 4: (G;?) = (f1g;)
This is the set with only the element 1, and the operation is normal multiplica-
tion. This is indeed a group, but it is extremely uninteresting, and is called the
Trivial Group .
Example 5: (G;?) = (Z3;+)
This is the set of integers mod 3, containing only the elements 0,1, and 2
(3 mod 3 is 0, 4 mod 3 is 1, 5 mod 3 is 2, etc.)
You can check yourself that this is a group.
2.1.2 Finite Discrete Groups and Their Organization
From the examples above, several things should be apparent about groups. One is that
there can be any number of objects in a group. We have a special name for the number of
objects in the group’s set. TheOrder of a group is the number of elements in it.
The order of (Z;+)is infinite (there are an infinite number of integers), as is the order of
(R;+)and(R f0g;). But, the order of (f1g;)is 1, and the order of (Z3;+)is 3.
If the order of a group is finite, the group is said to be Finite . Otherwise it is Infinite .
It is also clear that the elements of groups may be Discrete , or they may be Continuous .
For example, (Z;+),(f1g;), and (Z3;+)are all discrete, while (R;+)and(R f0g;)are
both continuous.
Now that we understand what a discrete finite group is, we can talk about how to orga-
nize one. Namely, we use what is called a Multiplication Table . A multiplication table is
a way of organizing the elements of a group as follows:
(G;?)eg1g2
ee?ee?g 1e?g 2
g1g1?eg1?g1g1?g2
g2g2?eg2?g1g2?g2
...............
We state the following property of multiplication tables without proof. A multiplication
table must contain every element of the group exactly one time in every row and every column . A
few minutes thought should convince you that this is necessary to ensure that the defini-
tion of a group is satisfied.
17
As an example, we will draw a multiplication table for the group of order 2. We won’t
look at specific numbers, but rather call the elements g1andg2. We begin as follows:
(G;?)eg1
e ??
g1 ??
Three of these are easy to fill in from the identity:
(G;?)eg1
eeg1
g1g1?
And because we know that every element must appear exactly once, the final question
mark must be e. So, there is only one possible group of order 2.
We will consider a few more examples, but we stress at this point that the temptation to
plug in numbers should be avoided. Groups are abstract things, and you should try to
think of them in terms of the abstract properties, not in terms of actual numbers.
We can proceed with the multiplication table for the group of order 3. You will find that,
once again, there is only one option. (Doing this is instructive and it would be helpful to
work this out yourself.)
(G;?)eg1g2
eeg1g2
g1g1g2e
g2g2eg1
You are encouraged to work out the possibilities for groups of order 4. (Hint: there are 4
possibilities.)
2.1.3 Group Actions
So far we have only considered elements of groups and how they relate to each other. The
point has been that a particular group represents nothing more than a structure. There
are a set of things, and they relate to each other in a particular way. Now, however, we
want to consider an extremely simple version of how this relates to nature.
18
Example 6
Consider three Easter eggs, all painted different colors (red, orange, and yel-
low), which we denote R, O, and Y. Now, assume they have been put into a
row in the order (ROY). If we want to keep them lined up, not take any eggs
away, and not add any eggs, what we can we do to them? We can do any of
the following:
1. Letebe doing nothing to the set, so e(ROY ) = (ROY ).
2. Letg1be a cyclic permutation of the three, g1(ROY ) = (OYR )
3. Letg2be a cyclic permutation in the other direction, g2(ROY ) = (YRO )
4. Letg3be swapping the first and second, g3(ROY ) = (ORY )
5. Letg4be swapping the first and third, g4(ROY ) = (YOR )
6. Letg5be swapping the second and third, g5(ROY ) = (RYO )
You will find that these 6 elements are closed, there is an identity, and each has
an inverse.2So, we can draw a multiplication table (you are strongly encour-
aged to write at least part of this out on your own):
(G;?)eg1g2g3g4g5
eeg1g2g3g4g5
g1g1g2eg5g3g4
g2g2eg1g4g5g3
g3g3g4g5eg1g2
g4g4g5g3g2eg1
g5g5g3g4g1g2e
There is something interesting about this group. Notice that g3?g1=g4, whereasg1?g3=
g5. So, we have the surprising result that in this group it is not necessarily true that
gi?gj=gj?gi.
This leads to a new way of classifying groups. We say a group is Abelian ifgi?gj=gj?gi
8gi;gj2G. If a group is not Abelian, it is Non-Abelian .
Another term commonly used is Commute . Ifgi?gj=gj?gi, then we say that giand
gjcommute. So, an Abelian group is Commutative , whereas a Non-Abelian group is
Non-Commutative .
2We should be very careful to draw a distinction between the elements of the group and the objects the
group acts on. The objects in this example are the eggs, and the permutations are the results of the group
action. Neither the eggs nor the permutations of the eggs are the elements of the group. The elements of
the group are abstract objects which we are assigning to some operation on the eggs, resulting in a new
permutation
19
The Easter egg group of order 6 above is an example of a very important type of group. It
is denoted S3, and is called the Symmetric Group . It is the group that takes three objects
to all permutations of those three objects.
The more general group of this type is Sn, the group that takes nobjects to all permu-
tations of those objects. You can convince yourself that Snwill always have order n!(n
factorial).
The idea above with the 3 eggs is that S3is the group , while the eggs are the objects that
the group acts on . The particular way an element of S3changes the eggs around is called
theGroup Action of that element. And each element of S3will move the eggs around
while leaving them lined up. This ties in to our overarching concept of “what doesn’t
change when something else changes”. The fact that there are 3 eggs with 3 particular
colors lined up doesn’t change. The order they appear in does.
2.1.4 Representations
We suggested above that you think of groups as purely abstract things rather than trying
to plug in actual numbers. Now, however, we want to talk about how to see groups,
or the elements of groups, in terms of specific numbers. But, we will do this in a very
systematic way. The name for a specific set of numbers or objects that form a group is
aRepresentation . The remainder of this section (and the next) will primarily be about
group representations.
We already discussed a few simple representations when we discussed (Z;+),(R f0g;),
and(Z3;+). Let’s focus on (Z3;+)for a moment (the integers mod 3, where e= 0,g1= 1,
g2= 2, with addition). Notice that we could alternatively define e= 1,g1=e2i
3, andg2=
e4i
3, and let?be multiplication. So, in the “representation” with (0;1;2)and addition, we
had for example
g1?g2= (1 + 2) mod 3 = 3 mod 3 = 0 =e
whereas now with the multiplicative representation we have
g1?g2=e2i
3e4i
3=e2i=e0= 1 =e
So the structure of the group is preserved in both representations.
We have two completely different representations of the same group. This idea of differ-
ent ways of expressing the same group is of extreme importance, and we will be using it
throughout the remainder of these notes.
We now see a more rigorous way of coming up with representations of a particular group.
We begin by introducing some notation. For a group (G;?)with elements g1;g2;:::, we
call the Representation of that group D(G), so that the elements of GareD(e),D(g1),
20
D(g2)(where each D(gi)is a matrix of some dimension). We then choose ?to be matrix
multiplication. So, D(gi)D(gj) =D(gi?gj).
It may not seem that we have done anything profound at this point, but we most defi-
nitely have. Remember above that we encouraged seeing groups as abstract things, rather
than in terms of specific numbers. This is because a group is fundamentally an abstract
object. A group is not a specific set of numbers, but rather a set of abstract objects with a
well-defined structure telling you how those elements relate to each other.
And the beauty of a representation Dis that, via normal matrix multiplication, we have
a sort of “lens”, made of familiar things (like numbers, matrices, or Easter eggs), through
which we can see into this abstract world. And because D(gi)D(gj) =D(gi?gj), we
aren’t losing any of the structure of the abstract group by using a representation.
So now that we have some notation, we can develop a formalism to figure out exactly
whatDis for an arbitrary group.
We will use Dirac vector notation, where the column vector
v=0
BBB@v1
v2
v3
...1
CCCA=jvi
and the row vector
vT=
v1v2v3
=hvj
So, the dot product between two vectors is
vu=
v1v2v30
BBB@u1
u2
u3
...1
CCCA=v1u1+v2u2+v3u3+hvjui
Now, we proceed by relating each element of a finite discrete group to one of the standard
orthonormal unit vectors:
e!jei=j^e1ig1!jg1i=j^e2ig2!jg2i=j^e3i
And we define the way an element in a representation D(G)acts on these vectors to be
D(gi)jgji=jgi?gji
Now, we can build our representation. We will (from now on unless otherwise stated)
represent the elements of a group Gusing matrices of various sizes, and the group opera-
tion?will be standard matrix multiplication. The specific matrices that represent a given
21
elementgkof our group will be given by
[D(gk)]ij=hgijD(gk)jgji (2.1)
As an example, consider again the group of order 2 (we wrote out the multiplication table
above on page 18). First, we find the matrix representation of the identity, [D(e)]ij,
[D(e)]11=hejD(e)jei=heje?ei=hejei= 1
[D(e)]12=hejD(e)jg1i=heje?g 1i=hejg1i= 0
[D(e)]21=hg1jD(e)jei=hg1je?ei=hg1jei= 0
[D(e)]22=hg1jD(e)jg1i=hg1je?g 1i=hg1jg1i= 1
So, the matrix representation of the identity is D(e) _ =1 0
0 1
. It shouldn’t be surprising
that the identity element is represented by the identity matrix.
Next we find the representation of D(g1):
[D(g1)]11=hejD(g1)jei=hejg1?ei=hejg1i= 0
[D(g1)]12=hejD(g1)jg1i=hejg1?g1i=hejei= 1
[D(g1)]21=hg1jD(g1)jei=hg1jg1?ei=hg1jg1i= 1
[D(g1)]22=hg1jD(g1)jg1i=hg1jg1?g1i=hg1jei= 0
So, the matrix representation of g1isD(g1) _ =0 1
1 0
. It is straightforward to check that
this is a true representation,
e?e =1 0
0 11 0
0 1
=1 0
0 1
=eX
e?g 1=1 0
0 10 1
1 0
=0 1
1 0
=g1X
g1?e=0 1
1 01 0
0 1
=0 1
1 0
=g1X
g1?g1=0 1
1 00 1
1 0
=1 0
0 1
=eX
Instead of considering the next obvious example, the group of order 3, consider the group
S3from above (the multiplication table is on page 19). The identity representation D(e)is
easy — it is just the 66identity matrix. We encourage you to work out the representation
ofD(g1)on your own, and check to see that it is
D(g1) _ =0
BBBBBB@0 0 1 0 0 0
1 0 0 0 0 0
0 1 0 0 0 0
0 0 0 0 1 0
0 0 0 0 0 1
0 0 0 1 0 01
CCCCCCA(2.2)
22
All 6 matrices can be found this way, and multiplying them out will confirm that they do
indeed satisfy the group structure of S3.
2.1.5 Reducibility and Irreducibility — A Preview
You have probably noticed that equation (2.1) will always produce a set of nnma-
trices, where nis the order of the group. There is actually a name for this particular
representation. Thennmatrix representation of a group of order nis called the Regular
Representation . More generally, themmmatrix representation of a group (of any order) is
called the m-dimensional representation.
But, as we have seen, there is more than one representation for a given group (in fact,
there are an infinite number of representations).
One thing we can immediately see is that any group that is Non-Abelian cannot have a
11matrix representation. This is because scalars ( 11matrices) always commute,
whereas matrices in general do not.
We saw above in equation (2.2) that we can represent the group Snbyn!n!matrices.
Or, more generally, we can represent any group using mmmatrices, were mequals
order(G). This is the regular representation. But it turns out that it is usually possible to
find representations that are “smaller” than the regular representation.
To pursue how this might be done, note that we are working with matrix representations
of groups. In other words, we are representing groups in linear spaces . We will therefore
be using a great deal of linear algebra to find smaller representations. This process, of
finding a smaller representation, is called Reducing a representation. Given an arbitrary
representation of some group, the first question that must be asked is “is there a smaller
representation?” If the answer is yes, then the representation is said to be Reducible . If
the answer is no, then it is Irreducible .
Before we dive into the more rigorous approach to reducibility and irreducibility, let’s
consider a more intuitive example, using S3. In fact, we’ll stick with our three painted
Easter eggs, R,O, andY:
1.e(ROY ) = (ROY )
2.g1(ROY ) = (OYR )
3.g2(ROY ) = (YRO )
4.g3(ROY ) = (ORY )
5.g4(ROY ) = (YOR )
6.g5(ROY ) = (RYO )
23
We will represent the set of eggs by a column vector jEi=0
@R
O
Y1
A.
Now, by inspection, what matrix would do to jEiwhatg1does to (ROY )? In other words,
how can we fill in the ?’s in
0
@? ? ?
? ? ?
? ? ?1
A0
@R
O
Y1
A=0
@O
Y
R1
A
to make the equality hold? A few moments thought will show that the appropriate matrix
is
0
@0 1 0
0 0 1
1 0 01
A0
@R
O
Y1
A=0
@O
Y
R1
A
Continuing this reasoning, we can see that the rest of the matrices are
D(e) _ =0
@1 0 0
0 1 0
0 0 11
A; D (g1) _ =0
@0 1 0
0 0 1
1 0 01
A; D (g2) _ =0
@0 0 1
1 0 0
0 1 01
A
D(g3) _ =0
@0 1 0
1 0 0
0 0 11
A; D (g4) _ =0
@0 0 1
0 1 0
1 0 01
A; D (g5) _ =0
@1 0 0
0 0 1
0 1 01
A
You can do the matrix multiplication to convince yourself that this is in fact a representa-
tion ofS3.
So, in equation (2.2), we had a 66matrix representation. Here, we have a new repre-
sentation of consisting of 33matrices. We have therefore “reduced” the representation.
In the next section, we will look at more mathematically rigorous ways of reducing rep-
resentations.
2.1.6 Algebraic Definitions
Before moving on, we must spend this section learning the definitions of several terms
which are used in group theory.
IfHis a subset of G, denotedHG, such that the elements of Hform a group, then we say that
Hforms a Subgroup ofG. We make this more precise with examples.
24
Example 7
Consider (as usual) the group S3, with the elements labeled as before:
1.g0(ROY ) = (ROY )
2.g1(ROY ) = (OYR )
3.g2(ROY ) = (YRO )
4.g3(ROY ) = (ORY )
5.g4(ROY ) = (YOR )
6.g5(ROY ) = (RYO )
(where we are relabeling g0efor later convenience). The multiplication
table is given on page 19.
Notice thatfg0;g1;g2gform a subgroup. You can see this by noticing that the
upper left 9 boxes in the multiplication table (the g0;g1;g2rows and columns)
all have only g0’s,g1’s, andg2’s. So, here is a group of order 3 contained in S3.
Example 8
Consider the subset of S3consisting of g0andg3only. Both g0andg3are their
own inverses, so the identity exists, and the group is closed. Therefore, we can
say thatfg0;g3gS3is a subgroup of S3.
In fact, if you write out the multiplication table for g0andg3only, you will
see that it is exactly equivalent to the group of order 2 considered above. This
means that we can say that S3contains the group of order 2 (and we know from
last time that there is only one such group, though there are an infinite number
of representations of it). The way we understand this is that the abstract entity
S3, of which there is only one, contains the group of order 2, of which there
is only one. However, the representations ofS3, of which there are an infinite
number, will each contain the group of order 2 (of which there are also an
infinite number of representations).
Example 9
Notice that the sets fg0;g3g,fg0;g4g, andfg0;g5g(allS3), are all the same as
the group of order 2. This means that S3actually contains exactly three copies
of the group of order 2 in addition to the single copy of the group of order 3.
Again, this is speaking in terms of the abstract entity S3. We can see this
through the “lens” of representation by the fact that anyrepresentation of S3
will contain three different copies of the group of order 2.
25
Example 10
As a final example of subgroups, there are two subgroups of anygroup, no
matter what the group. One is the subgroup consisting of only the identity,
fg0gG. All groups contain this, but it is never very interesting.
Secondly,8G,GG, and therefore Gis always a subgroup of itself. We call
these subgroups the “trivial” subgroups.
We now introduce another important definition.
IfGis a group, and His a subgroup of G(HG), then
The setgH=fg?hjh2Hgis called the Left Coset ofHinG
The setHg=fh?gjh2Hgis called the Right Coset ofHinG
There is a right (or left) coset for each element g2G, though they are not necessarily all
unique. This definition should be understood as follows; a coset is a setconsisting of the
elements of Hall multiplied on the right (or left) by some element of G.
Example 11
For the subgroup H=fg0;g1gS3discussed above, the left cosets are
g0fg0;g1g=fg0?g0;g0?g1g=fg0;g1g
g1fg0;g1g=fg1?g0;g1?g1g=fg1;g2g
g2fg0;g1g=fg2?g0;g2?g1g=fg2;g0g
g3fg0;g1g=fg3?g0;g3?g1g=fg3;g4g
g4fg0;g1g=fg4?g0;g4?g1g=fg4;g5g
g5fg0;g1g=fg5?g0;g5?g1g=fg5;g3g
So, the left cosets of fg0;g1ginS3arefg0;g1g,fg1;g2g,fg2;g0g,fg3;g4g,fg4;g5g,
andfg5;g3g.
We can now understand the following definition. His aNormal Subgroup ofGif8h2H,
g 1?h?g2H. Or, in other words, if Hdenotes the subgroup, it is a normal subgroup if
gH=Hg, which says that the left and right cosets are all equal.
As a comment, saying gHandHgare equal doesn’t mean that each individual element
in the coset gHis equal to the corresponding element in Hg, but rather that the two
cosets contain the same elements, regardless of their order. For example, if we had the
cosetsfgi;gj;gkgandfgj;gk;gig, they would be equal because they contain the same three
elements.
26
This definition means that if you take a subgroup Hof a group G, and you multiply the
entire set on the left by some element of g2G, the resulting set will contain the exact same
elements it would if you had multiplied on the right by the same element g2G. Here is
an example to illustrate.
Example 12
Consider the order 2 subgroup fg0;g3gS3. Multiplying on the left by, say,
g4, gives
g4?fg0;g3g=fg4?g0;g4?g3g=fg4;g2g
And multiplying on the right by g4givs
fg0;g3g?g4=fg0?g4;g3?g4g=fg4;g1g
So, because the final sets do not contain the same elements, fg4;g2g6=fg4;g1g,
we conclude that the subgroup fg0;g3gisnota normal subgroup of S3.
Example 13
Above, we found that fg0;g1;g2gS3is a subgroup of order 3 in S3. To use a
familiar label, remember that we previously called the group of order 3 (Z3;+).
So, dropping the `+0, we refer to the group of order 3 as Z3. Is this subgroup
normal? We leave it to you to show that it is.
Example 14
Consider the group of integers under addition, (Z;+). And, consider the sub-
groupZevenZ, the even integers under addition (we leave it to you to show
that this is indeed a group).
Now, take some odd integer noddand act on the left:
nodd+Zeven=fnodd+ 0;nodd2;nodd4;:::g
and then on the right:
Zeven+nodd=f0 +nodd;2 +nodd;4 +nodd;:::g
Notice that the final sets are the same (because addition is commutative). So,
ZevenZis a normal subgroup.
With a little thought, you can convince yourself that allsubgroups of Abelian groups are
normal.
IfGis a group and HGis normal, then the Factor Group ofHinG, denotedG=H (read “G
mod H”), is the group with elements in the set G=HfgHjg2Gg. The group operation ?is
understood to be
(giH)?(gjH) = (gi?gj)H
27
Example 15
Consider again Zeven. Notice that we can call Zeven = 2Zbecause 2Z=
2f0;1;23;:::g=f0;2;4;:::g=Zeven. We know that 2ZZis normal,
so we can build the factor group Z=2Zas
Z=2Z=f0 + 2Z;1 + 2Z;2 + 2Z;:::g
But, notice that
neven+ 2Z=Zeven
nodd+ 2Z=Zodd
So, the group Z=2Zonly has 2 elements; the set of all even integers, and the set
of all odd integers. And we know from before that there is only one group of
order 2, which we denote Z2. So, we have found that Z=2Z=Z2.
You can also convince yourself of the more general result
Z=nZ=Zn
Example 16
Finally, we consider the factor groups G=G andG=e.
G=G — The setG=fg0;g1;g2;:::gwill be the same coset for any element
ofGmultiplied by it. Therefore this factor group consists of only one
element, and therefore G=G =e, the trivial group.
G=e — The setfegwill be a unique coset for any element of G, and there-
foreG=e=G.
Something that might help you understand factor groups better is this: the factor group
G=H is the group that is left over when everything in His “collapsed” to the identity
element. Think about the above examples in terms of this picture.
IfGandHare both groups (not necessarily related in any way), then we can form the Product
Group , denotedKG
H, where an arbitrary element of Kis(gi;hj). If the group operation
ofGis?G, and the group operation of His?H, then two elements of Kare multiplied according
to the rule
(gi;hj)?K(gk;hl)(gi?Ggk;hj?Hhl)
28
2.1.7 Reducibility Revisited
Now that we understand subgroups, cosets, normal subgroups, and factor groups, we
can begin a more formal discussion of reducing representations. Recall that in deriving
equation (2.1), we made the designation
g0!j^e1ig1!j^e2ig2!j^e3i etc.
This was used to create an order( G)-dimensional Euclidian space which, while not having
any “physical” meaning, and while obviously not possessing any structure similar to the
group, was and will continue to be of great use to us.
We have an n-dimensional space spanned by the orthonormal vectors jg0i;jg1i;:::;jgn 1i,
whereg0is understood to always refer to the identity element. This brings us to the first
definition of this section. For a group G=fg0;g1;g2;:::g, we call the Algebra ofGthe set
C[G]n 1X
i=0cijgiici2C8i
In other words, C[G]is the set of all possible linear combinations of the vectors jgiiwith
complex coefficients.
We could have defined the algebra over ZorR, but we used Cfor generality at this point.
Addition of two elements of C[G]is merely normal addition of linear combinations,
n 1X
i=0cijgii+n 1X
i=0dijgii=n 1X
i=0(ci+di)jgii
This definition amounts to saying that, in the n-dimensional Euclidian space we have cre-
ated, withn= order(G), you can choose any point in the space with complex coefficients,
and this will correspond to a particular linear combination of elements of G.
Now that we have defined an algebra, we can talk about group actions. Recall that the
gi’s don’t act on the jgji’s, but rather the representation D(gi)does. We define the action
D(gi)on an element of C[G]as follows:
D(gi)n 1X
j=0cjjgji=D(gi)(c0jg0i+c1jg1i++cn 1jgn 1i)
=c0jgi?g0i+c1jgi?g1i++cn 1jgi?gn 1i=n 1X
j=0cjjgi?gji
Previously, we discussed how elements of a group act on each other, and we also talked
about how elements of a group act on some other object or set of objects (like three painted
29
eggs). We now generalize this notion to a set of qabstract objects a group can act on,
denotedM=fm0;m1;m2;:::;mq 1g. Just as before, we build a vector space, similar to
the one above used in building an algebra. The orthonormal vectors here will be
m0!jm0i; m 1!jm1i; ::: m q 1!jmq 1i
This allows us to understand the following definition.
The set
CMq 1X
i=0cijmiici2C8i
is called the Module ofM. (We don’t use the square brackets here to distinguish modules
from algebras). In other words, the space spanned by the jmiiis the module.
Example 17
Consider, once again, S3. However, we generalize from three eggs to three
“objects”m0;m1, andm2. So,CMis all points in the 3-dimensional space of
the formc0jm0i+c1jm1i+c2jm2iwithci2C8i.
Then, operating on a given point with, say, g1gives
g1(c0jm0i+c1jm1i+c2jm2i) = (c0jg1m0i+c1jg1m1i+c2jg1m2i)
and from the multiplication table on page 19, we know
g1m0=m1; g 1m1=m0; g 1m2=m2
So,
(c0jg1m0i+c1jg1m1i+c2jg1m2i) = (c0jm1i+c1jm0i+c2jm2i)
=c1jm0i+c0jm1i+c2jm2i
So, the effect of g1was to swap c1andc0. This can be visualized geometrically
as a reflection in the c0=c1plane in the 3-dimensional module space. We can
visualize every element of Gin this way. They each move points around the
module space in a well-defined way.
This allows us to give the following definition. IfCVis a module, and CWis
a subspace of CVthat is closed under the action of G, thenCWis an Invariant
Subspace ofCV.
Example 18
Working with S3, we know that S3acts on a 3-dimensional space spanned by
jm0i= (1;0;0)T;jm1i= (0;1;0)T; andjm2i= (0;0;1)T
30
Now, consider the subspace spanned by
c(jm0i+jm1i+jm2i) (2.3)
wherec2C, andcranges over all possible complex numbers. If we restrict
ctoR, we can visualize this more easily as the set of all points in the line
through the origin defined by (^i+^j+^k)(where2R). You can write out the
action of any element of S3on any point in this subspace, and you will see that
they are unaffected. This means that the space spanned by (2.3) is an invariant
subspace.
As a note, all modules CVhave two trivial invariant subspaces.
CVis a trivial invariant subspace of CV
Ceis a trivial invariant subspace of CV
Finally, we can give a more formal definition of reducibility. If a representation Dof a group
Gacts on the space of a module CM, then the representation Dis said to be Reducible ifCM
contains a non-trivial invariant subspace. If a representation is not reducible, it is Irreducible .
We encouraged you to write out the entire regular representation of S3above. If you have
done so, you may have noticed that every 66matrix appeared with non-zero elements
only in the upper left 33elements, and the lower right 33elements. The upper
right and lower left are all 0. This means that, for every element of S3, there will never
be any mixing of the first 3 dimensions with the last 3. So, there are two 3-dimensional
invariant subspaces in the module for this particular representation of S3(the regular
representation).
We can now begin to take advantage of the fact that representations live in linear spaces
with the following definition.
IfVis anyn-dimensional space spanned by nlinearly independent basis vectors, and UandW
are both subspaces of V, then we say that Vis the Direct Sum ofUandWif every vector v2V
can be written as the sum v= u+ w, where u2Uandw2W, and every operator Xacting on
elements of Vcan be separated into parts acting individually on UandW. The notation for this
isV=UW.
In order to make this clearer, ifXnis annnmatrix, it is the direct sum of mmmatrixAm
andkkmatrixBk, denotedXn=AmBk, iffXis in Block Diagonal form,
Xn=Am0
0Bk
wheren=m+k, andAm,Bk, and the 0’s are understood as matrices of appropriate dimension.
31
We can generalize the previous definition as follows,
Xn=An1Bn2Cnk=0
BBB@An10 0
0Bn2 0
.........
0 0...Cnk1
CCCA
wheren=n1+n2++nk.
Example 19
LetA3=0
@1 1 2
1 5
17 4 111
A, and letB2=1 2
3 4
. Then,
B2A3=0
BBBB@1 2 0 0 0
3 4 0 0 0
0 0 1 1 2
0 0 1 5
0 0 17 4 111
CCCCA
To take stock of what we have done so far, we have talked about algebras, which are the
vector spaces spanned by the elements of a group, and about modules, which are the vec-
tor spaces that representations of groups act on. We have also defined invariant subspaces
as follows: Given some space and some group that acts on that space, moving the points
around in a well-defined way, an invariant subspace is a subspace which always contains
the same points. The group doesn’t remove any points from that subspace, and it doesn’t
add any points to it. It merely moves the points around inside that subspace. Then, we
defined a representation as reducible if there are any non-trivial invariant subspaces in
the space that the group acts on.
And what this amounts to is the following: a representation of any group is reducible if it
can be written in block diagonal form.
But this leaves the question of what we mean when we say “can be written”. How can
you “rewrite” a representation? This leads us to the following definition. Given a matrix
Dand a non-singular matrix S, the linear transformation
D!D0=S 1DS
is called a Similarity Transformation .
Then, we can give the following definition. Two matrices related by a similarity transforma-
tion are said to be Equivalent .
Because similarity transformations are linear transformations, if D(G)is a representation
ofG, then so is S 1DSfor literally anynon-singular matrix S. To see this, if gi?gj=gk,
32
thenD(gi)D(gj) =D(gk), and therefore
S 1D(gi)SS 1D(gj)S=S 1D(gi)D(gj)S=S 1D(gk)S
So, if we have a representation that isn’t in block diagonal form, how can we figure out
if it is reducible? We must look for a matrix Sthat will transform it into block diagonal
form.
You likely realize immediately that this is not a particularly easy thing to do by inspection.
It turns out that there is a very straightforward and systematic way of taking a given rep-
resentation and determining whether or not it is reducible, and if so, what the irreducible
representations are.
However, the details of how this can be done, while very interesting, are not necessary
for the agenda of these notes. Therefore, for the sake of brevity, we will not pursue them.
What is important is that you understand not only the details of general group theory and
representation theory (which we outlined above), but also the concept of what it means
for a group to be reducible or irreducible.
2.2 Introduction to Lie Groups
In section 2.1, we considered groups which are of finite order and discrete, which allowed
us to write out a multiplication table.
Here, however, we examine a different type of group. Consider the unit circle, where
each point on the circle is specified by an angle , measured from the positive x-axis.
We will refer to the point at = 0 as the “starting point” (like ROY was for the Easter
eggs). Now, just as we considered all possible orientations of (ROY )that left the eggs
lined up, we consider all possible rotations the wheel can undergo. With the eggs there
33
were only 6 possibilities. Now however, for the wheel there are an infinite number of
possibilities for (any real number 2[0;2)).
And note that if we denote the set of all angles as G, then all the rotations obey closure
(1+2=32G;81;22G), associativity (as usual), identity (0 +=+ 0 =), and
inverse (the inverse of is ).
So, we have a group that is parameterized by a continuous variable . So, we are no longer
talking about gi’s, but about g().
Notice that this particular group (the circle) is Abelian, which is why we can (temporarily)
use addition to represent it. Also, note that we obviously cannot make a multiplication
table because the order of this group is 1.
One simple representation is the one we used above: taking and using addition. A
more familiar (and useful) representation is the Euler matrix g() _ =cossin
sincos
with
the usual matrix multiplication:
cos1sin1
sin1cos1cos2sin2
sin2cos2
(2.4)
=cos1cos2 sin1sin2 cos1sin2+ sin1cos2
sin1cos2 cos1sin2 sin1sin2+ cos1cos2
(2.5)
=cos(1+2) sin(1+2)
sin(1+2) cos(1+2)
(2.6)
This will prove to be a much more useful representation than with addition.
Groups that are parameterized by one or more continuous variables like this are called
Lie Groups . Of course, the true definition of a Lie group is much more rigorous (and com-
plicated), and that definition should eventually be understood. However, the definition
we have given will suffice for the purposes of these notes.
2.2.1 Classification of Lie Groups
The usefulness of group theory is that groups represent a mathematical way to make
changes to a system while leaving something about the system unchanged. For example,
we moved (ROY )around, but the structure “3 eggs with different colors lined up” was
preserved. With the circle, we rotated it, but it still maintained its basic structure as a
circle. It is in this sense that group theory is a study of Symmetry . No matter which of
“these” transformations you do to the system, “this” stays the same—this is symmetry.
To see the usefulness of this in physics, recall Noether’s Theorem (section 1.1.2). When
you do a symmetry transformation to a Lagrangian, you get a conserved quantity. Think
34
back to the Lagrangian for the projectile (1.2). The transformation x!x+was a sym-
metry because could take any value, and the Lagrangian was unchanged (note that
forms the Abelian group (R;+)).
So, given a Lagrangian, which represents the structure of a physical system, a symme-
try represents a way of changing the Lagrangian while preserving that structure. The
particular preserved part of the system is the conserved quantity jwe discussed in sec-
tions 1.1.2 and 1.1.6. And as you have no doubt noticed, nearly all physical processes are
governed by Conservation Laws : conservation of momentum, energy, charge, spin, etc.
So, group theory, and in particular Lie group theory, gives us an extremely powerful way
of understanding and classifying symmetries, and therefore conserved charges. And be-
cause it allows us to understand conserved charges, group theory can be used to under-
stand the entirety of the physics in our universe.
We now begin to classify the major types of Lie groups we will be working with in these
notes. To start, we consider the most general possible Lie group in an arbitrary number of
dimensions, n. This will be the group that, for any point pin then-dimensional space, can
continuously take it anywhere else in the space. All that is preserved is that the points in
the space stay in the space. This means that we can have literally any nnmatrix, or linear
transformation, so long as the matrix is invertible (non-singular). Thus, in ndimensions
the largest and most general Lie group is the group of all nnnon-singular matrices. We
call this group GL(n), or the General Linear group. The most general field of numbers
to take the elements of GL(n)from is C, so we begin with GL(n;C). This is the group of
allnnnon-singular matrices with complex elements. The preserved quantity is that all
points in Cnstay in Cn.
The most obvious subgroup of GL(n;C)isGL(n;R), or the set of all nninvertible
matrices with real elements. This leaves all points in RninRn.
To find a further subgroup, recall from linear algebra and vector calculus that in ndi-
mensions, you can take nvectors at the origin such that for a parallelepiped, we could
obtain
35
Then, if you arrange the components of the nvectors into the rows (or columns) of a
matrix, the determinant of that matrix will be the volume of the parallelepiped.
So, consider now the set of all General Linear transformations that transform all vectors
from the origin (or in other words, points in the space) in such a way that the volume of
the corresponding parallelepiped is preserved. This will demand that we only consider
General Linear matrices with determinant 1. Also, the set of all General Linear matrices
with unit determinant will form a group because of the general rule detjABj= detjAj
detjBj. So, if detjAj= 1 anddetjBj= 1, then detjABj= 1. We call this subgroup of
GL(n;C)theSpecial Linear group, orSL(n;C). The natural subgroup of this is SL(n;R).
This group preserves not only the points in the space (as GLdid), but also the volume, as
described above.
Now, consider the familiar transformations on vectors in n-dimensional space of gen-
eralized Euler angles. These are transformations that rotate all points around the ori-
gin. These rotation transformations leave the radius squared (r2)invariant. And, be-
cause r2= rTr, if we transform with a rotation matrix R, then r!r0=Rr, and
rT!r0T= rTRT, sor0Tr0= rTRTRr. But, as we said, we are demanding that the
radius squared be invariant under the action of R, and so we demand rTRTRr= rTr.
So, the constraint we are imposing is RTR=I, which implies RT=R 1. This tells us that
the rows and columns of Rare orthogonal. Therefore, we call the group of generalized
rotations, or generalized Euler angles in ndimensions, O(n), or the Orthogonal group.
We don’t specify CorRhere because it will be understood that we are always talking
aboutR.
Also, note that because detjRTRj= detjIj)(detjRj)2= 1)detjRj=1. We again
denote the subgroup with detjRj= +1 theSpecial Orthogonal group, orSO(n). To
understand what this means, consider an orthogonal matrix with determinant 1, such
as
M=0
@1 0 0
0 1 0
0 0 11
A
This matrix is orthogonal, and therefore is an element of the group O(3), but the determi-
nant is 1. This matrix will take the point (x;y;z )Tto the point (x;y; z)T. This changes
the handedness of the system (the right hand rule will no longer work). So, if we limit
ourselves to SO(n), we are preserving the space, the radius, the volume, and the handed-
ness of the space.
For vectors in Cspace, we do not define orthogonal matrices (although we could). In-
stead, we discuss the complex version of the radius, where instead of r2= rTr, we have
r2= ryr, where the dagger denotes the Hermitian conjugate, ry= (r?)T, where?denotes
complex conjugate.
So, with the elements in Rbeing in C, we have r!Rr, and ry!ryRy. So, ryr!ryRyRr,
and by the same argument as above with the orthogonal matrices, this demands that
36
RyR=I, orRy=R 1. We denote such matrices Unitary , and the set of all such nn
invertible matrices form the group U(n). Again, we understand the unitary groups to
have elements in C, so we don’t specify that. And, we will still have a subset of unitary
matricesRwith detjRj= 1calledSU(n), the Special Unitary groups.
We can summarize the hierarchy we have just described in the following diagram:
We will now describe one more category of Lie groups before moving on. We saw above
that the group SO(n)preserves the radius squared in real space. In coordinates, this
means that r2=x2
1+x2
2++x2
n, or more generally the dot product xy=x1y1+x2y2+
+xnynis preserved.
However, we can generalize this to form a group action that preserves not the radius
squared, but the value (switching to indicial notation for the dot product) xaya= x1y1
x2y2 xmym+xm+1ym+1++xm+nym+n. We call the group that preserves this quantity
37
SO(m;n). The space we are working in is still Rm+n, but we are making transformations
that preserve something different than the radius.
Note thatSO(m;n)will have an SO(m)subgroup and an SO(n)subgroup, consisting of
rotations in the first mand lastncomponents separately.
Finally, notice that the specific group of this type, SO(1;3), is the group that preserves the
values2= x1y1+x2y2+x3y3+x4y4, or written more suggestively, s2= c2t2+x2+y2+z2.
Therefore, the group SO(1;3)is the Lorentz Group . Any action that is invariant under
SO(1;3)is said to be a Lorentz Invariant theory (as all theories should be). We will find
that thinking of Special Relativity in these terms, rather than in the terms of Part I, will be
much more useful.
It should be noted that there are many other types of Lie groups. We have limited our-
selves to the ones we will be working with in these notes.
2.2.2 Generators
Now that we have a good “birds eye view” of Lie groups, we can begin to pick apart the
details of how they work.
As we said before, a Lie group is a group that is parameterized by a set of continuous
parameters, which we call ifori= 1;:::;n , wherenis the number of parameters the
group depends on. The elements of the group will then be denoted g(i).
Because all groups include an identity element, we will choose to parameterize them in
such a way that g(i)
i=0=e, the identity element. So, if we are going to talk about
representations, Dn(g(i))
i=0=I, where Iis thennidentity matrix for whatever
dimension ( n) representation we want.
Now, takeito be very small with i<<1. So,Dn(g(0 +i))can be Taylor expanded:
Dn(g(i)) =I+i@Dn(g(i))
@i
i=0+
The terms@Dn
@i
i=0are extremely important, and we give them their own expression:
Xi i@Dn
@i
i=0(2.7)
(we have included the iin order to make XiHermitian, which will be necessary later).
So, the representation for infinitesimal iis then
Dn(i) =I+iiXi+
38
(where we have switched our notation from Dn(g())toDn()for brevity).
TheXi’s are constant matrices which we will determine later.
Now, let’s say that we want to see what the representation will look like for a finite value
ofirather than an infinitesimal value. A finite transformation will be the result of an
infinite number of infinitesimal transformations. Or in other words, i=NiasN!1 .
So,i=i
N, and an infinite number of infinitesimal transformations is
lim
N!1(1 +iiXi)N= lim
N!1
1 +ii
NXiN
If you expand this out for several values of N, you will see that it is exactly
lim
N!1
1 +ii
NXiN=eiiXi
We call the Xi’s the Generators of the group, and there is one for each parameter required
to specify a particular element of the group. For example, consider SO(3), the group
of rotations in 3 dimensions. We know from vector calculus that an element of SO(3)
requires 3 angles, usually denoted ;, and . Therefore, SO(3)will require 3 generators,
which will be denoted X;X, andX . We will discuss how the generators can be found
soon.
In general, there will be several (in fact, infinite) different sets of Xi’s that define a given
group (just as there are an infinite number of representations of any finite group). What
we will find is that up to a similarity transformation, a particular set of generators defines
a particular representation of a group.
So,Dn(i) =eiiXifor any group (the iindex in the exponent is understood to be summed
over all parameters and generators). The best way to think of the parameter space for the
group is as a vector space, where the generators describe the behavior near the identity,
but form a basis for the entire vector space. By analogy, think of the unit vectors ^i,^j,
and ^kinR3. They are defined at the origin, but they can be combined with real num-
bers/parameters to specify any arbitrary point in R3. In the same way, the generators are
the “unit vectors” of the parameter space (which in general is a much more complicated
space than Euclidian space), and the parameters (like ;, and ) specify where in the
parameter space you are in terms of the generators. That point in the parameter space
will then correspond to a particular element of the group.
We call the number of generators of a group (or equivalently the number of parameters
necessary to specify an element), the Dimension of the group. For example, the dimen-
sion ofSO(3)is 3. The dimension of SO(2)(rotations in the plane) however is only 1 (only
is needed), so there will be only one generator.
39
2.2.3 Lie Algebras
In section 2.1 we discussed algebras. An algebra is a space spanned by elements of the
group with Ccoefficients parameterizing the Euclidian space we defined. Obviously we
can’t define an algebra in the same way for Lie groups, because the elements are continu-
ous. But, as discussed in the last section, a particular element of a Lie group is defined by
the values of the parameters in the parameter space spanned by the generators. We will
see that the generators will form the algebras for Lie groups.
Consider two elements of the same group with generators Xi, one with parameter values
iand the other with parameter values i. The product of the 2 elements will then be
eiiXieijXj. Because we are assuming this is a group, we know that the product must be
an element of the group (due to closure), and therefore the product must be specified by
some set of parameters k, soeiiXieijXj=eikXk. Note that the product won’t necessarily
simply be eiiXieijXj=ei(iXi+jXj)because the generators are matrices and therefore
don’t in general commute.
So, we want to figure out what iwill be in terms of iandi. We do this as follows.
iiXk= ln(eikXk) = ln(eiiXiejXj) = ln(1 +eiiXiejXj 1)ln(1 +x)
where we have defined xeiiXiejXj 1. We will proceed by expanding only to second
order iniandj, though the result we will obtain will hold at arbitrary order. By Taylor
expanding the exponential terms,
eiiXiejXj 1 = (1 + iiXi+1
2(iiXi)2+)(1 +ijXj+1
2(ijXj)2+) 1
= 1 +ijXj 1
2(jXj)2+iiXi iXijXj 1
2(iXi)2 1
=i(iXi+jXj) iXijXj 1
2
(iXi)2+ (jXj)2
Then, using the general Taylor expansion ln(1 +x) =x x2
2+x3
3 x4
4+, and again
40
keeping terms only to second order in and, we have
x x2
2=
i(iXi+jXj) iXijXj 1
2[(iXi)2+ (jXj)2]
1
2
i(iXi+jXj) iXijXj 1
2[(iXi)2+ (jXj)2]2
=i(iXi+jXj) iXijXj 1
2
(iXi)2+ (jXj)2
1
2
(iXi+jXj)(iXi+jXj)
=i(iXi+jXj) iXijXj 1
2
(iXi)2+ (jXj)2
+1
2
(iXi)2+ (jXj)2+ij(XiXj+XjXi)
=i(iXi+jXj) +1
2ij(XjXi XiXj)
=i(iXi+jXj) 1
2ij[Xi;Xj]
=i(iXi+jXj) 1
2[iXi;jXj]
So finally we can see
ikXk=i(iXi+jXj) 1
2[iXi;j;Xj]
or
eiiXieijXj=ei(iXi+jXj) 1
2[iXi;jXj](2.8)
Equation (2.8) is called the Baker-Campbell-Hausdorff formula, and it is one of the most
important relations in group theory and in physics. Notice that, if the generators com-
mute, this reduces to the normal equation for multiplying exponentials. You can think of
equation (2.8) as the generalization of the normal exponential multiplication rule.
Now, it is clear that the commutator [Xi;Xj]must be proportional to some linear combi-
nation of the generators of the group (because of closure). So, it must be the case that
[Xi;Xj] =ifijkXk (2.9)
for some set of constants fijk. These constants are called the Structure Constants of the
group, and if they are completely known, the commutation relations between all the gen-
erators are known, and so the entire group can be determined in any representation you
want.
The generators, under the specific commutation relations defined by the structure con-
stants, form the Lie Algebra of the group, and it is this commutation structure which
forms the structure of the Lie group.
41
2.2.4 The Adjoint Representation
We will talk about several representations for each group we discuss, but we will men-
tion a very important one now. We mentioned before that the structure constants fijk
completely determine the entire structure of the group.
We begin by using the Jacobi identity,
Xi;[Xj;Xk]
+
Xj;[Xk;Xi]
+
Xk;[Xi;Xj]
= 0 (2.10)
(if you aren’t familiar with this identity, try multiplying it out. You will find that it is
identically true — all the terms cancel exactly). But, from equation (2.9), we can write
Xi;[Xj;Xk]
=ifjka[Xi;Xa] =ifjkafiabXb
Plugging this into (2.10) we get
ifjkafiabXb+ifkiafjabXb+ifijafkabXb= 0
)(fjkafiab+fkiafjab+fijafkab)iXb= 0
)fjkafiab+fkiafjab+fijafkab= 0 (2.11)
So, if we define the matrices
[Ta]bc ifabc (2.12)
then it is easy to show that (2.11) leads to
[Ta;Tb] =ifabcTc
So, the structure constants themselves form a representation of the group (as defined by
(2.12). We call this representation the Adjoint Representation , and it will prove to be
extremely important.
Notice that the indices labeling the rows and columns in (2.12) each run over the same
values as the indices labeling the Tmatrices. This tells us that the adjoint representation
is made of nnmatrices, where nis the dimension of the group, or the number of
parameters in the group. For example, SO(3)requires 3 parameters to specify an element
(;; ), so the adjoint representation of SO(3)will consist of 33matrices.SO(2)on the
other hand is Abelian, and therefore all of the structure constants vanish. Therefore there
is no adjoint representation of SO(2).
We now go on to consider several specific groups in detail.
42
2.2.5SO(2)
We start by looking at an extremely simple group, SO(2). This is the group of rotations in
the plane that leaves r2=x2+y2=
x y
x
y
= vTvinvariant. So for some generator X
(which we will now find) of SO(2),v!R()v=eiXv, and vT!vTeiXT. So, expanding
to first order only, vTeiXTeiXv= vT(1 +iXT+iX)v= vTv+ vTi(X+XT)v. And
because we demand that r2be invariant, we demand that X+XT= 0)X= XT. So,
Xmust be antisymmetric. Therefore we take
X1
i0 1
1 0
(the1
iis included to balance the iwe inserted in equation (2.7) to ensure that Xis Hermi-
tian).
So, an arbitrary element of SO(2)will be
eiX=ei1
i0
@0 1
1 01
A
=e0
@0 1
1 01
A
=0 1
1 00
+0 1
1 01
+1
220 1
1 02
+
=1 0
0 1
+0 1
1 0
1
221 0
0 1
1
3!30 1
1 0
+
=1 1
22+ 1
3!3+
( 1
3!3+) 1 1
22+
=cossin
sincos
which is exactly what we would expect for a matrix describing rotations in the plane.
Also, notice that because SO(2)is Abelian, the commutation relations trivially vanish
([X;X ]0), and so all of the structure constants are zero.
Now that we have found an explicit example of a generator, and seen an example of how
generators relate to group elements, we move on to slightly more complicated examples.
2.2.6SO(3)
We could easily generalize the argument from the proceeding section and find the genera-
tors ofSO(3)in the same way, but in order to illustrate more clearly how generators work,
we will approach SO(3)differently by working backwards. Above, we found the gener-
ators and used them to calculate the group elements. Here, we begin with the known
43
group elements of SO(3), which are just the standard Euler matrices for rotations in 3-
dimensional space:
Rx() =0
@1 0 0
0 cossin
0 sincos1
A (2.13)
Ry( ) =0
@cos 0 sin
0 1 0
sin 0 cos 1
A (2.14)
Rz() =0
@cossin0
sincos0
0 0 11
A (2.15)
Now, recall the definition of the generators, equation (2.7). We can use it to find the
generators of SO(3), which we will denote Jx,Jy, andJz.
Jx=1
idRx()
d
=0=1
i0
@0 0 0
0 sincos
0 cossin1
A
=0=1
i0
@0 0 0
0 0 1
0 1 01
A
And similarly
Jy=1
i0
@0 0 1
0 0 0
1 0 01
A; Jz=1
i0
@0 1 0
1 0 0
0 0 01
A
You can plug these into the exponentials with the appropriate parameters ( ; , or) and
find thateiJx,ei Jy, andeiJzreproduce (2.13), (2.14), and (2.15), respectively.
Furthermore, you can multiply out the commutators to find
[Jx;Jy] =iJz; [Jy;Jz] =iJx; [Jz;Jx] =iJy
or
[Ji;Jj] =iijkJk
which tells us that the structure constants for SO(3)are
fijk=ijk (2.16)
whereijkis the totally antisymmetric tensor. The structure constants being non-zero is
consistent with SO(3)being a non-Abelian group.
44
2.2.7SU(2)
We will approach SU(2)yet another way: by starting with the structure constants. It turns
out they are the same as the structure constants for SO(3):
fijk=ijk (2.17)
To see why, recall that SU(2)are rotations in two complex dimensions. The most general
form of such a matrix U2SU(2)isU=a b
c d
. The “Special” part of SU(2)demands
that the determinant be equal to 1, or
ad bc= 1
and the “Unitary” part demands that U 1=Uy. So,
U 1=d b
c a
=Uy=a?c?
b?d?
or in other words,
U=a b
b?a?
where we demand jaj2+jbj2= 1.
Bothaandbare inC, and therefore have 2 real components each, so Uhas 4 real param-
eters. The constraint jaj2+jbj2= 1fixes one of them, leaving 3 real parameters, just like
inSO(3). This is a loose explanation of why SU(2)andSO(3)have the same structure
constants. They are both rotational groups with 3 real parameters.
This also tells us that SU(2)will have 3 generators.
2.2.8SU(2)and Physical States
The elements of any Lie group (in a d-dimensional representation consisting of ddma-
trices) will act on vectors, just like the 33matrices representing S3acted on
R O YT
in section 2.1.5. The most natural way to understand the space a Lie group acts on is to
study the eigenvectors and eigenvalues of the generators of the representation you are
using (the reason for this is beyond the scope of these notes at this point, but will be-
come more clear as we proceed). These eigenvectors will obviously form a basis of the
eigenspace of the physical space the group is acting on.
Using similarity transformations, one or more of the generators of a Lie group can be di-
agonalized. For now, trust us that with SU(2), it is only possible to diagonalize one of the
45
three generators at a time (you may convince yourself of this by studying the commuta-
tion relations). We will call the generators of SU(2)J1;J2, andJ3, and by convention we
takeJ3to be the diagonal one. So, consequently, the eigenvectors of J3will be the basis
vectors of the physical vector space upon which SU(2)acts.
Now, we know that J3(whatever it is ... we don’t know at this point) will in general
have more than one eigenvalue. Let’s call the greatest eigenvalue of J3(whatever it is)
j, and the eigenvectors of J3will be denotedjj;mi(the firstjis merely a label — the
second value describes the vector), where mis the eigenvalue of the eigenvector. The
eigenvector corresponding to the greatest eigenvalue jwill obviously then be jj;ji. So,
J3jj;ji=jjj;ji, or more generally J3jj;mi=mjj;mi.
Now let’s assume that we know jj;ji. There is a trick we can employ to find the rest of
the states. Define the following linear combinations of the generators:
J1p
2(J1iJ2) (2.18)
Now, using the fact that the SU(2)generators obey the commutation relations in equation
(2.16), it is easy to show the following relations,
[J2;J] =Jand [J+;J ] =J3(2.19)
Notice that, because by definition Jiare all Hermitian, we have
(J )y=J+(2.20)
Consider some arbitrary eigenvector jj;mi. We know the eigenvalue of this will be m, so
J3jj;mi=mjj;mi. But now let’s create some new state by acting on jj;miwith either of
the operators (2.18). The new state will be Jjj;mi, but what will the J3eigenvalue be?
Using the commutation relations in (2.19),
J3Jjj;mi= (J+JJ3)jj;mi= (m1)Jjj;mi
So, the vector J+jj;miis the eigenvector with eigenvalue m+ 1, and the vector J jj;mi
is the eigenvector with eigenvalue m 1.
If we have some arbitrary eigenvector jj;mi, we can use Jto move up or down to the
eigenvector with the next highest or lowest eigenvalue. For this reason, Jare called the
Raising and Lowering operators. They raise and lower the eigenvalue of the state by one.
Clearly, the eigenvector with the greatest eigenvalue j, with eigenvector jj;ji, cannot
be raised any higher, so we define J+jj;ji 0. We will see that there is also a lowest
eigenvalue j0, so we similarly define J jj;j0i0.
Now, considering once again jj;ji. We know that if we operate on this state with J , we
will get the eigenvector with the eigenvalue j 1. But, we don’t know exactly what this
46
state will be (knowing the eigenvalue doesn’t mean we know the actual state). But, we
know it will be proportional to jj;j 1i. So, we set J jj;ji=Njjj;j 1i, whereNjis the
proportionality constant. To find Nj, we take the inner product (and using (2.20)):
hj;jjJ+J jj;ji=jNjj2hj;j 1jj;j 1i
But we can also write
hj;jjJ+J jj;ji=hj;jj(J+J J J+)jj;ji=hj;jj
J+;J
jj;ji
=hj;jjJ3jj;ji=jhj;jjj;ji=j (2.21)
where we used the fact that J+jj;ji= 0to get the first equality, and (2.19) to get the third
equality. We also assumed that jj;jiis normalized.
So, (2.21) tells us
hj;j 1jj;j 1i= 1()Njp
j (2.22)
And our normalized state is therefore
J
Njjj;ji=J
pjjj;ji=jj;j 1i
Repeating this to find Nj 1, we have
jNj 1j2hj;j 2jj;j 2i=hj;j 1jJ+J jj;j 1i
=hj;jjJ+
pjJ+J J
pjjj;ji
=1
jhj;jjJ+J+J J jj;ji
=1
jhj;jjJ+(J3+J J+)J jj;ji
=1
jhj;jj(J+J3J +J+J J+J )jj;ji
=1
jhj;jj(J+( J +J J3) +J+J (J3+J J+))jj;ji
=1
j[hj;jj( J+J +jJ+J +jJ+J )jj;ji]
=1
jhj;jj( [J+;J ] + 2j[J+;J ])jj;ji
=1
jhj;jj( J3+ 2jJ3)jj;ji
=1
j(2j2 j) = 2j 1 (2.23)
So,jNj 1j2= 2j 1, orNj 1=p2j 1.
47
We can continue this process, and we will find that the general result is
Nj k=1p
2p
(2j k)(k+ 1) (2.24)
and the general states are defined by
jj;j ki=1
Nj k(J )kjj;ji
Notice that these expressions recover (2.22) and (2.23) for k= 0andk= 1, respectively.
Furthermore, notice that when k= 2j,
Nj 2j=1p
2p
(2j 2j)(2j+ 1)0
So, the statejj;j ki
k=2j=jj; jiis the state with the lowest eigenvalue, and by defini-
tionJ jj; ji0.
So, in a general representation of SU(2), we have 2j+ 1states:
fj;j 1;j 2;:::; j+ 2; j+ 1; jg
This therefore demands that j=n
2for some integer n. In other words, the highest eigen-
value of an SU(2)eigenvector can be 0;1
2;1;3
2;2, etc.
Furthermore, using these states, it is easy to show
hj;m0jJ3jj;mi=mm0;m
hj;m0jJ+jj;mi=1p
2p
(j+m+ 1)(j m)m0;m+1
hj;m0jJ jj;mi=1p
2p
(j+m)(j m+ 1)m0;m 1 (2.25)
2.2.9SU(2)forj=1
2
We will skip the j= 0 case because it is trivial (though we will discuss it later when we
return to physics).
Forj=1
2, the two eigenvalues of J3will be1
2and1
2 1= 1
2. So, denoting the J3generator
ofSU(2)whenj=1
2asJ3
1=2, we have
J3
1=2=1=2 0
0 1=2
48
Now, inverting (2.18) to get
J1=1p
2(J +J+) andJ2=ip
2(J J+)
and using the standard matrix equation [Ja
j]m0;m=hj;m0jJajj;mi, and the explicit prod-
ucts in (2.25), we can find (for example)
1
2; 1
2J11
2; 1
2
=1
2; 1
21p
2(J +J+)1
2; 1
2
== 0
So[J1]11= 0. Then,
1
2; 1
2J11
2;1
2
=1
2; 1
21p
2(J +J+)1
2;1
2
==1
2
So[J1]12=1
2.
We can continue this to find all the elements for each generator for j= 1=2. The final
result will be
J1
1=2=1
20 1
1 0
=1
2; J2
1=2=1
20 i
i0
=2
2; J3
1=2=1
21 0
0 1
=3
2(2.26)
where theimatrices are the Pauli Spin Matrices . This is no accident! We will discuss this
in much, much more detail later, but for now recall that we said that SU(2)is the group of
transformations in 2-dimensional complex space (with one of the real parameters fixed,
leaving 3 real parameters). We are going to see that SU(2)is the group which represents
quantum mechanical spin, where jis the value of the spin of the particle. In other words,
particles with spin 1=2are described by the j= 1=2representation (the 22representation
in (2.26)), and particles with spin 1 are described by the j= 1representation, and so on.
In other words, SU(2)describes quantum mechanical spin in 3 dimensions in the same
way thatSO(3)describes normal “spin” in 3 dimensions. We will talk about the physical
implications, reasons, and meaning of this later.
However, as a warning, be careful at this point not to think too much in terms of physics.
You have likely covered SU(2)in great detail in a quantum mechanics course (though
you may not have known it was called “ SU(2)”), but the approach we are taking here
has a different goal than what you have likely seen before. The properties of SU(2)we
are seeing here are actually very, very specific and simplified illustrations of much deeper
concepts in Lie groups, and in order to understand particle physics we must understand
Lie groups in this way. So for now, try to fight the temptation to merely understand
everything we are doing in terms of the physics you have seen before and learn this as
we are presenting it: pure mathematics. We will focus on how it applies to physics later,
in its fuller and more fundamental way than introductory quantum mechanics makes
apparent.
49
2.2.10SU(2)forj= 1
You can follow the same procedure we used above to find
J1
1=1p
20
@0 1 0
1 0 1
0 1 01
A; J2
1=1p
20
@0 i0
i0 i
0i01
A; J3
1=0
@1 0 0
0 0 0
0 0 11
A (2.27)
Notice that only J3
1is diagonal (as before), and that the eigenvalues are f1;0; 1g, or
fj;j 1;j 2 = jgas we’d expect.
2.2.11SU(2)for Arbitrary j
For any given j, we have 3 generators J1
j;J2
j, andJ3
j, and for whatever dimension ( d=
2j+ 1) the physical space we are working in, we have deigenvectors
jj;ji=0
BBBBB@1
0
0
...
01
CCCCCA;jj;j 1i=0
BBBBB@0
1
0
...
01
CCCCCA;jj;j 2i=0
BBBBB@0
0
1
...
01
CCCCCA; jj; ji=0
BBBBB@0
0
0
...
11
CCCCCA
with eigenvalues fj;j 1;j 2;; jg, respectively.
Then, for any j, we can form the linear combinations J
j1p
2(J1
jiJ2
j). For example, for
j= 1=2these are
J+
1=2=1p
21
20 1
1 0
+i
20 i
i0
=1
2p
20 2
0 0
=1p
20 1
0 0
and similarly
J
1=2==1p
20 0
1 0
So, the two j= 1=2eigenvectors will be1
2;1
2i=1
0
and1
2; 1
2i=0
1
. So,
J+
1=21
2;1
2
=1p
20 1
0 01
0
= 0
J
1=21
2; 1
2
=1p
20 0
1 00
1
= 0
50
and similarly
J
1=21
2;1
2
=1
2; 1
2
J+
1=21
2; 1
2
=1
2;1
2
which is exactly what we would expect.
The same calculation can be done for the j= 1 case and we will find the same results,
except that the j= 1 state (the first eigenvector) can be lowered twice . The first time
J
1=2acts it takes it to the state with eigenvalue 0, and the second time it acts it takes it
to the state with eigenvalue 1. Acting a third time will destroy the state (take it to 0).
Analogously, the lowest state, with eigenvalue j= 1can be raised twice.
We can do the same analysis for any j=integer or half integer.
As we said before, we interpret jas the quantum mechanical spin of a particle, and the
groupSU(2)describes that rotation. It is important to recognize that quantum spin is
not a rotation through spacetime (it would be described by SO(3)if it was), but rather
through the mathematically constructed spinor space. We will talk more about this space
later.
So for a given particle with spin, we can talk about both its rotation through physical
spacetime using SO(3), as well as its rotation through complex spinor space using SU(2).
Both values will be physically measurable and will be conserved quantities. The total
angular momentum of the particle will be the combination of both spin and spacetime
angular momentum. Again, we will talk much more about the spin of physical particles
when we return to a discussion of physics. We only mention this now to give a preview
of where this is going. However, spin is not the only thing SU(2)describes. We will also
find that it is the group which governs the weak nuclear force (whereas U(1)describes
the electromagnetic force, and SU(3)describes the strong force ... much, much more on
this later).
2.2.12 Root Space
As a comment before beginning this section, it is likely that you will find this to be the
most difficult section of these notes. The material here is both extremely difficult (espe-
cially the first time it is encountered), and extremely important to the development of
particle physics. In fact, this section is the most central to what will come later in these
notes. If the contents are not clear you are encouraged to read this section multiple times
until it becomes clear. It may also be helpful to study this section while looking closely
at the examples in the sections forming the remainder of this part of these notes. They
illustrate the point of where we are going with all of this.
51
We saw in the previous section that we can view the physical space that a group is acting
on by using the eigenvectors of the diagonal generators as a basis. These eigenvectors
can be arranged in order of decreasing eigenvalue. Then, the non-diagonal generators
can be used to form linear combinations that act as raising and lowering operators, which
transform one eigenvector to another, changing the eigenvalue by an amount defined by
the commutation relations of the generators.
We now see that this generalizes very nicely.
An arbitrary Lie group is defined in terms of its generators. As we said at the end of
section 2.2.2, it is best to think of the generators as being analogous to the basis vectors
spanning some space. Of course, the space the generators span is much more complicated
thanRnin general, but the generators span the space the same way. In this sense, the
generators form a linear vector space. So, we must define an inner product for them. For
reasons that are beyond the scope of these notes, we will choose the generators and inner
product so that, for generators TaandTb,
hTa;Tbi1
Tr (TaTb) =ab(2.28)
whereis some normalization constant.
Also, in the set of generators of a Lie group, there will be a closed subalgebra of generators
which all commute with each other, but not with generators outside of this subalgebra.
In other words, this is the set of generators which can be simultaneously diagonalized
through some similarity transformation. For SU(2), we saw that there was only one gen-
erator in this subalgebra which we chose to be J3
j(recall that a matrix will only commute
with all other matrices if it is equal to the identity matrix times a constant, whereas two
diagonal matrices will always commute regardless of what their diagonal elements are).
Let’s say that a particular Lie group has Ngenerators total, or is an N-dimensional group.
Then, let’s say that there are M < N generators in the mutually commuting subalgebra.
We call those Mgenerators the Cartan Subalgebra , and the generators in it are called
Cartan Generators . We define the number Mas the Rank of the group.
By convention we will label the Cartan generators Hi(i= 1;:::;M ) and the non-Cartan
generators Ei(i= 1;:::;N M).
For example, with SU(2)we hadH1=J3
j, andE1=J1
j,E2=J2
j.
Before moving on, we point out that this should seem familiar. If you think back to an
introductory class in quantum mechanics, recall that we always choose some set of vari-
ables that all commute with each other (usually we choose either position or momentum
because [x;p]6= 0). Then, we expand the physical states in terms of the position ormo-
mentum eigenvectors. Here, we are doing the exact same thing, only in a much more
general context.
Now, theHi’s are simultaneously diagonalized, so we will write the physical states in
52
terms of their eigenvalues. In an n-dimensional representation Dn, the generators are nn
matrices, so the eigenvectors are n-dimensional. So, there will be a total of neigenvectors,
and each will have one eigenvalue with each of the MCartan generators Hi. So, for each
of these eigenvectors, which we temporarily denote jji, forj= 1;:::;n , we have the
Meigenvalues with MCartan generators, which we call ti
j(wherej= 1;:::;n labels
the eigenvectors, and i= 1;:::;M labels the eigenvalues), and we form what is called a
Weight Vector
tj0
BBBBB@t1
j
t2
j
t3
j...
tM
j1
CCCCCA(2.29)
wherej= 1;:::;n . The individual components of these vectors, the ti
j’s, are called the
Weights .
So for a given representation Dn, we now denote the state jDn;tji(instead ofjji). So, our
eigenvalues will be
HijDn;tji=ti
jjDn;tji (2.30)
As we mentioned before, the adjoint representation is a particularly important represen-
tation. If you do not remember the details of the adjoint representation, go reread section
2.2.4. Here, the generators are defined by equation (2.12), [Ta]bc ifabc. Recall that each
index runs from 1 to N, so that the generators in the adjoint representation are NN
matrices, and the eigenvectors are N-dimensional.
Also, as a point of nomenclature, weights in the adjoint representation are called Roots ,
and the corresponding vectors (as in (2.29)) are called Root Vectors .
This means that there is exactly one eigenvector for each generator, and therefore one
root vector for each generator. So, in equation (2.29), j= 1;:::;N . We make this more
obvious by explicitly assigning each eigenvector to a generator as follows. First, because
we now have the same number of generators, eigenvectors, and root vectors, we label the
generators by the root vectors Ttjinstead ofTj. Also, we now refer to general eigenstates
asjAdj;Ttji, wherej= 1;:::;N andtjis theM-dimensional root vector corresponding
toTtj. And, we also divide the states jAdj;Ttjiinto two groups: those corresponding
to theMCartan generators jAdj;Hhji(wherej= 1;:::;M andhjis theM-dimensional
root vector corresponding to Hhj), and those corresponding to the N Mnon-Cartan
generatorsjAdj;Eeji(wherej= 1;:::N Mand ejis theM-dimensional root vector
corresponding to Eej).
Don’t be alarmed by the superscripts being vectors. We are using this notation for later
convenience, and Ttihere means the same thing Tjdid before (the jthgenerator). This
53
notation, which we use only for the adjoint representation, is simply taking advantage
of the fact that in the adjoint representation, the total number of generators, the number
of eigenvectors of the Cartan generators, the dimension of the representation, and the
number of weight/root vectors is the same.
Also, with the adjoint representation states jAdj;Ttji, we can use equation (2.28) to define
the inner product between states as
hAdj;TtjjAdj;Ttki=1
Tr (TtjTtk) =jk(2.31)
We will make use of this equation soon.
The matrix elements of a given generator will then be given by the familiar equation
ifabc= [Tta]bchAdj;TtbjTtajAdj;Ttci
We want to know what an arbitrary generator Ttawill do to an arbitrary state jAdj;Ttbi
in the adjoint representation. So,
TtajAdj;Ttbi=X
cjAdj;TtcihAdj;TtcjTtajAdj;Ttbi=X
cjAdj;Ttci[Tta]cb
=X
cjAdj;Ttci( ifacb) =X
cifabcjAdj;Ttci
And, because there is exactly one eigenvector for each generator, the state jAdj;Ttcicor-
responds to the generator Tc. And because we know that
ifabcTtc= [Tta;Ttb]
(wherecis understood to be summed) by definition of the structure constants, we can
infer that
TtajAdj;Ttbi=X
cifabcjAdj;Ttci=jAdj; [Tta;Ttb]i (2.32)
where [Tta;Ttb]is simply the commutator.
The derivation of equation (2.32) is extremely important, and it is vital that you under-
stand it. However, it is also one of the more difficult results of this already difficult sec-
tion. You are therefore encouraged (again) to read through this section, comparing it with
examples several times until it becomes clear.
So, let’s apply this to combinations of the two types of generators we have, Hha’s and
Eea’s. If we have a Cartan generator acting on a state corresponding to a Cartan generator,
we have (from equation (2.30))
HhajAdj;Hhbi=ha
bjAdj;Hhbi
54
But from (2.32) we have
HhajAdj;Hhbi=jAdj; [Hha;Hhb]i
By definition, the Cartan generators commute, so [Hta;Htb]0, and therefore
hb0 (2.33)
So we can drop them from our notation, leaving the eigentstates corresponding to non-
Cartan generators denoted jAdj;Hji.
On the other hand, if we have a Cartan generator acting on an eigenstate corresponding
to a non-Cartan generator, equation (2.30) gives
HajAdj;Eebi=ea
bjAdj; [Ha;Eeb]i (2.34)
And equation (2.32) gives
HajAdj;Eebi=jAdj; [Ha;Eeb]i (2.35)
Now, we don’t know a priori what [Ha;Eeb]is, but comparing (2.34) and (2.35), we see
jAdj;ea
bEebi=jAdj; [Ha;Eeb]i
And because we know that each of these vectors corresponds directly to the generators,
we have the final result
[Ha;Eeb] =ea
bEeb (2.36)
Now we want to know what a non-Cartan generator does to a given eigentstate. Consider
an arbitrary state jAdj;TtbiwithHceigenvalue tc
b. We can act on this with Eeato create
the new state EeajAdj;Ttbi. So what will the Hceigenvalue of this new state be? Using
(2.36),
HcEeajAdj;Ttbi= (HcEea EeaHc+EeaHc)jAdj;Ttbi= ([Hc;Eea] +EeaHc)jAdj;Ttbi
= (ec
aEea+Eeatc
b)jAdj;Ttbi= (tc
b+ec
a)EeajAdj;Ttbi
= (tb+ ea)cEeajAdj;Ttbi (2.37)
So, by acting on the one of the eigenstates with a non-Cartan generator Eea, we have
shifted the Hceigenvalue by one of the coordinates of the root vector. What this means
is that the non-Cartan generators play a role analogous to the raising and lowering op-
erators we saw in SU(2), except instead of merely shifting the state “up” and “down”, it
moves the states around through some M-dimensional space.
From this, we can also see that if there is an operator that can transform from one state
to another, there must be a corresponding operator that will make the opposite trans-
formation. Therefore, for every operator Eea, we expect to have the operator E ea, and
corresponding eigenstate jAdj;E eai.
55
Finally, consider the state EeajAdj;E eai. We know from (2.32) that EeajAdj;E eai=
jAdj; [Eea;E ea]i. The eigenvalue of this state can be found using equation (2.37):
HbEeajAdj;E eai= ( ea+ ea)bEeajAdj;E eai0
But according to equation (2.33), states with 0 eigenvalue are states corresponding to Car-
tan generators. Therefore we conclude that the state EeajAdj;E eaiis proportional to
some linear combination of the Cartan states,
EeajAdj;E eai=X
bNbjAdj;Hbi (2.38)
where theNb’s are the constants of proportionality. To find the constants Nb, we follow
an approach similar to the one we used in deriving (2.24). Taking the inner product and
using (2.32),
hAdj;HcjEeajAdj;E eai=X
bNbhAdj;HcjAdj;Hbi=X
bNbcb=Nc (2.39)
) hAdj;HcjAdj; [Eea;E ea]i=Nc
Then, using (2.31)
hAdj;HcjAdj; [Eea;E ea]i=1
Tr (Hc[Eea;E ea])
=1
Tr (E ea[Hc;Eea])
=1
ec
aTr (E eaEea)
=ec
aaa
=ec
a
So,
Nc=ec
a
And therefore equation (2.38) is now
EeajAdj;E eai=jAdj; [Eea;E ea]i=eb
ajAdj;Hbi
where the sum over bis understood. This leads to our final result,
[Eea;E ea] =eb
aHb(2.40)
Though we did all of this using the adjoint representation we have seen before, this struc-
ture is the same in any representation, and therefore everything we have said is valid
in anyDn. We worked in the adjoint simply because that makes the results easiest to
obtain. The extensive use we made of labeling the eigenvectors with the generators can
56
only be done in the adjoint representation because only in the adjoint does the number of
eigenvectors equal the number of eigenstates. However, this will not be a problem. The
important results from this section are (2.36) and (2.40), which are true in any representa-
tion. Part of what we will do later is find these structures in other representations.
The importance of the ideas in this section cannot be stressed enough. However, the
material is somewhat abstract. So, we consider a few examples of how all this works.
2.2.13 Adjoint Representation of SU(2)
We now illustrate what we did in section 2.2.12 with SU(2). We will work in the adjoint
representation to make the correspondence with section 2.2.12 as transparent as possible.
SU(2)has 3 generators, and therefore the adjoint representation will consist of 33ma-
trices. This is simply the j= 1representation, which we wrote out in equation (2.27).
First, it is easy to verify that (2.28) and (2.31) hold for = 2.
Next we look at the eigenstates. We know they will be the normal vectors
v1=0
@1
0
01
A; v 2=0
@0
1
01
A; v 3=0
@0
0
11
A
(we will relabel them to be consistent with section 2.2.12 shortly).
Obviously only J3
1is diagonal, so SU(2)has rankM= 1. We define
H1=J3
1=0
@1 0 0
0 0 0
0 0 11
A
E1=J1
1=1
20
@0 1 0
1 0 1
0 1 01
A; E2=J2
1=1
20
@0 i0
i0 i
0i01
A
Because the rank is 1, the root vectors will be 1-dimensional vectors, or scalars. We find
them easily by finding the eigenvalues of each eigenvector with H1:
H1v1= (+1)v1; H1v2= (0)v2; H1v3= ( 1)v3
So the root vectors are
t1=t1= +1 t2=t2= 0 t3=t3= 1 (2.41)
We can graph these on the real line as shown below,
57
Now our initial guess will be to associate v3withJ3
1=H1, and thenv1=E1andv2=E2.
But we want to exploit what we learned in section 2.2.12, and therefore we must make
sure that (2.36) and (2.40) hold.
Starting with (2.36), we check (leaving the tedious matrix multiplication up to you)
[H1;E1] ==1
20
@0 1 0
1 0 1
0 1 01
A (2.42)
[H1;E2] == i
20
@0 1 0
1 0 1
0 1 01
A (2.43)
But we have a problem. According to (2.36), [H1;Ei]should be proportional to Ei, but
this is not the case here. However notice that in (2.42),
1
20
@0 1 0
1 0 1
0 1 01
A=iE2
and in (2.43),
i
20
@0 1 0
1 0 1
0 1 01
A= iE1
58
Writing this more suggestively,
[H1;E1] =iE2;[H1;iE2] =E1(2.44)
So, if we take the linear combinations of equations (2.44), we get [H1;E1iE2] =
E1iE2, which has the correct form of equation (2.36) as long as =. Therefore we
are now working with the operators E(E1iE2).
Now we seek to impose (2.40). We start by evaluating
[E+;E ] =2[E1+iE2;E1 iE2]
=2
[E1;E1] i[E1;E2] +i[E2;E1] + [E2;E2]
= 2i2[E1;E2] =
= 2i2iH1= 22H1
Then, from equations (2.41) and the definition of E, we see thate1
1=(t1 t2) =
(1 0) =1. So we therefore set 2=1
2)=1p
2, and we find that the appropriate
non-Cartan generators (including the 1 to be consistent with the notation in section 2.2.12)
are
E1=1p
2(E1iE2) (2.45)
which is exactly what we had in equation (2.18) above. So, we have derived the trick used
to understand quantum mechanical spin in introductory quantum courses!
2.2.14SU(2)for Arbitrary j::: Again
Now that we have our operators in the adjoint representation, we can consider any arbi-
trary representation. As we saw in section 2.2.11, we can form the linear combinations in
equation (2.45) for any j=integer or half integer. The weight vectors will always look like
those in the diagram on page 58 (in other words, raising and lowering operators always
raise or lower their eigenvalue by 1).
59
The space of physical states, on the other hand, changes for each representation. For
j=1
2, we have
Forj= 1,
60
Forj= 3=2,
and so on.
Notice that the vectors graphed in the diagram on page 58 are the exact vectors required
to move from point to point in each of these graphs. This is obviously not a coincidence.
2.2.15SU(3)
Now that we have said pretty much everything we can about SU(2), which is only Rank 1
(and therefore not all that interesting), we move on to SU(3). However, we will expedite
the process by stating the structure constants up front. The non-zero structure constants
are
f123= 1; f 147=f165=f246=f257=f345=f376=1
2; f 458=f678=p
3
2
The most convenient representation is the Fundamental Representation (consisting of
61
33matrices). They are Ta=1
2afora= 1;:::; 8, where
1=0
@0 1 0
1 0 0
0 0 01
A; 2=0
@0 i0
i0 0
0 0 01
A; 3=0
@1 0 0
0 1 0
0 0 01
A; 4=0
@0 0 1
0 0 0
1 0 01
A
5=0
@0 0 i
0 0 0
i0 01
A; 6=0
@0 0 0
0 0 1
0 1 01
A; 7=0
@0 0 0
0 0 i
0i01
A; 8=1p
30
@1 0 0
0 1 0
0 0 21
A
(2.46)
Clearly, only two of these are diagonal, 3and8. So,SU(3)is a rank 2 group.
Before moving on, we summarize a few results (without proofs). An arbitrary SU(n)
group will always have n2 1generators, and will be rank n 1. An arbitrary SO(n)
group (for neven) will always haven(n 1)
2generators. We won’t worry about the rank of
the orthogonal groups.
Working in the adjoint representation of SU(3)would involve 88matrices, which would
obviously be very tedious. So, we exploit the fact that the techniques we developed in sec-
tion 2.2.12 are valid in any representation, and stick with the Fundamental Representation
defined by the generators in (2.46).
Proceeding as in section 2.2.13, we note that the eigenvectors will again be
v1=0
@1
0
01
A; v 2=0
@0
1
01
A; v 3=0
@0
0
11
A
(we will relabel them to be consistent with section 2.2.12 shortly).
Then, the Cartan generators are
H1=1
20
@1 0 0
0 1 0
0 0 01
A; H2=1
2p
30
@1 0 0
0 1 0
0 0 21
A
and the non-Cartan Generators are simply
E1=T1; E2=T2; E3=T4; E4=T5; E5=T6; E6=T7
So we have 6 eigenvalues to find,
H1v1=1
2
v1; H1v2=
1
2
v2; H1v3= (0)v3
H2v1=1
2p
3
v1; H2v2=1
2p
3
v2; H2v3=
1p
3
v3
62
So the weight vectors will be 2-dimensional (because the rank is 2). They are
t1=
1
21
2p
3T
;t2=
1
21
2p
3T
;t3=
0 1p
3T
(2.47)
We can graph these in R2as shown below,
Now, repeating nearly the identical argument we started with equation (2.42) and repeat-
ing it for all 6 non-Cartan generators, we find that in order to maintain (2.36) and (2.40),
we must work with the operators
1p
2(T1iT2) =1p
2(E1iE2)
1p
2(T4iT5) =1p
2(E3iE4)
1p
2(T6iT7) =1p
2(E5iE6) (2.48)
The weight vectors associated with these will be, respectively,
(t1 t2) =1
0
;(t1 t3) =1=2p
3=2
;and(t2 t3) = 1=2p
3=2
(2.49)
So, the non-Cartan generators are
E0
@1
01
A
; E0
@1=2p
3=21
A
; E0
@ 1=2p
3=21
A
63
We are no longer in the adjoint representation, so we had to be more deliberate about
choosing these linear combinations than we could be in section 2.2.13. What we did here is
more general; we chose them to be the differences in the three weight vectors in equation
(2.47), so that these vectors would naturally transform from one eigenvector to another
(just as the raising and lowering operators do, as we found for SU(2)and more generally
in section 2.2.12). The remarkable property of Lie groups is that this is always possible in
any representation.
We can graph the 6 vectors in (2.49), along with the two Cartan weight vectors, which we
know from (2.33) are 0:
And again, just as with SU(2), notice that the 6 non-zero vectors are the exact vectors
that would be necessary to move from point to point on the diagram on page 63. So once
again, we see that the non-Cartan generators act as raising and lowering operators which
transform between the eigenstates of the Cartan generators. Notice that there were 6 non-
Cartan generators, and they formed linear combinations to form 6 raising and lowering
operators.
2.2.16 What is the Point of All of This?
Before finally getting back to physics, we give a spoiler of how Lie theory is used in
physics. What we are going to find is that some physical interaction (electromagnetism,
weak force, strong force) will ultimately be described by a Lie group in some particular
64
representation. The particles that interact with that force will be described by the eigen-
vectors of the Cartan generators of the group, and the eigenvalues of those eigenvectors
will be the physically measurable charges. Clearly, the number of charges associated with
the interaction is equal to the number of dimensions of the representation. For example,
you likely are aware that the strong force has 3 charges, called “colors” (red, green, and
blue). So, the strong force (we will see) will be in a 3-dimensional representation of the
group that describes it.
We will find that all forces carrying particles (photons, gluons, WandZbosons) will be
described by the generators of their respective Lie group. The Cartan generators will be
force-carrying particles which can interact with any particle charged under that group
by transferring energy and momentum, but do not change the charge (photons and Z
bosons). This makes sense because Cartan generators are notraising or lowering opera-
tors. On the other hand, the non-Cartan generators will be force carrying particles which
interact with any particle charged under that group by not only transferring energy and
momentum, but also changing the charge ( Wbosons and gluons).
We won’t be able to come back to discussing how this works until much later, and until
examples are worked out, this may not be clear. We merely wanted to give an idea of
where we are going with this.
2.3 References and Further Reading
The material in section 2.1 came primarily from [9] and [30]. The material in 2.2 came
from [9], [10], and [15]. The sections on SU(2)also came from [31].
For further reading, we recommend [2], [8], [16], [17], and [28].
65
3 Part III — Quantum Field Theory
3.1 A Primer to Quantization
3.1.1 Quantum Fields
Our ultimate goal in the exposition that follows is to formulate a relativistic quantum me-
chanical theory of interactions . So, beginning with the fundamental equation of quantum
mechanics, Schroedinger’s equation,
H =i~@
@t
we know that for a non-interacting, non-relativistic particle, H=p2
2m= ~
2mr2, so
~
2mr2 =i~@
@t(3.1)
Of course, is in this case a Scalar Field , and therefore only has one state. So, it describes
a spin-0 particle (or, in the language we have learned in the previous sections, it sits in a
j= 0representation of SU(2), which is the trivial representation). And, since does not
have any spacetime indices, it also transforms trivially under the Lorentz group SO(1;3).
Notice, however, that we have a fundamental barrier in making a relativistic theory - the
spatial derivative in (3.1) acts quadratically ( r2), whereas the time derivative is linear.
Clearly, treating space and time differently in this way is unacceptable for a relativistic
theory. That is a hint of a much more fundamental problem with quantum mechanics;
space is always treated as an operator, but time is always treated as a parameter. This
fundamental asymmetry is what ultimately prevents a straightforward generalization to
relativistic quantum theory.
To fix this problem, we have two choices: either promote time to an operator along with
space, or demote space back to a parameter and quantize in a new way.
The first option would result in the Hermitian operators ^X;^Y;^Z, and ^T. It turns out that
this approach is very difficult and less useful as far as building a relativistic quantum
theory. So, we will take the second option.
In demoting position to a parameter along with time, we obviously have sacrificed the
operators which we imposed commutation relations on to get a “quantum” theory in the
first place. And because we obviously can’t impose commutation relations on parameters
(because they are scalars), quantization appears impossible. So, we are going to have to
make a fairly radical reinterpretation.
66
Rather than letting the coordinates be Hermitian operators that act on the state in the
Hilbert space representing a particle, we now interpret the particle as the Hermitian operator ,
and thisoperator (or particle) will be parameterized by the spacetime coordinates. The
physical state that the particle operators act on is then the vacuum itself, j0i. So, whereas
before you acted on the “electron” j iwith the operator ^x, now the “electron” (parame-
terized byx) (x)acts on the vacuum j0i, creating the state
(x)j0i
. In other words,
the operator representing an electron excites the vacuum (empty space) resulting in an
electron. We will see that all quantum fields contain appropriate raising and lowering
operators to do just this.
This approach, where the quantum mechanical entities are no longer the coordinates act-
ing on the fields, but the fields themselves, is called Quantum Field Theory (QFT).
So, whereas before, quantization was defined by imposing commutation relations on the
coordinate operators [x;p]6= 0, we now quantize by imposing commutation relations on
the field operators, [ 1; 2]6= 0.
Because we must still write down the equations of motion which govern the dynamics
of the fields, we will need to spend the rest of this section coming up with the classical
equations governing the fields we want to work with. We will quantize them in the next
section.
3.1.2 Spin-0 Fields
As we said above, Schroedinger’s equation (3.1) describes the time evolution of a spin-0
field, or a scalar field. Generalizing to higher spins will come later. Now, we see how to
make this description relativistic.
The most obvious guess for a relativistic form is to simply plug in the standard relativistic
Hamiltonian
H=p
p2c2+m2c4 (3.2)
Note that, using the standard Taylor expansionp
1 +x21 +1
2xforx << 1givesH
mc2+p2
2m, for p2<<c2, which is the standard non-relativistic form (plus a constant) we’d
expect from a low speed limit.
Plugging (3.2) into (3.1), we have
i~@
@t=p
~2c2r2+m2c4
But there are two problems with this:
1. The space and time derivatives are still treated differently, so this is inadequate as a
relativistic equation, and
67
2. Taylor expanding the square root will give an infinite number of derivatives acting
on, making this theory non-local.
One solution is to square the operator on both sides, giving
~2@2
@t2= ( ~2c2r2+m2c4)
)( @0@0+r2 m2c2
~2)= 0
Or, if we choose the so called “natural units” or “God units”, where c=~= 1, we have
(@2 m2)= 0 (3.3)
Equation (3.3) is called the Klein Gordon equation. It is nothing more than an operator
version of the standard relativistic relation E2=m2c4+ p2c2.
Note that because we will be quantizing fields and not coordinates, there is absolutely
nothing “quantum” about the Klein Gordon equation. It is, at this point, merely a rela-
tivistic wave equation for a classical, spinless, non-interacting field.
Finally, we note one major problem with the Klein Gordon equation. When we squared
the Hamiltonian H=p
m2c4+ p2c2to getH2=m2c4+ p2c2, the energy eigenvalues
becameE=p
m2c4+ p2c2. It appears that we have a negative energy eigenvalue! Ob-
viously this is unacceptable in a physically meaningful theory, because negative energy
means that we don’t have a true vacuum, and therefore a particle can cascade down for-
ever, giving off an infinite amount of radiation.
We will see that this problem plagues the spin- 1=2particles as well, so we wait to talk
about the solution until then.
3.1.3 Why SU(2)for Spin?
Because we are talking about particles “spinning”, a common question is why don’t we
useSO(3)instead ofSU(2)? The original answer to the question is historical. The experi-
ments done in the early days of quantum mechanics were not consistent with the particles
having a rotational degree of freedom in spacetime. Rather, the data indicated that, along
any given axis, the spin could have only one of two possible values, and SO(3)does not
explain this. Here, however, we consider a more mathematical explanation.
First, recall that spin is a purely quantum mechanical phenomenon, with no classical
analogue. Because the data demanded two possible spin states, the field describing the
particle had to have 2 spin components, = 1
2
. Now, if we seek a 2-dimensional
68
representation of SO(3), we find that there is only one: D0D0, the trivial representation
consisting of all 1’s. This means
0
1
0
2
=D0D0 1
2
=1 0
0 1 1
2
= 1
2
which is no transformation. This is the only 2-dimensional representation of SO(3)that
is possible.
The solution to the problem is found in one of the many peculiarities of quantum me-
chanics. The only physically measurable quantity in quantum theory is the probability
amplitude, which is proportional to the square of . Therefore, the state 1
2
is physi-
cally identical to 1
2
.
Now consider a general element of SO(3):ei(Jx+ Jy+Jz). On the other hand, a general
element of SU(2)will beei(x
2+ y
2+z
2).
Now consider rotating the system by an angle of 2around, say, the zaxis. TheSO(3)
element corresponding to this rotation will be ei2Jz, while the SU(2)element will be
eiz. The factor of 1/2 difference means that the spinor space rotates through only half
the angle of the SO(3)does. So, in the 2rotation,U2SU(2)! U, whereasR2
SO(3)!R. Therefore, both Uand Ucorrespond to R. There is a 2 to 1 correspondence
betweenSU(2)andSO(3).
And, as we said above, spin is a purely quantum mechanical effect and experimentally
only allows 2 values, but SO(3)has no such representation, whereas the j= 1=2rep-
resentation of SU(2)does. We therefore use SU(2). And, because SU(2)is2!1with
SO(3), but spin is quantum mechanical, both Uand Ucan consistently correspond to
the sameR2SO(3). The minus sign difference is not subject to measurement; only j j2
is physically measurable.
An important thing to understand is that “spin” is not a rotation through spacetime in any
meaningful way. It is a rotation in “spinor space”, which is an internal degree of freedom.
Like many things in quantum mechanics, spinor space is a mathematical structure. All
we can say for certain is what we can measure, or know ( j j2), not what “is”.
3.1.4 Spin1
2Particles
Finding equation (3.3) was easy because scalar fields have no spacetime indices and no
spinor indices, and they therefore transform trivially under SU(2)and the Lorentz group.
A particle of spin 1=2however, will have two complex components, one for spin +1=2,
69
and the other for spin 1=2. So, we describe such a particle as the two-component Spinor
= 1
2
where 1and 2are both2C. So, we want some differential operator in the form of 22
matrices to act on such a field to form the equation of motion.
Following Dirac’s approach, he reasoned that given such a 22operator, the equation of
motion should somehow “imply” the Klein Gordon equation (which merely makes the
theory relativistic). So his goal (and our goal) is to find an equation with a 22matrix
differential operator acting on that results in (3.3).
Dirac’s approach was to find an operator of the form
6D=
@=
0@0+
1@1+
2@2+
3@3
where the
’s are 22matrices, and the equation of motion is then 6D = im . The
challenge is in finding the appropriate 22
matrices. Dirac reasoned that, in order to
be properly relativistic, operating twice with 6Dshould give the Klein Gordon equation.
In other words,
6D= im ) 6D6D = im6D
)
@
@ = im( im )
)
@@ = m2
)
@@+m2
= 0 (3.4)
This will yield the Klein Gordon equation if
= I. Or, using the symmetry of the
sum in (3.4), it will yield the Klein Gordon equation if we demand1
2(
+
) = I.
Consider
f
;
g=
+
= 2I (3.5)
If the
matrices satisfy (3.5), then (3.4) gives
(
@@+m2) = 0)( @@+m2) = 0)(@2 m2) = 0
which is exactly the Klein Gordon equation (3.3).
So, we have the Dirac equation
6D+im
= 0 (3.6)
but we still have a problem. Namely, there does not exist a set of 22matrices that solve
(3.5). Nor does there exist a set of 33matrices. The smallest possible size where this is
possible is 44. Obviously, if we want to describe a spin- 1=2particle with exactly 2 spin
70
states, using 4 spin components does not seem right. But, we will accept the necessity of
44Dirac matrices and move on.
Instead of using = 1
2
, we will define the two 2-dimensional spinors
L 1
2
and R 3
4
(3.7)
and the 4-component spinor
L
R
(3.8)
Now it is possible to solve (3.5). Such a problem is actually very familiar to algebraists,
and we will not delve into the details of how this is done. Instead, we merely state one
solution (there are many, up to a similarity transformation). We define the 4 4 matrices
i=0 i
i0
and
0=00
00
(3.9)
where0is the 22 identity matrix, and iare the Pauli spin matrices. It should be no
surprise that they show up in attempting to describe spin- 1=2particles. What is interest-
ing is that we did not assume them—we derived them using (3.5).
Before moving on, notice that we have initiated a convention that will be used throughout
the rest of these notes. Whenever a greek index is used, it runs over all spacetime indices.
Whenever a latin index is used, it runs over only the spatial part. So in (3.9), iruns 1;2;3.
Now that we have an explicit form of the Dirac gamma matrices, we can write out (3.6)
explicitly:
0
BB@0 0 @0 @3 @1+i@2
0 0 @1 i@2@0+@3
@0+@3@1 i@2 0 0
@1+i@2@0 @3 0 01
CCA0
BB@ 1
2
3
41
CCA= im0
BB@ 1
2
3
41
CCA
Or, in terms of Land R,
i@ R= +m L
i@ L= +m R
where we have defined the 4-vectors = (0;12;3)and= (0; 1; 2; 3).
3.1.5 The Lorentz Group
This section is intended to give a deeper understanding of why we were unable to find a
22 matrix representation to solve (3.5).
71
Recall that the driving idea behind the derivation of the Dirac equation (3.6) was to make
it imply the Klein Gordon equation, or in other words to be a relativistic theory. Put
another way, it was to create a theory that was invariant under the Lorentz group SO(1;3).
So, let’s take a closer look at the Lorentz group.
We know from section 1.1.5 that the Lorentz group consists of 3 rotations and 3 boosts.
We gave the general forms of these transformations in equations (1.5) and (1.6). It is
easy, using those general expressions in addition to (2.7), to find all 6 generators, and
then multiply them out to get the commutation relations. We spare the (easy but tedious)
details and simply state the commutation relations. If we label the generators of rotation
Ji(i= 1;2;3), and the generators of boosts Ki(i= 1;2;3), then the commutation relations
are
[Ji;Jj] =iijkJk
[Ji;Kj] =iijkKk
[Ki;Kj] = iijkJk
In order to make the actual structure of this group more obvious, we define two new
linear combinations of these generators:
Ni=1
2(Ji iKi)Niy=1
2(Ji+iKi)
Writing out the commutation relations for NiandNiy, we get
[Ni;Nj] =iijkNk
[Niy;Njy] =iijkNky
[Ni;Njy] = 0
So, bothNiandNiyseparately form an SU(2). In more mathematical terms, we say that
SO(1;3)isIsomorphic toSU(2)
SU(2), which we denote SO(1;3)=SU(2)
SU(2).
While the idea of an isomorphism is a very rich mathematical idea, for now you can
simply think of it as a way of saying that two groups have the same group structure.
So, because a given representation of SU(2)is defined by the value of j, we can see that
a particular representation of the Lorentz group SO(1;3)=SU(2)
SU(2)is defined
bytwovalues ofj, or by the doublet (j;j0). The smallest possible representation then is
(j;j0) = (0;0). This has one state from j= 0and one state from j0= 0, and therefore has
11 = 1 state total. Therefore, this representation describes a scalar field.
Then, there is the state (0;1=2), which will have one state from j= 0, but two states from
j= 1=2, for a total of 12 = 2 states. Therefore, this describes a single spin- 1=2field. We
call this field L, and the (0;1=2)representation the Left-Handed Spinor Representation
of the Lorentz group.
72
Clearly, we will also have the representation (1=2;0), which also has 2 states, correspond-
ing to the Rfield. This is called the Right-Handed Spinor Representation of the Lorentz
group. This is the reason for the notation used in (3.7). The left-handed (0;1=2)represen-
tation acts on Land the right-handed (1=2;0)representation acts on R.
Next is the representation (1=2;1=2), which has two states from j= 1=2and two from the
j0= 1=2for a total of 22 = 4 states. It turns out that this representation is the space-
time vector representation we use to act on spacetime vectors for the standard Lorentz
transformations discussed in section 1.1.5.
Now, anSU(2)representation specified by some jis an irreducible representation, and
therefore the tensor products SU(2)
SU(2)specified by the doublet (j;j0)are irreducible.
This means that there are no irreducible subspaces, and so given a representation (j;j0),
there is a particular transformation taking the state (j;j0)to(j0;j). For the (0;0)and
(1=2;1=2)representations this doesn’t affect anything. However, this fact means that the
(0;1=2)and (1=2;0)representations must always appear together. To put this in more
mathematical language, our choices for representations of the Lorentz group are
(0;0); (1=2;1=2) and (1=2;0)(0;1=2)
which are 1, 4, and 4-dimensional representations, respectively. Furthermore, they are
the representations which transform Klein Gordon scalar/spinor-0 fields, spacetime 4-
vectors, and spin- 1=2spinors, respectively.
The physical meaning of this fact is that relativity demands that if you want a theory
with spin- 1=2particles, you cannot have them existing by themselves. They must come
in pairs, each transforming under an SU(2)representation of opposite handedness. In
the next two sections we will discuss ways of interpreting this fact, starting with Dirac’s
original approach which, while brilliant, didn’t ultimately work. Then we will consider
what appears to be the correct view.
3.1.6 The Dirac Sea Interpretation of Antiparticles
Initially, it may seem that the impossibility of finding a 2 2 matrix solution to (3.5)
means that we can’t have fields with 2 spinor states. However, we saw in the last section
that we aren’t limited to scalars and spacetime 4-component spinors. We can also have
two fields, Land R, which can be paired together to form two spin- 1=2fields in a 4-
component spinor = L
R
. So, Dirac was faced with the challenge of both interpreting
this, while at the same time dealing with the negative energy states mentioned in section
3.1.2.
Dirac’s solution, though today abandoned, was brilliant enough to mention. He sug-
gested that because spin- 1=2particles obey the Pauli Exclusion Principle , there could be
73
an infinite number of particles already in the negative energy levels, and so they are al-
ready occupied, preventing any more particles from falling down and giving off infinite
energy. Thus, the negative energy problem was solved.
Furthermore, he said that it is possible for one of the particles in this infinite negative sea
to be excited and jump up into a positive energy state, leaving behind a hole. This would
appear to us, experimentally, as a particle with the same mass, but the opposite charge.
He called such particles Antiparticles . For example, the antiparticle of the electron is the
antielectron, or the positron (same mass, opposite charge). The positron is not a particle
in the same sense as the electron, but rather is a hole in an infinite sea of electrons. And
where this negative charge is missing, all that is left is a hole which appears as a positively
charged particle.
So, Ldescribes a particle, and due to the infinite sea of negative particles, there can
always be a hole, which will be described by L. Everything about this worked out math-
ematically, and when antiparticles were detected about 5 years after Dirac’s prediction of
them, it appeared that Dirac’s suggestion was correct.
However, there were two major problems with Dirac’s idea, and they ultimately proved
fatal to the “Dirac Sea” interpretation:
1. This theory, which was supposed to be a theory of single particles, now requires an
infinite number of them.
2. Particles like photons, pions, mesons, or Klein-Gordon scalars don’t obey the Pauli
Exclusion Principle, but still have negative energy states, and therefore Dirac’s ar-
gument doesn’t work.
However, his labeling them “antiparticles” has stuck, and we therefore still refer to the
right-handed part of the spin- 1=2field as the antiparticle, whereas the left-handed part is
still the particle.
For these reasons, we must have some other way of understanding the existence of the
antiparticles.
3.1.7 The QFT Interpretation of Antiparticles
In presenting the problem of negative energy states, we have been somewhat intention-
ally sloppy. To take stock, we have two equations of motion: the Klein Gordon (3.3) for
scalar/spin-0 fields, and the Dirac equation (3.6) for spin- 1=2particles.
And in our discussion of negative energy states, we were “pretending” that the ’s and
’s are “states” with negative energy. But, as we said in section 3.1.1, QFT offers a different
interpretation of the fields. Namely, the fields are not states — they are operators. And
74
consequently they can’t have energy. A state is made by acting on the vacuum with either
of the operatorsor , and then the state j0ior j0ihas some energy.
So, QFT allows us to see the antiparticle as a real, actual particle, rather than the absence
of a particle. And, we do not need the conceptually difficult idea of an infinite sea of
negative energy particles. The vacuum j0i, with no particles in it, is now our state with
the lowest possible energy level. And, as we will see, there are never negative energy
states with these particles.
How exactlyj0iworks will become clearer when we quantize. The point to be understood
for now is that QFT solves the problem of negative energy by reinterpreting what is a state
and what is an operator. The fields and are operators, not states, and therefore they
do not have energy associated with them (any more than the operator ^xor^pxdid in non-
relativistic quantum mechanics). So, without any problems of negative energy, we merely
accept that nature, due to relativity, demands that particles come in particle/antiparticle
pairs, and we move on.
3.1.8 Lagrangians for Scalars and Dirac Particles
Now that we have the equations of motion (3.3) and (3.6), we want to know the actions
that lead to these equations of motion. In order to save time, we will merely write down
the answers and let you take the variations to see that they do indeed lead to the Klein
Gordon and Dirac equations of motion for and .
They are
LKG = 1
2@@ 1
2m2 (3.10)
LD=i y
L@ L+i y
R@ R m( y
L R+ y
R L) (3.11)
where the dagger represents the Hermitian conjugate, y
L= ( ?
1; ?
2), y
R= ( ?
3; ?
4), as
usual. You can actually take the variation of LDwith respect to either y
Land y
Rto get
the equations of motion for Land R, or you can take the variations with respect to
Land Rto get the equations for y
Land y
R. The two sets of equations are simply the
conjugates of each other, and therefore represent a single set of equations.
In order to simplify (3.11), the convention is to use the Dirac gamma matrices (3.9) to
define y
0(where here is the 4-component spinor in equation (3.8)). Using this,
all 4 terms in (3.11) can be summarized as
LD= (i
@ m) (3.12)
75
3.1.9 Conserved Currents
In Part I we discussed how symmetries and conserved quantities are related. Let’s con-
sider a few examples of this using the Lagrangians we have now defined.
Consider a massless Klein Gordon scalar particle, described by L= 1
2@@. Following
what we did starting with equation (1.2), consider the transformation !+, where
is a constant. Because @!@+@=@, the Lagrangian is invariant. So (using
= 1), our conserved quantity is
j=@L
@(@)= @
Or, consider the Klein Gordon Lagrangian with complex scalar fields andy, which
we write asL= @y m2y. We can make the transformation !eiand
y!ye i(whereis an arbitrary real constant). This type of transformation is called a
U(1)transformation, because eiis an element of the group of all 1 1 unitary matrices,
as discussed in section 2.2.1.
The conserved quantity associated with this U(1)symmetry is
j=@L
@(@)+@L
@(@y)y=i(@y y@)
Or consider the Dirac Lagrangian. Notice that it is invariant under the U(1)transforma-
tion !ei, with current
j=
(3.13)
In both of the previous examples, notice that the U(1)symmetry changes the field at all
points in space at once, and all in the same way. In other words, it is a single overall con-
stant phase ei. We therefore call such a symmetry a Global Symmetry . The implications
of this are likely not clear at this point. We merely wish to call your attention to the fact
thateihas no spacetime dependence.
3.1.10 The Dirac Equation with an Electromagnetic Field
Previously we found the Lagrangian for an electromagnetic field (1.13). Our goal now is
to find a Lagrangian that describes the electromagnetic field and a spin- 1=2particle that
couples to the electromagnetic field, and additionally the interaction between them. We
start by writing down a Lagrangian without any interaction. This will simply be the sum
of the two terms,
L=LD+LEM= (i
@ m) 1
4FF JA (3.14)
76
But, because the Dirac part has no terms in common with the electromagnetic part, the
equations of motion and the conserved quantities for both andAwill be exactly the
same, as if the other weren’t present at all. In other words, both fields go about their way
as if the other weren’t there—there is no interaction in this theory. Because this makes for
a boring universe (and horrible phenomenology), we need to find some way of coupling
the two fields together to produce some sort of interaction.
Interaction is added to a physical theory by adding another term to the Lagrangian called
theInteraction Term . So, the final Lagrangian will have the form L=LD+LEM+Lint.
Now, for reasons that will become clear in the next section (and even more clear when
we quantize), we do this by coupling the electromagnetic field Ato the current resulting
from theU(1)symmetry inLD, which we discussed in section 3.1.9, and wrote out in
equation (3.13). In other words, our interaction term will be proportional to Aj.
So, adding a constant of proportionality q(which we will see has the physical interpreta-
tion of a coupling constant, weighting the probability of an interaction to take place, or
equivalently the physical interpretation of electric charge), our Lagrangian is now
L= (i
@ m) 1
4FF JA qjA
= (i
@ m) 1
4FF (J+q
)A (3.15)
Notice thatLis still invariant under the global U(1)symmetry, and the U(1)current is
stillJ=
.
Also, notice that the Lagrangians in (3.14) and (3.15) are the same except for a shift in the
current term, J!J+qj. Recall that physically, J= (;J)represents the charge
and current creating the field. The fact that Jhas shifted in (3.15) simply means that the
spin- 1=2particle in this theory contributes to the field, which is exactly what we would
expect it to do.
If we setq=e, the electric charge, this Lagrangian becomes upon quantization the La-
grangian of Quantum Electrodynamics ( QED ), which to date makes the most accurate
experimental predictions ever.
In the next section, we will re-derive this Lagrangian in a more fundamental way.
3.1.11 Gauging the Symmetry
Physically speaking, this section is among the most important in these notes. Read this
section again and again until you understand every step.
Consider once again the Dirac Lagrangian (3.6). As we said in section 3.1.9, it is invariant
under the globalU(1)transformation !ei . It is global in that it acts on the field the
77
exact same way at every point in spacetime. The idea behind this section is that we are
going to make this symmetry Local , so thatdepends on spacetime ( =(x)), and then
try to force the Lagrangian to maintain its invariance under the localU(1)transformation.
Making a global symmetry local is referred to as Gauging the symmetry.
We start by making the local U(1)transformation:
L= (i
@ m) ! e (x)(i
@ m)ei(x)
and because the differential operators will now act on (x)as well as , we get extra
terms:
L! e (x)(i
@ m)ei(x) = (i
@ m)
@(x)
= (i
@ m
@(x))
If we want to demand that Lstill be invariant under this local U(1)transformation, we
must find a way of canceling the
@(x)term. We do this in the following way.
Define some arbitrary field Awhich under the U(1)transformation ei(x)transforms ac-
cording to
A!A 1
q@(x) (3.16)
We callAtheGauge Field for reasons that will be clear soon, and qis a constant we have
included for later convenience.
We introduce Aby replacing the standard derivative @with the Covariant Derivative
D@+iqA (3.17)
If you have studied general relativity or differential geometry at any point, you are famil-
iar with covariant derivatives. There is an incredibly rich geometric picture of all of this,
but it is beyond the scope of these notes. We will deal with it later in this series, however.
As a comment regarding vocabulary, to say that a particle “carries charge” mathemati-
cally means that it has the corresponding term in its covariant derivative. So, if a parti-
cle’s covariant derivative is equal to the normal differential operator @, then the particle
has no charge, and it will not interact with anything. But if it carries charge, it will have a
term corresponding to that charge in its covariant derivative. This will become clearer as
we proceed.
So, our Lagrangian is now
L= (i
D m) = (i
[@+iqA] m) = (i
@ m q
A)
78
And under the local U(1)we have
L ! e i(x)(i
@ m q
[A 1
q@(x)])ei(x)
= (i
@ m
@(x) q
A+
@(x))
= (i
@ m q
A) = (i
D m)
=L
So, the addition of the field Ahas indeed restored the U(1)symmetry. Notice that now it
is not only invariant under this local U(1), but also still under the global U(1)we started
with, with the same conserved U(1)currentj=
. This allows us to rewrite the
Lagrangian as
L= (i
D m) = (i
@ m) qjA (3.18)
But we have a problem. If we want to know what the dynamics of Awill be, we nat-
urally take the variation of the Lagrangian with respect to A. But because there are
no derivatives of A, the Euler-Lagrange equation is merely@L
@A= q
= 0. But
q
= qj. So the equation of motion for Asays that the current vanishes, or that
j= 0, and so the Lagrangian is reduced back to (3.12), which was not invariant under
the localU(1).
We can state this problem in another way. All physical fields have some sort of dynamics.
If they don’t then they are merely a constant background field that never changes and
does nothing. As it is written, equation (3.18) has a field AbutAhas no kinetic term,
and therefore no dynamics.
So, to fix this problem we must include some sort of dynamics, or kinetic terms, for A.
The way to do this turns out to involve a considerable amount of geometry which would
be out of place in these notes. We will cover the necessary ideas in a later paper in this
series and derive the following expressions. For now we merely give the results and ask
you for patience until we have the machinery to derive them.
For an arbitrary field A, the appropriate gauge-invariant kinetic term is
LKin;A = 1
4FF
where
Fi
q[D;D] (3.19)
andqis the constant of proportionality introduced in the transformation of Ain equation
(3.16).Dis the covariant derivative defined in (3.17).
79
Writing out (3.19) (and using an arbitrary test function f(x)),
Ff(x) =i
q[D;D]f(x)
=i
q
(@+iqA)(@+iqA) (@+iqA)(@+iqA
f(x)
=i
q
@@f(x) +iq@(Af(x)) +iqA@f(x) q2AAf(x)
@@f(x) +iq@(Af(x)) +iqA@f(x) q2AAf(x)
=i
q
iqf(x)@A+iqA@f(x) +iqA@f(x) q2AAf(x)
iqf(x)@A iqA@f(x) iqA@f(x) +q2AAf(x)
=
@A @A+iq[A;A]
f(x)
But for each value of ,Ais a scalar function, so the commutator term vanishes, leaving
(dropping the test function f(x))
F=i
q[D;D] =@A @A(3.20)
So, writing out the entire Lagrangian we have
L= (i
D m) 1
4FF
And finally, because Ais obviously a physical field, we can naturally assume that there
is some source term causing it, which we simply call J. This makes our final Lagrangian
L= (i
D m) 1
4FF JA
Comparing this to (3.15) we see that they are exactly the same. So what have we done?
We started with nothing but a Lagrangian for a spin- 1=2particle, which had a global U(1)
symmetry. Then, all we did was promote the U(1)symmetry to a local symmetry (we
gauged the symmetry), and then imposed what we had to impose to get a consistent
theory. The gauge field Awas forced upon us, and the form of the kinetic term for Ais
demanded automatically by geometric considerations we did not delve into.
In other words, we started with nothing but a non-interacting particle, and by specify-
ingnothing butU(1)we have created a theory with not only that same particle, but also
electromagnetism. The Afield, which upon quantization will be the photon, is a direct
consequence of the U(1).
This is what we meant at the end of section 2.2.11 when we said that electromagnetism is
described by U(1). We will talk more about the weak and strong forces later, as well as
the groups that give rise to them.
80
Theories of this type, where we generate forces by specifying a Lie group, are called
Gauge Theories , orYang-Mills Theories .
Finally, notice that (3.16) has exactly the same form as (1.15). This is why we call Aa
gauge field. The gauge symmetry in electromagnetism is a sort of remnant of the much
deeper and more fundamental U(1)structure of the theory.
3.2 Quantization
3.2.1 Review of What Quantization Means
In quantum mechanics (not QFT), quantization is done by taking certain dynamical quan-
tities and making use of the Heisenberg Uncertainty Principle . Normally we take posi-
tion xand momentum pand, according to Heisenberg, the measurement of the particle’s
position will effect its momentum and vice-versa.
To make this more precise, we promote xandpfrom merely being variables to being
Hermitian operators ^xand^p(which can be represented by matrices) acting on some vector
space. Calling a vector in this space j i, physically measurable quantities (like position
or momentum) become the eigenvalues of the operators ^xand^p,
^xj i=xj i
^pj i=pj i
Heisenberg Uncertainty says that measuring xwill affect the value of p, and vice-versa.
It is the act of measuring which enacts this effect. It is not an engineering problem in
the sense that there is no better measurement technique which would undo this. It is a
fundamental fact of quantum mechanics (and therefore the universe) that measurement
of one variable affects another.
So, if we measure x(using ^x) and then p(using ^p), we will in general get different values
for both than if we measured pand thenx. More mathematically, ^x^p6= ^p^x. Put another
way,
[^x;^p]^x^p ^p^x6= 0
For reasons learned in an introductory quantum course, the actual relation is
[^x;^p] =i~ (3.21)
where~is Planck’s Constant. We call (3.21) the Canonical Commutation Relation , and it
is this structure which allows us to determine the physical structure of the theory.
More generally, we choose some set of operators that all commute with each other, and
then label a physical state by its eigenvectors. For example ^x,^yand ^zall commute with
81
each other, so we may label a physical state by its eigenvectors j ri=jx;y;zi. Or,
because ^px,^py, and ^pzall commute, we may call the state j pi=jpx;py;pzi. We may
also include some other values like spin and angular momentum, to have (for example)
j i=jx;y;z;sz;Lz;:::i.
As discussed in section 3.1.1, when we make the jump to QFT, the fields are no longer the
states but the operators. We are therefore going to impose commutation relations on the
fields, not on the coordinates.
Furthermore, whereas before the states were eigenvectors of the coordinate operators, we
now will expand the fields in terms of the eigenvectors of the Hamiltonian.
3.2.2 Canonical Quantization of Scalar Fields
We begin with the Klein Gordon Lagrangian in equation (3.10), but we make the slight
modification of adding an arbitrary constant
,
LKG= 1
2@@ 1
2m22+
Note that
has absolutely no affect whatsoever on the physics.
Quantization then comes about by defining the field momentum and Hamiltonian (using
(1.7) and (1.8)) to get
=@L
@_(x)=_ (3.22)
H= _ L=1
22+1
2(r)2+1
2m22
(3.23)
Now, using the canonical commutation relations (3.21) as guides, we impose
[(t;x);(t0;x0)] = 0
[(t;x);(t0;x0)] = 0
[(t;x);(t0;x0)] =i(t t0)(x x0) (3.24)
(where we have set ~= 1).
We can see more clearly what this means if we expand the solutions of the Klein Gordon
equation. One solution is plane waves, eikxi!t, where
!= +pk2+m2 (3.25)
andkis the standard wave vector.
82
So, we write the field as
(t;x) =Zd3k
f(k)
a(k)eikx i!t+b(k)eikx+i!t
wheref(x)is a redundant function which we have included for later convenience. For
now, both a(k)andb(k)are merely arbitrary coefficients (integration constants) used to
expand(t;x)in terms of individual solutions.
We demand that (t;x)be Hermitian. This requires
y=)?=)b?(k) =a( k)
Then, changing the sign of the integration variable kon the second term in the integral
allows us to use 4-vector notation, so
(x) =Zd3k
f(k)
a(k)eikx+a?(k)e ikx
wherekx=kx.
Now notice that the integration measure, d3k, is not invariant under Lorentz transforma-
tions (because it integrates over the spatial part but not over the time part). We therefore
choosef(k)to restore Lorentz invariance.
We know that the measure d3kwould be invariant, as would functions and (step)
functions. So, consider the invariant combination
d4k(k2+m2)(k0) (3.26)
Thefunction merely requires that relativity hold ( k2+m2is simply the relativistic rela-
tion (3.2), and the function preserves causality. So this is a physically acceptable Lorentz
invariant integration measure.
Recall the general function identity,
Z1
1dx(g(x)) =X
i1dg(x)
dxjx=xi
where thexi’s are the zeros of the function g(x). We can do the k0integral over measure
(3.26), and using the fact that the zeros of k2+m2=k2 k0k0+m2in terms of k0are
k0k0=k2+m2=!2, we get
Z
d3kdk0(k2+m2)(k0) =Zd3k
2!
So, adding a factor of (2)3for later convenience, we take our invariant measure to be
d3k
(2)32!
83
So finally,
(x) =Z
fdk
a(k)eikx+a?(k)e ikx
(3.27)
where we have defined fdkd3k
(2)32!.
The commutation relations we defined in (3.24) will now hold provided we impose
[a(k);a(k0)] = 0
[ay(k);ay(k0)] = 0
[a(k);ay(k0)] = (2)32!3(k k0) (3.28)
(showing this is fairly tedious, but we encourage you to work it out). We are using y
instead of?to emphasize that, in the quantum theory, we are talking about Hermitian
operators. The operators a(k)anday(k)are scalars, so in this case a?=ay.
Furthermore, we can write the Hamiltonian Hin terms of (3.27):
H=Z
d3xH=Z
d3x1
22+1
2(r)2+1
2m22
=1
2Z
fdkfdk0d3x[( i!a(k)eikx+i!a?(k)e ikx)( i!0a(k0)eik0x+i!0a?(k0)e ik0x)
+(ika(k)eikx ika?(k)e ikx)(ik0a(k0)eik0x ik0a?(k0)e ik0x)
m2(a(k)eikx+a?(k)e ikx)(a(k0)eik0x+a?(k0)e ik0x)] Z
d3x
=1
2Z
fdkfdk0d3x[( !!0a(k)a(k0)ei(k+k0)x+!!0a(k)a?(k0)ei(k k0)x
+!!0a?(k)a(k0)e i(k k0) !!0a?(k)a?(k0)e i(k+k0)x)
+( kk0a(k)a(k0)ei(k+k0)x+kk0a(k)a?(k0)ei(k k0)x
+kk0a?(k)a(k0)e i(k k0)x kk0a?(k)a?(k0)e i(k+k0)x)
+m2(a(k)a(k0)ei(k+k0)x+a(k)a?(k0)ei(k k0)x
+a?(k)a(k0)e i(k k0)x+a?(k)a?(k0)e i(k+k0)x)
V
whereVis the volume of the space resulting from theR
d3xintegral. Then, from the fact
thatR
d3xeixy= (2)33(y), we have
H=1
2(2)3Z
fdkfdk0[3(k k0)(!!0+kk0+m2)(a?(k)a(k0)e i(! !0)t+a(k)a?(k0)e i(! !0)t)
+3(k+k0)( !!0 kk0+m2)(a(k)a(k0)e i(!+!0)t+a?(k)a?(k0)ei(!+!0)t)
V
=1
2Z
fdk1
2![(!2+k2+m2)(a?(k)a(k) +a(k)a?(k))
+( !2+k2+m2)(a(k)a( k)e 2i!t+a?(k)a?( k)e2i!t)] V
84
and finally, using the definition of !(equation (3.25)), this becomes
H=1
2Z
fdk!(a?(k)a(k) +a(k)a?(k)) V
And now, using (3.28), we can rewrite this as (switching from ?toyto emphasize the
Hermitian nature)
H=1
2Z
fdk!(ay(k)a(k) +a(k)ay(k)) V
=1
2Z
fdk!(ay(k)a(k) + (2)32!3(k k) +ay(k)a(k)) V
=Z
fdk!ay(k)a(k) +Z
fdk!(2)33(0) V
=Z
fdk!ay(k)a(k) +Zd3k
(2)32!!(2)33(0) V
=Z
fdk!ay(k)a(k) +1
23(0)Z
d3k V
Notice that both the second and third terms are infinite (assuming the volume Vof the
space we are in is infinite). This may be troubling, but remember that
is an arbitrary
constant we can set to be anything we want. So, let’s define
1
2V3(0)Z
d3k
leaving
H=Z
fdk!ay(k)a(k) (3.29)
Remember that measurement can only detect changes in energy, and therefore the infin-
ity we subtracted off does not affect the value we will measure experimentally. What
we have done here, by subtracting off the infinite part in a way that doesn’t change the
physics, is a very primitive example of Renormalization . Often, for various reasons,
measurable quantities in QFT are plagued by different types of infinities. However, it is
possible to subtract off those infinities in a well-defined way, leaving a finite part. It turns
out that this finite part is the correct value seen in nature. The reasons for this are very
deep, and we will not discuss them (or general renormalization theory) in much depth in
these notes. For correlating theoretical results with experiment, being able to renormalize
results correctly is vital. However, our goal is not to understand the subtleties of renor-
malization, but to understand the overall structure of particle physics. When you take a
course on QFT you will spend a great, great deal of time on renormalization, and a deeper
understanding of it will emerge.
85
So, we have our field expansion (3.27) and commutation relations (3.28). Notice that
(3.28) have the exact form of a simple harmonic oscillator, which you learned about in
introductory quantum mechanics. Therefore, because they have the same structure as
the harmonic oscillator, they will have the same physics. By doing nothing but imposing
relativity, we have found that scalar fields, which are Hermitian operators, act as raising
and lowering (or synonymously creation and annihilation) operators on the vacuum (just
like the simple harmonic oscillator).
Comparing (3.28) with the standard harmonic oscillator operators, it is clear that ay(k)
creates aparticle with momentum kand energy !, whereasa(k)annihilates a particle
with momentum kand energy !. A normalized state will be
jki=p
2!ay(k)j0i (3.30)
The entire spectrum of states can be studied by acting on j0iwith creation operators, and
probability amplitudes for one state to be found in another, hkfjkii, are straightforward
to calculate (and positive semi-definite). Naturally this theory does not discuss any inter-
actions between particles, and therefore we will have to do a great deal of modification
before we are done. But this simple exercise of merely imposing the standard commu-
tation relations (3.24) between the field and its momentum, we have gained complete
knowledge of the quantum mechanical states of the theory.
3.2.3 The Spin-Statistics Theorem
Notice that the states coming from (3.30) will include the two particle state
jk;k0i= 2p
!!0ay(k)ay(k0)j0i (3.31)
But the commutation relations (3.28) tell us that ay(k)ay(k0) =ay(k0)ay(k). So, this theory
also allows the state
jk0;ki= 2p
!0!ay(k0)ay(k)j0i (3.32)
Recall from a chemistry or modern physics course that particles with half-integer spin
obey the Pauli Exclusion Principle, whereas particles of integer spin do not. Our Klein
Gordon scalar fields are spinless ( j= 0), and therefore we would expect that they do not
obey Pauli exclusion. The fact that our commutation relations have allowed both states
(3.31) and (3.32) is therefore expected. This is an indication that we quantized correctly.
But notice that this statistical result (that the scalar fields do not obey Pauli exclusion)
is entirely a result of the commutation relations. Therefore, if we attempt to quantize a
spin- 1=2field in the same way, they will obviously not obey Pauli exclusion either. We
must therefore quantize spin- 1=2differently.
86
It turns out that the correct way to quantize spin- 1=2fields is to use, instead of commuta-
tion relations like we used for for scalar fields, anticommutation relations . If the operators
of our spin- 1=2fields obey
fay
1;ay
2g=ay
1ay
2= 0)ay
1ay
2= ay
2ay
1
then if we try to act twice with the same operator, we have
ay
1ay
1j0i= ay
1ay
1j0i)ay
1ay
1j0i= 0
In other words, if we quantize with anticommutation relations, it is not possible for two
particles to occupy the same state simultaneously.
This relationship between the spin of a particle and the statistics it obeys (which demands
that integer spin particles be quantized by commutation relations and half-integer spin
particles to be quantized with anticommutation relations) is called the Spin-Statistics
Theorem .
And, because particles obeying Pauli exclusion are said to have Bose-Einstein statistics,
and particles that do not obey Pauli exclusion are said to have Fermi-Dirac statistics, we
call particles with integer spin Bosons , and particles with half-integer spin Fermions .
3.2.4 Left-Handed and Right-Handed Fields
Recall that in the Dirac Lagrangian (3.12), our fundamental field was the 4-component
spinor = L
R
where Ltransforms under the left-handed (0;1=2)representation of
the Lorentz group, and Rtransforms under the right-handed (1=2;0)representation.
In general, we refer to these 2-component spinors as Weyl fields (usually pronounced
“vile”). So, the fermion is the spinor combination of two Weyl fields, one being the left-
handed particle, and the other being the right-handed antiparticle.
Also in (3.12) was the field we defined as = y
0= ( y
R; y
L). If we interpret as the
conjugate of (which the form of the Dirac Lagrangian implies we should), then we see
that the right-handed field is the conjugate of the left, and vice versa. Or, in other words,
y
L= R and y
R= L
We take advantage of the fact by writing all fields in terms of left-handed Weyl fields. For
example, given the two left-handed Weyl fields and, we can form the 4-component
spinor field =
y
, and so = (;y). We will refer to such a field as a Dirac Field ,
and denote it D.
87
On the other hand, we could define a 4-component spinor in terms of a single left-handed
Weyl field, or =
y
. But now notice that = (;y), which is equal simply to the
transpose of . We refer to such a field (whose conjugate is equal to its transpose) as a
Majorana Field , and denote it M.
Recall that an antiparticle has the same mass but opposite charge and opposite handed-
ness of its particle. So, working with the Dirac field D, we can change the charge by
merely swapping and, using the Charge Conjugation operatorCdefined by
C D=C
y
=
y
Also, consider the transpose of D(which is just returning the conjugate of Dto column
form), T
D=
y
. Acting on this with Cgives
C T
D=C
y
=
y
= D
So, we have
C D= T
DandC T
D= D
We therefore say that Dand T
DareCharge Conjugate to each other.
However, notice that with the Majorana field,
C M=C
y
=
y
= M
and
C T
M= T
M= M
So in summary, Dirac fields are not equal to their charge conjugate, while Majorana fields
are. By analogy with scalars (where the complex conjugate of a real number is equal
to itself, whereas the complex conjugate of a complex number is not), we often refer to
Majorana fields as Real , and to Dirac fields as Complex .
So, we can now write out the Lagrangian for Dirac and Majorana fields in terms of their
Weyl fields:
LD=iy@+iy@ m(+yy) (3.33)
LM=iy@ 1
2m(+yy) (3.34)
88
3.2.5 Canonical Quantization of Fermions
We first quantize the Dirac fermion. The general solution to the Dirac equation is
D(x) =2X
s=1Z
fdk[bs(k)us(k)eikx+dy
s(k)vs(k)e ikx]
wheres=1, 2 are the two spin states, bsanddy
sare (respectively) the lowering operator for
the particle and the raising operator for the antiparticle. The charge conjugate of Dwill
have the raising operator for the particle and the lowering operator for the antiparticle.
Theusandvsare constant 4-component vectors which act as a basis for all parti-
cle/antiparticle states in the spinor space (for our purposes, they are merely present to
make Da 4-component field).
We quantize, as we said in section 3.2.3, using anti-commutation relations. Writing only
the non-zero relation,
f (t;x); (t;x)g=3(x x0)(
0)
These imply that the only non-zero commutation relations in terms of the operators are
fbs(k);by
s0(k0)g= (2)33(k k0)2!ss0
fdy
s(k);ds0(k0)g= (2)33(k k0)2!ss0
Once again, these form the algebra of a simple harmonic oscillator, and we can therefore
find the entire spectrum of states by acting on j0iwithby
sanddy
s.
Then, following a series of calculations nearly identical to the ones in section 3.2.2, we
arrive at the Hamiltonian
H=2X
s=1Z
fdk![by
s(k)bs(k) +dy
s(k)ds(k)] (3.35)
whereis an infinite constant we can merely subtract off and therefore ignore.
Comparing (3.29) and (3.35), we see that they both have essentially the same form; !
(which is energy) to the left of the creation operator, which is to the left of the annihilation
operator. To understand the meaning of this, we will see how it generates energy eigen-
values. We will use equation (3.29) for simplicity. Consider acting with the Hamiltonian
89
operator on some arbitrary state jpiwith momentum p. Using (3.30),
Hjpi=Z
fdk!kay(k)a(k)jpi=Z
fdk!kay(k)a(k)p
2!pay(p)j0i
=Z
fdk!kp
2!pay(k)
(2)32!p3(k p) +ay(p)a(k)
j0i
=Z
fdk!kp
2!pay(k)(2)32!p3(k p)j0i
=Zd3k
(2)32!k!kp
2!pay(k)(2)32!p3(k p)j0i
=Z
d3kp
2!pay(k)!p3(k p)j0i
=!pp
2!payj0i=!pjpi
So,Hjpi=!pjpi, where!p= p2+m2, which is the relativistic equation for energy as in
equation (3.25). So, the Hamiltonian operator gives the appropriate energy eigenvalue on
our physical quantum states.
For the Dirac Hamiltonian the eigenvalue will be a linear combination of the energies of
each type of particle. If we denote the states as jpb; sb; pd; sdi, where the first two elements
give the state of a btype particle and the second of the dtype particle, we have
Hjpb; sb; pd; sdi== (!pb+!pd)jpb; sb; pd; sdi
For Majorana fields things are simpler. We only have one type of particle, so
M(x) =2X
s=1Z
fdk
bs(k)us(k)eikx+by
s(k)vs(k)e ikx
And quantization with anticommutation relations will give
H=2X
s=1Z
fdk!by
s(k)bs(k)
3.2.6 Insufficiencies of Canonical Quantization
While the Canonical Quantization procedure we have carried out in the past several sec-
tions has given us a tremendous amount of information (the entire spectrum of states
for bosons, Dirac fermions, and Majorana fermions), it is still lacking quite a bit. As we
said at the beginning of section 3.1.1, we ultimately want a relativistic quantum mechan-
ical theory of interactions. Canonical Quantization has provided a relativistic quantum
mechanical theory, but we aren’t close to being able to incorporate interactions into our
theory. While it is possible to incorporate interactions, it is very difficult, and in order to
simplify we will need a new way of quantizing.
90
3.2.7 Path Integrals and Path Integral Quantization
Perhaps the most fundamental experiment in quantum mechanics is the Double Slit ex-
periment. In brief, what this experiment tells us is that, when a single electron moves
through a screen with two slits, and no observation is made regarding which slit it goes
through, it actually goes through both slits, and until a measurement is made (for exam-
ple, when it hits the observation screen behind the double slit), it exists in a superposition
ofboth paths. As a result, the particle exhibits a wave nature, and the pattern that emerges
on the observation screen is an interference pattern—the same as if a classical wave was
passing through the double slit–all paths in the superposition of the single electron are
interfering with each other, both destructively and constructively. Once the electron is
observed on the observation screen, it collapses probabilistically into one of its possible
states (a particular location on the observation screen).
If, on the other hand, you set up some mechanism to observe which of the two slits the elec-
tron travels through, then the observation has been made before the observation screen,
and you no longer have the superposition, and therefore you no longer see any indica-
tion of an interference pattern. The electrons are behaving, in a sense, classically from the
double slit to the observation screen in this case.
The meaning of this is that a particle that has not been observed will actually take every
possible path at once. Once an observation has been made, there is some probability
associated with each path. Some paths are very likely, and others are less likely (some
are nearly impossible). But until observation, it actually exists in a superposition of all
possible states/paths.
So, to quantize, we will create a mathematical expression for a “sum over all possible
paths”. This expression is called a Path Integral , and will prove to be a much more useful
way to quantize a physical system.
We begin this construction by considering merely the amplitude for a particle at position
q1at timet1to propagate to q2at timet2. This amplitude will be given by
hq2;t2jq1;t1i=hq2jeiH(t2 t1)jq1i
To evaluate this, we begin by dividing the time interval Tt2 t1intoN+ 1equal inter-
vals of length t=T
N+1each. So, we can insert Ncomplete sets of position eigenstates,
hq2;t2jq1;t1i=Z1
1NY
i=1dQihq2je iHtjQNihQNje iHtjQN 1ihQ1je iHtjq1i (3.36)
Let’s look at a single one of these amplitudes. We know that in nearly all physical theories,
we can break the Hamiltonian up as H=P2
2m+V(Q). So, using the completeness of
91
momentum eigenstates,
hQi+1je iHtjQii=hQi+1je i
P2
2m+V(Q)
tjQii
=hQi+1je itP2
2me itV(Q)jQii
=Z
dP0hQi+1je itP2
2mjP0ihP0je itV(Q)jQii
=Z
dP0e itP02
2me itV(Qi)hQi+1jP0ihP0jQii
=Z
dP0e itP02
2me itV(Qi)eiP0Qi+1
p
2e iP0Qi
p
2
=ZdP0
2eiHteiP0(Qi+1 Qi)
=ZdP0
2ei
P0(Qi+1 Qi) Ht
=ZdP0
2eit
P0 Qi+1 Qi
t
H
And taking the limit as t!0,Qi+1 Qi
t!_Qi. So,
ZdP0
2eit
P0 Qi+1 Qi
t
H
=ZdP0
2eidti+1[P0_Qi H]
where the subscript on dtmerely indicates where the infinitesimal time interval “ends”.
So, we can plug this into (3.36) and taking the limit as t!0,
hq2;t2jq1;t1i=Z1
1NY
i=1dQihq2je iHtjQNihQNje iHtjQN 1ihQ1je iHtjq1i
= lim
N!1Z1
1NY
i=1dQiZdP0
i
2eidt2[P0
N_QN H]eidtN[P0
N 1_QN 1 H]eidt1[P0
1_Q1 H]
Z1
1DpDqe iRt2
t1dt(p_q H)
whereDp=Q1
i=1dpiandDq=Q1
i=1dqi.
And ifpshows up quadratically (as it always does;p2
2m), then we can merely do the Gaus-
sian integral over p, resulting in an overall constant which we merely absorb back into
the measure when we normalize. Then, recognizing that the integrand in the exponent is
p_q H=L, we have
hq2;t2jq1;t1i=Z
DqeiRt2
t1dtL=Z
DqeiS(3.37)
Formally, the measure of (3.37) has an infinite number of differentials, and therefore eval-
uating it would require doing an infinite number of integrals. This is to be expected, since
92
the point of the path integral is a sum over every possible path, of which there are an
infinite number. So, because we obviously can’t do an infinite number of integrals, we
will have to find a clever way of evaluating (3.37). But before doing so, we discuss what
the path integral means.
3.2.8 Interpretation of the Path Integral
Equation (3.37) says that, given an initial and final configuration (q1;t1)and(q2;t2), abso-
lutely anypath between them is possible. This is the content of the Dqpart: it is the sum
over all paths.
Then, for each of those paths, the integral assigns a statistical weight ofeiSto it, where the
actionSis calculated using that path (recall our comments in section 1.1.1 about Sbeing
a functional, not a function).
So, consider an arbitrary path q0, which receives statistical weight eiS[q0]. Now, consider
a pathq0very close to q0, only varying by a small amount: q0=q0+q0. This will have
statistical weight eiS[q0+q0]=eiS[q0]+iq0S[q0]
q, whereS
qis the Euler Lagrange derivative
(1.1)
S
q=d
dt@S
@_q @S
@q
To make our intended result more obvious, we do a Wick rotation, taking t!it, so
dt!idt, andS=R
dtL!iR
dtL=iS, andeiS!e S. Now, the path q0=q0+q0gets
weighteiS[q0]e iq0S[q0]
q.
So, ifS
qis very large, then the weight becomes exponentially small. In other words, the
larger the variation of the action is, the less probable that path is.
So the most probable path is the one for the smallest value ofS
t, or the path at which
S
q= 0. And as we discussed in 1.1.1, this is the path of Least Action . Thus, we have
recovered classical mechanics as the first order approximation of quantum mechanics.
So, the meaning of the path integral is that all imaginable paths are possible for the par-
ticle to travel in moving from one configuration to another. However, not all paths are
equally probable. The likelihood of a given path is given by the action exponentiated, and
therefore the most probable paths are the ones which minimize the action. This is the rea-
son that, macroscopically, the world appears classical. The likelihood of every particle in,
say, a baseball, simultaneously taking a path noticeably far from the path of least action
is negligibly small.
We will find that path integral quantization provides an extremely powerful tool with
which to create our relativistic quantum theory of interactions.
93
3.2.9 Expectation Values
Now that we have a way of finding hq2;t2jq1;t1i, the natural question to ask next is how do
we find expectation values like hq2;t2jQ(t0)jq1;t1iorhq2;t2jP(t0)jq1;t1i. By doing a similar
derivation as in the last section, it is easy to show that
hq2;t2jQ(t0)jq1;t1i==Z
DqQ(t0)eiS
We will find that evaluating integrals of this form is simplified greatly through making
use of Functional Derivatives . For some function f(x), the functional derivative is de-
fined by
f(y)f(x)(x y)
Next, we modify our path integral by adding an Auxiliary External Source function, so
that
L!L +f(t)Q(t) +h(t)P(t)
So we now have
hq2;t2jq1;t1if;h=Z
DqeR
dt(L+fQ+hP)
which allows us to write out expectation values in the simple form
hq2;t2jQ(t0)jq1;t1i=1
i
f(t0)hq2;t2jq1;t1if;h
f;h=0=Z
DqQ(t0)eiS+iR
dt(fQ+hP)
f;h=0
=Z
DqQ(t0)eiS
or
hq2;t2jP(t0)jq1;t1i=1
i
h(t0)hq2;t2jq1;t1if;h
f;h=0=Z
DqP(t0)eiS+iR
dt(fQ+hP)
f;h=0
=Z
DqP(t0)eiS
So, once we have hq2;t2jq1;t1i, we can find any expectation value we want simply by
taking successive functional derivatives.
94
3.2.10 Path Integrals with Fields
Because we can build whatever state we want by acting on the vacuum, the important
quantity for us to work with will be the Vacuum to Vacuum expectation value, or VEV ,
h0j0i, and the various expectation values we can build through functional derivatives
(h0jj0i,h0j j0i, etc.).
For simplicity let’s consider a scalar boson . The Lagrangian is given in equation (3.10).
Using this, we can write the path integral
h0j0i=Z
DeiR
d4x
1
2@@ 1
2m22
Z
DeiR
d4xL0
We will eventually want to find expectation values, so we introduce the auxiliary field J,
creating
h0j0iJ=Z
DeiR
d4x(L0+J)(3.38)
So, for example,h0jj0i=1
i
Jh0j0iJ
J=0.
Of course, we still have a path integral with an infinite number of integrals to evaluate.
But, we are finally able to discuss how we can do the evaluation.
We defineZ0(J)h0j0iJ. Then, making use of the Fourier Transform of ,
e(k) =Z
d4xe ikx(x)(x) =Zd4k
(2)4eikxe(k)
we begin with the L0part:
S0=Z
d4xL0=Z
d4x
1
2@@ 1
2m22
=Z
d4x
1
2@Zd4k
(2)4eikxe(k)
@Zd4k0
(2)4eik0xe(k0)
1
2m2Zd4k
(2)4eikxe(k)Zd4k0
(2)4eik0xe(k0)
=Z
d4x1
2Zd4kd4k0
(2)8eikxeik0xe(k)e(k0)(kk0
m2)
=1
2Zd4kd4k0
(2)8e(k)e(k0)(kk0
m2)Z
d4xei(k+k0)x
=1
2Zd4kd4k0
(2)8e(k)e(k0)(kk0
m2)(2)44(k+k0)
= 1
2Zd4k
(2)4e(k)(k2+m2)e( k)
95
Then, transforming the auxiliary field part,
Z
d4xJ(x)(x) =Z
d4xZd4k
(2)4eikxeJ(k)Zd4k0
(2)4eik0xe(k0)
=Zd4kd4k0
(2)8eJ(k)e(k0)Z
d4xei(k+k0)x
=Zd4kd4k0
(2)8eJ(k)e(k0)(2)44(k+k0)
=Zd4k
(2)4eJ(k)e( k)
And because the integral is over all k, we can rewrite this as
Zd4k
(2)4eJ(k)e( k) =1
2Zd4k
(2)4 eJ(k)e( k) +eJ( k)e(k)
(we did this to get the factor of 1=2out front in order to have the same coefficient as the
L0part from above).
So,
S=1
2Zd4k
(2)4
e(k)(k2+m2)e( k) +eJ(k)e( k) +eJ( k)e(k)
Now, we make a change of variables,
e(k)e(k) eJ(k)
k2+m2
(Note that this leaves the measure of the path integral unchanged: D!D.)
Plugging this in, we have,
S=1
2Zd4k
(2)4
e(k) +eJ(k)
k2+m2
(k2+m2)
e( k) +eJ( k)
k2+m2
+eJ(k)
e( k) +eJ( k)
k2+m2
+eJ( k)
e(k) +eJ(k)
k2+m2
=1
2Zd4k
(2)4
e(k)(k2+m2)e( k) +eJ(k)eJ( k)
k2+m2
(The point of all of this is that, in this form, we have all of the , or equivalently , depen-
dence in the first term, with no ordependence on the second term.)
Finally, our path integral (3.38) is
h0j0iJ=Z
Dei
2Rd4k
(2)4
e(k)(k2+m2)e( k)+eJ(k)eJ( k)
k2+m2
96
Now, using some clever physical reasoning, we can see how to evaluate the infinite num-
ber of integrals in this expression. Notice that if we set J= 0, we have a free theory in
which no interactions take place. This means that if we start with nothing (the vacuum),
the probability of having nothing later is 100%. Or,
h0j0iJ
J=0= 1 =Z
Dei
2Rd4k
(2)4
e(k)(k2+m2)e( k)
And if that part is 1, then we have
h0j0iJ=Z
Dei
2Rd4k
(2)4eJ(k)eJ( k)
k2+m2
And remarkably, the integrand has nodependence ! Therefore, the infinite number of
integrals over all possible paths becomes nothing more than a constant we can absorb
into the normalization, leaving
h0j0iJ=ei
2Rd4k
(2)4eJ(k)eJ( k)
k2+m2
We can Fourier Transform back to coordinate space to get
Z0(J) =h0j0iJ=ei
2R
d4xd4x0J(x)(x x0)J(x0)(3.39)
where
(x x0)Zd4k
(2)4eik(x x0)
k2+m2
is called the Feynman Propagator for the scalar field.
We can then find expectation values by operating on this with1
i
Jas described in section
3.2.9.
We can repeat everything we have just done for fermions, and while it is a great deal more
complicated (and tedious), it is in essence the same calculation. We begin by adding the
auxiliary function + , to get expectation values of and by using1
i
and1
i
,
respectively.
We then Fourier Transform every term in the exponent and find that we can separate out
the and dependence, allowing us to set the term which does depend on and equal
to 1. Fourier Transforming back then gives
Z0(;) =eiR
d4xd4x0(x)S(x x0)(x0)(3.40)
where
S(x x0) =Zd4k
(2)4(
k+m)eik(x x0)
k2+m2
97
is the Feynman propagator for fermion fields.
Recall that we are calling the auxiliary fields J,, and Source Fields . Comparing the
form of the Lagrangian in equation (3.38) to (1.13) reveals why. J,, and behave mathe-
matically as sources, giving rise to the field they are coupled to, in the same way that the
electromagnetic source Jgives rise to the electromagnetic field A. The meaning behind
equations (3.39) (and (3.40)) is that J(orand) act as sources for the fields, creating a
(or and ) at spacetime point x, and absorbing it at point x0. The terms (x x0)and
S(x x0)then represent the expression giving the probability amplitude h0j0ifor that par-
ticular event to occur. In other words, the propagator is the statistical weight of a particle
going from xtox0.
3.2.11 Interacting Scalar Fields and Feynman Diagrams
We can now consider how to incorporate interactions into our formalism, allowing us to
finally have our relativistic quantum theory of interactions.
Beginning with the free scalar Lagrangian (3.10), we can add an interaction term L1. At
this point, we only have one type of particle, , so we can only have ’s interacting with
other’s. Terms proportional to or2are either constant or linear in the equations
of motion, and therefore aren’t valid candidates for interaction terms. So, the simplest
expression we can have is
L1=1
3!g3
where1
3!is a conventional normalization, and gis a Coupling Constant . So our total
Lagrangian is
L=L0+L1= 1
2@@ 1
2m22+1
6g3
and the path integral is
Z(J) =h0j0iJ
=Z
DeiR
d4x[L0+L1+J]
=Z
DeiR
d4xL1eiR
d4x[L0+J]
=Z
DeiR
d4xL1Z0(J) =Z
DeiR
d4xL1h0j0iJ
But, recall that we can bring out a factor of fromh0j0iJusing the functional derivative
1
i
J. So, we can make the replacement
L1()!L 11
i
J
)1
6g3!g
61
i
J3
98
And notice that once this is done, there is no longer any dependence in Z(J). So, with
the free theory, we were able to remove the dependence, leading to (3.39). And here,
we were able to remove it from the interaction term as well. So, once again, the infinite
number of integrals in (3.37) will merely give a constant which we can absorb into the
normalization.
This leaves the result
Z(J) =ei
6gR
d4x
1
i
J(x)3
Z0(J)
=e 1
6gR
d4x
J(x)3
ei
2R
d4xd4x0J(x)(x x0)J(x0)
Now, we can do two separate Taylor expansions to these two exponentials,
Z(J) =1X
V=01
V!
g
6Z
d4x
J(x)3V
1X
P=01
P!i
2Z
d4yd4zJ(y)(y z)J(z)P
(3.41)
Now, recall that a functional derivative1
i
J, will remove a Jterm. Furthermore, after
taking the functional derivatives, we will set J= 0to get the physical result. So, for a term
to survive, the 2Psources must all be exactly removed by the 3Vfunctional derivatives.
So, using (3.41), we can expand in orders of g(the coupling constant), keeping only the
terms which survive, and after removing the sources, evaluate the integrals over the prop-
agators . The value of the integral will then be the physical amplitude for a particular
event.
In practice, a slightly different formalism is used to organize and keep track of each term
in this expansion. Note that there will be Ppropagators . We can represent each of these
terms diagrammatically, by making each source a solid dot, each propagator a line, and
let thegterms be vertices joining the lines together. There will be a total of Vvertices,
each joining 3 lines (matching the fact that we are looking at 3theory; there would be 4
lines at each vertex for 4theory, etc.).
For example, for V= 0andP= 1,
Z(J) =i
2Z
d4yd4zJ(y)(y z)J(z)
We have two sources, one located at zand the other located at y, so we draw two dots,
corresponding to those locations. Then, the propagator (y z)connects them together,
so we draw a line between the two dots. The diagram should look like this:
99
Of course, once we set J= 0, this will vanish because it contains two sources.
As another example, consider V= 0andP= 2. Now,
Z(J) =1
2!i
22Z
d4yd4zd4y0d4z0
J(y)(y z)J(z)
J(y0)(y0 z0)J(z0)
This corresponds to four sources, located at y;z;y0andz0, with propagator lines connect-
ingytoz, and connecting y0toz0. But, there are no lines connecting an unprimed source
to a primed source, so this results in two disconnected diagrams:
As another example, consider V= 1andP= 2,
Z(J) = g
6Z
d4x
J(x)3
1
2!i
22Z
d4yd4zd4y0d4z0
J(y)(y z)J(z)
J(y0)(y0 z0)J(z0)
=g
48Z
d4xd4yd4zd4y0d4z0(y x)(y z)(z x)(y0 x)(y0 z0)J(z0)
=g
48Z
d4xd4z0(x x)(x z0)J(z0)
This will correspond to
where the source Jis located at the dot, and the vertex joining the line to the loop is at x.
You can work out the following out, and see that there are multiple possible diagrams for
V= 3; P= 5
100
And forV= 2,P= 4,
And forV= 1,P= 3,
and so on.
Through a series of combinatoric and physical arguments, it can be shown that only con-
nected diagrams will contribute, and the1
P!and1
V!terms will always cancel exactly.
So, to calculate the amplitude for a particular interaction to happen (say N ’s in andM
’s out), draw every connected diagram that is topologically distinct, and has the correct
number of in and out particles. Then, through a set of rules which you will learn formally
in a QFT course, you can reconstruct the integrals which we started with in (3.41).
When you take a course on QFT, you will spend a tremendous amount of time learning
how to evaluate these integrals for low order (they cannot be evaluated past about second
order in most cases). While this is extremely important, it is not vital for the agenda of
these notes, and we therefore do not discuss how they are evaluated.
The idea is that each diagram represents one of the possible paths the particle can take,
along with the possible interactions it can be a part of. Because this is a quantum mechan-
ical theory, we know it is actually in a superposition of all possible paths and interactions.
101
We don’t make a measurement or observation until the particles leave the area in which
they collide, so we have no idea about what is going on inside the accelerator. We know
that if thisgoes in and thiscomes out, we can draw a particular set of diagrams which
have the correct input and output, and the nature of the interaction terms (which deter-
mines what types of vertices you can have) tells us what types of interactions we can
have inside the accelerator. Evaluating the integrals then tells us how much that partic-
ular event/diagram contributes towards the total probability amplitude. So, if you want
to know how likely a certain incoming/outgoing set of particles is, write down all the
diagrams, evaluate the corresponding integrals, and add them up.
And as we pointed out above, the classical behavior (which is more probable) is closer
to the first order approximation of the quantum behavior. Therefore, even though in
general we can’t evaluate the integrals past about second order, the first few orders tell us
to a reasonable (in fact, exceptional in most cases) degree of accuracy what the amplitude
is. If we want more accuracy, we can seek to evaluate higher orders, but usually lower
orders suffice for experiments at energy levels we can currently attain.
One of the difficulties encountered with evaluating these integrals is that you almost al-
ways find that they yield infinite amplitudes. Since an amplitude (which is a probability)
should be between 0 and 1, this is obviously unacceptable. The process of finding the
infinite parts and separating them from the finite parts of the amplitude is a very well
defined mathematical construct called Renormalization . The basic idea is that any infi-
nite term consists of a pure infinity and a finite part. For example (trust us for now) the
infinite sum:
1X
n=1n= lim
x!01
x2 1
12
There is a part which is a pure infinity (the first term on the right hand side), and a term
which is finite. While this may seem strange and extremely unfamiliar (and a bit like
hand waving), it is actually a very rigorous and very well understood mathematical idea.
Much of what particle physicists attempt to do is find theories (and types of theories) that
can be renormalized and theories that cannot. For example, the action which leads to
General Relativity leads to a quantum theory which cannot be renormalized. Renormal-
ization is a fascinating and deep topic, and will be covered in great depth in any standard
QFT text or course. Unfortunately, we will not discuss it further.
3.2.12 Interacting Fermion Fields
The analysis we performed above for scalar fields above is almost identical for fermions,
and we therefore won’t repeat it. The main difference is that the interaction terms will
have a field interacting with , and so the vertices will be slightly different. We won’t
bother with those details.
102
Finally, we can have a Lagrangian with both scalars and fermions. Then, naturally, you
could have interaction terms where the scalars interact with fermions. While there are
countless interaction terms of this type, the one that will be the most interesting to us is
theYukawa term,
LYuk=g (3.42)
If we represent by a dotted line, by a line with an arrow in the forward time direction,
and with an arrow going backwards in time, this interaction term will show up in a
Feynman diagram as
Once each diagram is drawn, there are well defined rules to write down an integral cor-
responding to each diagram.
3.3 Final Ingredients
The purpose of the previous section was merely to introduce the idea of Feynman Dia-
grams as a tool to calculate amplitudes for physical processes. In doing so, we have met
the goal set out in section 3.1.1, a relativistic quantum mechanical theory of interactions.
We achieve such a theory by finding a Lagrangian of a classical theory (both with and
without interaction terms), and using equation (3.41) (and the analogous equation for
fermions) to write down integrals which, when evaluated, give a contribution to a total
amplitude. It is important to remember that we will eventually set all sources Jto zero,
and a functional derivative (as contained in the interaction term L1) will set any term
withoutJ’s to zero. So, the only non-zero terms will be the ones where all of the J’s are
exactly removed by the functional derivatives.
A large portion of understanding QFT is learning how to set these integrals up in greater
detail, and learning several methods to evaluate them. We will not delve into those details
ofPerturbative Quantum Field Theory , where amplitudes are studied order by order,
here. The goal of these notes is merely to explain how, once given a Lagrangian, that
Lagrangian can be turned into a physically measurable quantity.
103
With this done, we now set out to find the Lagrangian for the Standard Model of Particle
Physics, the theory which seems to explain our universe (apart from gravity). Once this
Lagrangian has been explained, we trust you have a general concept of what to do with
it from the previous sections.
However, before we are able to explain the Standard Model Lagrangian, there are a few
final concepts we need. They will be the subject of this section. Namely, we will be study-
ing the ideas of Spontaneous Symmetry Breaking and Gauge Theories . In section 3.1.11,
we discussed the simple U(1)gauge theory, where we made a global U(1)symmetry of
the free Dirac Lagrangian a local U(1)symmetry, or a gauged symmetry, and showed that
consistency demanded the introduction of a gauge field A, and consequently a kinetic
term and a source term. Thus we recovered the entire electromagnetic force from nothing
butU(1). Later in this section, we generalize this to arbitrary Lie group. Because U(1)
is an Abelian group, we refer to the gauge theory of section 3.1.11 as an abelian gauge
theory. For a more general, non-Abelian group, we refer to the theory resulting as a
Non-Abelian Gauge Theory . Such theories introduce a great deal of complexity, and we
therefore consider them in detail in this section before moving on to the Standard Model.
However, we begin with the idea of spontaneous symmetry breaking.
3.3.1 Spontaneous Symmetry Breaking
Consider a complex scalar boson andy. The Lagrangian will be
L= 1
2@y@ 1
2m2y
Naturally we can write this as
L= 1
2@y@ V(y;) (3.43)
where
V(y;) =1
2m2y
This Lagrangian has the U(1)symmetry we discussed in 3.1.11.
Also, notice that we can graph V(y;), plottingVvs.jj,
104
We see a “bowl” with Vminimum atjj2= 0. The vacuum of any theory ends up being at the
lowest potential point, and therefore the vacuum of this theory is at = 0, as we would
expect.
Now, let’s change the potential. Consider
V(y;) =1
2m2(y 2)2(3.44)
whereandare real constants. Notice that the Lagrangian will still have the global
U(1)symmetry from before. But, now if we graph Vvs.jj, we get
where now the vacuum Vminimum is represented by the circle at jj= . In other words,
there are an infinite number of vacuums in this theory. And because the circle drawn
in the figure above represents a rotation through field space, this degenerate vacuum is
parameterized by ei, the global U(1). There will be a vacuum for every value of , located
atjj= .
In order to make sense of this theory, we must choose a vacuum by hand. Because the
theory is completely invariant under the choice of the U(1)ei, we can choose any and
105
define that as our true vacuum. So, we choose to make our vacuum at = , or where
is real and equal to . We have thus, in a sense, Gauged Fixed the symmetry in the
Lagrangian, and the U(1)symmetry is no longer manifest.
Now we need to rewrite this theory in terms of our new vacuum. We therefore expand
around the constant vacuum value to have the new field
++i
whereandare new real scalar fields (so y= + i). We can now write out the
Lagrangian as
L= 1
2@[ i]@[+i] 1
2m2[( + i)( ++i) 2]2
=
1
2@@ 1
24m222 1
2@@
1
2m2
43+ 42+4+22+4
(3.45)
This is now a theory of a massive real scalar field (with mass =p
4m22), amassless real
scalar field , and five different types of interactions (one allowing three ’s to interact,
the second allowing one and two’s, the third allowing four ’s, the fourth allowing
two’s and two’s, and the last allowing four ’s.) In other words, there are five different
types of vertices allowed in the Feynman diagrams for this theory.
Furthermore, notice that this theory has no obvious U(1)symmetry. For this reason, writ-
ing the field in terms of fluctuations around the vacuum we choose is called “breaking”
the symmetry. The symmetry is still there, but it can’t be seen in this form.
Finally, notice that breaking the symmetry has resulted in the addition of the massless
field. It turns out that breaking global symmetries as we have done always results in a
massless boson. Such particles are called Goldstone Bosons .
3.3.2 Breaking Local Symmetries
In the previous section, we broke a global U(1)symmetry. In this section, we will break
a localU(1)and see what happens. We begin with the Lagrangian for a complex scalar
field with a gauged U(1):
L= 1
2
@ iqA
y
@+iqA
1
4FF V(y;)
where we have taken the external source J= 0. Let’s once again assume V(y;)has the
form of equation (3.44), so the vacuum has the U(1)degeneracy atjj= .
Because our U(1)is now local, we choose (x)so that not only is the vacuum real, but
also so that is always real. We therefore expand
= +h (3.46)
106
wherehis a real scalar field representing fluctuations around the vacuum we chose.
Now,
L= 1
2
@ iqA
+h
@+iqA
+h)
1
4FF
1
2m2
+h
+h
22
=
= 1
2@h@h 1
24m22h2 1
4FF 1
2q22A2+Linteractions
where the allowed interaction terms include a vertex connecting an hand twoA’s, four
h’s, and three h’s.
So, before breaking, we had a complex scalar field and a massless vector field Awith
two polarization states (because it is a photon). Now, we have a single real scalar hwith
mass =p
4m22and a field Awith mass =q. In other words, our force-carrying
particleAhas gained mass! We started with a theory with no mass, and by merely
breaking the symmetry, we have introduced mass into our theory.
This mechanism for introducing mass into a theory, called the Higgs Mechanism , was
first discovered by Peter Higgs, and the resulting field his called the Higgs Boson .
So, whereas the consequence of global symmetry breaking is a massless boson called a
Goldstone boson, the consequence of a local symmetry breaking is that the gauge field,
which came about as a result of the symmetry being local, acquires mass.
3.3.3 Non-Abelian Gauge Theory
We are now ready to generalize what we did in section 3.1.11 to an arbitrary Lie group.
Consider a Lagrangian LwithNscalar (or spinor) fields i(i= 1;:::;N ) that is invariant
under a continuous SO(N)orSU(N)symmetry, i!Uijj, whereUijis anNNmatrix
ofSO(N)orSU(N).
In section 3.1.11, we saw that if the group is U(1), gauging it demands the introduction of
the gauge field Ato preserve the symmetry, which shows up in the covariant derivative
D=@ ieA. To say a field carried some sort of charge means that it has the corre-
sponding term in its covariant derivative. We then added a kinetic term for Aas well as
an external source J. Then, higher order interaction terms can be included in whatever
way is appropriate for the theory.
To generalize this, let’s say for the sake of concreteness that our Lie group is SU(N). An
arbitrary element of SU(N)iseiga(x)Ta, wheregis a constant we have added for later
convenience, aare theN2 1parameters of the group (cf. section 2.2.15), and the Taare
107
the generator matrices for the group. Notice that we have gauged the symmetry (in that
(x)is a function of spacetime).
By definition, we know that the generators Tawill obey the commutation relations
[Ta;Tb] =ifabcTc
(cf equation (2.9)), where fabcare the structure constants of the group.
When gauging the U(1)in section 3.1.11, the transformation of the gauge field was given
by equation (3.16). For the more general transformation i!Uijj, the gauge field trans-
forms according to
A!U(x)AUy(x) +i
gU(x)@Uy(x)
(where we have removed the indicial notation and it is understood that matrix multipli-
cation is being discussed). If U(x)is an element of U(1)(so it iseig(x)), then this transfor-
mation reduces to
A!eig(x)Ae ig(x)+i
geig(x)( ig@(x))e ig(x)=A+@(x)
which is exactly what we had in (1.15). For general SU(N), however, the U’s are elements
of a Non-Abelian group, and the A’s are matrices of the same size.
Generalizing, we find that a general element of the SU(N)is (changing notation slightly)
U(x)e ig a(x)Ta
withN2 1real parameters a. We then build the covariant derivative in the exact same
way as in equation (3.17) by adding a term proportional to the gauge field
D=INN@ igA
(Remember that each component of Ais anNNmatrix. They were scalars for U(1)
becauseU(1)is a11matrix.) Or, acting on the fields, the covariant derivative is
(D)j=@j(x) ig[A(x)]jkk(x) (3.47)
wherekis understood to be summed on the last term. It will be understood from now
on that the normal partial derivative term (the first term) has an NNidentity matrix
multiplied by it.
Then, just as in (3.19), we have the field strength
F(x)i
g[D;D] =@A @A ig[A;A] (3.48)
where the commutator term doesn’t vanish for arbitrary Lie group as it did for Abelian
U(1).
108
Recall from equation (1.16) that for U(1),Fis invariant under the gauge transforma-
tion (1.15) on its own, because the commutator term vanishes. In general, however, the
commutator term does not vanish, and we must therefore be careful in writing down the
correct kinetic term. It turns out that the correct choice is
LKin= 1
2Tr (FF) (3.49)
It may not be obvious, but this form is actually a consequence of (2.28). There is algebraic
machinery working under the surface of this that, while extremely interesting, is unfortu-
nately beyond the scope of what we are doing. We will discuss all of these ideas in much
greater depth later in this series.
So, starting with a non-interacting Lagrangian that is invariant under the global SU(N),
we can gauge the SU(N)to create a theory with a gauge field (or synonymously a “force
carrying” field) A, which is an NNmatrix. So, every Lie group gives rise to a particular
gauge field (which is a force carrying particle, like the photon), and therefore a particular
force.
For this reason, we discuss forces in terms of Lie groups, or synonymously Gauge
Groups . Each group defines a force. As we said at the very end of section 2.2.11, U(1)rep-
resents the electromagnetic force (as we have seen in section 3.1.11, while SU(2)describes
the weak force, and SU(3)describes the strong color force.
3.3.4 Representations of Gauge Groups
As we discussed in section 2.2, given a set of structure constants fabc, which define the
Lie algebra of some Lie group, we can form a representation of that group, which we
denoteR. So,Rwill be a set of D(R)D(R)matrices, where Dis the dimension of the
representation R. We then call the generators of the group (in the representation R)Ta
R,
and they naturally obey [Ta
R;Tb
R] =ifabcTc
R.
One representation which exists for any of the groups we have considered is the represen-
tation ofSO(N)orSU(N)consisting of NNmatrices. We denote this the Fundamental
Representation (also called Defining Representation in some books). Clearly, the funda-
mental representations of SO(2),SO(3),SU(2), andSU(3)are the 22,33,22, and
33matrix representations, respectively. We will denote the fundamental representation
for a given group by writing the number in bold. So, the fundamental representation of
SU(2)will be denoted 2, and the generators for SU(2)in the fundamental representation
will be denoted Ta
2. Obviously, the fundamental representation of SU(3)will be 3with
generators Ta
3.
Furthermore, let’s say we have some arbitrary representation generated by Ta
R, obeying
[Ta
R;Tb
R] =ifabcTc
R. We can take the complex conjugate of the commutation relations to get
[T?a
R;T?b
R] = ifabcT?c
R. So, notice that if we define the new set of generators T0a
R T?a
R,
109
then theT0a
Rwill obey the correct commutation relations, and will therefore form a rep-
resentation of the group as well. If it turns out that T0a
R= (Ta
R)?=Ta
R, or if there is
some unitary similarity transformation Ta
R!U 1Ta
RUsuch thatT0a
R= (Ta
R)?=Ta
R, then
we call the representation Real , and the complex conjugate of the Ta
R’s is the same repre-
sentation. However, if no such transformation exists, then we have a new representation,
called the Complex Conjugate representation to R, or the Anti-Rrepresentation, which
we denote R.
For example, there is the fundamental representation of SU(3), denoted 3, generated by
Ta
3, and then there is the anti-fundamental representation 3, generated by Ta
3.
The representations of a group which will be important to us are the fundamental, anti-
fundamental, and adjoint.
3.3.5 Symmetry Breaking Revisited
As we said in section 3.3.3, given a field transforming in a particular representation R, the
gauge fields Awill beD(R)D(R)matrices.
Once we know what representation we are working in, and therefore know the generators
Ta
R, it turns out that it is always possible to write the gauge fields in terms of the genera-
tors. Recall in sections 2.2.2 and 2.2.12, we encouraged you to think of the generators as
basis vectors which span the parameter space for the group. Because the gauge fields live
in theNNspace as well, we can write them in terms of the generators. That is, instead
of the gauge fields being NNmatrices on their own, we will use the NNmatrix
generators as basis vectors, and then the gauge fields can be written as scalar coefficients
of each generator:
A=A
aTa
R (3.50)
whereais understood to be summed, and each A
ais now a scalar function rather than a
D(R)D(R)matrix (the advantage of this is that we can continue to think of the gauge
fields as scalars with an extra index, rather than as matrices). As a note, we haven’t
done anything particularly profound here. We are merely writing each component of the
D(R)D(R)matrixAin terms of the D(R)D(R)generators, allowing us to work
with a scalar fieldA
arather than the matrix field A. We now actually view each A
aas a
separate field. So, if a group has Ngenerators, we say there are Ngauge fields associated
with it, each one having 4 spacetime components .
In matrix components, we will have
(A)ij= (A
aTa
R)ij
Then, the covariant derivative in (3.47) will be
(D)j=@j(x) ig[Aa(x)Ta
R]jkk(x) (3.51)
110
We may assume that the field strength Fcan also be expressed in terms of the genera-
tors, so that we have
F=F
aTa(3.52)
or
(F)ij= (F
aTa)ij (3.53)
Now, using (2.28) (and taking = 1=2by convention), we can write (3.49) in terms of the
new basis:
LKin= 1
2Tr (FF) = 1
2Tr (F
aTaFbTb)
= 1
2F
aFbTr(TaTb)
= 1
2F
aFbab
= 1
2F
aFa
= 1
4F
aFa
(3.54)
(we have raised the index aon the second field strength term in the last two lines simply
to explicitly imply the summation over it. The fact that it is raised doesn’t change its value
in this case; it is merely notational).
Furthermore, we can use (2.28) to invert (3.52):
F=F
aTa)FTb=F
aTaTb
)Tr (FTb) =F
aTr (TaTb)
)Tr (FTb) =F
aab
)Tr (FTb) =1
2F
b
)F
a= 2 Tr (FTa) (3.55)
In sections 3.3.1 and 3.3.2, we broke the U(1)symmetry, which only had one generator.
However, if we break larger groups we may only break part of it. For example, we will
see thatSU(3)has anSU(2)subgroup. It is actually possible to break only the SU(2)part
of theSU(3). So, three of the SU(3)generators are broken (the three corresponding to
theSU(2)subgroup/subalgebra), and the other five are unbroken. Because we are now
writing our gauge fields using the generators as a basis, this means that three of the gauge
fields are broken, while five of the gauge fields are not.
Finally, recall from section 3.3.2 that breaking a local symmetry results in a gauge field
gaining mass. We seek now to elucidate the relationship between breaking a symmetry
111
and a field gaining mass. First, we can summarize as follows: Gauge fields corresponding
to broken generators get mass, while those corresponding to unbroken generators do not. The
unbroken generators form a new gauge group that is smaller than the original group that was
broken .
In 3.3.2, we saw that breaking a symmetry gave the gauge field mass. Now, we see that
giving a gauge field mass will break the symmetry.
To make this clearer, we begin with a very simple example, then move on to a more
complicated example.
3.3.6 Simple Examples of Symmetry Breaking
Consider a theory with three real massless scalar fields i(i= 1;2;3) and with Lagrangian
L= 1
2@i@i
which is clearly invariant under the global SO(3)rotation
i!Rijj
whereRijis an element of SO(3), because the Lagrangian is merely a dot product in field
space, and we know that dot products are invariant under SO(3).
Now, let’s say that one of the fields, say 1, gains mass. The new Lagrangian will then be
L= 1
2@i@i 1
2m22
1
So this Lagrangian is no longer invariant under the full SO(3)group, which mixes any
two of the three fields. Rather, it is only invariant under rotations in field space that mix
2and3orSO(2). In other words, giving one field mass broke SO(3)to the smaller
SO(2).
As another simple example, if we started with five massless complex scalar fields i, with
Lagrangian
L= 1
2@y
i@i
This will be invariant under any SU(5)transformation.
Then let’s say we give two of the fields, 1and2(equal) mass. The new Lagrangian will
be
L= 1
2@y
i@i 1
2m(y
11+y
22)
112
So now, we no longer have the full SU(5)symmetry, but we do have the special unitary
transformations mixing 3,4, and5. This is an SU(3)subgroup. Furthermore, we can
do a special unitary transformation mixing 1and2. This is an SU(2)subgroup. So, we
have broken SU(5)!SU(3)
SU(2).
Before considering a more complicated example of this, we further discuss the connection
between symmetry breaking and fields gaining mass.
When we introduced spontaneous symmetry breaking in section 3.3.1, recall that we
shifted the potential minimum from Vminimum at= 0 toVminimum atjj= . But we
were discussing this in very classical language. We can interpret all of this in a more
“quantum” way in terms of VEV’s. As we said, the vacuum of a theory is defined as the
minimum potential field configuration. For the Vminimum at= 0 potential, the VEV of
the fieldwas at 0, or
h0jj0i= 0
However, for the Vminimum atjj= potential, we have
h0jj0i=
So, in quantum mechanical language, symmetry breaking occurs when a field, or some
components of a field, take on a non-zero VEV .
This seems to be what is happening in nature. At higher energies, there is some “Mas-
ter Theory” with some gauge group defining the physics, and all of the fields involved
have 0 VEV’s. At lower energies, for whatever reason (the reason for this is not well un-
derstood at the time of this writing), some of the fields take on non-zero VEV’s, which
break the symmetry into smaller groups, giving mass to certain fields through the Higgs
Mechanism discussed in section 3.3.2. We call the theory with the unbroken gauge sym-
metry at higher energies the more fundamental theory (analogous to equation (3.43)), and
the Lagrangian which results from breaking the symmetry (analogous to (3.45)) the Low
Energy Effective Theory .
And this is how mass is introduced into the Standard Model. It turns out that if a theory
is renormalizable one can prove that any lower energy effective theory that results from
breaking the original theory’s symmetry is also renormalizable, even if it doesn’t appear
to be. And, because the actions that appear to describe the universe at the energy level
we live at (and the levels attainable by current experiment) are not renormalizable when
they have mass terms, we work with a larger theory which has no massive particles but
can be renormalized, and use the Higgs Mechanism to give various particles mass. So,
whereas the physics we see at low energies may not appear renormalizable, if we can find
a renormalizable theory which breaks down to our physics, we are safe.
Now, we consider a slightly more complicated (and realistic) example of symmetry break-
ing.
113
3.3.7 A More Complicated Example of Symmetry Breaking
Consider the gauge group SU(N), acting on Ncomplex scalar fields i(i= 1;:::;N ) in
the fundamental representation N. Recall that in section 3.3.2, in order to get equation
(3.46), we made use of the U(1)symmetry to make the vacuum, or the VEV , real. We can
now do something similar: we make use of the SU(N)to not only make the VEV real, but
also to rotate it to a single component of the field, N. In other words, we do an SU(N)
rotation so that
h0jij0i= 0 fori= 1;:::;N 1
h0jNj0i=
So, we expand Naround this new vacuum:
i=i fori= 1;:::;N 1
N= +
This means that, in the vacuum configuration, the fields will have the form
0
BBB@1
2
...
N1
CCCA
vacuum=0
BBB@0
0
...
1
CCCA
So, how will the action of SU(N)be affected by this VEV? If we consider a general element
ofSU(N)acting on this,
0
BBB@U11U12U1N
U21U22U2N
.........
UN1UN2...UNN1
CCCA0
BBB@0
0
...
1
CCCA=0
BBB@U1N
U2N
...
UNN1
CCCA
So, only elements of SU(N)with non-zero elements in the last column will be affected by
this VEV . But the other N 1elements’ rows and columns are unaffected. This means
that we have an SU(N 1)symmetry left. Or in other words, we have broken SU(N)!
SU(N 1)with this VEV .
Let’s consider a specific example of this. Consider SU(3). The generators are written out
in (2.46). Notice that exactly three of them have all zeros in the last column; 1,2, and
3. We expect these three to give an SU(3 1) =SU(2)subgroup. And looking at the
upper left 22boxes in those three generators, we can see that they are the Pauli matrices,
the generators of SU(2). So, if we give a non-zero VEV to the fields transforming under
SU(3), we see that they do indeed break the SU(3)toSU(2). The other five generators of
SU(3)will be affected by the VEV , and consequently the corresponding fields will acquire
mass.
114
3.4 Particle Physics
3.4.1 Introduction to the Standard Model
We are finally ready to study the Standard Model of Particle Physics , which (except for
gravity) appears to be the theory which explains our universe. To state the Standard
Model in the simplest possible terms, it is
A Yang-Mills (Gauge) Theory with Gauge Group
SU(3)
SU(2)
U(1)
with left-handed Weyl fields fields in three copies of the representation
(1;2; 1=2)(1;1;1)(3;2;1=6)(3;1; 2;3)(3;1;1=3)
(where the last entry specifies the value of the U(1)hypercharge),
and a single copy of a complex scalar field in the representation
(1;2; 1=2)
Admittedly, our exposition will be somewhat cursory. This is largely because every con-
cept and tool we use in this section has been discussed in detail in the previous sections.
The purpose of these notes is to provide an introduction to the primary concepts and
mathematical tools used in Particle Physics, not to give the details of the theory. We will
cover the main points of the Standard Model, but there is tremendous detail we are skip-
ping over. A second reason the following section is cursory is that we will not be working
out every step in detail, as we have been doing. For nearly all calculations being done in
this section, we have worked out a similar tedious calculation previously. We will there-
fore frequently refer to previous sections/equations. It will be worthwhile to go back and
carefully study the parts which we refer to.
Because this section is slightly more experimental, or at least phenomenological, than the
rest, and because the general purpose of these notes is to develop the mathematical tools
and framework of particle physics (especially gauge theory), undue attention should not
be given to this section. The purpose is merely to show, as briefly as possible, where
everything we have done so far lines up with experiment. It will be useful to read through
this section, but do not spend too much time bogged down in the details.
Before diving into this in detail, look over the general structure of the Standard Model on
page 139.
115
3.4.2 The Gauge and Higgs Sector
We begin our exposition with the Electroweak part of the Standard Model gauge group,
theSU(2)
U(1)part, as well as the Higgs.
Beginning with the Higgs, a scalar field in the (2; 1=2)representation of SU(2)
U(1), the
first step is to write down the covariant derivative as in (3.51). We denote the generators
of the 2representation of SU(2)asTa
2=1
2a(the Pauli matrices) and the gauge fields
asAa
. The generator of U(1)isY=C1 0
0 1
whereCis the hypercharge ( 1=2in this
case), and the U(1)gauge field is B. So, the covariant derivative is
(D)i=@i i[g2Aa
Ta
2+g1BY]ijj (3.56)
whereg1andg2are coupling constants for the U(1)part and the SU(2)part, respectively.
If the reason we wrote it down this way isn’t clear, compare this expression to equation
(3.51), and remember that we are saying the field carries twocharges; one for SU(2)and
one forU(1). Therefore, it has two terms in its covariant derivative. And, as usual, is a
spacetime index.
Knowing that the generators of SU(2)are the Pauli matrices, we can expand the second
part of the covariant derivative in matrix form,
g2Aa
Ta
2+g1BY=g2
2(A1
1+A2
2+A3
3) g1
2BI22
=1
2g2A3
g1Bg2(A1
iA2
)
g2(A1
+iA2
) g2A3
g1B
So, the full covariant derivative is
(D)i_ =D1
D2
=@1+i
2(g2A3
g1B)1+ig2
2(A1
iA2
)2
@2+ig2
2(A1
+iA2
)1 i
2(g2A3
+g1B)2
(3.57)
Now, we know that the Lagrangian will have the kinetic term and some potential:
L= 1
2Dy
iDi V(y;) (3.58)
Let’s assume that the potential has a similar form as equation (3.44) (we add the factors
of one-half here for the sake of convention; they don’t amount to anything other than a
rescaling of and),
V(y;) =1
4
y 1
222
(3.59)
Clearly the minimum field configuration is not at = 0, but atjj=vp
2. So, following
what we did in section 3.3.7, we make a global SU(2)transformation to put the entire
116
VEV on the firstcomponent of , and then make a global U(1)transformation to make the
field real. So,
h0jj0i=1p
2v
0
(3.60)
and we expand around this new vacuum:
(x) =1p
2v+h(x)
0
(3.61)
Remember that we have chosen our SU(2)to keep the second component 0 and our U(1)
to keep the first component real. So, h(x)is a real scalar field.
Clearly, plugging this into the covariant derivative (3.57) will give the exact same expres-
sion as before, but with 1replaced by1p
2h(x)and2replaced by 0, plus an extra term
forv. When we plug this extra term into the kinetic term in the Lagrangian (3.58), we get
that it is
1
8v2
1 0g2A3
g1Bg2(A1
iA2
g2(A1
+iA2
g2A3
g1Bg2A3 g1Bg2(A1 iA2)
g2(A1+iA2) g2A3 g1B1
0
(3.62)
Before multiplying this out, we employ a trick. Define the Weak Mixing Angle
wtan 1g1
g2
and the shorthand notation
swsinwand c wcosw
And finally, we can define four new gauge fields as linear combinations of the four we
have been using:
W+
1p
2(A1
iA2
) (3.63)
W
1p
2(A1
+iA2
) (3.64)
ZcwA3
swB (3.65)
AswA3
+cwB (3.66)
These can easily be inverted to give the old fields in terms of the new fields,
A1
=1p
2(W+
+W
) (3.67)
A2
=ip
2(W+
W
) (3.68)
A3
=cwZ+swA (3.69)
B= swZ+cwA (3.70)
117
We make a few observations about these fields before moving on. First of all, they are
merely linear combinations of the gauge fields introduced in equation (3.56). Second,
notice that the two fields W
are both linear combinations of fields corresponding to non-
Cartan generators of SU(2), whereasZandAare both linear combinations of fields
corresponding to Cartan generators of SU(2)andU(1). So, according to our discussion in
section 2.2.16, we expect that ZandAwill interact but not change the charge, and that
W
will interact and change the charge. Incidentally, notice that W
has the exact form
of the raising and lowering operators defined in (2.18).
With these fields defined, we can now rewrite (3.62) as
1
8g2
2v2
1 01
c2Zp
2W+
p
2W
?21
0
= M2
wW+W
1
2M2
ZZZ
(the?is there because that matrix element will always be multiplied by 0, so we don’t
bother writing it), where we have defined
Mw=g2v
2andMZ=Mw
cw=g2v
2cw
So, we see that, by symmetry breaking, we have given mass to the W+
, theW
, and the
Zfields. However, the Ahas not gained mass.
These particles are the WandZvector bosons, which are the force carrying particles of
theWeak Force . Each of these particles has an extremely large mass ( MW80.4 GeV ,
andMZ91.2 GeV), which explains why they only act over a very short range ( 10 18
meters).
Also note that the Aremains massless, implying that it did not acquire a VEV , and be-
cause it is a single generator, we see that a single U(1)remains unbroken. This U(1)and
Aare the gauge group and field of Electromagnetism , as discussed in section 3.1.11.
The idea of all of this is that at very high energies (above the breaking of the SU(2)
U(1)),
we have only a Higgs complex scalar field, along with four identical massless vector
boson gauge fields ( A1
;A2
;A3
;B), each of which behave basically like a photon. At
low energies, however, the SU(2)
U(1)symmetry of the Higgs is broken, and the low
energy effective theory consists of a linear combination of the original four fields. Three
of those linear combinations have gained mass, and one remains massless, retaining the
photon-like properties from before symmetry breaking. The theory above the symmetry
breaking scale is called the Electroweak Theory (with four photon-like force carrying
particles), whereas below the breaking scale they become two separate forces; the Weak
and the Electromagnetic . This is the first and most basic example of unification we have
in our universe. At low energies, the electromagnetic and weak forces are separate. At
high energies, they unify into a single theory that is described by SU(2)
U(1).
We can express the new fields as simple Euler rotations of the old fields:
Z
A
=A3
cosw Bsinw
A3
sinw+Bcosw
)Z
A
=R(w)A3
B
118
So, theZis a massive linear combination of the A3
andB, while the photon Ais a
massless linear combination of the two.
We can do the same type of analysis for the W
, where they are both massive linear
combinations of A1
andA2
. TheZandAare both made up of a mixture of the SU(2)
andU(1)gauge groups, whereas the W
come solely from the SU(2)part.
Before moving on to include leptons (and then hadrons), we first write out the full La-
grangian for the effective field theory for h(x)and the gauge fields.
We start with the complete Lagrangian term for h(x). We have written the original field
as in equation (3.61). So, our potential in equation (3.59) is now
V(y;) =1
4(y 1
2v2)2==1
4v2h2+1
4vh3+1
16h4
The first term on the right hand side is clearly a mass term giving the mass of the Higgs
(=q
2v), and the second two terms are interaction vertices. The kinetic term for the
Higgs will be the usual 1
2@h@h.
Now, following loosely what we did in section 3.1.11, we want to find kinetic terms for
the gauge fields. We start by finding them for the original gauge fields before symmetry
breaking (A1
;A2
;A3
andB). Using (3.55), (3.48), and (3.50), and the SU(2)structure
constants given in equation (2.17), we have
F1
= 2 Tr (FT1)
= 2 Tr
(@A @A ig2[A;A])T1
= 2 Tr
(@Aa
Ta @Aa
Ta ig2Aa
Ab
[Ta;Tb])T1
= 2 Tr (@Aa
TaT1 @Aa
TaT1 ig2Aa
Ab
ifabdTcT1)
=@Aa
a1 @Aa
a1+g2Aa
Ab
fabcc1
=@A1
@A1
+g2Aa
Ab
fab1
=@A1
@A1
+g2Aa
Ab
ab1
=@A1
@A1
+g2(A2
A3
A2
A3
)
And similarly,
F2
=@A2
@A2
+g2(A3
A1
A3
A1
)
F3
=@A3
@A3
+g2(A1
A2
A1
A2
)
And the gauge field corresponding to the U(1)will be defined as in (3.20):
B=@B @B
So, we can now write the kinetic term for our fields according to equation (3.54):
LKin= 1
4F
aFa
1
4BB
119
We can then use (3.67–3.70) to translate these kinetic terms into the new fields. We will
spare the extremely tedious detail and skip right to the Lagrangian:
Leff= 1
4FF 1
4ZZ DyW DW+
+DyW DW+
+ie(F+ cotwZ)W+
W
1
2e2
sin2w
(W+W
W+W
W+W+
W W
)
(M2
WW+W
+1
2M2
ZZZ)
1 +h
v2
= 1
2@h@h 1
2m2
hh2 1
2m2
h
vh3 1
8m2
h
v2h4
where we have chosen the following definitions:
F=@A @A (Electromagnetic Field Strength)
Z=@Z @Z (Kinetic term for Z)
D=@ ie(A+ cotwZ)
and the rest of the terms were defined previously in this section.
3.4.3 The Lepton Sector
We now turn to the lepton sector (which is still in the SU(2)
U(1)part of the Standard
Model gauge group). A Lepton is a spin- 1=2particle that does notinteract with the SU(3)
color group (the strong force). There are six Flavors of leptons arranged into three Fami-
lies, orGenerations . The table on page 139 explains this. The first generation consists of
the electron ( e) and the electron neutrino ( e), the second generation the muon ( ) and the
muon neutrino ( ), and the third the tau ( ) and tau neutrino ( ). Each family behaves
exactly the same way, so we will only discuss one generation in this section ( eande). To
incorporate the physics of the other families, merely change the eto either a or, and
theeto aorin the following notes.
What we will see is that, in a sense, the neutrinos don’t really interact with anything on
their own (which is why they are incredibly difficult to detect). For this reason, neutrinos
don’t have their own place in a representation of SU(3)
SU(2)
U(1)(see table on 139).
Electrons on the other hand, do interact with other things on their own, and we therefore
see them in the (1;1)representation.
However, the neutrino does interact with other things as part of an SU(2)doublet with the
electron,
l=e
e
(3.71)
120
This is why it is arranged as it is on page 139 with the electron under the (2; 1=2)repre-
sentation of SU(2)
U(1).
This may seem confusing, but we hope the following will make it clear. We will proceed
in what we believe is the clearest way to see this (primarily following [32]). We start with
2 fields, eandl, where eis a single left-handed Weyl field (see section 3.2.4), and lis
defined in (3.71). As we have said, lis in the (2; 1=2)representation, eis in the (1;1)
representation, and ehas no representation of its own.
So, mimicking what we did in equation (3.56) in the previous section, we can write down
the covariant derivative for each field,
(Dl)i=@li ig2Aa
(Ta)ijlj ig1BYlli (3.72)
De=@e ig1BYee (3.73)
The field ehas noSU(2)term in its covariant derivative because the 1representation of
SU(2)is the trivial representation - this means it doesn’t carry SU(2)charge. Also, we
know that
Yl= 1
21 0
0 1
(3.74)
and
Ye= (1)1 0
0 1
(3.75)
Following the Lagrangian for the spin- 1=2fields we wrote out in equation (3.11), we can
write out the kinetic term for both (massless) fields:
LKin=ilyi(Dl)i+ieyDe (3.76)
At the end of section 3.2.2 and of section 3.2.11, we briefly discussed the idea of renormal-
ization. We said that certain theories can be renormalized and others cannot. It turns out
(for reasons beyond the scope of these notes) that while the theory we have outlined so
far is renormalizable, if we try to add mass terms for and landefields, the theory breaks
down. Therefore we cannot add a mass term. But, we know experimentally that electrons
and neutrinos have mass, so obviously something is wrong. We must incorporate mass
into the theory, but in a more subtle way than merely adding a mass term. It turns out
that we can use the Higgs mechanism as follows.
While adding mass terms renders the theory inconsistent, we can add a Yukawa term (cf.
equation (3.42)),
LYuk= yijilje+h.c.
121
whereyis another coupling constant, ijis the totally antisymmetric tensor, and h.c. is the
Hermitian Conjugate of the first term.
Now that we have added LYukto the Lagrangian, we want to break the symmetry exactly
as we did in the previous section. First, we replace 1with1p
2(v+h(x))and2with 0,
exactly as we did in equation (3.61). So,
LYuk = yijilje+h.c.
= y(1l2 2l1)e+h.c.
= 1p
2y(v+h)l2e+h.c.
= 1p
2y(v+h)ee 1p
2y(v+h)ee
= 1p
2y(v+h)EE (3.77)
whereE=e
ey
is the Dirac field for the electron ( eis the electron and eyis the anti-
electron, or positron). Comparing (3.77) with (3.12), we see that it is a mass term for the
electron and positron.
Now we want a kinetic term for the neutrino. It is believed that neutrinos are described
by Majorana fields (see section 3.2.4), so we begin with the field N0=e
y
e
. Now, we
employ a trick. Referring back to equations (3.33) and (3.34), the kinetic term for Majorana
fields has only one term (because Majorana fields have only one Weyl spinor), whereas
the Dirac field sums over both Weyl spinors composing it. So, instead of working with
the Majorana field N0, we can instead work with the Dirac field
N=e
0
So, the Dirac kinetic term iN
@Nwill clearly result in the correct kinetic term from
(3.76), oriy@.
Now, continuing with the symmetry breaking, we want to write the covariant derivative
(3.72) and (3.73) in terms of our low energy gauge fields (3.63–3.66).
We said in the previous section (which echoed our discussion in section 2.2.16) that the
gauge fields corresponding to Cartan generators ( AandZ) act as force carrying parti-
cles, but do not change the charge of the particles they interact with. On the other hand,
the non-Cartan generators’ gauge fields ( W
) are force carrying particles which dochange
the charge of the particle they interact with. Therefore, to make calculations simpler, we
will break the covariant derivative up into the non-Cartan part and the Cartan part.
122
The non-Cartan part of the covariant derivative (3.72) is
g2(A1
T1+A2
T2) =1
2g2
A1
0 1
1 0
+A2
0 i
i0
=1
2g20A1
iA2
A1
+iA2
0
=g2p
20W+
W
0
and the Cartan part is
g2A3
T3+g1BY=e
sw(swA+cwZ)T3+e
cw(cwA swZ)Y
=e(A+ cotwZ)T3+e(A tanwZ)Y
=e(T3+Y)A+e(cotwT3 tanwY)Z
We have noted before that Ais the photon, or the electromagnetic field, and eis the
electromagnetic charge. Therefore, the linear combination T3+Ymust be the generator
of electric charge. Notice that the electromagnetic generator is in a linear combination of
the two Cartan generators of SU(2)
U(1).
We know that T3=1
23, andYlandYeare defined in equations (3.74) and (3.75), so we
can write
T3l=1
21 0
0 1e
e
=1
2e
e
Yll= 1
21 0
0 1e
e
= 1
2e
e
And we know that ecarries noT3charge, so its T3eigenvalue is 0, while Yeis+1. So,
summarizing all of this,
T3e= +1
2eT3e= 1
2e T3e= 0
Ye= 1
2eYe= 1
2e Y e= +e
Then defining the generator of electric charge to be QT3+Y, we have
Qe= 0Qe= e Q e= +e
So the neutrino ehas no electric charge, the electron ehas negative electric charge, and
the antielectron, or positron, has plus one electric charge—all exactly what we would
expect.
We can now take all of the terms we have discussed so far and write out a complete
Lagrangian. However, doing so is both tedious and unnecessary for our purposes.
123
The primary idea is that electrons/positrons and neutrinos all interact with the SU(2)
U(1)gauge particles, the W,Z, andA. TheZandA(the Cartan gauge particles)
interact but do not affect the charge. On the other hand, the Wact asSU(2)raising
and lowering operators (as can easily be seen by comparing (3.63) and (3.64) to equation
(2.18)). The SU(2)doublet state acted on by these raising and lowering operators is the
doublet in equation (3.71). The W+interacts with a left-handed electron, raising its elec-
tric charge from minus one to zero, turning it into a neutrino. However W+does not
interact with left-handed neutrinos. On the other hand, W will lower the electric charge
of a neutrino, making it an electron. But W will not interact with an electron.3
3.4.4 The Quark Sector
A Quark is a spin- 1=2particle that interacts with the SU(3)color force. Just as with lep-
tons, there are six flavors of quarks, arranged in three families or generations (see the
table on page 139).
Following very closely what we did with the leptons, we work with only one generation.
Extending to the other generators is then trivial. To begin, define three fields: q,u, and d,
in the representations (3;2;+1=6),(3;1; 2=3), and (3;1;+1=2)ofSU(3)
SU(2)
U(1).
The fieldqwill be the SU(2)doublet
q=u
d
(3.78)
This is exactly analogous to equation (3.71).
Again, following what we did with the leptons, we can write out the covariant derivative
for all three fields:
(Dq)i=@qi ig3Aa
(Ta
2)
qi ig2Aa
(Ta
2)j
iqj ig11
6
Bqi (3.79)
(Du)=@u ig3Aa
(Ta
3)
u ig1
2
3
Bu(3.80)
(Dd)=@d ig3Aa
(T
3)
d ig11
3
Bd(3.81)
whereiis anSU(2)index andis anSU(3)index. The SU(3)index is lowered for the 3
representation and raised for the 3representation.
Just as with leptons, we cannot write down a mass term for these particles, but we can
include a Yukawa term coupling these fields to the Higgs:
LYuk= y0ijiqjd y00yiqiu+h.c.
3This does not mean that no vertex in the Feynman diagrams will include a W and an electron field,
but rather that if you collide an electron and a W , there will be no interaction
124
As with the leptons, we can break the symmetry according to equation (3.61), and writing
out this Yukawa term, we get
LYuk = 1p
2y0(v+h)(dd+dy
dy) 1p
2y00(v+h)(uu+ uy
uy)
= 1p
2y0(v+h)DD 1p
2y00(v+h)UU
where we have defined the Dirac fields for the up and down quarks:
Dddy
Uu
uy
Notice that, whereas both the up and down quarks were massless before breaking, they
have now acquired masses
md=y0vp
2mu=y00vp
2
Writing out the non-Cartan and Cartan parts of the covariant derivatives in terms of the
lower energy SU(2)
U(1)gauge fields, we get
g2A1
T1+g2A2
T2=g2p
20W+
W
0
g2A3
T3+g1BY=eQA+e
swcw(T3 s2
wQ)Z
And it is again straightforward to find the electric charge eigenvalue for each field:
Qu= +2
3u Qd = 1
3d Q u= 2
3u Q d= +1
3d
Again, we can collect all of these terms and write out a complete Lagrangian. But, doing
so is extremely tedious and unnecessary for our purposes.
The primary idea to take away is that the SU(2)doublet (3.78) behaves exactly as the lep-
ton doublet in (3.71) when interacting with the “raising” and “lowering” gauge particles
W. This is why the uanddare arranged in the SU(2)doubletqin (3.78), and why q
carries theSU(2)indexiin the covariant derivative (3.79), whereas uand dcarry only the
SU(3)index.
TheSU(3)index runs from 1 to 3, and the 3 values are conventionally denoted red, green ,
and blue(r;g;b ). These obviously are merely labels and have nothing to do with the colors
in the visible spectrum.
125
The eight gauge fields associated with the eight SU(3)generators are called Gluons , and
they are represented by the matrices in (2.48). We label each gluon as follows:
g
_ =0
@rr rg rb
gr gg gb
br bg bb1
A
so that the upper index is the anti-color index, and denotes the column of the matrix,
and the lower index is the color index denoting the row of the matrix. Then, from (2.48),
consider the gluon
gg
r/0
@0 1 0
0 0 0
0 0 01
A
and the quarks
qr=0
@1
0
01
Aqg=0
@0
1
01
Aqb=0
@0
0
11
A
It is easy to see that this gluon will interact as
gg
rqr= 0gg
rqg=qrgg
rqb= 0
Or in other words, the gluon with the anti-green index will only interact with a green
quark. There will be no interaction with the other quarks. Multiplying this out, and
looking more closely at the behavior of the SU(3)generators and eigentstates as discussed
in sections 2.2.12–2.2.15, you can work out all of the interaction rules between quarks and
gluons. You will see that they behave exactly according to the root space of SU(3).
3.5 References and Further Reading
The primary source for these notes is [32], which is an exceptionally clear introduction to
Quantum Field Theory. We also used a great deal of meterial from [3], [24], [29], and [35],
all of which are outstanding QFT texts. The derivation of the Dirac equation came from
[21], which is written mostly above the scope of these notes, but is an excellent survey of
some of the mathematical ideas of Non-Perturbative QFT and Gauge Theory.
The sections on the Standard Model come almost entirely from [32] with little change, in
that Srednicki’s exposition could hardly be improved upon for the scope of these notes.
For further reading, we also recommend [1], [11], [23], [26], and [27].
126
4 The Standard Model — A Summary
4.1 How Does All of This Relate to Real Life?
In the fifth century B.C., a Greek named Empedocles took the ideas of several others
before him and combined them to say that matter is made up of earth, wind, fire, and
water, and that there are two forces, Love and Strife, that govern the way they grow and
act. More scientifically, he was saying that matter is made of smaller substances that
interact with each other through repulsion and attraction. Democritus, a contemporary
of Empedocles, went a step further to say that all matter is made of fundamental particles
that are indestructible. He called these particles atoms, meaning “indivisible”4[14].
The field of particle physics seeks to continue studying these same concepts. Are there
fundamental, indivisible particles and if so, what are they? How do they behave? How
do they group together to form the matter that we see? How do they interact with each
other?
The current answer to these questions is called the Standard Model, the theory we spent
this paper developing. We have now spent more than one hundred pages expositing a se-
ries of mathematical tricks for various types of “fields”. In doing so, we talked about
“massless scalars with U(1)charge”, and about things “in a j=1
2representation of
SU(2)”. But one could easily be left wondering how exactly this relates to the things
we see in nature. We only discussed 25 particles in the previous section and in the table
on page 139 (particles and antiparticles), but you are likely aware that there are hundreds
of particles in nature. What about those? How does the mathematical framework detailed
so far form the building blocks for the universe?
While we wish to reiterate that the primary purpose of these notes is to provide the math-
ematical tools with which particle physics is done, and not to outline the phenomenolog-
ical details of the theory, we are physicists still—not mathematicians. Therefore, before
concluding this paper, we will take a brief hiatus from the mathematical rigor and look at
a qualitative summary of particle physics.
Throughout this section, the footnotes will provide brief explanations of the analogous
mathematical ideas from above. This section5can be read with or without paying atten-
tion to them. We provide them merely for those curious.
4Of course, our modern use of the word is different. At their discovery, it was thought that different
elements were the indivisible particles sought for, so the name atom seemed appropriate
5Nearly everything in this section is adapted from [4], including the tables on page 134
127
4.2 The Fundamental Forces
The two forces most familiar to people are Gravity and Electromagnetism . Just the act
of standing on the ground or sitting in a chair makes use of both, and every “Physics I”
student has drawn a free body diagram with a gravitational force going down and a
normal force (caused by the electromagnetic repulsion between the two objects) going
up. However, these two are only half of the four fundamental forces in our universe (that
we know of).
We can think about the third by first considering a compact nucleus which we know to be
made of protons and neutrons. From electromagnetism we know that the protons should
repel each other because of their like charge. But the nuclei of atoms somehow hold to-
gether, which is evidence for some stronger force that causes these particles to attract.
This force, which overcomes the electromagnetic repulsions and allows atomic nuclei to
remain stable, is called the Strong force.6Just as electrically-charged particles are sub-
ject to the electromagnetic force, some particles have a property similar to charge, called
Color , and are subject to the strong force. The field theory that describes this is called
Quantum Chromodynamics (QCD)7and was first proposed in 1965 by Han, Nambu,
and Greenberg [20]. This theory predicts the existence of the gluon, which is the mediator
of the strong force between two matter particles.
The fourth force is the one we have the least familiarity with. It is responsible for certain
types of radioactive decays; for example, permitting a proton to turn into a neutron and
vice versa. It is called the Weak force.8
In the 1960’s, Sheldon Glashow, Abdus Salum, and Steven Weinburg independently de-
veloped a gauge-invariant theory that unified the electromagnetic and weak force [20].
At sufficiently high energies it is observed that the difference between these two separate
forces is negligible and that they instead act together as the Electroweak force.9For pro-
cesses at lower energy scales, the symmetry between the electromagnetic and the weak
force is broken and we observe two different forces with different properties. Similar
to QCD, electroweak theory predicts four force-carrier particles10that mediate the force
between matter particles. The mediating particle for electromagnetism is the neutral pho-
ton, and those for the weak force are the W+(with +1electron charge), W (with 1
electron charge) and Z0(neutral) bosons.
The electromagnetic, weak, and strong forces forces described above form what is called
theStandard Model of Particle Physics . The Standard Model is an incomplete theory
in the sense that it fails to describe gravitation, the force that acts on matter. Physicists
6TheSU(3)color force
7Again, the study of the SU(3)color force
8TheSU(2)part that is left over when SU(2)
U(1)is broken
9The unbroken SU(2)
U(1)force
10Corresponding to the 3 generators of SU(2)and the 1 in U(1). Two of them are Cartan and are, there-
fore, uncharged, while two are non-Cartan and therefore carry charge
128
continue to work towards a theory that describes all four fundamental forces, with String
Theory currently the most promising. The papers later in this series will discuss these
ideas. For the rest of the sections in this review, however, it should be assumed that we are
talking about physics under the Standard Model only11which, despite the shortcoming
of not explaining gravity, has tremendous experimental support.
4.3 Categorizing Particles
In the last century, experimenters were surprised as they discovered new particle after
new particle. It seemed disorganized and overwhelming that there could be so many
elementary objects. Eventually, however, the properties of these particles became better
understood and it was found that there really is just a small, finite set of fundamental
particles, some of which can be grouped together to make up larger objects. In the next
two sections, we will introduce the elementary particles and then will discuss the types
of composite particles.
One property of the ”zoo” of discovered particles that helps in our organizing them is
their intrinsic spin.12Any particle, elementary or composite, that is of half-integer spin13
is aFermion . Those with integer spin are Bosons .14The spins govern the statistics of a set
of such particles, so fermions and bosons may also be defined according to the statistics
they obey.
Namely, fermions obey Fermi-Dirac statistics and therefore also obey the Pauli Exclusion
Principle. This means that no two identical fermions can be found in the same quantum
state at the same time. Furthermore, to accurately display this behavior it is found that
the wave function of a system with fermions must be antisymmetric; swapping any two
like fermions causes a change in sign of the overall wave function.
Bosons on the other hand obey Bose-Einstein statistics; any number of the same type of
particle can be in the same state at the same time. In contrast to fermions, the wavefunc-
tion of a system of bosons is symmetric.15
The Venn diagrams on page 138, the table provided on page 139, and the table below
should be referenced as you read through what follows.
11We are not assuming Supersymmetry in this paper, though we will consider Supersymmetry in a later
paper
12Or in other words, which representation of SU(2)they sit in
13Is in the j=1
2orj=3
2representation of SU(2)
14Is in the j= 0orj= 1representation of SU(2)
15Cf. section 3.2.3
129
4.4 Elementary Particles
The elementary particles are those that are considered fundamental, or in other words,
are not composed of smaller particles.16They can be divided into two groups: matter
particles and non-matter particles.
The elementary matter particles all have half-integer spin (so are fermions) and the ele-
mentary non-matter particles all have integer spin (so are bosons). We can then observe
that an equivalent grouping is made if we divide the elementary particles instead by their
intrinsic spin, which is commonly done. Then an elementary matter particle is the same
thing as an elementary fermion, and similarly for the bosons. The two terms are used
interchangeably in the discussion below.
4.4.1 Elementary Fermions
The elementary fermions are the building blocks of all other matter. For example, the
proton and neutron are made up of different combinations of three elementary quarks.
Electrons, which are also elementary, cloud around the protons and neutrons, and when
all three group together in a particular way, an atom is formed. Less familiar examples
include those that are unstable, such as the muon, which decay into something else fairly
quickly.
For every elementary (and sometimes composite) matter particle, there is also a corre-
sponding particle with the same mass but of different charge and magnetic moment.17
Generally the name of such a particle is the same as the corresponding “normal” matter
particle, but with the prefix “anti” in front of it (e.g. antiquark, antilepton, etc.). In this
paper, whenever we discuss matter and its properties, it is implied that the antimatter
counterparts have similar properties.
Now we further divide the elementary fermions into two groups, quarks and leptons. A
convenient way to distinguish these two sets is by whether or not they interact via the
strong force: quarks may interact via the strong force, while leptons do not.
Quarks
Experiments involving high energy collisions of electrons and protons led Murray Gell-
Mann to suggest in 1964 [25] that protons and neutrons are actually composite particles,
made of three point-like, spin- 1=2particles whose charges are either 1=3or+2=3units
of electron charge. He called these particles Quarks . Through further experiments it has
been found that there are six flavors of quarks total, grouped into three generations with
16These are the ones that are in some representation of SU(3)
SU(2)
U(1)on the table on page 139
17Cf. material on spin- 1=2particles in section 3.1
130
the first generation containing the up and down quarks, the second generation containing
the more massive charm and strange quarks, and the third generation containing the even
more massive top and bottom quarks.
As electrically charged particles are subject to the electromagnetic force, quarks have a
property similar to charge, called color, and any colored particle is subject to the strong
force. It is found that there are three different types of colors: (defined as) red, green,
and blue (plus three more for antiquarks: antired, antigreen, and antiblue). Quarks are
grouped together to make composite particles that are colorless (the color charges cancel
out), which is why the concept of color was only discovered after quarks themselves were
found. The addition of color to the quark model also ensures that any quarks contained
in a composite particle will not violate the exclusion principle since each has a different
color. Again, QCD is the field theory that describes these properties.
Another interesting feature of quarks is that they are never found alone, but rather always
inside of a composite particle. This phenomenon is called Confinement .18It is more
a property of the strong force, which increases in strength as two colored particles are
pulled away from each other, just as would happen when the ends of a piece of elastic are
pulled apart. We can consider reaching a distance between the two quarks where there
is sufficient potential energy built up that it can be converted to matter, creating a quark-
antiquark pair. The pair will separate and the resulting particles will recombine with
the original quarks. As this process repeats, and more quark-antiquark pairs are created,
the end result in the whole process is a multiplication of the number of quarks and of the
number of composite particles. In the opposite extreme, as two quarks get closer together,
the strong force between them becomes weaker until the quarks move around freely and
more independently. This is a called Asymptotic Freedom.
Quarks also interact with other particles via the weak force, which is the only force that
can cause a change of flavor (changing an up into a down, for example). When this
happens, a quark either turns into a heavier quark by absorbing a Wboson, or it emits a
Wboson and then decays to a lighter quark. Beta decay, a common radioactive process,
is caused by this mechanism. Instead of just thinking of beta decay as a neutron in the
nucleus of an atom decaying, or splitting, into a proton, electron, and antineutrino, we
can go a step further with our understanding of quarks subject to the weak force. We add
that, really, it is one of the down quarks in the neutron that emits a W boson and then
decays to the lighter up quark, keeping charge conserved in the process.19The neutron,
which used to have one up and two down quarks, now has one down and two up quarks,
which is the composition of a proton. The electron and antineutrino are created from the
decay of the W boson.
18We did not discuss confinement in the main body of this paper, though it can be derived from what we
did discuss
19Other conserved quantities are momentum, energy, quark number, lepton number, and (approxi-
mately) lepton generation number
131
Leptons
Leptons interact with other matter via the electromagnetic, the weak, and gravitational
forces, but not through the strong force.20There are three charged leptons, grouped, like
the quarks, into three different generations based on their masses.21The electron is the
lightest of the charged leptons, then the muon, and the tau. There are also three neutral
leptons, called neutrinos (“little neutral one”), one type for each of the charged leptons:
the electron neutrino, the muon neutrino, and the tau neutrino.
Some quantities in lepton events are found to be conserved.22If we define lepton num-
ber as the number of leptons minus the number of antileptons, then lepton number is
constant in all interactions. Additionally, the lepton number within each generation is also
approximately conserved. For example, the number of electrons and electron neutrinos
minus the number of antielectrons and electron antineutrinos is found to be constant in
most particle reactions.
An interesting exception is in neutrino oscillations, where a neutrino changes lepton fla-
vors as it travels. For example, we can take a measurement and observe an electron
neutrino, even though it was known to have been created as a muon neutrino. These
oscillations of flavor only occur if neutrinos have mass (even just very small mass), so the
fact that the Standard Model currently predicts them to be massless demonstrates that
there are some parameters in the theory that need to be adjusted.
4.4.2 Elementary Bosons
Throughout the development of the Standard Model it was found that some elementary
particles play a different role than the ordinary matter particles that make up the stuff of
the universe. Both the gauge bosons and the Higgs boson fall into this group.
Gauge Bosons
In the mathematical formulation of quantum field theory, the Lagrangian can be made
invariant under a local gauge transformation by the addition of a vector field called a
gauge field. As with the more familiar example of an electron, the quanta of the gauge
field is a type of particle, which in this case is called a Gauge Boson . There are three
types of gauge bosons described by the Standard Model.23They are the photon, which
20This means that leptons carry SU(2)
U(1)terms in their covariant derivatives, but not SU(3)terms
21This is equivalent to the statement above that there are three copies of the Standard Model Gauge Group
- Cf. page 115
22These conservation laws can all be derived from the rules we discussed above, though they are typically
treated separately because they are extremely useful when talking about specific interactions
23This is equivalent to saying that there are three gauge groups, each with their own set of generators
132
carries the electromagnetic force, the WandZbosons, which carry the weak force, and
the gluons which carry the strong force. Each of these bosons have been experimentally
detected.
Evidence for the neutral photon first came in 1905 when Einstein proposed an explanation
of the photoelectric effect, that light was quantized into energy packets [7]. Confirmation
of theW+,W , andZ0bosons came in 1983 through proton-proton collisions at the Eu-
ropean Organization for Nuclear Research (CERN) [5].
The gluons were first experimentally observed in 1979 in the electron-position collider
at the German Electron Synchroton (DESY) in Hamburg [5]. Further experiments have
demonstrated that the gluons have eight different color states and that, because they in-
teract via the strong force, they have properties similar to quarks, such as confinement.
Taking into account their possible charge or color, we find that there are 12 gauge bosons
in all, one for the electromagnetic force, three for the weak force, and eight for the strong
force.
The Higgs Boson
The Higgs boson is the only Standard Model particle that has not yet been observed. It is
also the only elementary boson that is not a gauge boson. Rather, it is the carrier particle
of the scalar Higgs field from which other particles acquire mass. The existence of the
Higgs would explain why some particles have mass and others do not. For example,
theWandZbosons are very massive, whereas the photon is massless. One of the main
goals of the Large Hadron Collider (LHC), located at CERN in Switzerland, is to provide
evidence for the Higgs. It is expected to be in full operation in 2009.
4.5 Composite Particles
Examples of composite particles include hadrons, nuclei, atoms, and molecules. The latter
three are well known and will not be described here.
Hadrons are made up of bound quarks and interact via the strong force. They can be
either fermions or bosons, depending on the number of quarks that make them up. An
odd number of bound quarks create a spin- 1=2or spin- 3=2hadron, which is called a
baryon, and an even number of quarks create spin-0 or spin-1 hadrons, called mesons.
Experimentally, only combinations of three quarks or two quarks have been found, so the
terms baryon and meson often just refer to three or two bound quarks, respectively.
You can understand why mesons and baryons have the spin that they do by considering
how many spin- 1=2quarks compose them. A meson has two quarks, and therefore the
total spin of a meson is the sum of an even number of half-integer spin particles, which
133
will be integer spin. And because there are only two of them, it is either spin 0 or 1.
Baryons, on the other hand, will have a linear combination of three particles with half-
integer spin, which will of course be half-integer: 1=2or3=2.
The most well known examples of baryons are protons and neutrons. Protons are made
of two up quarks and one down quark, or juudi, and neutrons are made of two down
and one up, orjuddi. The baryons are made of “normal” quarks only and their antimatter
counterparts are made of the corresponding antiquarks.
The mesons are made of a quark and an antiquark pair, though not necessarily of the
same generation. Examples include the +judiandK+jusi.
One of the reasons for the zoo of particles discovered in the past century is because of
the numerous possible combinations of six quarks put into a three-quark or two-quark
hadron. Additionally, each of these combinations can be in different quantum mechanical
states, thereby displaying different properties. For example, a rho meson has the same
combination of quarks as a pion , but theis spin-1 whereas the pion is spin-0.
4.6 Visualizing It All
Finally, we provide a few tables which should help you see all of this more clearly.
Interactions Acts On Strength Range
Strong Hadrons 1 10 15m
Electromagnetism Electric Charges 10 21(1=r2)
Weak Leptons and Hadrons 10 510 18m
Gravity Mass 10 391(1=r2)
where the relative strengths have been normalized to unity for the strong force.
Also, the four classes of force-carrying gauge bosons are shown below;24
Interaction Gauge Boson Spin Acts On
Strong Gluon 1 Hadrons
Electromagnetism Photon 1 Electric Charges
Weak W,Z , 1 Leptons and Hadrons
Gravity Graviton 2 Mass
24The graviton is a the hypothetical carrying particle for the gravitational force; it is not described by the
Standard Model
134
5 A Look Ahead
Now that we have completed our introduction to basic particle theory, we can begin our
uphill climb towards more fundamental concepts. As a preview, notice that everything
we have done so far has been an exposition of how gauge theories work. Our investiga-
tion into gauge theories has been purely algebraic (working entirely from group theory, as
Part II demonstrates). As gauge theory seems to be the correct approach to understand-
ing our universe, everything we do for the remainder of this series will be focused on a
more fundamental understanding of gauge theory, culminating in String Theory.
As we just stated, we have been treating gauge theory as a purely algebraic construct.
However String Theory, if true, must obviously be able to reproduce the same general
framework we have seen so far. But, String Theory is fundamentally a geometric con-
struct. As we will see, String Theory will reproduce literally everything we have seen
about gauge theory, but from a geometric framework.
This should not be entirely foreign, though. Recall that, for electromagnetism, the gauge
group isU(1). We can “draw” this geometrically as a circle in the complex plane. The
Weak force is represented by the gauge group SU(2), which we have seen is parame-
terized by three numbers, and therefore has three generators. As we discussed in these
notes, we should think of these spaces as vector spaces and the generators as basis vectors
spanning the entire space. The same is true of SU(3), though it is an eight-dimensional
space. So, because there is a space associated with each of these groups, it should be
somewhat obvious that there is a natural geometric picture associated with a Lie group.
While the idea of the parameter space of a Lie group having a geometric picture associ-
ated with it may seem straightforward, the geometry undergirding gauge theory can be
extremely complicated, and we therefore must spend a significant amount of time inves-
tigating it. Therefore, the next paper in this series will be an introduction to the geometric
structure of gauge theory. Just as we have built gauge theory from algebra, we will in a
sense start over and rebuild it using geometry. However, because we have already cov-
ered a great deal of detail in the physics and mathematics of gauge theory and particle
physics in general, we will move much more quickly to avoid being repetitive.
When we finally get to String Theory (later in this series), we will see that the geometric
and algebraic pictures come together beautifully, and that a thorough understanding of
both will be necessary to understand what may be the “ultimate” theory of our universe.
135
References
[1] D. Bailin and A. Love, “Introduction to Gauge Field Theory”, Taylor and Francis
(1993)
[2] R. Cahn, “Semi-Simple Lie Groups and Their Representations”, Dover (2006)
[3] W. N. Cottingham and D. A. Greenwood, “Introduction to the Standard Model”,
Cambridge University Press (2007)
[4] R. Dunlap, “An Introduction to the Physics of Nuclei and Particles”, Brooks Cole
(2003)
[5] V . Ezhela and B. Armstrong, ”Particle Physics: One Hundred Years of Discoveries:
An Annotated Chronological Bibliography”, Springer (1996)
[6] R. P . Feynman, R. B. Leighton, and M. Sands, “The Feynman Lectures Vols. I-III”,
Addison Wesley (2005)
[7] K. Ford, ”The Quantum World”, Harvard University Press (2005)
[8] J. B. Fraleigh, “A First Course in Abstract Algebra”, Addison Wesley (2002)
[9] H. Georgi, “Lie Algebras and Particle Physics”, Westview Press (1999)
[10] R. Gilmore, “Lie Groups, Lie Algebras, and Some of Their Applications”, Dover
(2006)
[11] R. Gilmore, “Lie Groups, Physics, and Geometry”, Cambridge University Press
(2008)
[12] H. Goldstein, C. P . Poole, and J. L. Safko, “Classical Mechanics”, Addison Wesley
(2001)
[13] D. J. Griffiths, “Introduction to Classical Electrodynamics”, Benjamin Cummings
(1999)
[14] J. Hakim, ”The Story of Science: Aristotle Leads the Way”, Smithsonian Books (2004)
[15] B. C. Hall, “Lie Groups, Lie Algebras, and Representations”, Springer (2004)
[16] J. E. Humphreys, “Introduction to Lie Algebras and Representation Theory”,
Springer (1994)
[17] T. W. Hungerford, “Algebra”, Springer (2003)
[18] J. D. Jackson, “Classical Electrodynamics”, Wiley (1998)
[19] J. V . Jose and E. J. Saletan, “Classical Dynamics: A Contemporary Approach”, Cam-
bridge University Press (1998)
136
[20] J. Mehra and H. Rechenberg, ”The Historical Development of Quantum Theory; v.6”,
Springer (2001)
[21] G. Naber, “Topology, Geometry, and Gauge Fields: Interactions”, Springer (2000)
[22] G. Naber, “The Geometry of Minkowski Spacetime”, Dover (2003)
[23] M. Nakahara, “Geometry, Topology, and Physics”, Taylor and Francis (2003)
[24] M. E. Peskin and D. V . Schroeder, “An Introduction to Quantum Field Theory”, West-
view Press (1995)
[25] A. Pickering, ”Constructing Quarks: A Sociological History of Particle Physics”, Uni-
versity of Chicago Press (1984)
[26] P . Ramond, “Field Theory: A Modern Primer”, Westview Press (2001)
[27] P . Ramond, “Journeys Beyond the Standard Model”, Westview Press (2003)
[28] J. Rotman, “An Introduction to the Theory of Groups”, Springer (1999)
[29] L. H. Ryder, “Quantum Field Theory”, Cambridge University Press (1996)
[30] B. Sagan, “The Symmetric Group”, Springer (2001)
[31] J. J. Sakurai, “Modern Quantum Mechanics”, Addison Wesley (1993)
[32] M. Srednicki, “Quantum Field Theory”, Cambridge University Press (2007)
[33] J. Schwinger, “Classical Electrodynamics”, Westview Press (1998)
[34] N. M. J. Woodhouse, “Special Relativity”, Springer (2007)
[35] A. Zee, “Quantum Field Theory in a Nutshell”, Princeton University Press (2003)
137
Particles
Fermions
Fundamental
Leptons (Spin-1/2) Quarks (Spin-1/2)
e
e u c t
d s b
Composite
Baryons
Spin-1/2 Spin-3/2
proton =juudi
neutron =juddi++=juuui
=jdddi
Bosons
Fundamental
Spin-0 Gauge
HiggsSpin-1 Spin-2
AZ
WgiGraviton
Composite
Mesons
Spin-0 Spin-1
+=judi
K+=jusi+=judi
K?+=jusi
138
Leptons Hadrons Higgs
(1,2, –1/2) (1, 1, 1) (3,2, 1/6) (3, 1, –2/3) (3, 1, 1/3) (1,2, –1/2)
Generation 1electron neutrino
electron
electronup
down
up down 1 Generation Only
Generation 2muon neutrino
muon
muoncharm
strange
charm strange
Generation 3tau neutrino
tau
tautop
bottom
top bottom
139