warren siegel fields book
PDF · 731 pages · 4.1 MB
Open PDF file
A full-length textbook by Warren Siegel of the C. N. Yang Institute at Stony Brook, dated December 1999, kept in a folder of downloaded physics books. Its three parts cover symmetry (Lorentz, spin, supersymmetry, Yang-Mills, the standard model), quanta (path integrals, BRST, gauges, loops, anomalies), and higher spin (general relativity, supergravity, strings, and mechanics). The preface criticizes traditional QFT texts and explains the book's unified approach.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
arXiv:hep-th/9912205 21 Dec 1999YITP-99-67
FIELDSFIELDS
Warren Siegel
C. N. Yang Institute for Theoretical Physics
State University of New York at Stony Brook
Stony Brook, New York 11794-3840 USA
mailto:[email protected]
http://insti.physics.sunysb.edu/~siegel/plan.html
i
CONTENTS
Preface::::::::::::::::::::::::::::: ivSome eld theory texts ::::::::: xviii
:::::::::::::::::: ::::::::::::::::::
:::::::::::::::::: PART ONE: SYMMETRY ::::::::::::::::::
I. Global
A.Coordinates
1.Nonrelativity :::::::::::::: 3
2.Fermions:::::::::::::::::: 7
3.Lie algebra ::::::::::::::: 11
4.Relativity:::::::::::::::: 15
5.Discrete: C, P, T ::::::::: 20
6.Conformal::::::::::::::: 23
B.Indices
1.Matrices::::::::::::::::: 28
2.Representations :::::::::: 30
3.Determinants :::::::::::: 35
4.Classical groups :::::::::: 38
5.Tensor notation :::::::::: 40
C.Representations
1.More coordinates ::::::::: 45
2.Coordinate tensors ::::::: 47
3.Young tableaux :::::::::: 51
4.Color and
avor :::::::::: 53
5.Covering groups :::::::::: 58
II. Spin
A.Two components
1.3-vectors::::::::::::::::: 61
2.Rotations:::::::::::::::: 64
3.Spinors:::::::::::::::::: 66
4.Indices::::::::::::::::::: 67
5.Lorentz:::::::::::::::::: 70
6.Dirac:::::::::::::::::::: 76
7.Chirality/duality ::::::::: 78
B.Poincar e
1.Field equations ::::::::::: 80
2.Examples:::::::::::::::: 83
3.Solution:::::::::::::::::: 84
4.Mass::::::::::::::::::::: 89
5.Foldy-Wouthuysen ::::::: 92
6.Twistors::::::::::::::::: 96
7.Helicity:::::::::::::::::: 98
C.Supersymmetry
1.Algebra::::::::::::::::: 103
2.Supercoordinates :::::::: 104
3.Supergroups :::::::::::: 106
4.Superconformal ::::::::: 109
5.Supertwistors ::::::::::: 110III. Local
A.Actions
1.General::::::::::::::::: 115
2.Fermions:::::::::::::::: 119
3.Fields::::::::::::::::::: 120
4.Relativity::::::::::::::: 122
5.Constrained systems ::::128
B.Particles
1.Free:::::::::::::::::::: 132
2.Gauges::::::::::::::::: 136
3.Coupling:::::::::::::::: 137
4.Conservation :::::::::::: 138
5.Pair creation :::::::::::: 141
C.Yang-Mills
1.Nonabelian :::::::::::::: 144
2.Lightcone::::::::::::::: 148
3.Plane waves ::::::::::::: 152
4.Self-duality ::::::::::::: 153
5.Twistors:::::::::::::::: 156
6.Instantons:::::::::::::: 159
7.ADHM::::::::::::::::: 163
8.Monopoles:::::::::::::: 165
IV. Mixed
A.Hidden symmetry
1.Spontaneous breakdown :171
2.Sigma models ::::::::::: 173
3.Coset space ::::::::::::: 176
4.Chiral symmetry :::::::: 177
5.Stuckelberg::::::::::::: 180
6.Higgs::::::::::::::::::: 182
B.Standard model
1.Chromodynamics :::::::: 185
2.Electroweak ::::::::::::: 189
3.Families::::::::::::::::: 193
4.Grand Unied Theories ::195
C.Supersymmetry
1.Chiral:::::::::::::::::: 200
2.Actions::::::::::::::::: 202
3.Covariant derivatives ::::204
4.Prepotential ::::::::::::: 207
5.Gauge actions ::::::::::: 209
6.Breaking:::::::::::::::: 211
7.Extended::::::::::::::: 214
ii
:::::::::::::::::::: ::::::::::::::::::::
:::::::::::::::::::: PART TWO: QUANTA ::::::::::::::::::::
V. Quantization
A.General
1.Path integrals ::::::::::: 221
2.Semiclassical expansion ::225
3.Propagators ::::::::::::: 229
4.S-matrices:::::::::::::: 231
5.Wick rotation ::::::::::: 235
B.Propagators
1.Particles:::::::::::::::: 239
2.Properties::::::::::::::: 242
3.Generalizations :::::::::: 245
4.Wick rotation ::::::::::: 248
C.S-matrix
1.Path integrals ::::::::::: 253
2.Graphs::::::::::::::::: 257
3.Semiclassical expansion ::262
4.Feynman rules :::::::::: 266
5.Semiclassical unitarity :::272
6.Cutting rules :::::::::::: 274
7.Cross sections ::::::::::: 277
8.Singularities ::::::::::::: 280
9.Group theory ::::::::::: 282
VI. Quantum gauge theory
A.Becchi-Rouet-Stora-Tyutin
1.Hamiltonian :::::::::::: 288
2.Lagrangian :::::::::::::: 292
3.Particles:::::::::::::::: 295
4.Fields::::::::::::::::::: 296
B.Gauges
1.Radial:::::::::::::::::: 300
2.Lorentz::::::::::::::::: 303
3.Massive::::::::::::::::: 305
4.Gervais-Neveu ::::::::::: 307
5.Super Gervais-Neveu ::::310
6.Spacecone::::::::::::::: 313
7.Superspacecone ::::::::: 317
8.Background-eld :::::::: 319
9.Nielsen-Kallosh ::::::::: 325
10.Super background-eld ::327
C.Scattering
1.Yang-Mills :::::::::::::: 331
2.Recursion::::::::::::::: 335
3.Fermions:::::::::::::::: 337
4.Masses:::::::::::::::::: 339
5.Supergraphs :::::::::::: 345VII. Loops
A.General
1.Dimensional renormaliz'n350
2.Momentum integration ::353
3.Modied subtractions :::357
4.Optical theorem ::::::::: 361
5.Power counting :::::::::: 363
6.Infrared divergences ::::: 367
B.Examples
1.Tadpoles:::::::::::::::: 371
2.Eective potential ::::::: 374
3.Dimensional transmut'n :377
4.Massless propagators ::::378
5.Massive propagators ::::: 381
6.Renormalization group ::385
7.Overlapping divergences :388
C.Resummation
1.Improved perturbation ::395
2.Renormalons :::::::::::: 400
3.Borel::::::::::::::::::: 403
4.1/N expansion :::::::::: 406
VIII. Gauge loops
A.Propagators
1.Fermion::::::::::::::::: 412
2.Photon::::::::::::::::: 415
3.Gluon::::::::::::::::::: 416
4.Grand Unied Theories ::422
5.Supermatter :::::::::::: 425
6.Supergluon :::::::::::::: 427
7.Bosonization :::::::::::: 432
8.Schwinger model :::::::: 435
B.Low energy
1.JWKB:::::::::::::::::: 441
2.Axial anomaly :::::::::: 444
3.Anomaly cancelation ::::447
4.0!2
:::::::::::::::: 450
5.Vertex:::::::::::::::::: 452
6.Nonrelativistic JWKB :::455
C.High energy
1.Conformal anomaly ::::: 460
2.e+e !hadrons:::::::: 463
3.Parton model ::::::::::: 465
iii
:::::::::::::: ::::::::::::::
:::::::::::::: PART THREE: HIGHER SPIN ::::::::::::::
IX. General relativity
A.Actions
1.Gauge invariance :::::::: 473
2.Covariant derivatives ::::476
3.Conditions :::::::::::::: 481
4.Integration :::::::::::::: 484
5.Gravity::::::::::::::::: 488
6.Energy-momentum :::::: 491
7.Weyl scale:::::::::::::: 494
B.Gauges
1.Lorentz::::::::::::::::: 500
2.Geodesics::::::::::::::: 502
3.Axial::::::::::::::::::: 505
4.Radial:::::::::::::::::: 507
5.Weyl scale:::::::::::::: 512
C.Curved spaces
1.Self-duality ::::::::::::: 517
2.De Sitter:::::::::::::::: 518
3.Cosmology :::::::::::::: 521
4.Red shift:::::::::::::::: 523
5.Schwarzschild ::::::::::: 525
6.Experiments :::::::::::: 531
7.Black holes :::::::::::::: 534
X. Supergravity
A.Superspace
1.Covariant derivatives ::::538
2.Field strengths :::::::::: 543
3.Compensators ::::::::::: 547
4.Scale gauges :::::::::::: 549
B.Actions
1.Integration :::::::::::::: 555
2.Ectoplasm:::::::::::::: 558
3.Component transform'ns 561
4.Component approach ::::562
5.Duality::::::::::::::::: 565
6.Superhiggs :::::::::::::: 569
7.No-scale:::::::::::::::: 571
C.Higher dimensions
1.Dirac spinors :::::::::::: 574
2.Wick rotation ::::::::::: 577
3.Other spins ::::::::::::: 580
4.Supersymmetry ::::::::: 582
5.Theories:::::::::::::::: 585
6.Reduction to D=4 ::::::: 588XI. Strings
A.Scattering
1.Regge theory :::::::::::: 595
2.Classical mechanics ::::: 598
3.Gauges::::::::::::::::: 601
4.Quantum mechanics ::::: 606
5.Anomaly:::::::::::::::: 609
6.Tree amplitudes ::::::::: 611
B.Symmetries
1.Massless spectrum ::::::: 618
2.Reality and orientation ::620
3.Supergravity :::::::::::: 621
4.T-duality::::::::::::::: 622
5.Dilaton::::::::::::::::: 624
6.Superdilaton :::::::::::: 626
7.Conformal eld theory ::628
8.Triality::::::::::::::::: 633
C.Lattices
1.Spacetime lattice :::::::: 637
2.Worldsheet lattice ::::::: 641
3.QCD strings :::::::::::: 643
XII. Mechanics
A.OSp(1,1j2)
1.Lightcone::::::::::::::: 649
2.Algebra::::::::::::::::: 652
3.Action:::::::::::::::::: 655
4.Spinors::::::::::::::::: 657
5.Examples::::::::::::::: 659
B.IGL(1)
1.Algebra::::::::::::::::: 664
2.Inner product ::::::::::: 665
3.Action:::::::::::::::::: 667
4.Solution::::::::::::::::: 670
5.Spinors::::::::::::::::: 673
6.Masses:::::::::::::::::: 674
7.Background elds ::::::: 675
8.Strings:::::::::::::::::: 677
9.Relation to OSp(1,1 j2)::682
C.Gauge xing
1.Antibracket ::::::::::::: 685
2.ZJBV::::::::::::::::::: 688
3.BRST:::::::::::::::::: 692
AfterMath :::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::: 696
iv
::::::::::::::::::::::::::::::: :::::::::::::::::::::::::::::::
::::::::::::::::::::::::::::::: PREFACE :::::::::::::::::::::::::::::::
Scientic method
Although there are many ne textbooks on quantum eld theory, they all have
various shortcomings. Instinct is claimed as a basis for most discussions of quantum
eld theory, though clearly this topic is too recent to aect evolution. Their subjectiv-
ity more accurately identies this as fashion :( 1 )T h e old-fashioned approach justies
itself with the instinct of intuition . However, anyone who remembers when they rst
learned quantum mechanics or special relativity knows they are counter-intuitive;quantum eld theory is the synthesis of those two topics. Thus, the intuition in thiscase is probably just habit : Such an approach is actually historical ortraditional ,
recounting the chronological development of the subject. Generally the rst half (orvolume) is devoted to quantum electrodynamics, treated in the way it was viewed in
the 1950's, while the second half tells the story of quantum chromodynamics, as it
was understood in the 1970's. Such a \dualistic" approach is n ecessarily redundant,
e.g., using canonical quantization for QED but path-integral quantization for QCD,contrary to scientic principles, which advocate applying the same \unied" methodsto all theories. While some teachers may feel more comfortable by beginning a topicthe way they rst learned it, students may wonder why the course didn't begin with
the approach that they will wind up using in the end. Topics that are unfamiliar
to the author's intuition are often labeled as \formal" (lacking substance) or even\mathematical" (devoid of physics). Recent topics are usually treated there as ad-vanced: The opposite is often true, since explanations simplify with time, as the topicis better understood. On the positive side, this approach generally presents topicswith better experimental verication.
(2) In contrast, the fashionable approach is described as being based on the in-
stinct of beauty . But this subjective beauty of artis not the instinctive beauty of
nature, and in science it is merely a consolation. Treatments based on this approachare usually found in review articles rather than textbooks, due to the shorter life ex-pectancy of the latest fashion. On the other hand, this approach has more imagination
than the traditional one, and attempts to capture the future of the subject.
A related issue in the treatment of eld theory is the relative importance of con-
cepts vs.calculations : (1) Some texts emphasize the concepts, including those which
have not proven of practical value, but were considered motivational historically (in
the traditional approach) or currently (in the artistic approach). However, many ap-proaches that were once considered at the forefront of research have faded into oblivionnot because they were proven wrong by experimental evidence or lacked conceptual
v
attractiveness, but because they were too complex for calculation, or so vague they
lacked predicitive ability. Some methods claimed total generality, which they used to
prove theorems (though sometimes without examples); but ultimately the only usefulproofs of theorems are by construction. Often a dualistic, two-volume approach isagain advocated (and frequently the author writes only one of the two volumes): Likethe traditional approach of QED volume + QCD volume, some prefer concept volume+ calculation volume. Generally, this means that gauge theory S-matrix calculationsare omitted from the conceptual eld theory course, and left for a \particle physics"course, or perhaps an \advanced eld theory" course. Unfortunately, the particlephysics course will nd the specialized techniques of gauge theory too technical tocover, while the advanced eld theory course will frighten away many students by itstitle alone.
(2) On the other hand, some authors express a desire to introduce Feynman graphs
as quickly as possible: This suggests a lack of appreciation of eld theory outside ofdiagrammatics. Many essential aspects of eld theory (such as symmetry breakingand the Higgs eect) can be seen only from the action, and its analysis also leads to
better methods of applying perturbation theory than those obtained from a xed set
of rules. Also, functional equations are often simpler than pictorial ones, especiallywhen they are nonlinear in the elds. The result of over-emphasizing the calculationsis a cookbook, of the kind familiar from some lower-division undergraduate courses
intended for physics majors but designed for engineers.
The best explanation of a theory is the one that ts the principles of scientic
method : simplicity, generality, and experimental verication. In this text we thus
take a more economical orpragmatic approach, with methods based on eciency
and power. Unattractiveness or counter-intuitiveness of such methods become ad-vantages, because they force one to accept new and better ways of thinking aboutthe subject: The eciency of the method directs one to the underlying idea. Forexample, although some consider Einstein's original explanation of special relativityin terms of relativistic trains and Lorentz transformations with square roots as be-ing more physical, the concept of Minkowski space gave a much simpler explanationand deeper understanding that proved more useful and led to generalization. Manytheories have \miraculous cancelations" when traditional methods are used, which
led to new methods (background eld gauge, supergraphs, spacecone, etc.) that not
only incorporate the cancelations automatically (so that the \zeros" need not be cal-culated), but are built on the principles that explain them. We place an emphasison such new concepts, as well as the calculational methods that allow them to becompared with nature. It is important not to neglect one for the sake of the other,articial and misleading to try to separate them.
vi
As a result, many of our explanations of the standard topics are new to textbooks,
and some are completely new: For example, (1) we derive the Foldy-Wouthuysen
transformation by dimensional reduction from an analogous one for the massless case(subsections IIB3,5). (2) We derive the Feynman rules in terms of background eldsrather than sources (subsection VC1); this avoids the need for amputation of exter-nal lines for S-matrices or eective actions, and is more useful for background-eldgauges. (3) We obtain the nonrelativistic QED eective action, used in modern treat-ments of the Lamb shift (because it makes perturbation easier than the older Bethe-Salpeter methods), by eld redenition of the relativistic eective action (subsectionVIIIB6), rather than tting parameters by comparing Feynman diagrams from therelativistic and nonrelativistic actions. (In general, manipulations in the action areeasier than in diagrams.) (4) We present a somewhat new method for solving for
the curvature in general relativity that is slightly easier than all previous methods
(subsections IXA2,C5). There are also some completely new topics, like: (1) the anti-Gervais-Neveu gauge, where spin in U(N) Yang-Mills is treated in almost the sameway as internal symmetry | with Chan-Paton factors (subsection VIB4); (2) thesuperspacecone gauge, the simplest gauge for QCD (subsection VIB7); and (3) a new\(almost-)rst-order" superspace action for supergravity, analogous to the one forsuper Yang-Mills (subsection XB1).
We try to give the simplest possible calculational tools, not only for the above
reasons, but also so group theory (internal and spacetime) and integrals can be per-formed with the least eort and memory. (Some traditionalists may claim that theold methods are easy enough, but their arguments are less convincing when the orderof perturbation is increased. Even computer calculations are more ecient when leftas a last resort.) We give examples of (and excercises on) these methods, but notexhaustively. We also include more recent topics (or those more recently appreciatedin the particle physics community) that might be deemed non-introductory, but are
commonly used, and are simple and important enough to include at the earliest level.
For example, the related topics of (unitary) lightcone gauge, twistors, and spinorhelicity are absent from all eld theory texts, and as a result no such text performsthe calculation of as basic a diagram as the 4-gluon tree amplitude. Another missingtopic is the relation of QCD to strings through the random worldsheet lattice andlarge-color (1/N) expansion, which is the only known method that might quantita-tively describe its high-energy nonperturbative behavior (bound states of arbitrarilylarge mass).
This text is meant to cover all the eld theory every high energy theorist should
know, but not all that any particular theorist might need to know. It is not meant asan introduction to research, but as a preliminary to such courses: We try to ll in the
vii
cracks that often lie between standard eld theory courses and advanced specialized
courses. For example, we have some discussion of string theory, but it is more oriented
toward the strong interactions, where it has some experimental justication, ratherthan quantum gravity and unication, where its usefulness is still under investigation.We do not mention statistical mechanics, although many of the eld theory methodswe discuss are useful there. Also, we do not discuss any experimental results in detail;phenomenology and analysis of experiments deserve their own text. We give and applythe methods of calculation and discuss the qualitative features of the results, but donot make a numerical comparison to nature: We concentrate more on the \forest"than the \trees".
Unfortunately, our discussions of the (somewhat related) topics of infrared-diver-
gence cancelation, Lamb shift, and the parton model are sketchy, due to our inabilityto give fully satisfying treatments | but maybe in a Second Edition?
Unlike all previous texts on quantum eld theory, this one is available for free over
the Internet (as usual, from xxx.lanl.gov and its mirrors), and may be periodically
updated. Errata, additions, and other changes will be posted on my web page at
http://insti.physics.sunysb.edu/~siegel/plan.html until enough are accumulated fora new edition. Electronic distribution not only makes it more available, but easierto update to new editions, and you don't have to bother to lug your copy with youwhen you go anywhere (like home) that has a computer (so you can access it froma cartridge/diskette or the internet). It also oers the option of reading it on thecomputer, which saves space and trees (and the gures are nicer in color). ThePDF version allows searches that are more general than the Index, and includes an\outline" window with clickable \bookmarks" that is more convenient than the Tableof Contents, as well as the usual web links to xxx.lanl.gov (and a couple of otherplaces).
Highlights
This text also diers from others in most of the following ways: (1) We place a
greater emphasis on mechanics in introducing some of the more elementary physical
concepts of eld theory: (a) Some basic ideas, such as antiparticles, can be more sim-ply understood already with classical mechanics. (b) Some interactions can also betreated through rst-quantization: This is sucient for evaluating certain tree andone-loop graphs as particles in external elds. Also, Schwinger parameters can beunderstood from rst-quantization: They are useful for performing momentum inte-grals (reducing them to Gaussians), studying the high-energy behavior of Feynmangraphs, and nding their singularities in a way that exposes their classical mechanics
viii
interpretation. (c) Quantum mechanics is very similar to free classical eld the-
ory, by the usual \semiclassical" correspondence between particles (mechanics) and
waves (elds). They use the same wave equations, since the mechanics Hamiltonianor Becchi-Rouet-Stora-Tyutin operator is the kinetic operator of the correspondingclassical eld theory, so the free theories are equivalent. In particular, (relativistic)quantum mechanical BRST provides a simple explanation of the o-shell degreesof freedom of general gauge theories, and introduces concepts useful in string theory.As in the nonrelativistic case, this treatment starts directly with quantum mechanics,rather than by (rst-)quantization of a classical mechanical system. Since supersym-metry and strings are so important in present theoretical research, it is useful to havea text that includes the eld theory concepts that are prerequisites to a course onthese topics. (For the same reason, and because it can be treated so similarly to
Yang-Mills, we also discuss general relativity.)
(2) We also emphasize conformal invariance . Although a badly broken sym-
metry, the fact that it is larger than Poincar e invariance makes it useful in many
ways: (a) General classical theories can be described most simply by rst analyzing
conformal theories, and then introducing mass scales by various techniques. This
is particularly useful for the general analysis of free theories, and for constructingactions for supergravity theories. (b) Quantum theories that are well-dened withinperturbation theory are conformal (\scaling") at high energies. (A possible excep-tion is string theories, but the supposedly well understood string theories that arenite perturbatively have been discovered to be hard-to-quantize membranes in dis-guise nonperturbatively.) This makes methods based on conformal invariance usefulfor nding classical solutions, as well as studying the high-energy behavior of thequantum theory, and simplifying the calculation of amplitudes. (c) Theories whoseconformal invariance is not (further) broken by quantum corrections avoid certainproblems at the nonperturbative level. Thus conformal theories ultimately may be
required for an unambiguous description of high-energy physics.
( 3 )W em a k ee x t e n s i v eu s eo f two-component (chiral) spinors , which are ubiqui-
tous in particle physics: (a) The method of twistors (more recently dubbed \spinorhelicity") greatly simplies the Lorentz algebra in Feynman diagrams for massless(or high-energy) particles with spin, and it's now a standard in QCD. (Twistors
are also related to conformal invariance and self-duality.) On the other hand, most
texts still struggle with 4-component Dirac (rather than 2-component Weyl) spinornotation, which requires gamma-matrix and Fierz identities, when discussing QCDcalculations. (b) Chirality and duality are important concepts in all the interactions:Two-component spinors were rst found useful for weak interactions in the days of4-fermion interactions. Chiral symmetry in strong interactions has been important
ix
since the early days of pion physics; the related topic of instantons (self-dual solutions)
is simplied by two-component notation, and general self-dual solutions are expressed
in terms of twistors. Duality is simplest in two-component spinor notation, even whenapplied to just the electromagnetic eld. (c) Supersymmetry still has no convincingexperimental verication (at least not at the moment I'm typing this), but its the-oretical properties promise to solve many of the fundamental problems of quantumeld theory. It is an element of most of the proposed generalizations of the StandardModel. Chiral symmetry is built into supersymmetry, making two-component spinorsunavoidable.
(4) The topics are ordered in a more pedagogical manner: (a) Abelian and non-
abelian gauge theories are treated together using modern techniques. (Classical grav-ity is treated with the same methods.) (b) Classical Yang-Mills theory is discussed be-fore any quantum eld theory. This allows much of the physics, such as the StandardModel (which may appeal to a wider audience), of which Yang-Mills is an essentialpart, to be introduced earlier. In particular, symmetries and mass generation in theStandard Model appear already at the classical level, and can be seen more easily from
the action (classically) or eective action (quantum) than from diagrams. (c) Only
the method of path integrals is used for second-quantization. Canonical quantizationis more cumbersome and hides Lorentz invariance, as has been emphasized even byFeynman when he introduced his diagrams. We thus avoid such spurious concepts asthe \Dirac sea", which supposedly explains positrons while being totally inapplica-ble to bosons. However, for quantum physics of general systems or single particles,operator methods are more powerful than any type of rst-quantization of a classicalsystem, and path integrals are mainly of pedagogical interest. We therefore \review"quantum physics rst, discussing various properties (path integrals, S-matrices, uni-tarity, BRST, etc.) in a general (but simpler) framework, so that these properties neednot be rederived for the special case of quantum eld theory, for which path-integral
methods are then sucient as well as preferable.
(5)Gauge xing is discussed in a way more general and ecient than older meth-
ods: (a) The best gauge for studying unitarity is the (unitary) lightcone gauge. Thisrarely appears in eld theory texts, or is treated only half way, missing the importantexplicit elimination of all unphysical degrees of freedom. (b) Ghosts are introduced
by BRST symmetry, which proves unitarity by showing equivalence of convenient
and manifestly covariant gauges to the manifestly unitary lightcone gauge. It can beapplied directly to the classical action, avoiding the explicit use of functional determi-nants of the older Faddeev-Popov method. It also allows direct introduction of moregeneral gauges (again at the classical level) through the use of Nakanishi-Lautrupelds (which are omitted in older treatments of BRST), rather than the functional
x
averaging over Landau gauges required by the Faddeev-Popov method. (c) For non-
abelian gauge theories the background eld gauge is a must. It makes the eective
action gauge invariant, so Slavnov-Taylor identities need not be applied to it. Beta
functions can be found from just propagator corrections.
(6)Dimensional regularization is used exclusively (with the exception of one-loop
axial anomaly calculations): (a) It is the only one that preserves all possible sym-metries, as well as being the only one practical enough for higher-loop calculations.
(b) We also use it exclusively for infrared regularization, allowing all divergences to
be regularized with a single regulator (in contrast, e.g., to the three regulators used
for the standard treatment of Lamb shift). (c) It is good not only for regularization,
but renormalization (\dimensional renormalization"). For example, the renormaliza-
tion group is most simply described using dimensional regularization methods. More
importantly, renormalization itself is performed most simply by a minimal prescrip-
tion implied by dimensional regularization. Unfortunately, many books, even amongthose that use dimensional regularization, apply more complicated renormalization
procedures that require additional, nite renormalizations as prescribed by Slavnov-
Taylor identities. This is a needless duplication of eort that ignores the manifestgauge invariance whose preservation led to the choice of dimensional regularization in
the rst place. By using dimensional renormalization, gauge theories are as easy to
treat as scalar theories: BRST does not have to be applied to amplitudes explicitly,
since the dimensional regularization and renormalization procedure preserves it.
(7) Perhaps the most fundamental omission in most eld theory texts is the
expansion of QCD in the inverse of the number of colors : (a) It provides a gauge-
invariant organization of graphs into subsets, allowing simplications of calculations
at intermediate stages, and is commonly used in QCD today. (b) It is useful as a
perturbation expansion, whose experimental basis is the Okubo-Zweig-Iizuka rule.(c) At the nonperturbative level, it leads to a resummation of diagrams in a way that
can be associated with strings, suggesting an explanation of connement.
Notes for instructors
This text is intended for reference and as the basis for a full-year course on rela-
tivistic quantum eld theory for second-year graduate students. A preliminary version
of the rst two parts was used for a one-year course I taught at Stony Brook. Thechapter on gravity and pieces of early chapters cover a one-semester graduate rela-
tivity course I gave several times here | I used most of the following: IA, IB3, IC2,
IIA, IIIA-C5, VIB1, IX, XIA2-4, XIB4-5. The prerequisites (for the quantum eld
xi
theory course) are the usual rst-year courses in classical mechanics, classical elec-
trodynamics, and quantum mechanics. For example, the student should be familiarwith Hamiltonians and Lagrangians, Lorentz transformations for particles and elec-
tromagnetism, Green functions for wave equations, SU(2) and spin, and Hilbert space.
Unfortunately, I nd that many second-year graduate students (especially many who
got their undergraduate training in the USA) still have only an undergraduate level
of understanding of the prerequisite topics, lacking a working knowledge of actionprinciples, commutators, creation and annihilation operators, etc. While most such
topics are brie
y reviewed here, they should be learned elsewhere.
There is far more material here than can be covered comfortably in one year,
mostly because of included material that should be covered earlier, but rarely is.
Ideally, a modern curriculum for eld theory students should include: (1) courses on
classical mechanics, nonrelativistic quantum mechanics, and classical electrodynamicsin the rst semester of graduate study, without overly reviewing aspects that should
have been covered in undergraduate study (and in particular avoiding the enormous
overlap of the last two subjects due to both covering primarily the solution of wave
equations); (2) in the second semester, statistical mechanics as the sequel to clas-
sical, relativistic quantum mechanics as the sequel to nonrelativistic, and classicalnonabelian eld theory (Yang-Mills and gravity) as the sequel to classical electrody-
namics; (3) in the second year, one year of quantum eld theory, and at least one
semester on \phenomenology" (model-building and direct comparison with observa-
tions, including those for general relativity and cosmology); and (4) in the third year,
more specialized courses, such as a semester on supersymmetry and strings. Unfor-
tunately, in practice little of relativistic quantum mechanics and classical eld theory
(other than electromagnetism) will have been covered previously, which means theywill comprise half of the \quantum" eld theory course, while the true quantum eld
theory will be squeezed into the last half.
One way to cut the material to t a one-year course is to omit Part Three, which
can be left for a third semester on \advanced quantum eld theory"; then the rstsemester (Part One) is classical while the second (Part Two) is quantum. Further-
more, the ordering of the chapters is somewhat
exible: The \
ow" is indicated by
the following \3D" plot:
8
<
:lower spin
&.
higher spinclassical ! quantum
symmetry elds quantize loop
Bose I III V VII
# IX XI
XX I I
Fermi II IV VI VIII
xii
where the 3 dimensions are spin (\ j"), quantization (\ h"), and statistics (\ s"): The
three independent
ows are down the page, to the right, and into the page. (The thirddimension has been represented as perpendicular to the page, with \higher spin" insmaller type to indicate perspective, for legibility.) To present these chapters in the1 dimension of time we have classied them as jhs, but other orderings are possible:
jhs: I II III IV V VI VII VIII IX X XI XII
jsh: I III V VII II IV VI VIII IX XI X XII
hjs: I II III IV IX X V VI XI XII VII VIII
hsj: I II III IX IV X V XI VI XII VII VIII
sjh: I III V VII IX XI II IV VI VIII X XII
shj: I III IX V XI VII II IV X VI XII VIII
(However, the spinor notation of II is used for discussing instantons in III, so some
rearrangement would be required, except in the jhs,hjs,a n dhsjcases.) For exam-
ple, the rst half of the course can cover all of the classical, and the second quantum,dividing Part Three between them ( hjsor hsj). Another alternative ( jsh)i sao n e -
semester course on quantum eld theory, followed by a semester on the Standard
Model, and nishing with supergravity and strings. Although some of these (espe-
cially the rst two) allow division of the course into one-semester courses, this shouldnot be used as an excuse to treat such courses as complete: Any particle physicsstudent who was content to sit through another entire year of quantum mechanics in
graduate school should be prepared to take at least a year of eld theory.
Notes for students
Field theory is a hard course. (If you don't think so, name me a harder one at this
level.) But you knew as an undergraduate that physics was a hard major. Students
who plan to do research in eld theory will nd the topic challenging; those with less
enthusiasm for the topic may nd it overwhelming. The main dierence between eldtheory and lower courses is that it is not set in stone: There is much more variation instyle and content among eld theory courses than, e.g., quantum mechanics courses,
since quantum mechanics (to the extent taught in courses) was pretty much nished
in the 1920's, while eld theory is still an active research topic, even though it has hadmany experimentally conrmed results since the 1940's. As a result, a eld theorycourse has the
avor of research: There is no set of mathematically rigorous rules to
solve any problem. Answers are not nal, and should be treated as questions: One
should not be satised with the solution of a problem, but consider it as a rst steptoward generalization. The student should not expect to capture all the details of
xiii
eld theory the rst time through, since many of them are not yet fully understood
by people who work in the area. (It is far more likely that instead you will discover
details that you missed in earlier courses.) And one reminder: The only reason for
lectures (including seminars and conferences) is for the attendees to ask questions
(and not just in private), and there are no stupid questions (except for the infamous
\How many questions are on the exam?"). Only half of teaching is the responsibility
of the instructor.
Outline
The preceding Table of Contents lists the three parts of the text: Symmetry,
Quanta, and Higher Spin. Each part is divided into four chapters, each of which has
three sections, divided further into subsections. Each section is followed by references
to reviews and original papers. Excercises appear throughout the text, immediately
following the items they test: This purposely disrupts the
ow of the text, forcing
the reader to stop and think about what he has just learned. These excercises areinteresting in their own right, and not just examples or memory tests. This is not a
crime for homeworks and exams, which at least by graduate school should be about
more than just grades.
The rst part of the text focuses on symmetry: The Poincar e group is special
relativity, and is sucient to nd all free equations of motion for particles and elds.
Internal symmetries include both global ones, used for classifying particles, and local
ones, which describe the interactions of elds.
The rst chapter discusses global symmetry, both spacetime and internal. Space-
time symmetries covered include not only Poincar e but also Galilean (i.e., nonrel-
ativistic, used as an introduction) and conformal (broken in nature, but still very
useful). Parity, time reversal, and even charge conjugation are described simply in
terms of classical mechanics. Lightcone bases are introduced. Some general proper-
ties of Lie algebras are summarized, including fermions and anticommuting numbers.
Classical groups are described using tensor methods and index notation, including
Young tableaux. Dirac gamma matrices appear as coordinates for orthogonal groups.
The color and
avor symmetries of the particles of the Standard Model, and observed
light hadrons, are given as examples.
The second chapter extends the rst chapter's treatment of spacetime symmetry
to include spin. The methods introduced in this chapter are the most ecient ones for
handling Lorentz indices in QCD (or even pure Yang-Mills theory). Two-component
spinor notation is introduced by the rotation group in three (space) dimensions: Ten-
sor notation avoids Clebsch-Gordan-Wigner coecients. The study of the simple
xiv
algebraic properties of 2 2 matrices is extended straightforwardly from three dimen-
sions to four, and applied to simple examples in free eld theory. The conformal
group gives an easy and unied way to nd massless free eld equations in general,and dimensional reduction does the same for masses. As an application, we discussthe Foldy-Wouthuysen transformation (and its massless analog) for arbitrary spin,with minimal electromagnetic coupling to spin 1/2 as an example. (The case of non-minimal coupling will be useful in chapter VIII for the Lamb shift.) Twistors, relatedto conformal invariance and self-duality, yield a convenient and covariant method tosolve the massless equations, and also explain helicity. The chapter concludes with adiscussion of the general properties of supersymmetry and its representations, usingsuperspace and supertwistors.
Local symmetries are covered in the third chapter. It begins with a discussion
of the action principle, including fermions and constrained systems, applied to thesimple free examples given earlier (spins 1/2 and 1). The concepts of gauge invarianceand gauge xing are introduced through the simple case of the (spinless) relativisticparticle, whose classical mechanics will prove useful later in understanding severalfeatures of Feynman diagrams. As an example, classical pair creation and annihilation
is examined. Finally, pure Yang-Mills theory is analyzed, including some solutions to
the classical eld equations. Twistors are used to study self-duality and instantons.The lightcone gauge is used as a unitary gauge, and in combination with self-duality.
Gauge symmetry is coupled to lower spins in chapter four. Following an intro-
duction to spontaneous symmetry breakdown of global symmetries, nonlinear sigma
models are considered as low-energy theories, particularly in the study of chiral sym-metry, and gauge invariance is used in their general construction. The use of scalarsto generate mass for vectors is illustrated rst by the free case of the St uckelberg
model, generalized to the Higgs model, and applied to the Standard Model. Familiesand Grand Unied Theories are also described, as well as the basics for construct-ing actions for supersymmetric theories in superspace (including a brief discussion ofextended supersymmetry).
The second part of the text covers the quantum aspects of eld theory, as revealed
through perturbation theory. Although some have conjectured that nonperturbativeapproaches might solve the renormalization diculties found in perturbation, all ev-idence indicates these features survive in the complete theory.
Chapter ve focuses on the method of quantization of classical theories based
on path integrals. The chapter begins by considering various properties of quantumphysics in a general context | relation to canonical quantization, Wick rotation,unitarity, and causality | so that these items need not be repeated in the more
xv
specialized and complicated cases of eld theory. Path integrals are then applied
to classical mechanics to explain the St uckelberg-Feynman propagator. Path integra-
tion of eld theory produces a generating functional of background elds for Feynmandiagrams, as well as its connected and one-particle-irreducible parts. We use back-grounds elds instead of sources exclusively: All uses of Feynman diagrams involveeither the S-matrix or the eective action, both of which require the removal of ex-ternal propagators, which is equivalent to replacing sources with elds. The classical(tree) graphs are shown to give the perturbative solution to the classical eld equa-tions. Also described are the properties of the classical action needed for unitarity, thediagrammatic translation of unitarity and causality, the denition of cross sections,the relation of Landau singularities to classical mechanics, and the use of the quarkline rules for dealing with group theory in graphs easily. Throughout the chapter
simple examples are given from scalar theories.
Complications that arise from quantization of gauge theories are described in the
sixth chapter. BRST symmetry is the easiest way to gauge x, and makes unitarityclear by relating general gauges to unitary gauges. Again a general discussion is
given in the framework of quantum physics and canonical quantization, including
the relation of Hamiltonian and Lagrangian approaches, so that later eld theorycan be addressed covariantly with path integrals. Various gauges are considered forYang-Mills elds: radial, Lorentz, Landau, Fermi-Feynman, unitary, renormalizable,Gervais-Neveu, anti-Gervais-Neveu, super Gervais-Neveu, spacecone, superspacecone,background-eld, and Nielsen-Kallosh. The spacecone gauge is used as the simplestmethod to calculate graphs in massless theories, with examples given from (massless)QCD, including the 4-gluon and 5-gluon tree amplitudes. The Fermi-Feynman gaugeis used to calculate all the 4-point tree amplitudes of QED, and their dierentialcross sections. The supergraph rules are derived for supersymmetric theories, andthe locality of the eective action in the anticommuting coodinates is shown to imply
nonrenormalization theorems.
General features of higher orders in perturbation theory due to momentum inte-
gration are examined in chapter seven. Renormalization is explained (but not proven),and dimensional regularization is applied. Tadpole integrals are used to explain di-mensional transmutation through the example of the eective potential. The running
of couplings with energy is shown through the evaluation of one-loop massless and
massive propagator corrections. Some simple overlapping two-loop divergences areused to illustrate renormalization of subdivergences. The renormalization group equa-tions are introduced via dimensional regularization. The problems solved perturba-tively by renormalization are shown to reappear upon resummation of the expansion.Instantons and IR and UV renormalons are analyzed through a Borel transform in
xvi
the coupling, and the resultant ambiguities are related to nonperturbative vacuum
values of composite elds. The expansion in the inverse of the number of colors is an
approach to this problem, related to string theory, that also has uses at nite ordersof perturbation.
Chapter eight applies these methods to gauge theories. Propagator corrections in
QED and QCD are used to analyze the one-loop conformal anomaly and its relation to
asymptotic freedom. Finite N=1 supersymmetric theories are considered as a solutionto the renormalon problem. The Schwinger model is given as another example fromtwo dimensions, illustrating interesting features at one loop such as bound states,bosonization, and the axial anomaly. The axial anomaly is then evaluated moregenerally, and considered in four dimensions in relation to constraints on electroweakmodels and electromagnetic pion decay. The nonrelativistic form of the eectiveaction useful for nding the Lamb shift (including the anomalous magnetic moment) isgiven as an example of vertex corrections. Finally, the production of hadrons throughelectron-positron annihilation, deep inelastic scattering, and Drell-Yan scattering arebrie
y described as applications of perturbative QCD.
Part Three treats general spins, particularly spin 2, which ultimately must be
included in any complete theory of nature. Such spins are observed experimentallyfor bound states, but may be required also as fundamental elds.
Gravity is described through the theory of general relativity in chapter nine.
The classic experimental tests are described, including cosmology. The treatmentused is closely related to that applied to Yang-Mills theory, and diers from thatof most texts on gravity: (1) We emphasize the action for deriving eld equationsfor gravity (and matter), rather than treating it as an afterthought. (2) We makeuse of local (Weyl) scale invariance for cosmological solutions, gauge xing, eldredenitions, and studying conformal properties. In particular, other texts neglect the
(unphysical) dilaton, which is crucial in such treatments (especially for generalization
to supergravity and strings). (3) While most gravity texts leave spinors till the end,and treat them brie
y, our discussion of gravity is based on methods that can beapplied directly to spinors, and therefore to supergravity and superstrings. (4) Ourmethod of calculating curvatures for purposes of solving the classical eld equationsis somewhat new, but probably the simplest, and is directly related to the simplestmethods for super Yang-Mills theory and supergravity.
The approach of the previous chapter is generalized straightforwardly to super-
gravity in chapter ten. Actions with matter are constructed, and are analyzed insuperspace, and in terms of component elds using component expansion and sep-aration of superconformal breaking (\compensator") terms. The spin-3/2 particle
xvii
is given mass by the superhiggs eect, and no-scale supergravity provides a model
whereby the would-be resultant cosmological constant vanishes naturally. Extendedsupergravity is analyzed through general properties of extended supersymmetry andby reduction from higher dimensions.
Strings are proposed in chapter eleven as an approach to studying the most im-
portant yet least understood property of QCD: connement. Other methods havebeen proposed to study this phenomenon, but none have achieved explicit results formore than low hadron energy, which no more exhibits connement than chemistrydisproves the existence of free nuclei. Known string theories are not suitable for de-scribing hadrons quantitatively, but are useful models of observed properties, suchas Regge behavior. The classical and quantum theory of the simplest such model is
analyzed, and qualitative features expected of general theories are described. The dis-
cretization of the worldsheet of the string into a sum of Feynman diagrams is shownto exhibit features relevant to a string theory of hadrons.
The nal chapter gives a general derivation of free actions for any gauge theory,
based on adding equal numbers of commuting and anticommuting ghost dimensionsto the lightcone formulation of the Poincar e group. The usual ghost elds appear as
components of the gauge elds in anticommuting directions, as do necessary auxiliaryelds like the determinant of the metric tensor in gravity. Gauge xing to the Fermi-Feynman gauge is automatic. The \antields" and \antibracket" of the Zinn-Justin-Batalin-Vilkovisky method appear naturally from the anticommuting coordinate that
is the rst-quantized ghost of the Klein-Gordon equation.
Following the body of the text (and preceding the Index) is the AfterMath, con-
taining conventions and some of the more important equations.
Acknowledgments
I thank everyone with whom I have discussed eld theory, especially Gordon
Chalmers, Marc Grisaru, Marcelo Leite, Martin Ro cek, Jack Smith, George Sterman,
and Peter van Nieuwenhuizen. More generally, I thank the human race, without
whom this work would have been neither possible nor necessary.
December 20, 1999
xviii
::::::::::::: :::::::::::::
::::::::::::: SOME FIELD THEORY TEXTS :::::::::::::
Comprehensive, traditional
Complete texts; use canonical quantization for QED, then path integrals for QCD
1S. Weinberg, The quantum theory of elds , 3 v. (Cambridge University, 1995,6,9?)
609+489+c.500 pp.:First volume just QED; second volume contains many interesting topics; thirdvolume supersymmetry. By one of the developers of the Standard Model.
2M.E. Peskin and D.V. Schroeder, An introduction to quantum eld theory
(Addison-Wesley, 1995) 842 pp.:
Many applications; style similar to Bjorken and Drell.
3M. Kaku, Quantum eld theory: a modern introduction (Oxford University, 1993)
785 pp.:Includes introduction to supergravity and superstrings.
4C. Itzykson and J.-B. Zuber, Quantum eld theory (McGraw-Hill, 1980) 705 pp.
(but with lots of
small print ):
Emphasis on QED.
Somewhat specialized
Basics, plus thorough treatment of an advanced topic
5J. Zinn-Justin, Quantum eld theory and critical phenomena , 3rd ed. (Clarendon,
1996) 1008 pp.:
First 1/2 is basic text, with interesting treatments of many topics, but no S-matrix
examples or discussion of cross sections; second 1/2 is statistical mechanics.
6G. Sterman, An introduction to quantum eld theory (Cambridge University,
1993) 572 pp.:First 3/4 can be used as basic text, including S-matrix examples; last 1/4 hasextensive treatment of perturbative QCD, emphasizing factorization.
Basic; few S-matrix examples
Should be supplemented with a \QED/particle physics text"
7L.H. Ryder, Quantum eld theory , 2nd ed. (Cambridge University, 1996) 487 pp.:
Includes introduction to supersymmetry.
8D. Bailin and A. Love, Introduction to gauge eld theory ,2 n de d .( I n s t i t u t eo f
Physics, 1993) 364 pp.:
All the fundamentals.
9P. Ramond, Field theory: a modern primer , 2nd ed. (Addison-Wesley, 1989)
329 pp.:Short text on QCD: no weak interactions or Higgs.
xix
Classics
Older but unconventional treatments from their originators; no Yang-Mills or Higgs
10N.N. Bogoliubov and D.V. Shirkov, Introduction to the theory of quantized elds ,
3rd ed. (Wiley, 1980) 620 pp.:Ahead of its time (1st English ed. 1959); early treatments of path integrals, causal-ity, background elds, and renormalization of all eld theories (not just QED).
11R.P. Feynman, Quantum electrodynamics: a lecture note and reprint volume (Ben-
jamin, 1961) 198 pp.:
Original treatment of quantum eld theory as we know it today, but from me-
chanics; includes reprints of original articles (1949).
QED/particle physics
Numerous Feynman diagram calculations; no Yang-Mills or Higgs
12B. de Wit and J. Smith, Field theory in particle physics , v. 1 (Elsevier Science,
1986) 490 pp.:Oriented toward experimentalists (but wait till v. 2, due any day now...).
13A.I. Akhiezer and V.B. Berestetskii, Quantum electrodynamics (Wiley, 1965) 868
pp.:
Extensive examples of lower-order QED diagrams.
Advanced topics
For further reading; including brief reviews of some standard topics
14Theoretical Advanced Study Institute in Elementary Particle Physics (TASI) pro-
ceedings, University of Colorado, Boulder, CO (World Scientic):Annual collection of summer school lectures on recent research topics.
15W. Siegel, Introduction to string eld theory (World Scientic, 1988) 244 pp.:
Reviews lightcone, BRST, gravity, rst-quantization, spinors, twistors, strings;
besides, I like the author.
16S.J. Gates, Jr., M.T. Grisaru, M. Ro cek, and W. Siegel, Superspace: or one thou-
sand and one lessons in supersymmetry (Benjamin/Cummings, 1983) 548 pp.:
Covers supersymmetry, spinor notation, lightcone, St uckelberg elds, gravity,
Weyl scale, gauge xing, background-eld method, regularization, and anoma-lies; same author as previous, plus three other guys whose names sound familiar.
May soon be available for free where you found this book.
1
PART ONE: SYMMETRY
The rst four chapters present a one-semester course on \classical eld theory".
Perhaps a more accurate description would be \everything you should know before
learning quantum eld theory". This is basically a study of global and local symme-
tries: Classical dynamics represents only a certain limit of quantum dynamics, and not
the one usually emphasized, but most of the symmetries of classical physics survive
quantization. The phenomenon of symmetry breaking, and the related mechanisms ofmass generation, can also be seen at the classical level. In perturbative quantum eld
theory, classical eld theory is simply the leading term in the perturbation expansion.
Continuous symmetry is one of the most fundamental and important concepts
of physics. In the framework of an action principle (which is required in quantumphysics), it is equivalent to conservation laws, which have been a cornerstone of physics
since Newton. From a practical viewpoint, it simplies calculations by relating dier-
ent solutions to equations of motion, and allowing these equations to be written moreconcisely by treating independent degrees of freedom as a single entity. In particular,
local (\gauge") symmetries, which allow independent transformations at each coordi-
nate point, are basic to all the fundamental interactions: All the fundamental forcesare mediated by particles described by Yang-Mills theory and its generalizations.
Symmetries are the result of a redundant, but useful, description of a theory.
(Note that here we refer to symmetries of a theory, not of a solution to the theory.) For
example, translation invariance says that only dierences in position are measurable,not absolute position: We can't measure the position of the \origin". There are two
ways to deal with this: (1) Choose an origin; i.e., make a \choice of coordinates".
For example, place an object at the origin; i.e., choose the position of an object ata certain time to be the origin. (2) Work only in terms of dierences of coordinates,
which are \translationally invariant". Although the latter choice is more physical,
the former is usually more convenient: The use of redundant variables, together with
symmetry, often gives a simpler description of a theory. Another example is quantum
mechanics, where the arbitrariness of the phase of the wave function can be considereda symmetry: Although quantum mechanics can be reformulated in terms of phase-
invariant probabilities, currents, or density matrices instead of wave functions, and
this can be useful for some purposes of exposing physical properties, formulating andsolving the Schr odinger equation is simpler in terms of the wave function. The same
applies to \local" symmetries, where there is an independent symmetry at each point
of space and time: For example, quarks and gluons have a local \color" symmetry,
2
and are not (yet) observed independently in nature, but are simpler objects in terms
of which to describe strong interactions than the observed hadrons (protons, neutrons,
etc.), which are described by color-invariant products of quark/gluon wave functions,
in the same way that probabilities are phase-invariant products of wave functions.
(Note that in quantum mechanics there is a subtle distinction between observed andobserver that can obscure this symmetry if the observer is not invariant under it. This
can always be avoided by choosing to dene the observer as invariant: For example,
the detection apparatus can be included as part of the quantum mechanical system,while the observer can be dened as some \remote" recorder, who may be abstracted
as even being translationally invariant. In practice we are less precise, and abstract
even the detection apparatus to be invariant: For example, we describe the scatteringof particles in terms of the coordinates of only the particles, and deal with the origin
problem as above in terms of just those coordinates.)
Note that \global" (time-, and usually space-independent) symmetries can elim-
inate a variable, but not its time derivative. For example, translation invarianceallows us to x (i.e., eliminate) the position of the center of mass of a system at some
initial time, but not its time derivative, which is just the total momentum, whose
conservation is a consequence of that same symmetry. A local symmetry, being timedependent, may allow the elimination of a variable at all times: The existence of this
possibility depends on the dynamics, and will be discussed later.
Of particular intrerest are ways in which symmetries can be made manifest. Fre-
quently in the literature \manifest" is used vacuously; a \manifest symmetry" is anobvious one: If you know the group, the representation under consideration doesn't
need to be stated, but can be seen from just the notation. (In fact, one of the main
uses of index notation is just to manifest the symmetry.) Formulations where globaland local symmetries are manifest simplify calculations and their results, as well as
clarifying their meaning.
One of the main uses of manifest symmetry is rarely needing to explicitly perform
a specic symmetry transformation. For example, one might need to examine a rela-tivistic problem in dierent Lorentz frames. Rather than starting with a description
of the problem in one frame, and then explicitly transforming to another, it is much
simpler to start with a manifestly covariant description, make one choice of frame,then make another choice of frame. One then never uses the messy square roots of the
familiar Lorentz contraction factors (although they may appear at the end from kine-
matic constraints). A more extreme example is the corresponding situation for local
A. COORDINATES 3
symmetries, where such transformations are intractable in general, and one always
starts with the manifestly covariant form.
I. GLOBAL
In the rst chapter we study symmetry in general, concentrating primarily on
spacetime symmetries, but also discussing general properties that will have other
applications in the following chapter.
::::::::::::::::::::::: :::::::::::::::::::::::
::::::::::::::::::::::: A. COORDINATES :::::::::::::::::::::::
In this section we discuss the Poincar e (and conformal) group as coordinate trans-
formations. This is the simplest way to represent it on the physical world. In later
sections we nd general representations by adding spin.
1. Nonrelativity
We begin by reviewing some general properties of symmetries, including as an
example the symmetry group of nonrelativstic physics. In the Hamiltonian approach
to mechanics, both symmetries and dynamics can be expressed conveniently in terms
of a \bracket": the Poisson bracket for classical mechanics, the commutator for quan-tum mechanics. In this formulation, the fundamental variables (operators) are some
set of coordinates and their canonically conjugate momenta, as functions of time.
The (Heisenberg) operator approach to quantum mechanics then is related to classi-cal mechanics by identifying the semiclassical limit of the commutator as the Poisson
bracket: For any functions AandBofpandq, the quantum mechanical commutator
is
AB BA= ih@A
@pm@B
@qm @B
@pm@A
@qm
+O(h2)
In other words, the true classical limit of AB BAis zero, since classically functions
commute; thus the semiclassical limit is dened by
lim
h!01
h(AB BA)
(which is really a derivative with respect to h). We therefore dene the bracket for
the two cases by
[A;B]8
><
>: i@A
@pm@B
@qm @B
@pm@A
@qm
semiclassically
AB BA quantum mechanically
4I . G L O B A L
The semiclassical denition of the bracket then can be applied to classical physics
(where it was originally discovered). Classically AandBare two arbitrary functions
of the coordinates qand momenta p; in quantum mechanics they can be arbitrary
operators. We have included an \ i" in the classical normalization so the two agree
in the semiclassical limit. We generally use (natural/Planck) units h=1 ,s om a s si s
measured as inverse length, etc.; when we do use an explicit h, it is a dimensionless
parameter, and appears only for dening Jeries-Wentzel-Kramers-Brillouin (JWKB)
expansions or (semi)classical limits.
Our indices may appear either as subscripts or superscripts, with preferences to
be explained later: For nonrelativistic purposes we treat them the same. We also use
the Einstein summation convention, that any repeated index in a product is summed
over (\contracted"); usually we contract a superscript with a subscript:
AmBmX
mAmBm
The denition of the bracket is equivalent to using
[pm;qn]= in
m
(wheren
mis the \Kronecker delta function": 1 if m=n,0i fm6=n) together with
the general properties of the bracket
[A;B]= [B;A]; [A;B]y= [Ay;By]
[[A;B];C]+[ [B;C];A]+[ [C;A];B]=0
[A;BC ]=[A;B]C+B[A;C]
The rst set of identities exhibit the antisymmetry of the bracket; next are the \Ja-
cobi identities". In the last identity the ordering is important only in the quantummechanical case: In general, the dierence between classical and quantum mechanics
comes from the fact that in the quantum case operator reordering after taking the
commutator results in multiple commutators.
Innitesimal symmetry transformations are then written as
A=i[G;A];A
0=A+A
whereGis the \generator" of the transformation. More explicitly, innitesimal gen-
erators will contain innitesimal parameters: For example, for translations we have
G=ipi)xi=i[G;xi]=i
A. COORDINATES 5
whereiare innitesimal numbers.
The most evident physical symmetries are those involving spacetime. For nonrel-
ativistic particles, these symmetries form the \Galilean group": For the free particle,those innitesimal transformations are linear combinations of
M=m; P
i=pi;Jij=x[ipj]xipj xjpi;E =H=p2
i
2m;Vi=mxi pit
in terms of the position xi(i=1;2;3), momenta pi, and (nonvanishing) mass m,
where [ij] means to antisymmetrize in those indices, by summing over all permuta-
tions (just two in this case), with plus signs for even permutations and minus for odd.(In three spatial dimensions, one often writes J
i=1
2ijkJjkto makeJi n t oav e c t o r .
This is a peculiarity of three dimensions, and will lose its utility once we consider rela-tivity in four spacetime dimensions.) These transformations are the space translations(momentum) P, rotations (angular momentum | just orbital for the spinless case)
J, time translations (energy) E, and velocity transformations (\Galilean boosts")
V.( T h e m a s s Mis not normally associated with a symmetry, and is not conserved
relativistically.)
Excercise IA1.1
Let's examine the Galilean group more closely. Using just the relations for[x;p]a n d[A;BC ] (and the antisymmetry of the bracket):
aFind the action on x
iof each kind of inntesimal Galilean transformation.
bShow that the nonvanishing commutation relations for the generators are
[Jij;Pk]=ik[iPj]; [Jij;Vk]=ik[iVj]; [Jij;Jkl]=i[k
[iJj]l]
[Pi;Vj]= iijM; [H;Vi]= iPi
For more than one free particle, we introduce an m,xi,a n dpifor each particle
(but the same t), and the generators are the sums over all particles of the above ex-
pressions. If the particles interact with each other the expression for His modied, in
such a way as to preserve the commutation relations. If the particles also interact withdynamical elds, eld-dependent terms must be added to the generators. (External,nondynamical elds break the invariance. For example, a particle in a Coulomb po-tential is not translation invariant since the potential is centered about some point.)Note that for N particles there are 3N coordinates describing the particles, but stillonly 3 translations: The particles interact in the same 3-dimensional space. We canuse translational invariance to x the position of any one particle at a given time, butnot the rest: The dierences in position are translationally invariant. On the other
6I . G L O B A L
hand, it is often useful not to x the position of any particle, since keeping this invari-
ance (and the corresponding redundant variables) allows all particles to be treated
equally. We might also consider using the dierences of positions themselves as thevariables, allowing a symmetric treatment of the particles in terms of translationally
invariant variables: However, this would require applying constraints on the variables,
since there are 3N(N-1)/2 dierences, of which only 3(N-1) are independent. We willnd similar features later for \local" invariances: In general, the most convenientdescription of a theory is with the invariance; the invariance can then be xed, or
invariant combinations of variables used, appropriately for the particular application.
The rotations (or at least their \orbital" parts) and space translations are exam-
ples of coordinate transformations. In general, generators of coordinate transforma-tions are of the form
G=
i(x)pi)(x)=i[G;]=i@i
where@i=@=@xiand(x) is a \scalar eld" (or \spin-0 wave function"), a function
of only the coordinates.
In classical mechanics, or quantum mechanics in the Heisenberg picture, time
development also can be expressed in terms of the Hamiltonian using the bracket:
d
dtA=@
@t+iH;A
=@
@tA+i[H;A]
(The middle expression with the commutator of @=@t makes sense only in the quantum
case, and is not dened for the Poisson bracket.) Again, this general relation is
equivalent to the special cases, which in the classical limit are Hamilton's equationsof motion:dq
m
dt=i[H;qm]=@H
@pm;dpm
dt=i[H;pm]= @H
@qm
The Hamiltonian has no explicit time dependence in the absence of time-dependent
nondyamical elds (external potentials whose time dependence is xed by hand,
rather than by introducing the elds and their conjugate variables into the Hamilto-nian). Consequently, time development is itself a symmetry: Time translations are
generated by the Hamiltonian; the @=@t term ind=dt term can be dropped when
acting on operators without explicit time dependence.
Invariance of the theory under a symmetry means that the equations of motion
are unchanged under the transformation:
dA
dt0
=dA0
dt
A. COORDINATES 7
To apply our above translation of innitesimal transformations into bracket language,
we dene(d=dt)b y
d
dtA
=d
dt
A+d
dtA
In the quantum case we can write
d
dt
=
iG;@
@t+iH
;
which follows from the Jacobi identity using B=iGandC=@=@t +iH, and inserting
Ainto the blank spaces of the commutators above. (The classical case can be treated
similarly, except that the time derivatives are not written as brackets.) We then ndthat the generator of a symmetry transformation is conserved (constant), since
0=d
dt
=
i@G
@t [G;H ];
= idG
dt;
Excercise IA1.2
Show that the generators of the Galilean group are conserved, using the rela-tiond=dt =@=@t +i[H;] for the hamiltonian Hof a free particle. Solve the
equations of motion for x(t)a n dp(t) in terms of initial conditions, and substi-
tute into the expression for the generators to give an independent derivationof their time independence.
In the cases where time dependence is not involved, symmetries can be treated
in almost exactly the same way either classically or quantum mechanically using thecorresponding bracket (Poisson or commutator), by using the properties that theyhave in common. In particular, the fact that a symmetry generator G=
m(x)pm
is conserved means that we can solve for a component of pin terms of the constant
G, and substitute the result into the remaining equations of motion. For example,
translation invariance of a potential in a particular direction means that componentof the momentum is a constant ( dp
1=dt= @H=@q1= 0), rotational invariance
about some axis means that component of angular momentum is a constant ( dJ=dt =
@H=@ = 0), etc.
2. Fermions
As we know experimentally, and we will see follows from relativistic eld theory,
particles with half-integral spins obey Fermi-Dirac statistics. Let's therefore considerthe classical limit of fermions: This will lead to generalizations of the concepts ofbrackets and coordinates. Bosons obey commutation relations, such as [ x;p]=ih;
8I . G L O B A L
in the classical limit they just commute. Fermions obey anticommutation relations,
such asf;yg=hfor a single fermionic harmonic oscillator, where
fA;Bg=AB+BA
is the anticommutator. So, in the truly classical (not semiclassical) limit they an-
ticommute, y+y= 0. Actually, the simplest case is a single real (hermitian)
fermion: Quantum mechanically, or semiclassically, we have
h=f;g=22
while classically 2= 0. There is no analog for a single boson: [ x;x]=x2 x2=0 .
This means that classical fermionic elds must be \anticommuting": Two such objects
get a minus sign when pushed past each other. As a result, the product of two
fermionic quantities is bosonic, while fermionic times bosonic gives fermionic.
Excercise IA2.1
Show
[B;C]=[A;D]=0) [AB;CD ]=1
2fA;Cg[B;D ]+1
2[A;C]fB;Dg
To work with wave functions that are functions of anticommuting numbers, we
rst must understand how to dene general properties of functions of anticommuting
variables. For instance, given a single anticommuting variable , we need to be able
to Taylor expand functions in , e.g., to nd a basis for the states. We therefore have
an anticommuting derivative @=@ , satisfying
@
@ 2
=0
from either anticommutativity or the fact functions of terminate at rst order in .
We also need a integral to dene the inner product; indenite integration turns out
to be enough. The most important property of the integral is integration by parts;
then, when acting on any function of ,
Z
d @
@ =0)Z
d =@
@
where the normalization is xed for convenience. This also implies
( )=
Excercise IA2.2
Prove this is the most general possibility for anticommuting integration by
A. COORDINATES 9
considering action of integration and dierentiation on the most general func-
tion of (which has only two terms).
In general, when Taylor expanding a function of anticommuting variables we must
preserve the statistics: If we Taylor expand a quantity that is dened to be commuting(bosonic), then the coecients of even powers of anticommuting variables will also be
commuting, while the coecients of odd powers will be anticommuting (fermionic),
to maintain the commuting nature of that term (the product of the variables and
coecient). Similarly, when expanding an anticommuting quantity the coecients of
even powers will also be anticommuting, while for odd powers it will be commuting.
We can now consider operators that depend on both commuting (
m)a n da n t i -
commuting ( ) classical variables,
M=(m; )
Classically they satisfy the \graded" commutation relations (anticommutation if both
elements are fermionic, commutation otherwise), not to be confused with the Poisson
bracket,
classically [M;Ng=0 :mn nm=m m= + =0
This relation is then generalized to the the graded quantum mechanical commutator
or Poisson bracket by
[M;Ng=h
MN;
MN
PN=M
P
where
is constant, hermitian, and \graded antisymmetric":
(MN]=0:
(mn)=
[]=
m+
m=0
For the standard normalization of canonically conjugate pairs of bosons
m=i=(qi;pi)
and self-conjugate fermions, we choose
=;
i;j=ijC;C= 0
ii
0
Because of signs resulting from ordering anticommuting quantities, we dene
derivatives unambiguously by their action from the left:
@
@MN=N
M
10 I. GLOBAL
The general Poisson bracket then can be written as
semiclassically [A;Bg A
@
@M
NM@
@NB
Since derivatives are normally dened to act from the left, there is a minus sign from
pushing the rst derivative to the left if Aand that particular component of @=@M
are both fermionic.
Excercise IA2.3
Let's examine some properties of fermionic oscillators:
aFor a single set of harmonic oscillators we have
fa;ayg=1;fa;ag=fay;ayg=0
Show that the \number operator" ayahas the property
fa;eiayag=0
(Hint: Since this system has only 2 states, the easiest way is to check the
action on those states.)
bDene eigenstates of the annihilation operator (\coherent states") by
aji=ji
whereis anticommuting. Show that this implies
ayji= @
@ji;ji=eayj0i;e0ayji=j+0i;xayaji=jxi;
hj0i=e*0;1=Z
d*d e *jihj
Dene wave functions in this space, ( )=hj i.T a y l o r e x p a n d t h e m i n
, and compare this to the usual two-component representation using j0iand
ayj0ias a basis.
cDene the \supertrace" by
str(A)=Z
d*d e *hjAji
Find the relation between any operator in this space and a 2 2 matrix, and
nd the expression for the supertrace in terms of this matrix.
dRepeat part bfor the bosonic oscillator ([ a;ay] = 1), where the Hilbert space
is innite-dimensional, paying attention to signs, etc. Show that the analog
of part cdenes the ordinary trace.
A. COORDINATES 11
eFortwosets of fermionic oscillators, we dene
fa1;ay
1g=fa2;ay
2g=1;o t h e rf;g=0
Show that the new operators
~a1=a1; ~a2=eiay1a1a2
(and their Hermitian conjugates) are equivalent to the original ones except
that one set of the new oscillators commutes (not anticommutes) with the
other ([~a1;~ay2] = 0, etc.), even though each set satises the same anticom-
mutation relations with itself ( f~a1;~ay1g= 1, etc.). Thus, choice of statistics
is relevant only for particles in the same state: at most one fermion, but
unlimited bosons. (This change of oscillator basis is called a \Klein trans-formation". It can be useful for discrete sets of oscillators, but not for thoselabeled by a continuous parameter, because of the discontinuity in the com-mutation relations when the two labels are equal.)
3. Lie algebra
Since the same symmetries can be expressed in terms of dierent kinds of brackets
for classical and quantum theories, it can be useful to work with just those propertiesthat the Poisson bracket and commutator have in common, i.e., those that involveonly the bracket of two operators, not just their ordinary product:
[A+B;C ]=[A;C]+[B;C]f o r n u m b e r s ; (distributivity)
[A;B]= [B;A] (antisymmetry)
[A;[B;C]] + [B;[C;A]] + [C;[A;B]] = 0 (Jacobi identity)
with similar expressions (diering only by signs) for anticommutators or mixed com-
mutators and anticommutators.
Excercise IA3.1
Find the generalizations of the Jacobi identity using also anticommutators,corresponding to the cases where 2 or 3 of the objects involved are consideredas fermionic instead of bosonic.
These properties also give an abstract denition of a form of multiplication, the
\Lie bracket", which denes a \Lie algebra". (The rst property is true of algebrasin general.) Other Lie brackets include those dened by another, associative, form of
12 I. GLOBAL
multiplication, such as matrix multiplication, or operator (innite matrix) multipli-
cation as in quantum mechanics: In those cases we can write [ A;B]=AB BA,a n d
use the usual properties of multiplication (distributivity and associativity) to derive
the properties of the Lie bracket. (Another familiar example in physics is the \cross"
product for three-vectors; however, this can also be expressed in terms of matrixmultiplication.) The most important use of Lie algebras for physics is for describing
(continuous) innitesimal transformations, especially those describing symmetries.
Excercise IA3.2
Using only the commutation relations of the generators of the Galilean group(excercise IA1.1), check all the Jacobi identities.
For describing transformations, we can also think of the bracket as a derivative:
The \Lie derivative" of Bwith respect to Ais dened as
L
AB=[A;B]
As a consequence of the properties of the Lie bracket, this derivative satises the
usual properties of a derivative, including the Leibniz rule. (In fact, for coordinatetransformations the Lie derivative is really a derivative with respect to the coordi-
nates.)
We can now dene nite transformations by exponentiating innitesimal ones:
A
0(1 +iLG)A)A0= lim
!0(1 +iLG)1=A=eiLGA
In cases where we have [ A;B]=AB BA, we can also write
eiLGA=eiGAe iG
This follows from replacing Gon both sides with Gand taking the derivative with
respect to, to see that both satisfy the same dierential equation with the same
initial condition. We then can recognize this as the way transformations are performed
in quantum mechanics: A linear transformation that preserves the Hilbert-space innerproduct must be unitary, which means it can be written as the exponential of an
antihermitian operator.
Just as innitesimal transformations dene a Lie algebra with elements A, nite
ones dene a \Lie group" with elements
g=e
iG
The multiplication law of two group elements follows from the fact the product of
two exponentials can be expressed in terms of multiple commutators:
eAeB=eA+B+1
2[A;B]+:::
A. COORDINATES 13
We now have the mathematical properties that dene a group, namely: (1) a product,
so that for two group elements g1andg2, we can dene g1g2, which is another element
of the group (closure), (2) an identity element, so gI=Ig=g, (3) an inverse, where
gg 1=g 1g=I, and (4) associativity, g1(g2g3)=(g1g2)g3. In this case the identity
is 1 =e0, while the inverse is ( eA) 1=e A.
Since the elements of a Lie algebra form a vector space (we can add them and
multiply by numbers), it's useful to dene a basis:
G=iGi)g=eiiGi
The parameters ithen also give a set of (redundant) coordinates for the Lie group.
(Previously they were required to be inntesimal, for inntesimal transformations;
now they are nite, but may be periodic, as determined by topological considerations
that we will mostly ignore.) Now the multiplication rules for both the algebra andthe group are given by those of the basis:
[G
i;Gj]= ifijkGk
for the (\structure") constants fijk= fjik, which dene the algebra/group (but are
ambiguous up to a change of basis). They satisfy the Jacobi identity
[[G[i;Gj];Gk]]=0)f[ijlfk]lm=0
A familiar example is SO(3) (SU(2)), 3D rotations, where fijk=ijkif we useGi=
1
2ijkJjk.
Another useful concept is a \subgroup": If some subset of the elements of a group
also form a group, that is called a \subgroup" of the original group. In particular, for
a Lie group the basis of that subgroup will be a subset of some basis for the original
group. For example, for the Galilean group Jijgenerate the rotation subgroup.
Excercise IA3.3
Let's examine the subgroup of the Galilean group describing (spatial) coor-
dinate transformations | rotations and spatial translations:
aShow that the innitesimal transformations are given by
xi=xjji+^i;ij= ji
where the's are constants.
bExponentiate to nd the nite transformations
x0i=xjji+^i
14 I. GLOBAL
cShow that ijmust satisfy
ikjlkl=ij
both to preserve the scalar product, and as a consequence of exponentiating.
(Hint: Use matrix notation, and nd the equivalent relation between and
1.)
dShow that the last equation implies det=1, while exponentiating can
give onlydet = 1 (since +1 can't change continuously to 1). What is the
physical interpretation of a transformation with det= 1? (Hint: Consider
a simple example.)
These results can be generalized to include anticommutators: When some of the
basis elements Giare fermionic, the corresponding parameters iare anticommuting
numbers, the structure constants are dened by [ Gi;Gjg,e t c . . T h e n G=iGiis
bosonic term by term, as is g, so bosons transform into bosons and fermions into
fermions, but Taylor expansion in the 's will have both bosonic and fermionic co-
ecients. (For example, for A=B,i fAis bosonic, then so is B, but if also is
fermionic, then Bwill also be fermionic.)
For some purposes it is more convenient to absorb the \ i" in the innitesimal
transformation into the denition of the generator:
G! iG)A=[G;A]=LGA; g =eG;[Gi;Gj]=fijkGk
This aects the reality properties of G: In particular, if gis unitary ( ggy=I), as
usually required in quantum mechanics, g=eiGmakesGhermitian ( G=Gy), while
g=eGmakesGantihermitian ( G= Gy). In some cases anithermiticity can be
an advantage: For example, for translations we would then have Pi=@iand for
rotationsJij=x[i@j], which is more convenient since we know the i's in these (and
any) coordinate transformations must cancel anyway. On the other hand, the U(1)transformations of electrodynamics (on the wave function for a charged particle) arejust phase transformations g=e
i(whereis a real number), so clearly we want the
expliciti; then the only generator has the representation Gi= 1. In general we'll nd
that for our purposes absorbing the i's into the generators is more convenient for just
spacetime symmetries, while explicit i's are more convenient for internal symmetries.
A. COORDINATES 15
4. Relativity
The Hamiltonian approach singles out the time coordinate. In relativistic theories
time can be treated on equal footing with space, and it is useful to take advantage of
this fact, so that the full Poincar e invariance is manifest. So, we treat the time tand
spatial position xitogether as a four-vector (or D-vector in D 1s p a c ea n d1t i m e
dimension)
xm=(x0;xi)=(t;xi)
wherem=0;1;:::;3( o rD 1),i=1;2;3. Since the energy Eand three-momentum
piare canonically conjugate to them,
[pi;xj]= iij; [E;t]=+i
we dene the 4-momentum as
pm=(E;pi)=mnpn;pm=mnpn;[pm;xn]= imn; [pm;xn]= in
m
where we raise and lower indices with the \Minkowski metric", in an \orthonormal
basis",
mn=0
BBBB@01 2 3
0 1000
101 0 0200 1 0
300 0 11
CCCCA)p
0= p0= E
in four spacetime dimensions, with obvious generalizations to higher dimensions.
(Sometimes the metric with signs + is used; we prefer + ++ because it
is more convenient for quantum calculations.) Therefore, we now distinguish upperand lower indices in general: At least for position and momentum, the upper-indexed
x
mandpmhave the usual physical interpretation (so xmandpmhave extra signs).
This is consistent with our previous nonrelativistic notation, since 3-vector indices donot change sign upon raising or lowering.
Of course, we could have done that much nonrelativistically. Relativity is a
symmetry of kinematics and dynamics: In particular, a free, spinless, relativistic
particle is completely described by the constraint
p
2+m2=0
where we dene the covariant square
p2=pmpm=pmpnmn= (p0)2+(p1)2+(p2)2+(p3)2
16 I. GLOBAL
Our relativistic symmetry must leave this constraint invariant: Thus the metric de-
nes the norm of a vector (and an invariant inner product). Therefore, to preserve
Lorentz invariance it is important that we contract only an upper index with a lower
index. For similar reasons, we have
@m=@
@xm;@mxn=n
m
so quantum mechanically pm= i@m.
Unlike the positive-denite nonrelativistic norm of a 3-vector Vi, for an arbitrary
4-vectorVmwe can have
V28
<
:<
=
>9
=
;0:8
<
:timelike
lightlike=null
spacelike
In particular, the 4-momentum is timelike for massive particles ( m2>0) and lightlike
for massless ones (while \tachyons", with spacelike momenta and m2<0, do not exist,
for reasons that are most clear from quantum eld theory).
The quantum mechanics will be described later, but the result is that this con-
straint can be used as the wave equation. The main qualitative distinction from the
nonrelativistic case in the constraint
nonrelativistic : 2mE+~p2=0
relativistic : E2+m2+~p2=0
is that the equation for the energy Ep0is now quadratic, and thus has two
solutions:
p0=!; ! =p
(pi)2+m2
Later we'll see how the second solution is interpreted as an \antiparticle". We also
use (natural/Planck) units c= 1, so length and duration are measured in the same
units;cthen appears only as a parameter for dening nonrelativistic expansions and
limits.
The translations and Lorentz transformations make up the Poincar e group, the
symmetry that denes special relativity. (The Lorentz group in D 1s p a c ea n d1
time dimension is the \orthogonal" group \O(D 1,1)". The \proper" Lorentz group
\SO(D 1,1)", where the \S" is for \special", transforms the coordinates by a matrix
whose determinant is 1. The Poincar eg r o u pi sI S O ( D 1,1), where the \I" stands
for \inhomogeneous".) For the spinless particle they are generated by coordinate
transformations GI=(Pa;Jab):
Pa=pa;Jab=x[apb]
A. COORDINATES 17
(where also a;b=0;:::;3). Then the fact that the physics of the free particle is
invariant under Poincar e transformations is expressed as
[Pa;p2+m2]=[Jab;p2+m2]=0
Writing an arbitrary innitesimal transformation as a linear combination of the gen-
erators, we nd
xm=xnnm+^m;mn= nm
where the's are constants. Note that antisymmetry of mndoes not imply antisym-
metry ofmn=mppn, because of additional signs. (Similar remarks apply to Jab.)
Exponentiating to nd the nite transformations, we have
x0m=xnnm+^m; mpnqpq=mn
The same Lorentz transformations apply to pm, but the translations do not aect
it. The condition on follows from preservation of the Minkowski norm (or inner
product), but it is equivalent to the antisymmetry of mnby exponentiating = e
(compare excercise IA3.3).
Sincedxapais invariant under the coordinate transformations dened by the Pois-
son bracket (the chain rule, since eectively pa@a), it follows that the Poincar e
invariance of p2is equivalent to the invariance of the line element
ds2= dxmdxnmn
which denes the \proper time" s. Spacetime with this indenite metric is called
\Minkowski space", in contrast to the \Euclidean space" with positive denite metric
used to describe nonrelativistic length measured in just the three spatial dimensions.
For the massive case, we also have
pa=mdxa
ds
For the massless case ds= 0: Massless particles travel along lightlike lines. However,
we can dene a new parameter such that
pa=dxa
d
is well-dened in the massless case. In general, we then have
s=m
While this xes =s=m in the massive case, in the massless case it instead restricts
s= 0. Thus, proper time does not provide a useful parametrization of the world
18 I. GLOBAL
line of a classical massless particle, while does: For any piece of such a line, dis
given in terms of (any component of) paanddxa. Later we'll see how this parameter
appears in relativistic classical mechanics, and is useful for quantum mechanics andeld theory.
Excercise IA4.1
The relation between xandpis closely related to the Poincar e conservation
laws:
aShow that
dP
a=dJab=0)p[adxb]=0
and use this to prove that conservation of PandJimply the existence of a
parametersuch thatpa=dxa=d.
bConsider a multiparticle system (but still without spin) where some of the
particles can interact only when at the same point (i.e., by collision; theyact as free particles otherwise). Dene P
a=PpaandJab=Px[apb]as the
sum of the individual momenta and angular momenta. Show that momentum
conservation implies angular momentum conservation,
Pa=0) Jab=0
where \" refers to the change from before to after the collision(s).
Special relativity can also be stated as the fact that the only physically observable
quantities are those that are Poincar e invariant. (Other objects, such as vectors,
depend on the choice of reference frame.) For example, consider two spinless particles
that interact by collision, producing two spinless particles (which may dier from theoriginals). Without loss of generality, we can describe this process in terms of justthe momenta. (Quantum mechanically, this is automatically a complete description;classically, the position is found by p=dx=d .) All invariants can be expressed in
terms of the masses and the \Mandelstam variables" (not to be confused with time
and proper time)
s= (p
1+p2)2;t = (p1 p3)2;u = (p1 p4)2
where we have used momentum conservation, which shows that even these three
quantities are not independent:
p2
I= m2
I;p 1+p2=p3+p4)s+t+u=4X
I=1m2
I
A. COORDINATES 19
(The explicit index now labels the particle, for the process 1+2 !3+4.) The simplest
reference frame to describe this interaction is the center-of-mass frame (actually the
center of momentum, where the two 3-momenta cancel). In that Lorentz frame, using
also rotational invariance, momentum conservation, and the mass-shell conditions, the
momenta can be written in terms of these invariants as
p1=1ps(1
2(s+m2
1 m2
2);12;0;0)
p2=1ps(1
2(s+m2
2 m2
1); 12;0;0)
p3=1ps(1
2(s+m2
3 m2
4);34cos; 34sin; 0)
p4=1ps(1
2(s+m2
4 m2
3); 34cos; 34sin; 0)
cos =s2+2st (Pm2
I)s+(m2
1 m2
2)(m2
3 m2
4)
41234
2
IJ=1
4[s (mI+mJ)2][s (mI mJ)2]
The \physical region" of momentum space is then given by s(m1+m2)2and
(m3+m4)2,a n djcos j1.
Excercise IA4.2
Derive the above expressions for the momenta in terms of invariants in the
center-of-mass frame.
Excercise IA4.3
Find the conditions on s;tanduthat dene the physical region in the case
where all masses are equal.
For some purposes it will prove more convenient to use a \lightcone basis"
p=1p
2(p0p1))mn=0
BBBB@+ 23
+0 100
100 0
200 1 0
300 0 11
CCCCA;p
2= 2p+p +(p2)2+(p3)2
and similarly for the \lightcone coordinates" ( x;x2;x3). (\Lightcone" is an unfor-
tunate but common misnomer, having nothing to do with cones in most usages.) In
this basis the solution to the mass-shell condition p2+m2= 0 can be written as
p= p=(pi)2+m2
2p
(where now i=2;3), which more closely resembles the nonrelativistic expression.
(Note the change on indices + $ upon raising and lowering.) A special lightcone
basis is the \null basis",
p=1p
2(p0p1);pt=1p
2(p2 ip3);pt=1p
2(p2+ip3)
20 I. GLOBAL
)mn=0
BBBB@+ tt
+0 100
100 0
t 00 0 1
t 00 1 01
CCCCA;p
2= 2p+p +2ptpt
where the square of a vector is linear in each component. (We often use \ "t o
indicate complex conjugation.)
Excercise IA4.4
Show that for p2+m2=0(m20,pa6= 0), the signs of p+andp are
always the same as the sign of the canonical energy p0.
Excercise IA4.5
Consider the Poincar e group in 1 extra space dimension (D space, 1 time)
for a massless particle. Interpret p+as the mass, and p as the energy.
Show that the constraint p2= 0 gives the usual nonrelativistic expression
for the energy. Show that the subgroup of the Poincar e group generated by
all generators that commute with p+is the Galilean group (in D 1s p a c e
and 1 time dimensions). Now nonrelativistic mass conservation is part ofmomentum conservation, and all the Galilean transformations are coordinate
transformations. Also, positivity of the mass is related to positivity of the
energy (see excercise IA4.4).
5 .D i s c r e t e :C ,P ,T
By considering only symmetries than can be obtained continuously from the iden-
tity (Lie groups), we have missed some important symmetries: those that re
ect some
of the coordinates. It's sucient to consider a single re
ection of a spacelike axis,and one of a timelike axis; all other re
ections can be obtained by combining these
with the continuous (\proper, orthochronous") Lorentz transformations. (Spacelike
and timelike vectors can't be Lorentz transformed into each other, and re
ection of
a lightlike axis won't preserve p
2+m2.) Also, the re
ection of one spatial axis can
be combined with a rotation about that axis, resulting in re
ection of all three
spatial coordinates. (Similar generalizations hold for higher dimensions. Note that
the product of an even number of re
ections about dierent axes is a proper rotation;
thus, for even numbers of spatial dimensions re
ections of all spatial coordinates areproper rotations, even though the re
ection of a single axis is not.) The reversal of the
spatial coordinates is called \parity (P)", while that of the time coordinate is called
\time reversal" (\T"; actually, for historical reasons, to be explained shortly, this is
A. COORDINATES 21
usually labeled \CT".) These transformations have the same eect on the momen-
tum, so that the denition of the Poisson bracket is also preserved. These \discrete"
transformations, unlike the proper ones, are not symmetries of nature (except in cer-tain approximations): The only exception is the transformation that re
ects all axes(\CPT").
While the metric
mnis invariant under all Lorentz transformations (by deni-
tion), the \Levi-Civita tensor"
mnpqtotallyantisymmetric; 0123= 0123=1
is invariant under only proper Lorentz transformations: It has an odd number of
space indices and of time indices, so it changes sign under parity or time reversal.
Consequently, we can use it to dene \pseudotensors": Given \polar vectors", whose
signs change as position or momentum under improper Lorentz transformations, andscalars, which are invariant, we can dene \axial vectors" and \pseudoscalars" as
V
a=abcdBbCcDd; =abcdAaBbCcDd
which get an extra sign change under such transformations (P or CT, but not CPT).
There is another such \discrete" transformation that is dened on phase space,
but which does not aect spacetime. It changes the sign of all components of themomentum, while leaving the spacetime coordinates unchanged. This transformationis called \charge conjugation (C)", and is also only an approximate symmetry innature. (Quantum mechanically, complex conjugation of the position-space wave
function changes the sign of the momentum.) Furthermore, it does not preserve the
Poisson bracket, but changes it by an overall sign. (The misnomer \CT" for timereversal follows historically from the fact that the combination of reversing the timeaxis and charge conjugation preserves the sign of the energy.) The physical meaningof this transformation is clear from the spacetime-momentum relation of relativisticclassical mechanics p=md x = d s : It is proper-time reversal, changing the sign of s.
The relation to charge follows from \minimal coupling": The \covariant momentum"
md x = d s =p+qA(for charge q) appears in the constraint ( p+qA)
2+m2=0i na n
electromagnetic background; p! pthen has the same eect as q! q.
In the previous subsection, we mentioned how negative energies were associated
with \antiparticles". Now we can better see the relation in terms of charge conjuga-
tion. Note that charge conjugation, since it only changes the sign of but does not
eect the coordinates, does not change the path of the particle, but only how it isparametrized. This is also true in terms of momentum, since the velocity is given by
22 I. GLOBAL
pi=p0. Thus, the only observable property that is changed is charge; spacetime prop-
erties (path, velocity, mass; also spin, as we'll see later) remain the same. Another
way to say this is that charge conjugation commutes with the Poincar e group. One
way to identify an antiparticle is that it has all the same kinematical properties (mass,
spin) as the corresponding particle, but opposite sign for internal quantum numbers
(like charge). (Another way is pair creation and annihilation: See subsection IIIB5
below.) Quantum mechanically, we can identify a particle with its antiparticle by
requiring the wave function or eld to be invariant under charge conjugation: For
example, for a scalar eld (spinless particle), we have the reality condition
(x)=*(x)
or in momentum space, by Fourier transformation,
~(p)=[ ~( p)]*
which implies the particle has charge zero (neutral).
All these transformations are summarized in the table:
CCT PTC PPT CPT
s ++ +
t+ + +
~x++ +
E ++ +
~p + ++
(The upper-left 33 matrix contains the denitions, the rest is implied.)
However, from the point of view of the \particle" there issome kind of kinematic
change, since the proper time has changed sign: If we think of the mechanics of a
particle as a one-dimensional theory in space (the worldline), where x()( a sw e l l
as any such variables describing spin or internal symmetry) is a wave function or eld
on that space, then ! is T on that one-dimensional space. (The fact we don't
get CT can be seen when we add additional variables: For example, if we describe
internal U(N) symmetry in terms of creation and annihilation operators ayiandai,
then C mixes them on both the worldline and spacetime. So, on the worldline we
have the \pure" worldline geometric symmetry CT times C = T.) Thus, in terms of
\zeroth quantization",
worldlineT$spacetimeC
A. COORDINATES 23
On the other hand, spacetimePandCTare simply internal symmetries with respect
to the worldline (as are proper, orthochronous Poincar e transformations).
6. Conformal
Although Poincar e transformations are the most general coordinate transforma-
tions that preserve the mass condition p2+m2= 0, there is a larger group, the
\conformal group", that preserves this constraint in the massless case. Transforma-
tionsthat satisfy
[a(x)pa;p2]=(x)p2
for somealso preserve p2= 0, although they don't leave p2invariant. Equivalently,
we can look for coordinate transformations that scale
dx02=(x)dx2
Excercise IA6.1
Find the conformal group explicitly in two dimensions, and show it's innite
dimensional (not just the SO(2,2) described below). (Hint: Use lightconecoordinates.)
This symmetry can be made manifest by starting with a space with one extra
space and time dimension:
y
A=(ya;y+;y ))y2=yAyBAB=(ya)2 2y+y
where (ya)2=yaybabuses the usual D-dimensional Minkowski-space metric ab,
and the two additional dimensions have been written in a lightcone basis (not to
be confused for the similar basis that can be used for the Minkowski metric itself).
With respect to this metric, the original SO(D 1,1) Lorentz symmetry has been
enlarged to SO(D,2). This is the conformal group in D dimensions. However, rather
than also preserving (D+2)-dimensional translation invariance, we instead impose the
constraint and invariance
y2=0; yA=yA
This reduces the original space to the \projective" (invariant under the scaling)
lightcone (which in this case really is a cone).
These two conditions can be solved by
yA=ewA;wA=(xa;1;1
2xaxa)
24 I. GLOBAL
Projective invariance then means independence from e(y+), while the lightcone con-
dition has determined y .y2= 0 implies ydy= 0, so the simplest conformal
invariant is
dy2=(edw+wde)2=e2dw2=e2dx2
w h e r ew eh a v eu s e d w2=0)wdw= 0. This means any SO(D,2) transformation
onyAwill simply scale dx2,a n ds c a l e e2in the opposite way:
dx02=e2
e02
dx2
in agreement with the previous denition of the conformal group.
The explicit form of conformal transformations on xa=ya=y+now follows from
their linear form on yA, using the generators
GAB=y[ArB]; [rA;yB]= iB
A
of SO(D,2) in terms of the momentum rAconjugate to yA. (These are dened the
same way as the Lorentz generators Jab=x[apb].) For example, G+ just scalesxa.
(Scale transformations are also known as \dilatations".) We can also recognize G+aas
generating translations on xa. The only complicated transformations are generated
byG a, known as \conformal boosts" (acceleration transformations). Since they
commute with each other (like translations), it's easy to exponentiate to nd thenite transformations:
y
0=eGy; G =vay[ @a]
for some constant D-vector va(where@A@=@yA). Since the conformal boosts act
as \lowering operators" for scale weight (+ !a! ), only the rst three terms in
the exponential survive:
Gy =0;G ya=vay ;G y+=vaya)
y0 =y ;y0a=ya+vay ;y0+=y++vaya+1
2v2y )
x0a=xa+1
2vax2
1+vx+1
4v2x2
usingxa=ya=y+,y =y+=1
2x2.
Excercise IA6.2
Make the change of variables to xa=ya=y+,e=y+,z=1
2y2.E x p r e s s
rAin terms of the momenta ( pa;n;s) conjugate to ( xa;e;z). Show that the
conditionsy2=yArA=r2= 0 become z=en=p2= 0 in terms of the new
variables.
A. COORDINATES 25
Excercise IA6.3
Find the generator of innitesimal conformal boosts in terms of xaandpa.
We actually have the full O(D,2) symmetry: Besides the continuous symmetries,
and the discrete ones of SO(D 1,1), we have a second \time" reversal (from our
second time dimension):
y+$ y )xa$ xa
1
2x2
This transformation is called an \inversion".
Excercise IA6.4
Show that a nite conformal boost can be obtained by performing a transla-tion sandwiched between two inversions.
Excercise IA6.5
The conformal group for Euclidean space (or any spacetime signature) can beobtained by the same construction. Consider the special case of D=2 for theseSO(D+1,1) transformations. (This is a subgroup of the 2D superconformalgroup: See excercise IA6.1.) Use complex coordinates for the two \physical"dimensions:
z=
1p
2(x1+ix2)
aShow that the inversion is
z$ 1
z*
bShow that the conformal boost is (using a complex number also for the boost
vector)
z!z
1+v*z
Excercise IA6.6
Any parity transformation (re
ection in a spatial axis) can be obtained from
any other by a rotation of the spatial coordinates. Similarly, when there ismore than one time dimension, any time reversal can be obtained from another(but time reversal can't be rotated into parity, since a timelike vector can't berotated into a spacelike one). Thus, the complete orthogonal group O(m,n)can be obtained from those transformations that are continuous from theidentity by combining them with 1 parity transformation and 1 time reversaltransformation (for mn 6=0). For the conformal group, nd the rotation (in
terms of an angle) that rotates between the two time directions, and expressits action on x
a. Show that for angle it produces a transformation that is
26 I. GLOBAL
the product of time reversal and inversion. Use this to show that inversion is
related to time reversal by nding the continuum of conformal transformationsthat connect them.
Although conformal symmetry is not observed in nature, it is important in all
approaches to eld theory: (1) First of all, it is useful in the construction of free the-ories (see subsections IIB1-4 below). All massive elds can be described consistentlyin quantum eld theory in terms of coupling massless elds. Massless theories are asubset of conformal theories, and some conditions on massless theories can be found
more easily by nding the appropriate subset of those on conformal theories. This is
related to the fact that the conformal group, unlike the Poincar e group, is \simple":
It has no nontrivial subgroup that transforms into itself under the rest of the group(like the way translations transform into themselves under Lorentz transformations).(2) In interacting theories at the classical level, conformal symmetry is also important
in nding and classifying solutions, since at least some parts of the action are confor-
mally invariant, so corresponding solutions are related by conformal transformations(see subsections IIIC5-7). Furthermore, it is often convenient to treat arbitrary theo-ries as broken conformal theories, introducing elds with which the breaking is asso-ciated, and analyze the conformal and conformal-breaking elds separately. This is
particularly true for the case of gravity (see subsections IXA7,B5,C2-3,XA3-4,B5-7).
(3) Within quantum eld theory at the perturbative level, the only physical quantumeld theories are ones that are conformal at high energies (see subsection VIIIC1).The quantum corrections to conformal invariance at high energy are relatively sim-
ple. (4) Beyond perturbation theory, the only quantum theories that are well dened
may be just the ones whose breaking of conformal invariance at low energy is onlyclassical (see subsections VIIC2-3,VIIIA5-6). Furthermore, the largest possible sym-metry of a nontrivial S-matrix is conformal symmetry (or superconformal symmetryif we include fermionic generators). (5) Self-duality (a generalization of a condition
that equates electric and magnetism elds) is useful for nding solutions to classical
eld equations as well as simplifying perturbation theory, and is closely related to\twistors" (see subsections IIB6-7,C5,IIIC4-7). In general, self-duality is related toconformal invariance: For example, it can be shown that the free conformal theoriesin arbitrary even dimensions are just those with (on-mass-shell) eld strengths on
which self-duality can be imposed. (In arbitrary odd dimensions the free conformal
theories are just the scalar and spinor.)
REFERENCES
1
F.A. Berezin, The method of second quantization (Academic, 1966):
A. COORDINATES 27
calculus with anticommuting numbers.
2P.A.M. Dirac, P r o c .R o y .S o c . A126 (1930) 360:
antiparticles.
3E.C.G. St uckelberg, Helv. Phys. Acta 14(1941) 588, 15(1942) 23;
J.A. Wheeler, 1940, unpublished:
the relation of antiparticles to proper time.
4S. Mandelstam, Phys. Rev. 112(1958) 1344.
5P.A.M. Dirac, Ann. Math. 37(1936) 429;
H.A. Kastrup, Phys. Rev. 150(1966) 1186;
G. Mack and A. Salam, Ann. Phys. 53(1969) 174;
S. Adler, Phys. Rev. D6(1972) 3445;
R. Marnelius and B. Nilsson, Phys. Rev. D22 (1980) 830:
conformal symmetry.
6S. Coleman and J. Mandula, Phys. Rev. 159(1967) 1251:
conformal symmetry as the largest (bosonic) symmetry of the S-matrix.
7W. Siegel, Int. J. Mod. Phys. A 4(1989) 2015:
equivalence between conformal invariance and self-duality in all dimensions.
28 I. GLOBAL
::::::::::::::::::::::::::::: :::::::::::::::::::::::::::::
::::::::::::::::::::::::::::: B. INDICES :::::::::::::::::::::::::::::
In the previous section we saw various spacetime groups (Galilean, Poincar e,
conformal) in terms of how they acted on coordinates. This not only gave them a
simple physical interpretation, but also allowed a direct relation between classical
and quantum theories. However, as we know from studying rotations in quantum
theory in terms of spin, we will often need to study symmetries of quantum theories
for which the classical analog is not so useful or perhaps even nonexistent.
We therefore now consider some general results of group theory, mostly for con-
tinuous groups. We use tensor methods, rather than the slightly more powerful but
greatly less convenient Cartan-Weyl-Dynkin methods. Much of this section should
be review, but is included here for completeness; it is not intended as a substitute for
a group theory course, but as a summary of those results commonly useful in eld
theory.
1. Matrices
Matrices are dened by the way they act on some vector space; an n n matrix
takes one n-component vector to another. Given some group, and its multiplication
table (which denes the group completely), there is more than one way to represent
it by matrices. Any set of matrices we nd that has the same multiplication table as
the group elements is called a \representation" of that group, and the vector space onwhich those matrices act is called the \representation space." The representation of
the algebra or group in terms of explicit matrices is given by choosing a basis for the
vector space. If we include innite-dimensional representations, then a representation
of a group is simply a way to write its transformations that is linear:
0=M is
linear in . More generally, we can also have a \realization" of a group, where the
transformations can be nonlinear. These tend to be more cumbersome, so we usually
try to make redenitions of the variables that make the realization linear. A precise
denition of \manifest symmetry" is that all the realizations used are linear. (One
possible exception is \ane" or \inhomogeneous" transformations 0=M1 +M2,
such as the usual coordinate representation of Poincar e transformations, since these
transformations are still very simple, because they are really still linear, though not
homogeneous.)
For convenience, we write matrices with a Hilbert-space-like notation, but unlike
Hilbert space we don't necessarily associate bras directly with kets by Hermitian
conjugation, or even transposition. In general, the two spaces can even be dierent
B. INDICES 29
sizes, to describe matrices that are not square; however, for group theory we are
interested only in matrices that take us from some vector space into itself, so they
are square. Bras have an inner product with kets, but neither necessarily has a norm
(inner product with itself): In general, if we start with some vector space, written
as kets, we can always dene the \dual" space, written as bras, by dening such aninner product. In our case, we may start with some representation of a group, in
terms of some vector space, and that will give us directly the dual representation. (If
the representation is in terms of unitary matrices, we have a Hilbert space, and thedual representation is just the complex conjugate.)
So, we dene column vectors j iwith a basisj
Ii, and row vectors h jwith a
basishIj,w h e r eI=1;:::;n to describe nn matrices. The two bases have a relative
normalization dened so that the inner product gives the usual component sum:
j i=jIi I;hj=IhIj;hIjJi=J
I)hj i=I I;hIj i= I;hjIi=I
These bases then dene not only the components of vectors, but also matrices:
M=jIiMIJhJj;hIjMjJi=MIJ
where as usual the Ion the component (matrix element) MIJlabels the row of the
matrixM,a n dJthe column. This implies the usual matrix multiplication rules,
inserting the identity in terms of the basis,
I=jKihKj) (MN)IJ=hIjMjKihKjNjJi=MIKNKJ
Closely related is the denition of the trace,
tr M =hIjMjIi=MII)tr(MN)=tr(NM)
(We'll discuss the determinant later. The bra-ket notation is really just matrix nota-
tion written in a way to clearly distinguish column vectors, row vectors, and matrices.)
Thus, for example, we can easily translate transformation laws from matrix no-
tation into index notation just by using a basis for the representation space:
gjIi=jJigJI;GjIi=jJiGJI
G=iGi;j i=iGj i=jIiii(Gi)IJ J) I=ii(Gi)IJ J
The dual space isn't needed for this purpose. However, for any representation of a
group, the transpose
(MT)I
J=MJI
30 I. GLOBAL
of the inverse of those matrices also gives a representation of the group, since
g1g2=g3) (g1)T 1(g2)T 1=(g3)T 1
[G1;G2]=G3) [ GT
1; GT
2]= GT
3
This is the dual representation, which follows from dening the above inner product
to be invariant under the group:
h ji=0) I= i Ji(Gi)JI
The complex conjugate of a complex representation is also a representation, since
g1g2=g3)g1*g2*=g3*
[G1;G2]=G3) [G1*;G2*] =G3*
From any given representation, we can thus nd three others from taking the dual
and the conjugate: In matrix and index notation,
0=g : 0
I=gIJ J
0=(g 1)T : 0I=g 1
JI J
0=g* : 0.
I=g*.
I.
J .
J
0=(g 1)y : 0.
I=g* 1.
J.
I .
J
since (g 1)T,g*, and (g 1)y(but notgT, etc.) satisfy the same multiplication algebra
asg, including ordering. We use up/down and dotted/undotted indices to denote
the transformation law of each type of index; contracting undotted up indices with
undotted down indices preserves the transformation law as indicated by the remaining
indices, and similarly for dotted indices. These four representations are not necessarily
independent: Imposing relations among them is how the classical groups are dened
(see subsections IB4-5 below).
2. Representations
For example, we always have the \adjoint" representation of a Lie group/algebra,
which is how the algebra acts on its own generators:
G=iGi;A =iGi)A=i[G;A]=jifijkGk
)i= ikj(Gj)ki;(Gi)jk=ifijk
B. INDICES 31
This gives us two ways to represent the adjoint representation space: as either the
usual vector space, or in terms of the generators. Thus, we either use the matrix
A=iGi(for arbitrary representation of the matrices Gi, or treating Gias just
abstract generators), and write A=i[G;A], or we can write Aas a row vector,
hAj=ihij)hAj= ihAjG)ihij= ikj(Gj)kihij
The adjoint representation also provides a convenient way to dene a (symmetric)
group metric invariant under the group, the \Cartan metric":
ij=trA(GiGj)= fiklfjlk
For \Abelian" groups the structure constants vanish, and thus so does this metric.
\Semisimple" groups are those where the metric is invertible (no vanishing eigenval-
ues). A \simple" group has no nontrivial subgroup that transforms into itself under
the rest of the group: Semisimple groups can be written as \products" of simple
groups. \Compact" groups are those where it is positive denite (all eigenvalues pos-
itive); they are also those for which the invariant volume of the group space is nite.For simple, compact groups it's convenient to choose a basis where
ij=cAij
for some constant cA(the \Dynkin index" for the adjoint representation). For some
general irreducible representation Rof such a group the normalization of the trace is
trR(GiGj)=cRij=cR
cAij
Now the proportionality constant cR=cAis xed by the choice of R(only), since we
have already xed the normalization of our basis.
In general, the cyclicity property of the trace implies, for any representation, that
0=tr([Gi;Gj]) = ifijktr(Gk)
sotr(Gi) = 0 for semisimple groups. Similarly, we nd
fijkfijllk=it rA([Gi;Gj]Gk)
is totally antisymmetric: For semisimple groups, this implies the total antisymmetry
of the structure constants fijk, up to factors (which are absent for compact groups in
ab a s i sw h e r e ijij). This also means the adjoint representation is its own dual.
32 I. GLOBAL
(For example, for the compact group SO(3), we have ij= ikljlk=2ij.) Thus,
we can write Ain a third way, as a column vector
jAi=jiiijiijji
We can also do this for Abelian groups, by dening an invertible metric unrelated to
the Cartan metric: This is trivial for Abelian groups, since the generators themselves
are invariant, and thus so is anymetric on them.
An identity related to the trace one is the normalization of the value kRof the
\Casimir operator" for any particular representation,
ijGiGj=kRI
Its proportionality to the identity follows from the fact that it commutes with each
generator:
[jkGjGk;Gi]= ifj
ikfGj;Gkg=0
using the antisymmetry of the structure constants. (Thus it takes the same value on
any component of an irreducible representation, since they are all related by grouptransformations.) By tracing this identity, and contracting the trace identity,
c
R
cAdA=trR(ijGiGj)=kRdR
)kR=cRdA
cAdR
wheredRtrR(I) is the dimension of that representation.
Although quantum mechanics is dened on Hilbert space, which is a kind of com-
plex vector space, more generally we want to consider real objects, like spacetime
vectors. This restricts the form of linear transformations: Specically, if we absorbi's asg=e
G, then in such representations Gitself must be real. These represen-
tations are then called \real representations", while a \complex representation" is
one whose representation isn't real in any basis. A complex representation space can
have a real representation, but a real representation space can't have a complex rep-
resentation. In particular, coordinate transformations (of real coordinates) have onlyreal representations, which is why absorbing the i's into the generators is a useful
convention there. For semisimple unitary groups, hermiticity of the generators of the
adjoint representation implies (using total antisymmetry of the structure constantsand reality of the Cartan metric) that the structure constants are real, and thus the
adjoint representation is a real representation. More generally, any real unitary rep-
resentation will have antisymmetric generators ( G=G*= G
y)G= GT). If
B. INDICES 33
the complex conjugate representation is the same as the original (same matrices up
to a similarity transformation g*=MgM 1), but the representation is not real, then
it is called \pseudoreal". (An example is the spinor of SU(2), to be described in thenext section.)
For any representation gof the group, a transformation g!g
0gg 1
0on every
group element gfor some particular group element g0clearly maps the algebra to
itself, and preserves the multiplication rules. (Similar remarks apply to applying thetransformation to the generators.) However, the same is true for complex conjugation,
g!g*: Not only are the multiplication rules preserved, but for any element g
of that representation of the group, g* is also an element. (This can be shown,
e.g., by dening representations in terms of the values of all the Casimir operators,contructed from various powers of the generators.) In quantum mechanics (wherethe representations are unitary), the latter is called an \antiunitary transformation".Although this is a symmetry of the group, it cannot be reproduced by a unitarytransformation, except when the representation is (pseudo)real.
A very simple way to build a representation from others is by \direct sum". If we
have two representations of a group, on two dierent spaces, then we can take theirdirect sum by just putting one column vector on top of the other, creating a biggervector whose size (\dimension") is the sum of that of the original two. Explicitly, ifwe start with the basis j
ifor the rst representation and j0ifor the second, then
the union (ji;j0i) is the basis for the direct sum. (We can also write jIi=(ji;j0i),
where=1;:::;m ;0=1;:::;n ;I=1;:::;m;m +1;:::;m +n.) The group then acts
on each part of the new vector in the obvious way:
=ji ; =j0i0;gji=jig;gj0i=j0ig00
)j i=ji j0i0=j ijior( ) =
gj i=jig j0ig000or(g)=g0
0g00
(We can replace the with an ordinary + if we understand the basis vectors to be
now in a bigger space, where the elements of the rst basis have zeros for the newcomponents on the bottom while those of the second have zeros for the new compo-nents on top.) The important point is that no group element mixes the two spaces:The group representation is block diagonal. Any representation that can be writtenas a direct sum (after an appropriate choice of basis) is called \reducible". For exam-ple, we can build a reducible real representation from an irreducible complex one by
34 I. GLOBAL
just taking the direct sum of this complex representation with the complex conjugate
representation. Similarly, we can take direct sums of more than two representations.
A more useful way to build representations is by \direct product". The idea there
is to take a colummn vector and a row vector and use them to construct a matrix,
where the group element acts simultaneously on rows according to one representation
and columns according to the other. If the two original bases are again jiandj0i,
the new basis can also be written as jIi=j0i(I=1;:::;mn ). Explicitly,
j i=ji
j0i 0;g(ji
j0i)=ji
j0igg00)g00=gg00
or in terms of the algebra
G00=G00+G00
A familar example from quantum mechanics is rotations (or Lorentz transformations),
where the rst space is position space (so is the continuous index x), acted on by
the orbital part of the generators, while the second space is nite-dimensional, and isacted on by the spin part of the generators. Direct product representations are usuallyreducible: They then can be written also as direct sums, in a way that depends onthe particulars of the group and the representations.
Consider a representation constructed by direct product: In matrix notation
^G
i=Gi
I0+I
G0
i
Usingtr(A
B)=tr(A)tr(B), and assuming tr(Gi)=tr(G0
i)=0 ,w eh a v e
tr(^Gi^Gj)=tr(I0)tr(GiGj)+tr(I)tr(G0
iG0j)
For example, for SU(N) (see subsection IB4 below) we can construct the adjoint rep-
resentation from the direct product of the N-dimensional, \dening" representation
and its complex conjugate. (We also get a singlet, but it will not aect the result forthe adjoint.) In that case we nd
tr
A(GiGj)=2NtrD(GiGj))cD
cA=1
2N
For most purposes, we use trD(GiGj)=ij(cD= 1) for SU(N), so cA=2N.
B. INDICES 35
3. Determinants
We now \review" some properties of determinants that will prove useful for the
group analysis of the following subsections. Determinants can be dened in terms of
the Levi-Civita tensor . As a consequence of its antisymmetry,
totally antisymmetric; 12:::n=12:::n=1)J1:::JnI1:::In=I1
[J1In
Jn]
since each possible numerical index value appears once in each ,s ot h e yc a nb e
matched up with 's. By similar reasoning,
1
m!K1:::KmJ1:::Jn mK1:::KmI1:::In m=I1
[J1In m
Jn m]
where the normalization compensates for the number of terms in the summation.
This tensor is used to dene the determinant:
det MIJ=1
n!J1:::JnI1:::InMI1J1MInJn)J1:::JnMI1J1MInJn=I1:::IndetM
since anything totally antisymmetric in nindices must be proportional to the tensor.
This yields an explicit expression for the inverse:
(M 1)J1I1=1
(n 1)!J1:::JnI1:::InMI2J2MInJn(detM ) 1
From this follows a useful expression for the variation of the determinant:
@
@MIJdet M =(M 1)JIdet M
which is equivalent to
lnd e tM =tr(M 1M)
ReplacingMwitheMgives the often-used identity
l nd e teM=tr(e MeM)=tr M)det eM=etrM
where we have used the boundary condition for M= 0. Finally, replacing Min
the last identity with ln(1 +L) and expanding both sides to order Lngives general
expressions for determinants of nnmatrices in terms of traces:
det(1 +L)=etrln(1+L))det L =1
n!(tr L)n 1
2(n 2)!(tr L2)(tr L)n 2+
Excercise IB3.1
Use the denition of the determinant (and not its relation to the trace) to
show
det(AB)=det(A)det(B)
36 I. GLOBAL
These identities can also be derived by dening the determinant in terms of a
Gaussian integral. We rst collect some general properties of (indenite) Gaussian
integrals. The simplest such integral is
Zd2x
2e x2=2=Z2
0d
2Z1
0dr re r2=2=Z1
0du e u=1
)ZdDx
(2)D=2e x2=2=Zdxp
2e x2=2D
=Zd2x
2e x2=2D=2
=1
The complex form of this integral is
ZdDz*dDz
(2i)De jzj2=1
by reducing to real parameters as z=(x+iy)=p
2. These generalize to integrals
involving a real, symmetric matrix Sor a Hermitian matrix Has
ZdDx
(2)D=2e xTSx=2=(detS ) 1=2;ZdDz*dDz
(2i)De zyHz=(det H ) 1
by diagonalizing the matrices, making appropriate redenitions of the integration
variables, and identifying the determinant of a diagonal matrix. Alternatively, wecan use these integrals to dene the determinant, and derive the previous denition.
The relation for the symmetric matrix follows from that for the Hermitian one by
separatingzinto its real and imaginary parts for the special case H=S. If we treat
zandz* as independent variables, the determinant can also be understood as the
Jacobian for the (dummy) variable change z!H
1z,z*!z*. More generally, if
we dene the integral by an appropriate limiting procedure or analytic continuation
(for convergence), we can choose zandz* to be unrelated (or even separate real
variables), and SandHto be complex.
Excercise IB3.2
Other properties of determinants can also be derived directly from the integral
denition:
aFind an integral expression for the inverse of a (complex) matrix Mby using
the identity
0=Z@
@zI(zJe zyMz)
bDerive the identity l nd e tM =tr(M 1M) by varying the Gaussian de-
nition of the (complex) determinant with respect to M.
B. INDICES 37
An even better denition of the determinant is in terms of an anticommuting
integral (see subsection IA2), since anticommutativity automatically gives the anti-symmetry of the Levi-Civita tensor, and we don't have to worry about convergence.We then have, for anymatrixM,
Z
d
DydDe yM=det M
whereycan be chosen as the Hermitian conjugate of or as an independent variable,
whichever is convenient. From the denition of anticommuting integration, the onlyterms in the Taylor expansion of the exponential that contribute are those with theproduct of one of each anticommuting variable. Total antisymmetry in and in
y
then yields the determinant; we dene \ dDydD" to give the correct normalization.
(The normalization is ambiguous anyway because of the signs in ordering the d's.)
This determinant can also be considered a Jacobian, but the inverse of the commutingresult follows from the fact that the integrals are now really derivatives.
Excercise IB3.3
Divide up the range of a square matrix into two (not necessarily equal) parts:
In block form,
M=AB
CD
and do the same for the (commuting or anticommuting) variables used in
dening its determinant. Show that
detAB
CD
=det Ddet(A BD
1C)=det Adet(D CA 1B)
aby integrating over one part of the variables rst (this requires o-diagonal
changes of variables of the form y!y+Ox, which have unit Jacobian), or
bby rst proving the identity
AB
CD
=IB D 1
0IA BD 1C0
0DI 0
D 1CI
We then have, for any antisymmetric (even-dimensional) matrix A,
Z
d2De TA=2=PfA; (PfA)2=det A
by the same method as the commuting case (again with appropriate denition of the
normalization of d2D; the determinant of an odd-dimensional antisymmetric matrix
vanishes, since detM =detMT). However, there is now an important dierence: The
38 I. GLOBAL
\Pfaan" is not merely the square root of the determinant, but itself a polynomial,
since we can evaluate it also by Taylor expansion:
PfAIJ=1
D!2DI1:::I2DAI1I2AI2D 1I2D
which can be used as an alternate denition. (Normalization can be checked by
examining a special case; the overall sign is part of the normalization convention.)
4. Classical groups
The rotation group in three dimensions can be expressed most simply in terms
of 22 matrices. This description is the most convenient for not only spin 1/2, but
all spins. This result can be extended to orthogonal groups (such as the rotation,
Lorentz, and conformal groups) in other low dimensions, including all those relevantto spacetime symmetries in four dimensions.
There are an innite number of Lie groups. Of the compact ones, all but a nite
number are among the \classical" Lie groups. These classical groups can be dened
easily in terms of (real or complex) matrices satisfying a few simple constraints. (Theremaining \exceptional" compact groups can be dened in a similar way with a little
extra eort, but they are of rather specialized interest, so we won't cover them here.)
These matrices are thus called the \dening" representation of the group. (Sometimesthis representation is also called the \fundamental" representation; however, this term
has been used in slightly dierent ways in the literature, so we will avoid it.) These
constraints are a subset of:
volume: Special:det(g)=1
metric:8
<
:hermitian: Unitary:
(anti)symmetric:Orthogonal:
Symplectic:gg
y=
ggT=
g
gT=
(y= )
(T=)
(
T=
)
reality:Real:
pseudoreal (*):g*=g 1
g*=
g
1
wheregis any matrix in the dening representation of the group, while ;;
a r e
group \metrics", dening inner products (while the determinant denes the volume,
as in the Jacobian). For the compact cases and can be chosen to be the identity,
but we will also consider some noncompact cases. (There are also some uninteresting
variations of \Special" for complex matrices, setting the determinant to be real or its
magnitude to be 1.)
B. INDICES 39
Excercise IB4.1
Write all the dening constraints of the classical groups (S, U, O, Sp, R,
pseudoreal) in terms of the algebra rather than the group.
Note the modied denition of unitarity, etc. Such things are also encountered
in quantum mechanics with ghosts, since the resulting Hilbert space can have an
indenite metric. For example, if we have a nite-dimensional Hilbert space where
the inner product is represented in terms of matrices as
h ji= y
then \observables" satisfy a \pseudohermiticity" condition
h jHi=hH ji) H=Hy
and unitarity generalizes to
hU jUi=h ji)UyU=
Similar remarks apply when replacing the Hilbert-space \sesquilinear" (vector times
complex conjugate of vector) inner product with a symmetric (orthogonal) or anti-symmetric (symplectic) bilinear inner product. An important example is when the
wave function carries a Lorentz vector index, as expected for a relativistic description
of spin 1; then clearly the time component is unphysical.
The groups of matrices that can be constructed from these conditions are then:
GL(n,C) [SL(n,C)] U: [S]U(n
+,n )
O: [S]O(n,C)
Sp: Sp(2n,C)R: GL(n) [SL(n)]
*: [S]U*(2n)U R *
O [S]O(n +,n ) SO*(2n)
Sp Sp(2n) USp(2n +,2n )
Of the non-determinant constraints, in the rst column we applied none (\GL" means
\general linear", and \C" refers to the complex numbers; the real numbers \R" areimplicit); in the second column we applied one; in the third column we applied three,
since two of the three types (unitarity, symmetry, reality) imply the third. (The
corresponding groups with unit determinant, when distinct, are given in brackets.)T h e s es q u a r em a t r i c e sa r eo fs i z en ,n
++n , 2n, or 2n ++2n , as indicated. n +and
n refer to the number of positive and negative eigenvalues of the metric or .
O(n) diers from SO(n) by including \parity"-type transformations, which can't be
40 I. GLOBAL
obtained continuously from the identity. (SSp(2n) is the same as Sp(2n).) For this
reason, and also for studying \topological" properties, for nite transformations itis sometimes more useful to work directly with the group elements g, rather than
parametrizing them in terms of algebra elements as g=e
iG. U(n) diers from SU(n)
(and similarly for GL(n) vs. SL(n)) only by including a U(1) group that commutes
with the SU(n): Although U(1) is noncompact (it consists of just phase transforma-tions), a compact form of it can be used by requiring that all \charges" are integers
(i.e., all representations transform as
0=eiq for group parameter ,w h e r eqis an
integer dening the representation).
Of these groups, the compact ones are just SU(n), SO(n) (and O(n)), and USp(2n)
(all with n =0). The compact groups have an interesting interpretation in terms of
various number systems: SO(n) is the unitary group of n n matrices over the real
numbers, SU(n) is the same for the complex numbers, and USp(2n) is the same for
the quaternions. (Similar interpretations can be made for some of the noncompactgroups.) The remaining compact Lie groups that we didn't discuss, the \exceptional"
groups, can be interpreted as unitary groups over the octonions. (Unlike the classical
groups, which form innite series, there are only ve exceptional compact groups,because of the restrictions following from the nonassociativity of octonions.)
5. Tensor notation
Although historically group representations have usually been taught in the no-
tation where an m-component representation of a group dened by n n matrices is
represented by an m-component vector, carrying a single index with values 1 to m,
a much more convenient and transparent method is \tensor notation", where a gen-
eral representation carries many indices ranging from 1 to n, with certain symmetries(and perhaps tracelessness) imposed on them. (Tensor notation for a covering group
is generally known as \spinor notation" for the corresponding orthogonal group: See
subsection IC5.) This notation takes advantage of the property described above forexpressing arbitrary representations in terms of direct products of vectors. In termsof transformation laws, it means we need to know only the dening representation,
since the transformation of this representation is applied to each index. There are
at most four vector representations, by taking the dual and complex conjugate; weuse the corresponding index notation. Then the group constraints simply state the
invariance of the group metrics (and their complex conjugates and inverses), which
thus can be used to raise, lower, and contract indices:
B. INDICES 41
volume: Special:I1:::In
metric:8
<
:hermitian: Unitary:
(anti)symmetric:Orthogonal:
Symplectic:.
IJ
IJ
IJ
reality:Real:
pseudoreal (*):.
IJ
.
IJ
As a result, we have relations such as
hIjJi=IJor
IJ;h.
IjJi=.
IJ
We also dene inverse metrics satisfying
KIKJ=
KI
KJ=.
KI.
KJ=I
J
(and similarly for contracting the second index of each pair). Therefore, with uni-
tarity/(pseudo)reality we can ignore complex conjugate representations (and dottedindices), converting them into unconjugated ones with the metric, while for orthogo-nality/symplecticity we can do the same with respect to raising/lowering indices:
Unitary: .
I=.
IJ J
Orthogonal: I=IJ J
Symplectic: I=
IJ J
Real: .
I=.
IJ J
pseudoreal (*): .
I=
.
IJ J
For the real groups there is also the constraint of reality on the dening representation:
.
I( I)* = .
I.
IJ J
Excercise IB5.1
As an example of the advantages of index notation, show that SSp is the sameas Sp. (Hint: Write one in the denition of the determinant in terms of
's
by total antisymmetrization, which then can be dropped because it is enforcedby the other . One can ignore normalization by just showing detM =detI .)
For SO(n
+,n ), there is a slight modication of a sign convention: Since then
indices can be raised and lowered with the metric, I:::is usually dened to be the
result of raising indices on I:::,w h i c hm e a n s
12:::n=1)12:::n=det =( 1)n
42 I. GLOBAL
ThenI:::should be replaced with ( 1)n I:::in the equations of subsection IB3: For
example,
J1:::JnI1:::In=( 1)n I1
[J1In
Jn]
We now give the simplest explicit forms for the dening representations of the
classical groups. The most convenient notation is to label the generators by a pair
of fundamental indices, since the adjoint representation is obtained from the directproduct of the fundamental representation and its dual (i.e., as a matrix labeled by
row and column). The simplest example is GL(n), since the generators are arbitrary
matrices. We therefore choose as a basis matrices with a 1 as one entry and 0'severywhere else, and label that generator by the row and column where the 1 appears.
Explicitly,
GL(n): (G
IJ)KL=L
IJ
K)GIJ=jJihIj
This basis applies for GL(n,C) as well, the only dierence being that the coecients
inG=IJGJIare complex instead of real. The next simplest case is U(n): We can
again use this basis, although the matrices GIJare not all hermitian, by requiring
thatIJbe a hermitian matrix. This turns out to be more convenient in practice
than using a hermitian basis for the generators. A well known example is SU(2),where the two generators with the 1 as an o-diagonal element (and 0's elsewhere)
are known as the \raising and lowering operators" J
, and are more convenient than
their hermitian parts for purposes of contructing representations. (This generalizesto other unitary groups, where all the generators on one side of the diagonal are
raising, all those on the other side are lowering, and those along the diagonal give the
maximal Abelian subalgebra, or \Cartan subalgebra".)
Representations for the other classical groups follow from applying their deni-
tions to the GL(n) basis. We thus nd
SL(n): (G
IJ)KL=L
IJ
K 1
nJ
IL
K)GIJ=jJihIj 1
nJ
IjKihKj
SO(n): (GIJ)KL=K
[IL
J])GIJ=j[IihJ]j
Sp(n): (GIJ)KL=K
(IL
J))GIJ=j(IihJ)j
As before, SL(n,C) and SU(n) use the same basis as SL(n), etc. For SO(n) and Sp(n)
we have raised and lowered indices with the appropriate metric (so SO(n) includes
SO(n +,n )). For some purposes (especially for SL(n)), it's more convenient to impose
tracelessness or (anti)symmetry on the matrix , and use the simpler GL(n) basis.
Excercise IB5.2
Our normalization for the generators of the classical groups is the simplest,
and independent of n (except for subtracting out traces):
B. INDICES 43
aFind the commutation relations of the generators (structure constants) for the
dening representation of GL(n) as given in the text. Note that the values of
all the structure constants are 0, i. Show that
cD=1
(see subsection IB2).
bConsider the GL(m) subgroup of GL(n) (m <n) found by restricting the range
of the index of the above dening representation. Show the structure con-
stants are the same as those given by starting with the above representationof GL(m).
cFind the structure constants for SO(n) and Sp(n).
dDirectly evaluate k
DcA(=ijGiGj) for SL(n), SO(n), and Sp(n), and compare
withcDdA=dD.
Excercise IB5.3
At e n s o rt h a tp o p su pi nv a r i o u sc o n t e x t si s
dijk=tr(GifGj;Gkg)
It takes a very simple form in terms of dening indices:
aShow that for SU(n) this tensor is determined to be, up to an overall normal-
ization (that depends on the representation),
tr
GI1J1
GI2J2;GI3J3
[(231)+(312)] 2
n[(132)+(213)+(321)]+4
n2(123)
(ijk)J1
IiJ2
IjJ3
Ik
from just the total symmetry of dijk(andGII= 0), since the only invariant
tensor available is J
I.( I fIJ:::were used, IJ:::would also be required, to
balance the number of subscripts and superscripts; but their product can be
expressed in terms of just 's also.)
bCheck this result by using the explicit G's for the dening representation, and
determine the proportionality constant for that representation.
With the exception of the \spinor" representations of SO(n) (to be discussed
in subsection IC5, section IIA, and subsection XC1), general representations can be
obtained by reducing direct products of the dening representations. This means theycan be described by objects with multiple indices (up/down, dotted/undotted), where
each index is that of a dening representation, and satisfying various (anti)symmetry
and tracelessness conditions on the indices.
44 I. GLOBAL
Excercise IB5.4
Consider the representations of SU(n) obtained from the symmetric and an-
tisymmetric part of the direct product of two dening representations. For
simplicity, one can work with the U(n) generators, since the U(1) pieces will
appear in a simple way.
aUsing tensor notation for the generators ( GIJ)KLMN, nd their explicit rep-
resentation for these two representations.
bBy evaluating the trace, show that the Dynkin index for the two cases is
ca=n 2;cs=n+2
cShow the sum of these two is consistent with the argument at the end of
subsection IB2. Show each case is consistent with n=2, and the antisymmetriccase with n=3, by relating those cases to the singlet, dening, and adjoint
representations.
REFERENCES
1
H. Georgi, Lie algebras in particle physics: from isospin to unied theories
(Benjamin/Cummings, 1982):best book on Lie groups; unlike other texts, covers not only more powerful Cartan-Weyl methods, but also more useful tensor methods; also has useful applications tononrelativistic quark model.
2S. Helgason, Dierential geometry and symmetric spaces (Academic, 1962);
R. Gilmore, Lie groups, Lie algebras, and some of their applications (Wiley, 1974):
noncompact classical Lie groups.
C. REPRESENTATIONS 45
::::::::::::::::::: :::::::::::::::::::
::::::::::::::::::: C. REPRESENTATIONS :::::::::::::::::::
We now consider some of the more useful representations, as explicit examples
of the results of the previous section. In particular, we consider symmetries of the
quark model.
1. More coordinates
We began our \review" of group theory by looking at how symmetries were rep-
resented on coordinates. We now return to coordinates as a special case (particular
representation) of the general results of the previous section. The idea is that the
coordinates themselves are already a representation of the group, and the wave func-tions are functions of these coordinates. For example, for ordinary rotations we use
wave functions that depend on position or momentum, which transforms as a vec-
tor. (This is not always the case: For example, in our description of the conformal
group the usual space and time coordinates transformed nonlinearly, and not just
by multiplication by constant matrices unless the extra two coordinates were intro-duced.) This is the basic distinction between classical mechanics and classical eld
theory: Mechanics uses the coordinates themselves as the basic variables, while eld
theory uses functions of the coordinates. (Similarly, in quantum mechanics the wavefunctions are functions of the coordinates, while in quantum eld theory the wave
functions are \functionals" of functions of the coordinates.)
In general, the construction of such a \coordinate representation" starts with a
given matrix representation (usually nite dimensional) ( G
i)IJand then denes a
new representation
^Gi=qI(Gi)IJpJ;[pI;qJg=J
I;[q;qg=[p;pg=0
for some objects qandp, which are interpreted as either coordinates and their con-
jugate momenta (up to a factor of i), or as creation and annihilation operators: The
latter nomenclature is used when the boundary conditions allow the existence of a
statej0icalled the \vacuum", satisfying pj0i= 0, so we can dene the other states
as functions of qacting onj0i. (If the coordinates are fermionic, the distinction is
moot, since by the usual Taylor expansion the Hilbert space is nite dimensional. See
excercise IA2.3.) It is easy to check that ^Gisatisfy the same commutation relations
asGi. In particular, if the matrices are in the adjoint representation, qican be inter-
preted as the group coordinates themselves: This follows from considering the action
of an innitesimal transformation on the group element g(q)=eiqiGi(or just the Lie
algebra element G(q)=qiGi).
46 I. GLOBAL
If we write these results in bra/ket notation, since
[^Gi;qI]=qJ(Gi)JI;[^Gi;pI]= (Gi)IJpJ
it is more natural to look at the action on bras:
hqj=qIhIj;jpi=jIipI) ^Gihqj=hqjGi;^Gijpi= Gijpi
Note that this vector space is coordinate space itself, not the space of functions of
the coordinates; it is the same space on which Giis dened. (Of course, ^Giis dened
on arbitrary functions of the coordinates; it has a reducible representation bigger
than (Gi)IJ. Eectively, ( Gi)IJis represented on the space of functions linear in the
coordinates.) Then, for example
^G1^G2hqj=^G1hqjG2=hqjG1G2
is obviously equivalent, while (ignoring any extra signs for fermions)
^G1^G2jpi= ^G1G2jpi= G2^G1jpi=G2G1jpi
at least gives an equivalent result for the commutator algebra [ ^G1;^G2]. This is the
expected result for the dual representation Gi! GT
i.
Interesting examples are given by using the dening representation for G.F o r
example, the commonly used oscillator representation for U(n) is
U(n):GIJ=ayJaI;[aI;ayJg=J
I
where the oscillators can be bosonic or fermionic. For the SO and Sp cases, because
we can raise and lower indices, and because of the (anti)symmetry on the indices, the
interesting possibility arises to identify the coordinates with their momenta, with the
statistics appropriate to the symmetry:
Sp(n):GIJ=1
2z(IzJ);[zI;zJ]=
IJ
SO(n):GIJ=1
2
[I
J];f
I;
Jg=IJ
For SO(n) the representation is nite dimensional because of the Fermi-Dirac statis-
tics, and is called a \Dirac spinor" (and
the \Dirac matrices"). If the opposite
statistics are chosen, the coordinates and momenta can't be identied: For example,
bosonic coordinates for SO(n) give the usual spatial rotation generators GIJ=x[I@J].
Excercise IC1.1
Use this bosonic oscillator representation for U(2)=SU(2)
U(1), and use the
C. REPRESENTATIONS 47
SU(2) subgroup to describe spin. Show that the spin s(the integer or half-
integer number that denes the representation) itself has a very simple ex-
pression in terms of the U(1) generator. Show this holds in the quantummechanical case (by interpreting the bracket as the quantum commutator),giving the usual s(s+1) for the sum of the squares of the generators (with ap-
propriate normalization). Use this result to show that these oscillators, actingon the vacuum state, can be used to construct the usual states of arbitraryspins.
Excercise IC1.2
Considering SO(2n), divide up
Iinto pairs of canonical (and complex) con-
jugatesa1=(
1+i
2)=p
2, etc., sofa;ayg= 1. Write the SO(2n) generators
in terms of aa,ayay,a n daya. Show that the aya's by themselves generate a
U(n) subgroup. Decompose the Dirac spinor into U(n) representations. Show
that the product of all the
's is related to the U(1) generator, and commutes
with all the SO(2n) generators. Show that the states created by even or oddnumbers of a
y's on the vacuum don't mix with each other under SO(2n), so
the Dirac spinor is reducible into two \Weyl spinors".
2. Coordinate tensors
Many groups can be represented on coordinates. Depending on the choice of
coordinates, the coordinates may transform nonlinearly (i.e., as a realization, not arepresentation), as for the D-dimensional conformal group in terms of D (not D+2)coordinates. However, given the nonlinear transformation of the coordinates, thereare always representations other than the dening one (scalar eld) that we can im-mediately write down (such as the adjoint). We now consider such representations:These are useful not only for the spacetime symmetries we have already considered,
but also for general relativity, where the symmetry group consists of arbitrary coor-
dinate transformations. Furthermore, these considerations are useful for describingcoordinate transformations that are not symmetries, such as the change from Carte-sian to polar coordinates in nonrelativistic theories.
When applied to quantum mechanics, we write the action of a symmetry on a
state as =iG (or
0=eiG ), but on an operator as A=i[G;A]( o rA0=
eiGAe iG). In classical mechanics, we always write A=i[G;A] (since classical
objects are identied with quantum operators, not states). However, if G=m@mis
a coordinate transformation (e.g., a rotation) and is a scalar eld, then in quantum
48 I. GLOBAL
notation we can write
(x)=[G;]=G=m@m (0=eGe G=eG)
since the derivatives in Gjust dierentiate . (For this discussion of coordinate
transformations we switch to absorbing the i's into the generators.) The coordinate
transformation Ghas the usual properties of a derivative:
[G;f(x)] =Gf)Gf1f2=[G;f 1f2]=(Gf1)f2+f1Gf2
eGf1f2=eGf1f2e G=(eGf1e G)(eGf2e G)=(eGf1)(eGf2)
and similarly for products of more functions.
The adjoint representation of coordinate transformations is a \vector eld" (in the
sense of a spatial vector), a function that has general dependence on the coordinates
(like a scalar eld) but is also linear in the momenta (as are the Poincar e generators):
G=m(x)@m;V =Vm(x)@m)V=[G;V]=(m@mVn Vm@mn)@n
)Vm=n@nVm Vn@nm
The same result follows if we use the Poisson bracket instead of the quantum me-
chanical commutator, replacing @mwithipmin bothGandV.
Finite transformations can also be expressed in terms of transformed coordinates
themselves, instead of the transformation parameter:
(x)=e m@m0(x)=0(e m@mx)
as seen, for example, from a Taylor expansion of 0,u s i n ge G0=e G0eG.W e t h e n
dene
0(x0)=(x))x0=e m@mx
This is essentially the statement that the active and passive transformations cancel.
However, in general this method of dening coordinate transformations is not con-
venient for applications: When we make a coordinate transformation, we want toknow
0(x). Working with the \inverse" transformation on the coordinates, i.e., our
originale+G,
~xe+m@mx)0(x)=eG(x)=(~x(x))
So, for nite transformations, we work directly in terms of ~ x(x), and simply plug this
intoin place ofx(x!~x(x)) to nd0as a function of x.
C. REPRESENTATIONS 49
Similar remarks apply for the vector, and for derivatives in general. We then use
x0=e Gx)@0=e G@eG
where@0=@=@x0,s i n c e@0x0=@x=. This tells us
Vm(x)@m=e GV0m(x)@meG=V0m(x0)@0
m
orV0(x0)=V(x). Acting with both sides on x0m,
V0m(x0)=Vn(x)@x0m
@xn
On the other hand, working in terms of ~ xis again more convenient: Since for ~@=@=@~x
~@m=eG@me G="@~x(x)
@x 1#
mn
@n
we have, for example,
(Vm@m)0(x)=eG(Vm@m)(x)=Vm(~x(x))"@~x(x)
@x 1#
mn
@n(~x(x))
which yields an explicit expression for the transformed elds.
A \dierential form" is dened as an innitesimal W=dxmWm(x). Its transfor-
mation law under coordinate transformations, like that of scalar and vector elds, is
dened byW0(x0)=W(x). For any vector eld V=Vm(x)@m,VmWmtransforms as
a scalar, as follows from the \chain rule" d=dx0m@0
m=dxm@m. Explicitly,
W0
m(x0)=Wn(x)@xn
@x0m
or in innitesimal form
Wm=n@nWm+Wn@mn
Thus a dierential form is dual to a vector, at least as far as the matrix part of coor-
dinate transformations is concerned. They transform the same way under rotations,because rotations are orthogonal; however, more generally they transform dierently,
and in the absence of a metric there is not even a way to relate the two by raising or
lowering indices.
Higher-rank dierential forms can be dened by antisymmetric products of the
above \one-forms". These are useful for integration: Just as the line integralRW=Rdx
mWmis invariant under coordinate transformations by denition (as long as we
choose the curve along which the integral is performed in a coordinate-independent
50 I. GLOBAL
way), so is a totally antisymmetric Nth-rank tensor (\ N-form")Wm1mNintegrated
on anN-dimensional subspace as
Z
dxm1dxmNWm1mN:W0
m1mN(x0)=Wp1pN(x)@xp1
@x0m1@xpN
@x0mN
w h e r et h es u r f a c ee l e m e n t dxm1dxmNis interpreted as antisymmetric. (The signs
come from switching initial and nal limits of integration, as prescribed by the \ori-entation" of the hypersurface.) This is clear if we rewrite the integral more explicitlyin terms of coordinates
ifor the subspace: Then
Z
dxm1dxmNWm1mN(x)=Z
di1diNcWi1iN()=Z
dNi1iNcWi1iN()
where
cWi1iN()=@xm1
@i1@xmN
@iNWm1mN(x)
is the result of a coordinate transformation that converts Nof thex's to's, an
interpretation of the functions x() that dene the surface. Then any coordinate
transformation on x!x0(not on) will leavecW() invariant. In particular, if the
subspace is the full space, so we can look directly atR
dNxm1mNWm1mN,w es e e
that a coordinate transformation generates from WanN-dimensional determinant
exactly canceling the Jacobian resulting from changing the integration measure dNx.
Excercise IC2.1
For all of the following, use the exponential form of the nite coordinate trans-formation: Show that any (local) function of a scalar eld (without explicit x
dependence additional to that in the eld) is also a scalar eld (i.e., satisesthe same coordinate transformation law). Show that the transformation lawof a vector eld or dierential form remains the same when multiplied by ascalar eld (at the same x). Show that V=V
m@mis a scalar eld for any
scalar eld and vector eld V. Show that [ V;W ] is a vector eld for any
vector elds VandW.
Excercise IC2.2
Examine nite coordinate transformations for integrals of dierential formsin terms of ~ xrather than x
0. Find the explicit expression for W0(x)i nt e r m s
ofW(~x(x)), etc., and use this to show invariance:
Z
dxm1dxmNW0
m1mN(x)=Z
d~xm1d~xmNWm1mN(~x)
=Z
dxm1dxmNWm1mN(x)
C. REPRESENTATIONS 51
where in the last step we have simply substituted ~ x!xas a change of integra-
tion variables. Note that, using the ~ xform of the transformation rather than
x0, the transformation generates the needed Jacobain, rather than canceling
one.
From the above transformation law, we see that the curl of a dierential form is
also a dierential form:
@0
[m1W0
m2mN](x0)=@0
[m1(@0
m2xp2)(@0
mN]xpN)Wp2pN(x)
=[@[p1Wp2pN](x)](@0
m1xp1)(@0
mNxpN)
because the curl kills @0@0xterms that would appear if there were no antisymmetriza-
tion. Objects that transform \covariantly" under coordinate transformations, without
such higher derivatives of x(orin the other notation), like scalars, vectors, dieren-
tial forms and their products, are called (coordinate) \tensors". Getting derivatives of
tensors to come out covariant in general requires special elds, and will be discussed
in chapter IX. An important application of the covariance of the curl of dierentialforms is the generalized Stokes' theorem (which includes the usual Stokes' theorem
and Gauss' law as special cases):
Z
dx
m1dxmN+1 1
(N+1)!@[m1Wm2mN+1]=I
dxm1dxmNWm1mN
where the second integral is over the boundary of the space over which the rst is
integrated. (We use the symbol \H
" to refer to boundary integrals, including those
over contours, which are closed boundaries of 2D surfaces.)
3. Young tableaux
We now return to our discussion of nite-dimensional representations. In the
previous section we gave the machinery for describing them using index notation,but examined only the dening representation in detail. Now we analyze general
irreducible representations.
All the irreducible nite-dimensional representations of the groups SU(N) can be
described by tensors with lower N-valued indices with various (anti)symmetrizations.(An upper index can be replaced with N 1 lower indices by using the Levi-Civita
tensor.) Although detailed calculations require explicit use of these indices, three
properties can be more conveniently discussed pictorially: (1) the (anti)symmetriesof the indices, (2) the dimension (number of independent components) of the repre-
sentation, and (3) the reduction of the direct product of two representations (which
irreducible representations result, and how many of each).
52 I. GLOBAL
A \Young tableau" is a picture representing an irreducible representation in terms
of boxes arranged in a regular grid into rows and columns, such that the columns are
aligned at the top, and their depths are nonincreasing to the right: for example,
Each box represents an index, with antisymmetry among indices in any column, and
symmetry among indices in any row. More precisely, since one can't simultaneously
have these symmetries and antisymmetries, it corresponds to the result of taking anyarbitrary tensor with that many indices, rst symmetrizing the indices in each row,
and then antisymmetrizing the indices in each column (or vice versa; symmetrizing
and then antisymmetrizing and then symmetrizing again gives the same result asskipping the rst symmetrization, etc.). This gives a simple way to classify and
symbolize each representation. (We can denote the singlet representation, which has
no boxes, by a dot.) Note that the deepest column should have no more than N 1
boxes for SU(N) because of the antisymmetry.
To calculate the dimension of the representation for a given tableau, we use the
\factors over hooks" rule: (1) Write an \N" in the box in the upper-left corner, and
ll the rest of the boxes with numbers that decrease by 1 for each step down and
increase by 1 for each step to the right. (2) Draw (or picture in your mind) a \hook"
for each box | a \ " with its corner in the box and lines extending right and downout of the tableau. (3) The dimension is then given by the formula
dimension =Y
each boxinteger written there
# boxes intersected by its hook
For the previous example, we nd (listing boxes rst down and then to the right)
N
8N 1
6N 2
3N 3
1N+1
6N
4N 1
1N+2
4N+1
2N+3
3N+2
1N+4
1
The direct product of two Young tableaux A
B is analyzed by the following rules:
First, label all the boxes in B by putting an \a" in each box in the top row, \b" in
the second row, etc. Then, take the following steps in all possible ways to nd theYoung tableaux resulting from the direct product: (1) Add all the \a" boxes from B
to the right side and bottom of A, then \b" to the right and bottom of that, etc., to
make a new Young tableaux. Any two tableaux constructed in this way with the samearrangement of boxes but dierent assignment of letters are considered distinct, i.e.,
multiple occurences of the same representation in the direct product. (2) No more
than 1 \a" can be in any column, and similarly for the other letters. (3) Reading from
C. REPRESENTATIONS 53
right to left, and then from top to bottom (i.e., like Hebrew/Arabic), the number of
a's read should always be the number of b's, b's c's, etc. For example,
a
ba=baaaa
ba
ba
Note that A
B always gives the same result as B
A, but one way may be simpler
than the other. For a given value of N, a column of N boxes is equivalent to none(again by antisymmetry), while more than N boxes in a column gives a vanishing
tableau.
Excercise IC3.1
Calculate
Check the result by nding the dimensions of all the representations and
adding them up.
These SU(N) tableaux also apply to SL(N): Only the reality properties are dif-
ferent. Similar methods can be applied to USp(2N) (or Sp(2N)), but tracelessness(with respect to the symplectic metric) must be imposed in antisymmetrized indices,so these trace pieces must be separated out when considering the above rules. (I.e.,consider USp(2N) SU(2N).) Similar remarks apply to SO(N), which has a symmet-
ric metric, but there are also \spinor" representations (see below). The additional
irreducible representations then can be constructed from taking direct products ofthe above with the smallest spinors, and removing the \gamma-matrix" traces.
4. Color and
avor
We now consider the application of these methods to \internal symmetries" (those
that don't act on the coordinates) in particle physics. The symmetries with experi-
mental conrmation involve only the unitary groups (U and SU) of small dimension.However, we will nd later that larger unitary groups can be useful for approxima-tion schemes. (Also, larger unitary and other groups continue to be investigated forunication and other purposes, which we consider in later chapters.)
The \Standard Model" describes all of particle physics that is well conrmed
experimentally (except gravity, which is not understood at the quantum level). Itincludes as its \fundamental" particles: (1) the spin-1/2 quarks that make up the
observed strongly interacting particles, but do not exist as asymptotic states, (2) the
weakly interacting spin-1/2 leptons, (3) the spin-1 gluons that bind the quarks to-gether, which couple to the charges associated with SU(3) \color" symmetry, but also
54 I. GLOBAL
are not asymptotic, (4) the spin-1 particles that mediate the weak and electromag-
netic interactions, which couple to SU(2)
U(1) \
avor", and (5) the yet unobserved
spin-0 Higgs particles that are responsible for all the masses of these weakly interact-ing particles. These particles, along with their masses (in GeV) and (electromagnetic)charges (Q=Q+Q), are:
s=1
2
color:! quark (3) lepton (1)
avor (Q)(Q=1
6) (Q= 1
2)
+1
2u(.005)e(<10 8)
1
2d(.009)e(.0005109991)
+1
2c(1.4)(<.00017)
1
2s(.18)(.10565839)
+1
2t(174)(<.0182)
1
2b(4.7)(1.7770)s=1
color:! gluon electroweak
avor (Q) (8) (1)
0g(0)
(<210 25)
0 Z(91.187)
1 W(80.4)
s=0
(Q=0 )H (>77.5)
The quark masses we have listed are the \current quark masses", the eective masses
when the quarks are relativistic with respect to their hadron, and act as almost free.Nonrelativistic quark models use instead the \constituent quark masses", which in-clude potential energy from the gluons. This extra potential energy is about .30GeV per quark in the lightest mesons, .35 GeV in the lightest baryons; there is alsoa contribution to the binding energy from spin-spin interaction. Unlike electrody-namics, where the potential energy is negative because the electrons are free at largedistances, where the potential levels o (the top of the \well"), in chromodynamicsthe potential energy is positive because the quarks are free at high energies (shortdistances, the bottom of the well), and the potential is innitely rising. Masslessnessof the gluons is implied by the fact that no colorful asymptotic states have ever beenobserved. We have divided the spin-1/2 particles into 3 \families" with the same
quantum numbers (but dierent masses). Within each family, the quarks are similar
to the leptons, except that: (1) the masses and average charges ( Q) are dierent,
(2) the quarks come in 3 colors, while the leptons are colorless, and (3) the neutrinos,to within experimental error, are massless, so they have half as many componentsas the massive fermions (1 helicity state each, instead of 2 spin states each). Thismeans that each lepton family has 1 SU(2) doublet and 1 SU(2) singlet. For symme-try (and better, quantum mechanical, reasons to be explained later), we also assumethe quarks have 1 SU(2) doublet, but therefore 2 SU(2) singlets. (Some experimentshave indicated small masses for neutrinos: This would require generalization of the
C. REPRESENTATIONS 55
Standard Model, such as models with parity broken by interactions. Some examples
of such theories will be discussed in subsection IVB4.)
We rst look at the color group theory of the physical states, which are color
singlets. The fundamental unobserved particles are the spin-1 \gluons", described by
the Yang-Mills gauge elds, and the spin-1/2 quarks. Suppressing all but color indices,
we denote the quark states by qi, and the antiquarks by qyi, where the indices are those
of the dening representation of SU(n), and its complex conjugate. The quarks alsocarry a representation of a \
avor" group, unlike the gluons. The simplest
avorfulstates are those made up of only (anti)quarks, with indices completely contracted byone factor of an SU(n) group metric: From the \U" of SU(n), we can contract dening
indices with their complex conjugates, giving the \mesons", described by q
yiqi(quark-
antiquark), which are their own antiparticles. From the \S" of SU(n), we have the\baryons", described by
i1:::inqi1:::qin(n-quark), and the antibaryons, described by
the complex conjugate elds. All other colorless states made of just (anti)quarkscan be written as products of these elds, and therefore considered as describingcomposites of them. Thus, we can approximate the ground states of the mesons by
q
yi(x)qi(x), which describe spins 0 and 1 because of the various combinations of spins
(from1
2
1
2=01). The rst excited level will then be described by qy$
@q(where
A$
@BA@B (@A)B
and picks out the relative momentum of the two quarks): It includes spins 0, 1, and 2,
etc., where each derivative introduces orbital angular momentum 1. (Similar remarksapply to baryons.) We can also have
avorless states made from just gluons, called
\glueballs": The ground states can be described by F
ijFji,w h e r ee a c h Fis a gluon
state (in the adjoint representation of SU(n)), and includes spins 0 and 2 (from thesymmetric part of 1
1). Because of their
avor multiplets and (electroweak) inter-
actions, many mesons and baryons corresponding to such ground and excited stateshave been experimentally identied, while the glueballs' existence is still uncertain.Actually, quarks and gluons can almost be observed independently at high energies,where the \strong" interaction is weak: The energetic particle appears as a \jet" |
a particle of high energy accompanied by particles of much lower energy (perhaps
too small to detect) in color-singlet combinations. (Depending on the available decaymodes, the jet might not be observed until after decaying, but still within a smallangle of spread.)
We now look at the
avor group theory of the physical hadronic states. In contrast
to the previous paragraph, we now suppress all but the
avor indices. Mesons M
ij=
56 I. GLOBAL
qyiqjare thus in the adjoint representation of
avor U(m) ( m
m,w h e r emis the
dening representation and mits complex conjugate), for both the spin-0 and the
spin-1 ground states. The baryons are more complicated: For simplicity we considerSU(3) color, which accurately describes physics at observed energies. Then the colorstructure described above results in total symmetry in combined
avor and Lorentzindices (from the antisymmetry in the color indices, and the overall antisymmetry forFermi-Dirac statistics). Thus, for the 3-quark baryons, the Young tableaux
for SU(m)
avor are accompanied by the same Young tableaux for spin indices: Innonrelativistic notation, the rst tableau, being totally antisymmetric in
avor in-dices, is also totally antisymmetric in the three two-valued spinor indices, and thusvanishes. Similarly, the last tableau describes spin 3/2 (total symmetry in both types
of indices), while the middle one describes spin 1/2. Since only 3
avors of quarks
have small masses compared to the hadronic mass scale, hadrons can be most conve-niently grouped into
avor multiplets for SU(3)
avor: The ground states are then,in terms of SU(3)
avor multiplets, 8 1 for the pseudoscalars, 8 1 for the vectors, 8
for spin 1/2, and 10 for spin 3/2.
Excercise IC4.1
What SU(
avor) Young tableaux, corresponding to what spins, would we havefor mesons and baryons if there were 2 colors? 4 colors?
However, the diering masses of the dierent
avors of quarks break the SU(3)
avor symmetry (as does the weak interaction). In particular, the mass eigenstatestend to be pure states of the various combinations of the dierent
avors of quarks,rather than the linear combinations expected from the
avor symmetry. Specically,the linear combinations predicted by an 8 1 separation for mesons (trace and traceless
pieces of a 3
3 matrix) are replaced with particles that are more accurately described
by a particular
avor of quark bound to a particular
avor of antiquark. (This isknown as \ideal mixing".) The one exception is the lighest mesons (pseudoscalars),which are more accurately described by the 8 1 split, for this restriction to the 3
lighter
avors of quarks, but the mass of the singlet diers from that naively expected
from group theory or nonrelativistic quark models. (This is known as the \U(1)
problem".) The solution is probably that the singlet mixes strongly with the lightestpsuedoscalar glueball (described by tr
abcdFabFcd); the mass eigenstates are linear
combinations of these two elds with the same quantum numbers. In any case, themost convenient notation for labeling the entries of the matrix M
ijrepresenting the
C. REPRESENTATIONS 57
various meson states for any particular spin and angular momentum of the quark-
antiquark combination is that corresponding to the choice we gave earlier for the
generators of U(n): Label each entry by a separate name, where the complex conjugateappears re
ected across the diagonal. These directly correspond to the combinationof a particular quark with a particular antiquark, and to the mass eigenstates, withthe possible exception of the entries along the diagonal for the 3 lightest
avors,where the mass eigenstates are various linear combinations. (However, the SU(2) ofthe 2 lightest
avors is only slightly broken by the quark masses, so in that case thecombinations are very close to the 3 1 split of SU(2).)
For example, for the lightest multiplet of mesons (spin 0, and relative angular mo-
mentum 0 for the quark and antiquark, but not all of which have yet been observed),we can write the U(6) matrix (for the 6
avors of the 3 known families)
M
ij=0
BBBBBBBB@uu ud ucus utub
du dd dc ds dt db
cu cd cccsctcb
su sd scssstsb
tu td tctstttb
bu bdbcbsbtbb1
CCCCCCCCA
=0
BBBBBBBB@ud c s t b
u
u (:1395700)D0(1:8646)K (:49368)T0B (5:279)
d+(")d D+(1:8693) K0(:49767)T+B0(5:279)
c D0(")D (")c(2:980)D
s(1:9685)T0
cB
c
sK+(")K0(")D+
s(")s T+
sB0
s
t T0T T0
c T
s tT
b
bB+(")B0(")B+
c B0
s T+
bb1
CCCCCCCCA
where (approximately)
u=1p
20(:1349764) +1
2[0(:9578) +(:5473)]
d= 1p
20+1
2(0+);s=1p
2(0 )
in terms of the mass eigenstates (observed particles), with masses again in GeV,
and ditto marks refer to the transposed entry. (We have neglected the important
58 I. GLOBAL
contribution from the glueball.) For the corresponding spin-1 multiplet,
fMij=0
BBBBBBBB@ud c s t b
u!
u (:7700)D*0(2:0067)K* (:8917)T*0B* (5:325)
d+(")!dD*+(2:0100) K*0(:8961)T*+B*0(5:325)
c D*0(")D* (")J= (3:09688)D*
sT*0
cB*
c
sK *+(")K*0(")D*+
s (1:019413)T*+
sB*0
s
t T*0T* T*0
c T*
s T *
b
bB *+(")B*0(")B*+
c B*0
sT*+
b(9:4604)1
CCCCCCCCA
where
!
u=1p
2[!(:7819) +0(:7700)];!d=1p
2(! 0)
(with ss=, ideal mixing, also approximate).
5. Covering groups
The orthogonal groups O(n +,n ) are of obvious interest for describing Lorentz
symmetry in spacetimes with n +space and n time dimensions, or conformal sym-
metry in spacetimes with n + 1s p a c ea n dn 1 time dimensions. This means we
should be interested in O(n) for n 6, and their \Wick rotations": transformations
that put in extra factors of ito change some signs on the metric. Coincidentally, these
are just the cases where the Lie algebras of the orthogonal groups are equivalent to
those of some algebras for smaller matrices. The smaller representation then can be
identied as the \spinor" representation of that orthogonal group. Since the \vec-tor", or dening representation space of the orthogonal group, itself is represented asa matrix with respect to the other group (i.e., the state carries two spinor indices),the other group may include certain phase transformations (such as 1) that cancel
in the transformation of the vector. The other group is then called the \covering"group for that orthogonal group, since it includes those missing transformations inits dening representation. (As a result, its group space also has a more interesting
topology, which we won't discuss here.)
One way to discover these covering groups is to rst count generators, then try
to construct explicitly the orthogonal metric on matrices. SO(n) has n(n 1)/2 gen-
erators (antisymmetric matrices), Sp(n) has n(n+1)/2 (symmetric), and SU(n) has
n
2 1 (traceless). (These are hermitian generators, since we applied reality or her-
miticity.) So, for some group SO(n), we look for another group that has the samenumber of generators. Then, if the new group is dened on m m matrices, we look
for conditions to impose on an m m matrix (not necessarily the adjoint) to get an
C. REPRESENTATIONS 59
n-component representation. This is easy to do by inspection for small n; for large n
it's easy to see that it can't work, since m will be of the order of n, and the simpleconstraints will give of the order of n
2components instead of n. We then construct
the norm of this matrix Mastr(MyM), which is just the sum of the absolute value
squared of the components, for SO(n), and the other orthogonal groups by Wick
rotation. (Wick rotation aects mainly the reality conditions on M.)
The identications for the Lie algebras are then:
SO(2) = U(1), SO(1,1) = GL(1)
SO(3) = SU(2) = SU*(2) = USp(2), SO(2,1) = SU(1,1) = SL(2) = Sp(2)SO(4) = SU(2)
SU(2), SO(3,1) = SL(2,C) = Sp(2,C), SO(2,2) = SL(2)
SL(2)
SO(5) = USp(4), SO(4,1) = USp(2,2), SO(3,2) = Sp(4)
SO(6) = SU(4), SO(5,1) = SU*(4), SO(4,2) = SU(2,2), SO(3,3) = SL(4)
Note that the Euclidean cases are all unitary, while the ones with (almost) equal
numbers of space and time dimensions are all real. There are also some similarrelations for the pseudoreal orthogonal groups:
SO*(2) = U(1), SO*(4) = SU(2)
SL(2), SO*(6) = SU(3,1), SO*(8) = SO(6,2)
The norm and conditions for an m-spinor of SO(n
+,n )a r e :
n ) 0 1 2 3
mnnorm symmetry :zT= reality :z*=
12z0z z0z(z0*=z0)
23zz
z zz
4z0z
0
00 zzTz
45zz
z(z
=0 )1
2z1
2(z)z
6zz
z1
2z
z
1
2(z)z
Note that in all but the 2D cases the norms are associated with determinants: For
D=3 and 4 the norm is given by the determinant, while for D=5 and 6 we use thefact that the determinant of an antisymmetric matrix is the square of the Pfaan.
Excercise IC5.1
Show that for D=5 zzandzz
give the same norm. (Hint: Consider
[
].)
Unfortunately, for SO(n) for larger n, the spinor is as least as large as, and usu-
ally larger than, the vector. In general, the spinor is like the \square root" of the
vector, in that the vector can be found by taking the direct product of two spinors.
60 I. GLOBAL
It is impossible to nd the spinor representation by taking direct products of vec-
tors. This situation occurs only for orthogonal groups: In all other classical groups,
all (nite-dimensional) representations are among those obtained from multiple di-
rect products of vectors. Furthermore, in those cases the \irreducible" representa-
tions (those that can't be divided into smaller representations) can be picked out by(anti)symmetrization, and by separating trace and traceless pieces (where traces are
taken with the group metrics). Fortunately, for the above cases of orthogonal groups,
we can perform the same construction starting with the spinor representations, sincethose are the \vectors" of non-orthogonal groups.
REFERENCES
1
J. Schwinger, On angular momentum, Quantum theory of angular momentum: a collec-
tion of reprints and original papers , eds. L.C. Biedenharn and H. Van Dam (Academic,
1965) p. 229:spin using spinor oscillators.
2Georgi, loc. cit. (IB).
3P.A.M. Dirac, P r o c .R o y .S o c . A117 (1928) 610.
4H. Weyl, Z. Phys. 56(1929) 330.
5M. Hamermesh, Group theory and its application to physical problems (Addison-Wesley,
1962):detailed discussion of Young tableaux.
6M. Gell-Mann, Phys. Lett. 8(1964) 214;
G. Zweig, preprints CERN-TH-401 and 412 (1964); Fractionally charged particles andSU
6,i nSymmetries and elementary particle physics , proc. Int. School of Physics \Ettore
Majorana", Erice, Italy, Aug.-Sept., 1964, ed. A. Zichichi (Academic, 1965) p. 192:quarks.
7Particle Data Group (in 1998, C. Caso, G. Conforto, A. Gurtu, M. Aguilar-Benitez, C.Amsler, R.M. Barnett, P.R. Burchat, C.D. Carone, O. Dahl, M. Doser, S. Eidelman, J.L.F e n g ,M .G o o d m a n ,C .G r a b ,D . E .G r o o m ,K .H a g i w a r a ,K . G .H a y e s ,J . J .H e r n andez,
K. Hikasa, K. Honscheid, F. James, M.L. Mangano, A.V. Manoha, K. M onig, H. Mu-
rayama, K. Nakamura, K.A. Olive, A. Piepke, M. Roos, R.H. Schindler, R.E. Shrock,M. Tanabashi, N.A. T ornqvist, T.G. Trippe, P. Vogel, C.G. Wohl, R.L. Workman, and
W.-M. Yao), http://pdg.lbl.gov/pdg.html (published biannually, including European
Phys. J. C3(1998) 1):
Review of Particle Physics, a.k.a. Review of Particle Properties, a.k.a. Rosenfeld tables;tables of masses, decay rates, etc., of all known particles, plus useful brief reviews.
A. TWO COMPONENTS 61
II. SPIN
Special relativity is simply the statement that the laws of nature are symmetric
under the Poincar e group. Free relativistic quantum mechanics or eld theory is then
equivalent to a study of the representations of the Poincar e group. Since the con-
formal group is a classical group, while its subgroup the Poincar e group is not, it is
easier to rst study the conformal group, which is sucient for nding the masslessrepresentations of the Poincar e group. The massive ones then can be found by di-
mensional reduction, which gives them in the same form as occurs in interacting eldtheories. In four spacetime dimensions we use the covering group of the conformalgroup, which is the easiest way to include spinors. These methods extend straight-forwardly to supersymmetry, a symmetry between fermions and bosons that includesthe Poincar e group.
:::::::::::::::::: ::::::::::::::::::
::::::::::::::::::
A. TWO COMPONENTS ::::::::::::::::::
Although we have already specialized to spacetime symmetries, we have consid-
ered arbitrary spacetime dimensions. We have also noted that many of the lower-dimensional Lie groups have special properties, especially with regard to coveringgroups. In this section we will take advantage of those features; specically, we ex-amine the physical case D=4, where the rotation group is SO(3)=SU(2), the Lorentzgroup is SO(3,1)=SL(2,C), and the conformal group is SO(4,2)=SU(2,2).
1. 3-vectors
The most important nontrivial Lie group in physics is the rotations in three
dimensions. It is also the simplest nontrivial example of a Lie group. This makesit the ideal example to illustrate the properties discussed in the previous chapter,as well as lay the groundwork for later discussions. We have already mentioned theorbital part of rotations, i.e., the representation of rotations on spatial coordinates.In this chapter we discuss the spin part; this is really the same as nding all (nitedimensional, unitary) representations.
A useful way to understand spin, or general representations of the rotation group
in three dimensions of space, is to consider properties of 2 2 matrices. (This way
also generalizes in a very simple way to relativity, in three space and one time dimen-sions.) Consider such matrices to be hermitian, which is natural from the quantummechanical point of view. Then they have four real components, one too many for a
62 II. SPIN
three-vector (but just right for a relativistic four-vector), so we restrict them to also
be traceless:
V=Vy;t r V =0
The simplest way to get a single number out of a matrix, besides taking the trace, is
to take the determinant. By expanding a general matrix identity to quadratic order
we nd an identity for 2 2 matrices
det(I+M)=etrln(I+M)) 2detM =tr(M2) (tr M )2
It is then clear that in our case det V is positive denite, as well as quadratic, so
we can dene the norm of this 3-vector as
jVj2= 2detV =tr(V2)
This can be compared easily with conventional notation by picking a basis:
V=1p
2V1V2 iV3
V2+iV3 V1
=~V~)det V = 1
2(Vi)2
where~are the Pauli matrices, up to normalization. As usual, the inner product
follows from the norm:
jV+Wj2=jVj2+jWj2+2VW
)VW=detV +det W det(V+W)=tr(VW)
Applying our previous identities for determinants to 2 2 matrices, we have
MCMTC=Id e tM ; M 1=CMTC(detM ) 1
where we now use the imaginary, hermitian matrix
C= 0
i i
0
If we make the replacement M!eMand expand to linear order in M, we nd
M+CMTC=It rM
This implies
tr V =0, (VC)T=VC
i.e., the tracelessness of Vis equivalent to the symmetry of VC. Furthermore, the
combination of the trace and determinant identities tell us
M2=Mt rM Id e tM)V2= Id e tV =I1
2jVj2
A. TWO COMPONENTS 63
Here by \V2" we mean the square of the matrix, while \ jVj2"= (Vi)2is the square
of the norm (neither of which should be confused with the component V2=Vi2
i.)
Again expressing the inner product in terms of the norm, we then nd
fV;Wg=(VW)I
Furthermore, since the trace of a commutator of two nite matrices is traceless, and
picks up a minus sign under hermitian conjugation, we can dene an outer product
(vectorvector = vector) by
[V;W ]=p
2iVW
Combining these two results,
VW =1
2(VW)I+1p
2iVW
In other words, the product of two traceless hermitian 2 2 matrices gives a real trace
piece, symmetric in the two matrices, plus an antihermitian traceless piece, antisym-
metric in the two. Thus, we have a simple relation between the matrix product, the
inner (\dot") product and the outer (\cross") product. Therefore, the cross product
is a special case of the Lie bracket, or commutator. This way of treating vectors
(except for factors of \ i") is basically Hamilton's \quaternions".
Excercise IIA1.1
Check this result in two ways:
aShow the normalization agrees with the usual outer product. Using only the
above denition of VW,a l o n gw i t hfV;Wg=(VW)I,s h o w
IjVWj2=( [V;W ])2= I[jVj2jWj2 (VW)2]
bUse components, with the above basis.
Excercise IIA1.2
Write an arbitrary two-dimensional vector in terms of a complex number as
V=1p
2(vx ivy).
aShow that the phase (U(1)) transformation V0=Veigenerates the usual
rotation. Show that for any two vectors V1andV2,V1*V2is invariant, and
identify its real and imaginary parts in terms of well known vector prod-
ucts. What kind of transformation is V!V*, and how does it aect these
products?
64 II. SPIN
bConsider two-dimensional functions in terms of z=1p
2(x+iy)a n dz*=
1p
2(x iy). Show by the chain rule that @z=1p
2(@x i@y). Write the real
and imaginary parts of the equation @z*V= 0 in terms of the divergence and
curl. (Then Vis a function of just z.)
cConsider the complex integral
Idz
2iV
where \H
" is a \contour integral": an integral over a closed path in the
complex plane dened by parametrizing dz=du(dz=du ) in terms of some
real parameter u.T h i s i s u s e f u l i f Vcan be Laurent expanded as V(z)=P1
n= 1cn(z z0)ninside the contour about a point z0there, since by con-
sidering circles z=z0+reiwe nd only the 1 =(z z0) term contributes.
Show that this integral contains as its real and imaginary parts the usual lineintegral and \surface" integral. (In two dimensions a surface element diersfrom a line element only by its direction.) Use this fact to solve Gauss' lawin two dimensions for a unit point charge as E=1=z.
Excercise IIA1.3
Consider electromagnetism in 2 2 matrix notation: Dene the eld strength
as a complex vectorF=p
2(E+iB). Write partial derivatives as the sum
of a (rotational) scalar plus a (3-)vector as @=1p
2I@t+r,w h e r e@t=@=@t
is the time derivative and ris the partial space derivatives written as a
traceless matrix. Do the same for the charge density and (3-)current jas
J= 1p
2I+j. Using the denition of dot and cross products in terms of
matrix multiplication as discussed in this section, show that the simple matrixequation@F= J, when separated into its trace and traceless pieces, and
its hermitian and antihermitian pieces, gives the usual Maxwell equations
rB=0;rE=;rE+@
tB=0;rB @tE=j
(Note: Avoid the Pauli -matrices and explicit components.)
2. Rotations
One convenience of representing three-vectors as 2 2 instead of 31 is that rota-
tions are easier to write. Since vectors are hermitian, we expect their transformationsto be unitary:
V
0=UVUy;Uy=U 1
A. TWO COMPONENTS 65
It is easily checked that this preserves the properties of these matrices:
(V0)y=(UVUy)y=V0;t r (V0)=tr(UVU 1)=tr(U 1UV)=tr(V)=0
Furthermore, it also preserves the norm (and thus the inner product):
det(V0)=det(UVU 1)=det(U)det(V)(det U ) 1=det V
Unitary 22 matrices have 4 parameters; however, we can elimimate one by the
condition
det U =1
This eliminates only the phase factor in U, which cancels out in the transformation
law anyway. Taking the product of two rotations now involves multiplying only 2 2
matrices, and not 3 3 matrices.
We can also write Uin exponential notation, which is useful for going to the
innitesimal limit:
U=eiG)Gy=G; tr G =0
This means that Gitself can be considered a vector. Rotations can be parametrized
by a vector whose direction is the axis of rotation, and whose magnitude is (1 =p
2)
the angle of rotation:
V0=eiGVe iG)V=i[G;V]= p
2GV
We also now see that the Lie bracket we previously identied as the cross product is
the bracket for the rotation group.
Excercise IIA2.1
Evaluate the elements of the matrix eiGin closed form for a diagonal generator
G. Generalize this result to arbitrary G. (Hint: Use rotational invariance.)
The hermiticity condition on Vcan also be expressed as a reality condition:
V=Vyand tr V =0)V*= CVC; (VC)* =C(VC)C
where \ * " is the usual complex conjugate. A similar condition for Uis
Uy=U 1and det U =1)U*=CUC
which is also a consequence of the fact that we can write Uin terms of a vector as
U=eiV. As a result, the transformation law for the vector can be written in terms
ofVCin a simple way, which manifestly preserves its symmetry:
(VC)0=UVU 1C=U(VC)UT
66 II. SPIN
Excercise IIA2.2
Write an arbitrary rotation in two dimensions in terms of the slope (dy=dx )o f
the rotation (the slope to which the x-axis is rotated) rather than the angle.
(This is actually more convenient to measure if you happen to have a ruler,
which you need to measure lengths anyway, but not a protractor.) This avoids
trigonometry, but introduces ugly square roots. Show that the square roots
can be eliminated by using the slope of half the angle of transformation as
the variable. Show the relation to the variables used in writing 3D rotations
in terms of 22m a t r i c e s .( H i n t :C o n s i d e r UandVCdiagonal.)
3. Spinors
Note that the mapping of SU(2) to SO(3) is two-to-one: This follows from the
factV0=VwhenUis a phase factor. We eliminated continuous phase factors from
Uby the condition det U = 1, which restricts U(2) to SU(2). However,
det(Iei)=e2i=1)ei=1
for 22 matrices. More generally, for any SU(2) element U, Uis also an element of
SU(2), but acts the same way on a vector; i.e., these two SU(2) transformations give
the same SO(3) transformation. Thus SU(2) is called a \double covering" of SO(3).
However, this second transformation is not redundant, because it acts dierently on
half-integral spins, which we discuss in the following subsections.
The other convenience of using 2 2 matrices is that it makes obvious how to
introduce spinors | Since a vector already transforms with two factors of U,w e
dene a \square root" of a vector that transforms with just one U:
0=U ) y0= yU 1
where is a two-component \vector", i.e., a 2 1 matrix. The complex conjugate of
a spinor then transforms in essentially the same way:
(C *)0=CU* *=U(C *)
Note that the antisymmetry of Cimplies that must be complex: We might think
that, sinceC * transforms in the same way as , we can identify the two consistently
with the transformation law. But then we would have
=C *=C(C *)* =CC* =
A. TWO COMPONENTS 67
Thus the representation is pseudoreal. The fact that C * transforms the same way
under rotations as leads us to consider the transformation
0=C *
Since a vector transforms the same way under rotations as y, under this transfor-
mation we have
V0=CV*C= V
which identies it as a re
ection.
Another useful way to write rotations on (like looking at VCinstead ofV)i s
( TC)0=( TC)U 1
This tells us how to take an invariant inner product of spinors:
0=U ; 0=U) ( TC)0=( TC)
In other words, Cis the \metric" in the space of spinors. An important dierence of
this inner product from the familiar one for three-vectors is that it is antisymmetric.
Thus, if andare anticommuting spinors,
TC= TCT =TC
where one minus sign comes from anticommutativity and another is from the anti-
symmetry of C. Thus, it makes sense to take the norm of an anticommuting spinor
as TC , which would vanish if were commuting. Of course, since rotations are
unitary, we also have the usual y as an invariant, positive denite, inner product.
4. Indices
The best way to discuss general spins is to use index notation, rather than matrix
notation. Then a spinor rotates as
0
=U
with two-valued indices =; . The inner product is dened by
= C=
where we have dened raising and lowering of indices by
= C; =C
68 II. SPIN
C= C= C=C= 0
i i
0
paying careful attention to signs. (In general, we x signs by using a convention of
contracting indices from upper-left to lower-right.) Then objects with many indicestransform as the product of spinors:
A
0
:::
=UU:::U
A:::
An innitesimal transformation is then a sum:
iA:::
=GA:::
+GA:::
+:::+G
A:::
This is also true for C, even though it is an invariant constant:
C0
=U
UC
=Cdet U =C
A more interesting case is the vector: The transformation law is
V0
=U
UV
whereVis the symmetric VCconsidered earlier (in contrast to the antisymmetric
C).
There is basically only one identity in index notation, namely
0=1
2C[C
]=CC
+C
C+C
C
The expression vanishes because it is antisymmetric in those indices, and thus the
indices must all have dierent values, but there are three two-valued indices. Anotherway to write this identity is to use the denition of C
as the inverse of C:
C
C
=
)CC
=
[
]
This tells us that antisymmetrizing in any pair of indices automatically contracts
(sums over) them: Contracting this identity with an arbitrary tensor A
,
A[]=CC
A
= CA
That means that we need to consider only objects that are totally symmetric in their
free indices. This gives all spins: Such a eld with 2s indices describes spin s; we have
already seen spins 0, 1/2, and 1.
We have dened the transformation law of all elds with lower indices by con-
sidering the direct product of spinors. Transformations for upper indices follow from
multiplication with C: They all follow from
0= (U 1)
A. TWO COMPONENTS 69
Since the vertical position of the index indicates the form of the transformation law,
we dene
=( )*
where the \ " indicates complex conjugation. Thus, a hermitian matrix is written
as
M=(M)* =M
)M=M
So, for a vector we have
V=V=V
Spin s is usually formulated in terms of a (2s+1)-component \vector". Then
one needs to calculate Clebsch-Gordan-Wigner coecients to construct Hamiltonians
relating dierent spins. For example, to couple two spin-1/2 objects to a spin-1 object,
one might write something like ~V y~. The matrix elements of the Pauli matrices
~are the CGW coecients for the spin-1 piece of1
2
1
2=10. This method gets
progressively messier for higher spins. On the other hand, in spinor notation such a
term would be simply V ; no special coecients are necessary, only contraction
of indices. Similarly the decomposition of products of spins involves only the picking
out of the various symmetric and antisymmetric pieces: For example, for1
2
1
2,
=1
2( ()+ [])=1
2 () C
=V+CS
where () means to symmetrize in those indices, by adding all permutations with
plus signs. We have thus explicitly separated out the spin-1 and spin-0 parts Vand
Sof the product. The square roots of various integers that appear in the CGW
coecients come from permutation factors that appear in the normalizations of the
various elds/wave functions that appear in the products: For example,
A
A
=jAj2+3jA j2+3jA j2+jA j2
In the spinor index method, the square roots never appear explicitly, only their squares
appear in normalizations: For example, in calculating a probability for A
B!C,
we evaluate
hA
BjCihCjA
Bi
hAjAihBjBihCjCi
whereA,B,a n dCeach have 2s indices for spin s, and hjimeans contracting all
indices (with the usual complex conjugation).
70 II. SPIN
5. Lorentz
Consider now the general 2 2 hermitian matrix
(V)
.
=Vy=V+Vt*
VtV
=1p
2V0+V1V2 iV3
V2+iV3V0 V1
=Va(a)
.
where we distinguish the right spinor index by a dot because it will be chosen to
transform dierently from the left one. For comparison, raising both spinor indices
with the matrix Cas for SU(2), and the vector indices with the Minkowski metric (in
either the orthonormal or null basis, as appropriate | see subsection IA4), we ndanother hermitian matrix
(V)
.
=V+Vt*
VtV
=1p
2V0+V1V2+iV3
V2 iV3V0 V1
=Va(a).
In the orthonormal basis, aare the Pauli matrices and the identity, up to normal-
ization. They are also the Clebsch-Gordan-Wigner coecients for spinor
spinor =
vector. In the null basis, they are completely trivial: 1 for one element, 0 for the rest,
the usual basis for matrices. In other words, they are simply an arbitrary way (ac-
cording to choice of basis) to translate a 2 2 (hermitian) matrix into a 4-component
vector.
Examining the determinant of (either version of) V, we nd the correct Minkowski
norms:
2detV = 2V+V +2VtVt*= (V0)2+(V1)2+(V2)2+(V3)2=V2
Thus Lorentz transformations will be those that preserve the hermiticity of this matrix
and leave its determinant invariant:
V0=gVgy;d e t g =1
(detg could also have a phase, but that would cancel in the transformation.) Thus g
is an element of SL(2,C). In terms of the representation of the Lie algebra,
g=eG;t r G =0
Thus the group space is 6-dimensional ( Ghas three independent complex compo-
nents), the same as SO(3,1) (where ggT=)(G)T= G).
Excercise IIA5.1
SL(2,C) also can be seen (less conveniently) from vector notation:
aConsider the generators
J()
ab=1
2(Jabi1
2abcdJcd)
A. TWO COMPONENTS 71
of SO(3,1). Find their commutation relations, and in particular show
[J(+);J( )]=0 . E x p r e s s J()
0iin terms of J()
ij. ShowJ()
ijhave the same
commutation relations as Jij. Finally, take a general innitesimal Lorentz
transformation in terms of Jaband rewrite it in terms of J()
ij, paying special
attention to the reality properties of the coecients. This demonstrates thatthe algebra of SO(3,1) is the same as that of SU(2)
SU(2), but Wick rotated
to SL(2,C).
bApply the same procedure to SO(4) and SO(2,2) to derive their covering
groups.
Excercise IIA5.2
Consider relativity in two dimensions (one space, one time):
aShow that SO(1,1) is represented in lightcone coordinates by
x
0+=x+;x0 = 1x
for some (nonvanishing) real number , and therefore SO(1,1) = GL(1). Write
this one Lorentz transformation, in analogy to excercise IIA1.2a on rotationsin two space dimensions, in terms of an analog of the angle (\rapidity") forthose transformations that can be obtained continuously from the identity.Do the relativistic analog of excercise IIA2.2.
bStill using lightcone coordinates, nd the parity and time reversal transfor-
mations. Note that writing as an exponential, so it can be obtained con-tinuously from the identity, restricts it to be positive, yielding a subgroup ofGL(1). Explicitly, what are the transformations of O(1,1) missing from thissubgroup? Which of P, (C)T, and (C)PT are missing from these transforma-tions, and which are missing from GL(1) itself?
In index notation, we write for this vector
V
0
.
=g
g*.
.
V
.
while for a (\Weyl") spinor we have
0
=g
The metric of the group SL(2,C) is the two-index antisymmetric symbol, which is
also the metric for Sp(2,C): In our conventions,
C= C= C=C..
= 0
i i
0
72 II. SPIN
We also have the identities
det L=1
2CC
L
L=1
2(tr L)2 1
2tr(L2);(L 1)=CC
L
(det L ) 1
A[]=CC
A
;A [
]=0
discussed earlier in this section. As there, we use the metric to raise, lower, and
contract indices:
= C; .= .
C.
.
VW=V.
W
.
These results for SO(3,1) = SL(2,C) generalize to SO(4) = SU(2)
SU(2) and
SO(2,2) = SL(2)
SL(2). As described earlier, the reality conditions change, so now
SO(4) : (V0)* =V
0C
C00;S O (2;2) : (V0)* =V
0
consistent with the (pseudo)reality properties of spinors for SU(2) and SL(2), where
we now use unprimed and primed indices for the two independent group factors(V!gVg
0).
Excercise IIA5.3
Take the explicit 2 2 representation for a vector given above, change the
factors ofito satisfy the new reality conditions for SO(4) and SO(2,2), and
show the determinant gives the right signatures for the metrics.
A common example of index manipulation is to use antisymmetry whenever pos-
sible to give vector products. For example, from the fact that V.
V
.
C.
.
is antisym-
metric in
we have that
V.
V
.
=1
2
V2
where the normalization follows from tracing both sides. Similarly,
V.
W
.
+W.
V
.
=
VW
It then follows that
V.
W
.
V
.
=(
VW W.
V
.
)V
.
=VWV.
1
2V2W.
Antisymmetry in vector indices also implies some antisymmetry in spinor indices.
For example, the antisymmetric Maxwell eld strength Fab= Fba, after translating
vector indices into spinor, can be separated into its parts symmetric and antisymmet-ric in undotted indices; antisymmetry in vector indices (now spinor index pairs) then
implies the opposite symmetry in dotted indices:
F
.
;.
= F
.
;.
=1
4(F
()[.
.
]+F
[](.
.
))=C.
.
f+Cf.
.
;f=1
2F.
;.
A. TWO COMPONENTS 73
Thus, an antisymmetric tensor also can be written in terms of a (complex) 2 2
matrix.
We also need to dene complex (hermitian) conjugates carefully because Cis
imaginary, and uses indices consistent with transformation properties:
.=( )*) .= ( )*;( )y= . .
V.
=(V.)*) x.
=x.
where we have used the spacetime coordinates as an example of a real vector (her-
mitian 22 matrix). In general, hermitian conjugation properties for any Lorentz
representation are dened by the corresponding product of spinors: For example,
( ())y=(. .
)= (..
)) f..
(f)*
More generally, we nd
(T(1:::j)(.
1:::.
k))y( 1)j(j 1)=2+k(k 1)=2T(1:::k)(.1:::.j)
As we'll see later, most spinor algebra involves, besides spinors, just vectors and
antisymmetric tensors, which carry only two spinor indices, so matrix algebra is often
useful. When using bra-ket notation for 2-component spinors, it is often convenientto distinguish undotted and dotted spinors. Furthermore, since spinor indices can
be raised and lowered, we can always choose the bras to carry upper indices and the
kets lower, consistent with our index-contraction conventions, to avoid extra signsand factors of C. We therefore dene (see subsection IB1)
h j=
hj;j i=ji ;[ j= .[.j;j ]=j.] .
V=jiV.
[.
j)V*= j.]V.hj;f =jifhj)f*=j.]f..
[.
j
As a result, we also have
h i=h i= ;[ ]= ..;h iy=[ ]
h jVj]= V.
.
;h jfji= f
VW*+WV*=(VW)I
where we have used the anticommutativity of the spinor elds. From now on, we use
this notation for the matrix representing a vector V(V.
), rather than the one with
which we started ( V.
).
74 II. SPIN
Excercise IIA5.4
Consider the generators
G=x.
@.
+jihj
and their Hermitian conjugates, where @
.
=@=@x.
. Show their algebra
closes. What group do they generate? Find a subset of these generators thatcan be identied with (a representation of) the Lorentz group.
Since we have exhausted all possible linear transformations on spinors (except for
scale, which relates to conformal transformations), the only way to represent discrete
Lorentz transformations is as antilinear ones:
0
=p
2n.
.
( 0= p
2n *)
From its index structure we see that nis a vector, representing the direction of the
re
ection. The product of two identical re
ections is then, in matrix notation
00=2n(n *)* =n2 )n2=1
where we have required closure on an SL(2,C) transformation ( 1). Thusnis a unit
vector, either spacelike or timelike. Applying the same transformation to a vector,
whereV.
transforms like .
, we write in matrix notation
V0= 2nV*n=n2V 2(nV)n
(The overall sign is ambiguous, and depends on whether it is a polar or axial vector.)
This transformation thus describes parity (actually CP, because of the complex con-jugation). In particular, to describe purely CP without any additional rotation (i.e.,
exactly re
ection of the 3 spatial axes), in our basis we must choose a unit vector in
the time direction,
p
2n.
=.
) 0= .( 0.= )
)V0.
= V.
which corresponds to the usual in vector notation, since in our basis
.
a=a
.
To describe time reversal, we need a transformation that does not preserve the com-
plex conjugation properties of spinors: For example, CPT is
0= ; 0.
= .
)V0= V
A. TWO COMPONENTS 75
(The overall sign on Vis unambiguous.)
In principle, whenever we work on a problem with both spinors and vectors we
could use a mixed vector-spinor notation, converting between the usual basis for
vectors and the spinor-index basis with identities such as
a
.
.
b=b
a;a
.
.
a=
.
.
However, in practice it's much simpler to use spinor indices exclusively, since then
one needs no -matrix identities at all, but only the trivial identities for the matrix
Cthat follow from its antisymmetry. For example, converting the vector index on
thematrices themselves into spinor indices ( a!.
), they become trivial:
(.
)
.
=
.
.
(This is the same as saying an orthonormal basis of vectors has the components
(Va)b=a
bwhen the components are dened with respect to the same basis.)
Thus, the most general irreducible (nite-dimensional) representation of SL(2,C)
(and thus SO(3,1)) has an arbitrary number of dotted and undotted indices, and istotally symmetric in each: A
(1:::m)(.
1:::.
n). Treating a vector index directly as a
dotted-undotted pair of indices (e.g., a=., which is just a funny way of labeling
a 4-valued index), we can translate into spinor notation the two constant tensors ofSO(3,1): Since the only constant tensor of SL(2,C) is the antisymmetric symbol, theyc a nb ee x p r e s s e di nt e r m so fi t :
.;.
=CC..
;
.;.
;
.
;.
=i(CC
C..
C.
.
CC
C..
C.
.
)
When we work with just vectors, these can be expressed in matrix language:
VW=tr(VW*)
abcdVaWbXcYd(V;W;X;Y )=it r(VW*XY* Y*XW*V)
Excercise IIA5.5
Prove this expression for the tensor (in either index or matrix version)
by (1) showing total antisymmetry, (2) explicitly evaluating a nonvanishing
component.
76 II. SPIN
6. Dirac
The Dirac spinor we encountered earlier is a 4-component reducible representation
in D=4: in terms of two (\left" and \right") two-component spinors,
= L
R.
The Hermitian metric that denes the (Lorentz-invariant) Dirac spinor inner prod-
uct
= y =
L R+h:c:; y=(
R .
L)
takes the simple form
=0 C..
C0
=p
2
0
The Dirac matrices are given by
V=
V=0V.
V.0
=0V
V*0
where the indices have been chosen to insure that the
matrices always take a Dirac
spinor to the same type of spinor. Since fV=;W=g= VW,t h e
matrices satisfy
f
a;
bg= ab
The extra sign is the result of normalizing the
's to be pseudohermitian with respect
to the metric:
y 1=+
. This Dirac spinor can be made irreducible by imposing
a reality condition that relates Land R: The resulting \Majorana spinor" is then
=1p
2
.
The product of all the
's is a pseudoscalar:
1=2p
2
4!abcd
a
b
c
d=1p
2
i
0
0i.
.!
(This is usually called \
5" in the literature for D=4 ,o r\
D"f o rD6=4 . W eh a v e
renamed it for consistency with dimensional reduction.) It can be used to project aDirac or Majorana spinor onto its two two-component spinors:
=1
2(Ip
2i
1)= I
00
0; 0
00
I
Various identities for these matrices can be derived directly from the anticommu-
tation relations: For example,
a
a= 2;
aa=
a=a=;
aa=b=
a=ab;
aa=b=c=
a=c=b=a=
A. TWO COMPONENTS 77
tr(I)=4;t r (a=b=)= 2ab; tr (a=b=c=d=)=abcd+adbc acbd
The trace identities follow from the fact that the only way to get a nonvanishing trace
out of a product of
matrices is when there are terms proportional to the identity;
sincef
a;
bg= ab, this only happens when the indices are pairwise identical. The
above results then follow from examination of relevant special cases. (Traces of odd
numbers of
matrices vanish. An exception is
1, until it is rewritten in terms of
its denition as the product of the other
-matrices.)
Although use of the anticommutation relations is convenient for generalization of
such identities to arbitrary dimensions, 2-spinor bra-ket notation is easier for deriving
4D identities. Since a Dirac spinor is the direct sum of a Weyl spinor and its complexconjugate, we write
=j
i L+j.] R.; =
Rhj+ .
L[.j
In this notation, there is no need to use a spinor metric , just as in Minkowski
4-vector bra-ket notation there is no need for an explicit matrix to represent theMinkowski metric: It is included implicitly in the denition of the inner product
for the basis elements ( h
ajbi=aborhji=C). Thus hermitian conjugation is
automatically pseudohermitian conjugation, etc.: is y, from the eect of hermitian
conjugating the basis vectors along with the components of the spinors they multiply.
(See subsections IB4-5.) We then have simply
.
= ji[.
j j.
]hj; +=jihj; =j.ih.j
where we have replaced the vector index a!.
on
a.
Excercise IIA6.1
Use this representation for the
matrices and projection operators for all
of the following:
aDerive
.
a=1a=2n+1
.
=a=2n+1a=1
.
a=1a=2n
.
= 1
2It r(a=1a=2n)
1tr(
1a=1a=2n)
bRederive the trace identities above. (Hint: For the last identity, use the
identityC[C
]= 0 repeatedly.)
cShow that
tr[(+ )
.
.
.
.
]= i
.;.
;
.
;.
by comparison with the expression of the previous subsection for .
78 II. SPIN
7. Chirality/duality
are often called \chiral projectors"; 2-component spinors (not paired into
Dirac spinors) are often called \chiral spinors", and appear in \chiral theories"; the
two 2-component spinors of a Dirac spinor are often labeled as having left and right
\chirality"; etc. When these two halves decouple, a theory can have a \chiral sym-
metry"
0
=ei
Since chirality is closely related to parity (chiral spinors can represent CP, but need to
be doubled to allow C, and thus P), Dirac spinors are often used to describe theories
where parity is preserved, or softly broken, or to analyze parity violation specically,
using
1to identify it.
A similar feature appears in electrodynamics. We rst translate the theory into
spinor notation: The Maxwell eld strength Fabis expressed in terms of the vector po-
tential (\gauge eld") Aa, with a \gauge invariance" in terms of a \gauge parameter"
with spacetime dependence. The gauge transformation Aa= @abecomes
A0
.
=A
.
@
.
where@
.
=@=@x.
. It leaves invariant the eld strength Fab=@[aAb]:
F
.
;.
=@.
A
.
@
.
A.
=1
4(F
()[.
.
]+F
[](.
.
))
=C.
.
f+Cf.
.
;f=1
2@(.
A).
Maxwell's equations are
@.
fJ.
They include both the eld equations (the hermitian part) and the \Bianchi identi-
ties" (the antihermitian part).
Excercise IIA7.1
We already saw VW*+WV* gave the dot product; show how VW* WV*
is related to the \cross product" V[aWb].
Excercise IIA7.2
Write Maxwell's equations, and the expression for the eld strength in termsof the gauge vector, in 2 2 matrix notation, without using C's. Combine
them to derive the wave equation for A.
Maxwell's equations now can be easily generalized to include magnetic charge by
allowing the current Jto be complex. (However, the expression for Fin terms of Ais
A. TWO COMPONENTS 79
no longer valid.) This is because the \duality transformation" that switches electric
and magnetic elds is much simpler in spinor notation: Using the expression given
above for the 4D Levi-Civita tensor using spinor indices,
F0
ab=1
2abcdFcd)f0
= if
More generally, Maxwell's equations in free space (but not the expression for Fin
terms ofA) are invariant under the continuous duality transformation
f0
=eif
(andJ0
.
=eiJ
.
in the presence of both electric and magnetic charges).
Excercise IIA7.3
Prove the relation between duality in vector and spinor notation. Show that
Fab+i1
2abcdFcdcontains only fand notf..
.
Excercise IIA7.4
How does complexifying J
.
modify Maxwell's equations in vector notation?
In even time dimensions, Wick rotation kills the i(or i) in the spinor-index
expression for 0;
0;0;0. Since the (discrete and continuous) duality transformation
now contains no i, we can impose self-duality or anti-self-duality; i.e., that forf00
vanishes, since they are now independent and real instead of complex conjugates.
These continuous chirality and duality symmetries on the eld strengths generalize
to the free eld equations for arbitrary massless elds in four dimensions. For reasons
to be explained in the following section, they distinguish the two polarizations of the
waves described by such elds. They are closely related to conformal invariance: In
higher dimensions, where not all free, massless theories are conformal (even on the
mass shell), these symmetries exist exactly for those that are conformal.
REFERENCES
1L.D. Landau and E.M. Lifshitz, Quantum mechanics, non-relativstic theory ,2 n de d .
(Pergamon, 1965) ch. VIII:review of SU(2) spinor notation.
2B. van der Waerden, Gottinger Nachrichten (1929) 100:
SL(2,C) spinor notation.
3E. Majorana, Nuo. Cim. 14(1937) 171.
4A. Salam, Nuo. Cim. 5(1957) 299;
L. Landau, JETP 32(1957) 407, Nucl. Phys. 3(1957) 27;
T.D. Lee and C.N. Yang, Phys. Rev. 105(1957) 1671:
identication of neutrino with Weyl spinor.
80 II. SPIN
:::::::::::::::::::::::::::: ::::::::::::::::::::::::::::
:::::::::::::::::::::::::::: B. POINCAR E::::::::::::::::::::::::::::
The general procedure for nding arbitrary representations of the Poincar eg r o u p
relevant to physics is to: (1) Describe spin 0. As we have seen, this means startingwith the coordinate representation, which is reducible, and apply the constraint p
2+
m2= 0 to get an irreducible one. (2) Find arbitrary, nite-dimensional, irreducible
representations of the Lorentz group. This we have done in the previous section.(3) Take the direct product of these two representations of the Poincar e group, which
give the orbital and spin parts of the generators. (The spin part of translationsvanishes.) We then need a further constraint to pick out an irreducible piece of thisproduct, which is the subject of this section.
1. Field equations
We rst show how the derivation above of the massless particle from the confor-
mal particle for spin 0 can be generalized to all \spins", i.e., all representations ofthe Poincar e group in arbitrary dimensions. There is a way to do this in terms of
classical mechanics for all representations of the conformal group, by generalizing the
description of the classical spinning particle. However, by analyzing the conformalparticle quantum mechanically instead, applying a set of constraints, it will be clearhow to generalize from conformal particles to general massless particles by weakeningthe constraints. The general idea is that the symmetry group for massive particles isthe Poincar e group, while that for massless particles includes also scale transforma-
tions, and nally conformal particles have also conformal boosts. So, starting withthe conformal group and dropping anything to do with conformal boosts will givemassless particles. Massive particles then follow from dimensional reduction: addinga further spatial dimension and xing its component of momentum to a constant,the mass, so p
2!p2+m2. The equation p2+m2= 0, as an operator equation
acting on a eld or wave function is the \Klein-Gordon (or relativistic Schr odinger)
equation". States or elds that satisfy this (and the other) eld equation are called\on-(mass-)shell", while those that don't (or for which the equations haven't beenimposed) are \o-shell".
We begin with a general representation of the conformal group SO(D,2) in terms
of generators G
AB,w h e r eA;B are D+2-component vector indices. We then impose
constraints that are the conformally covariant form of p2= 0: Identifying
(G+a;Gab;G+ ;G a)=(Pa;Jab;;Ka)
B. POINCAR E8 1
(whereA=(;a)) as the generators for translations, Lorentz transformations, di-
latations, and conformal boosts, we see that
GAB=1
2GC(AGCB) 1
D+2ABGCDGCD=0
is an irreducible piece of the product GG(symmetric and traceless) and includes:
(G++;G+a;Gab;G+ ;G a;G )=(P2;1
2fJab;Pbg+1
2f;Pag;:::)
where \:::" all have terms containing Ka.
Excercise IIB1.1
Work out all theG's in terms of P,J, ,a n dK.
In general theories, even massless ones, it is not always possible to have invariance
under conformal boosts. (We'll see examples of this insubsection IXA7.) However, all
massless theories are scale invariant, at least at the free level. (In D=4, free masslesstheories can always be made conformal on shell. However, the fact that even these
theories can have actions that are not invariant under conformal boosts proves that it
is sucient to add just dilatations to the Poincar e group. Furthermore, the fact that
conformal boosts are not always an invariance in D >4 means that dropping them will
give results in a dimension-independent form.) Therefore, only G
++andG+acan be
dened in general massless theories, but we'll see that these are sucient to dene
the kinematics. The former is just the masslessness condition, which we used to pick
the constraints in the rst place.
As we saw earlier, just scales xa: We can therefore write the relevant generators
as
Pa=@a;Jab=x[a@b]+Sab; =1
2fxa;@ag+w 1=xa@a+w+D 2
2
(We have used the antihermitian form of the generators.) The \scale weight" w+D 2
2
is the real \spin" part of , just as Sabis the spin part of the angular momentum
Jab. To preserve the algebra it must commute with everything, and thus we can
set it equal to a constant on an irreducible representation. We'll see shortly that
its value is actually determined by the spin Sab. It is the engineering dimension of
the corresponding eld. It has been normalized for later convenience; the value ofwdepends on the representation of S
ab, but is independent of D. The dilatation
generator is not exactly antihermitian because the integration measure dDxisn't
invariant under scaling. This is another reason wis determined, by the free action.
The form we have given preserves reality of elds. The commutation relations for the
spin parts, and the total generators, are the same as those for the orbital parts; e.g.,
[Sab;Scd]= [c
[aSb]d]
82 II. SPIN
(A convenient mnemonic for evaluating this commutator in general is to use Sab!
x[a@b]instead.)
Excercise IIB1.2
Find an expression for Kain terms of x,@,S,a n dwthat preserves the
commutation relations. Evaluate all the constraints G, and express the inde-
pendent ones in terms of just @,S,a n dw(nox).
Substituting the explicit representation of the generators into the constraint G+a,
and using the former constraint P2= 0 (when acting on wave functions on the right),
we nd that all xdependence drops out, leaving for G+athe condition
Sab@b+w@a=0
(paying careful attention to quantum mechanical ordering). This equation is the
general eld equation for all spins (acting on the eld strength), in addition to theKlein-Gordon equation (which is redundant except for spin 0).
Excercise IIB1.3
Dene spin for the conformal group by starting in D+2 dimensions: In terms
of the (D+2)-dimensional coordinates y
Aand their derivatives @A,
GAB=y[A@B]+SAB
Besides the previous conditions
y2=@2=fyA;@Ag=0
impose the constraints, in analogy to the D-dimensional eld equations, and
taking into account the symmetry between yand@,
SAByB+wyA=SAB@B+w@A=0
Show that the algebra of constraints closes, if we include the additional con-
straint
1
2S(ACSB)C+w(w+D
2)AB=0
Solve all the constraints with explicit y's for everything with an upper \ "
index, reducing the manifest symmetry to SO(D 1,1), in analogy to the way
y2= 0 was solved to nd y . Write all the conformal generators in terms of
xa,@a,Sab,a n dw.
B. POINCAR E8 3
2. Examples
We now examine the constraints Sab@b+w@a= 0 in more detail. We begin by
looking at some simple (but useful) examples. The simplest case is spin 0:
Sab=0)w=0
The next simplest case (for arbitrary dimension) is the Dirac spinor:
Sab= 1
2[
a;
b])Sab@b+w@a=
a
b@b+(w 1
2)@a
)
a@a=0;w =1
2
where we have separated out the pieces of the constraint that are irreducible with
respect to the Lorentz group (e.g., by multiplying on the left with
a). This gives
the (massless) \Dirac equation" @= = 0. The next case is the vector: In terms of the
basisjVi=Vajai, the spin is
Sab=j[aihb]j
However, the vector yields just another description of the scalar:
Excercise IIB2.1
Apply the eld equations for general eld strengths to the case of a vectoreld strength.
aFind the independent eld equations (assuming the eld strength is not just
a constant)
@
[aFb]=0;@aFa=0;w =1
Note that solving the rst equation determines the vector in terms of a scalar,
while the second then gives the Klein-Gordon equation for that scalar, and
the third xes the weight of the scalar to be the same as that found by startingwith a scalar eld strength.
bSolve the second equation rst to nd a gauge eld that is not a scalar.
All other representations can be built up from the spinor and vector. As our nal
example, we consider the case where the eld is a 2nd-rank antisymmetric tensor,
(S
abF)cd=[c
[aFb]d])
(Sab@b+w@a)Fcd=1
2@[aFcd] a[c@bFd]b+(w 1)@aFcd
)@[aFbc]=@bFab=0;w =1
which are Maxwell's equations, again separating out irreducible pi eces (e.g., by tracing
and antisymmetrizing).
84 II. SPIN
Excercise IIB2.2
Verify the representation of Lorentz spin given above for Fabby nding the
commutation relations implied by this representation.
Excercise IIB2.3
Consider the eld equations in 4D spinor notation for a general eld strength,
totally symmetric in its mundotted indices and ndotted indices,
S@.
m@.
=S..
@
.
n@
.=0;w =1
2(m+n)
where the spin acts on spinors as
(S )
=C
( ); (S..
).
=C.
(. .
)((S ).
=(S..
)
=0 )
aShow this implies
@.
:::.
:::=@
.
:::.
:::=0
bTranslate the eld equations into vector notation (in terms of Sab), nding
Sab@b+w@a= 0 and an axial vector equation.
cShow that the two equations are equivalent by deriving the equations of part
afromSab@b+w@a= 0 alone, and from the axial equation alone.
In each case, choosing the wrong scale weight wwould imply the eld was con-
stant. Note that we chose the eld strengthFabto describe electromagnetism: The
arguments we used to derive eld equations were based on physical degrees of free-dom, and did not take gauge invariance into account. In chapter XII we use more
powerful methods to nd the gauge covariant eld equations for the gauge elds, and
their actions.
3. Solution
Free eld equations can be solved easily in momentum space. Then the simplest
way to do the algebra is in the \lightcone frame". This is a reference frame, obtained
by a Lorentz transformation, where a massless momentum takes the simple form
pa=a
+p+
(using only rotations), or the even simpler form pa=a
+(using also a Lorentz
boost), where again is the sign of the energy. In that frame the general eld
equationSab@b+w@a= 0 reduces to
S i=0;w =S+
B. POINCAR E8 5
The constraint S i= 0 determines S+ to take its maximum possible value within
that irreducible representation, since S iis the raising operators for S+ : For any
eigenstate of S+ ,
S+ jhi=hjhi)S+ (S ijhi)=(S iS+ +[S+ ;S i])jhi=(h+1 ) (S ijhi)
The remaining constraint then determines w: It is the maximum value of S+ for
that representation. By parity (+ $ ), wis the minimum, so
w0;w=0,Sab=0
since ifS+ = 0 for all states then Sab= 0 by Lorentz transformation. As we have
seen by other methods (but can easily be derived by this method), w=1
2for the
Dirac spinor and w= 1 for the vector; since general representations can be built from
reducing direct products of these, we see that wis an integer for bosons and half-
integer for fermions. If we describe a general irreducible representation by a Young
tableau for SO(D 1,1) (with tracelessness imposed), or a Young tableau times a
spinor (with also
-tracelessness
a a:::b= 0), then it is easy to see from the results
for the spinor and vector, and antisymmetry in rows, that wis simply the number
of columns of the tableau (its \width"), counting a spinor index as half a column:
S+ just counts the maximum number of \ " indices that can be stuck in the boxes
describing the basis elements. (In fact, Dirac spinor
Dirac spinor gives just all
possible 1-column representations.)
This leaves undetermined only SijandS+i.H o w e v e r , S+i(\creation operator")
is canonically conjugate to S i(\annihilation operator"), so its action has also been
xed:
[S i;S+j]=ijS+ +Sij
(Sijvanishes for i=j,s oS+iandS iare conjugate, though not \orthonormal". The
constantS+ was xed above to be nonvanishing, except for the trivial case of spin
0.)
Thus only the \little group" SO(D 2) spinSijremains nontrivial: The original
irreducible representation of SO(D 1,1) Lorentz spin Sabwas a reducible representa-
tion of SO(D 2) spinSij; the irreducible SO(D 2) representation with the highest
value ofS+ is picked out of this SO(D 1,1) representation. This solution also gives
the eld strength in terms of the gauge eld: Working with just the highest- S+ -
weight states is equivalent to working with the gauge eld, up to factors of @+.
86 II. SPIN
As an explicit example, for spin 1/2 we have simply
= 0, which kills half the
components, leaving the half given by
+ . For spin 1, we nd
pbFab=0)F a=0
p[aFbc]=0)only F+a6=0
)only F+i6=0
In the \lightcone gauge" A+=0 ,w eh a v e F+i=@+Ai, so the highest-weight part of
Fabis the transverse part of the gauge eld. The general pattern, in terms of eld
strengths, is then to keep only pieces with as many as possible upper + indices andno upper indices (and thus highest S
+ weight). In terms of the vector potential,
we have
Fabp[aAb])only Ai6=0
The general rule for the gauge eld is to drop indices, so the eld becomes an
irreducible representation of SO(D 2). All + indices on the eld strength are picked
up by the momenta, which also account for the scale weight of the eld strength: Allgauge elds have w= 0 for bosons and w=
1
2for fermions.
Excercise IIB3.1
Using only the anticommutation relations f
a;
bg= ab, construct projec-
tion operators from
: These are operators Ithat satisfy
IJ=IJI
noX
;X
I=1
Because of time reversal symmetry
+$
(or parity
+$
), these
project onto two subspaces equal in size.
A method equivalent to using the lightcone frame is to perform a unitary trans-
formationUon the spin that is the inverse of the transformation on the coordi-
nates/momentum that would take us to the lightcone frame: We want a Lorentztransformation
abon the eld equations, which are of the form
Oabpb=0;Oab=Sab+wb
a
that has the eect
UOabU 1=acOcdb
d; b
apb=p0
a;p0a=a
+p+)
0=UOabpbU 1=acOcdb
dpb=acOcdp0
d)Oabp0
b=0
Ifj isatises the original constraint, then Uj iwill satisfy the new one. If we like,
we can always transform back at the end. This is equivalent to a gauge transformationin the eld theory.
B. POINCAR E8 7
It is easy to check that the appropriate operator is
U=eS+ipi=p+
Any operator Vathat transforms as a vector under Sab,
[Sab;Vc]=V[ab]c
but commutes with p, is transformed by UintoUVU 1=V0as
V0+=V+;V0i=Vi+V+pi
p+;V0 =V +Vipi
p++V+(pi)2
2(p+)2
as follows from explicit Taylor expansion, which terminates because S+iact as low-
ering operators (as for conformal boosts in subsection IA6). This yields the desiredresult
V
0apa=Vap0
a+V+
2p+p2
when we impose the eld equation p2=0 .
Excercise IIB3.2
Check this result by performing the transformation explicitly on the con-
straint. Before the transformation, the lightcone decomposition of the con-
straint is
( S+ +w)p++S+ipi=0
Si p++Sijpj+wpi Si+p =0
S ipi+( S ++w)p =0
Show that after this transformation, the constraint becomes
( S+ +w)p+=0
Si p++( S+ +w)pi 1
2S+ip2
p+=0
S ipi+( S+ +w)p S+ p2
p+ 1
2S+ipip2
p+2=0
Clearly these imply
w=S+ ;S i=0
withp2=0 .
On the other hand, if instead of using the lightcone identication of x+as \time",
we choose to use the usual x0for purposes of nding the evolution of the system, then
we want to consider transformations that do not involve p0, instead of not involving
88 II. SPIN
the \energy" p . Thus, by p0-independent rotations alone, the best we can do is to
choose
pi=0;p1=!
i.e., we can x the value of the spatial momentum, but not in a way that relates to
the sign of the energy. The result is then
p0>0:pa=a
+p+
p0<0:pa=a
p
The result is similar to before, but now the positive and negative energy solutions are
separated: In this frame the eld equations reduce to
p0>0:S i=0;S+ =w
p0<0:S+i=0;S+ = w
Thus, while wtakes the same value as before, now the positive-energy states are
associated with the highest weight of S+ , while the negative-energy ones go with
the lowest weight (and nothing between). The unitary transformation that achieves
this result is a spin rotation that rotates Sabin the eld equations with the same
eect as an orbital transformation that would rotate ( p1;pi)!(!;0). By looking at
the special case D= 3 (where there is only one rotation generator), we easily nd
the explicit transformation
U=exp
tan 1jpij
p1
S1ipi
jpij
Excercise IIB3.3
Perform this transformation:
aFind the action of the above transformation on an arbitrary vector Va.( H i n t :
Look atD= 3 to get the transformation on the \longitudinal" part of the
vector.) In particular, show that
V0apa=Vap0
a;p0a=a
0p0+a
1!
bShow the eld equations are transformed as
S0apa+wp0) !S10+wp0=p0(w p0
!S10) 1
!S10p2
S1apa+wp1)1
!pi(!S1i p0S0i)+p1(w p0
!S10)
Siapa+wpi) [ij 1
!(!+p1)pipj](!S1j p0S0j)+pi(w p0
!S10)
B. POINCAR E8 9
Note that the rst equation gives the time-dependent Schr odinger equation,
with Hamiltonian
H=1
w(S10p1 S0ipi)!1
wS10!
This diagonalizes the Hamiltonian H(in a representation where S10is diag-
onal). Thus the only independent equations are
p2=0;S10=(p0)w; S1i (p0)S0i=0
leading to the advertised result.
cFind the transformation that rotates to the pidirection instead of the 1 di-
rection, so
H! 1
wS0ipi
jpjj!
4. Mass
So far we have considered only massless theories. We now introduce masses
by \dimensional reduction", identifying mass with the component of momentum inan extra dimension. As with the extra dimensions used for describing conformal
symmetry, this extra dimension is just a mathematical construct used to give a simple
derivation. (Theories have been postulated with extra, unseen dimensions that arehidden by \compactication": Space curls up in those directions to a size too small
to detect with present experiments. However, no compelling reason has been given
for why the extra dimensions should want to compactify.)
The method is to: (1) extend the range of vector indices by one additional spatial
direction, which we call \ 1"; (2) set the corresponding component of momentum to
equal the mass,
p
1=m
and (3) introduce extra factors of ito restore reality, since @ 1=ip 1=im,b y
a unitary transformation. Since all representations can be constructed by directproducts of the vector and spinor, it's sucient to dene this last step on them. For
the scalar this method is trivial, since then simply p
2!p2+m2.
For the spinor, since any transformation on the spinor index can be written in
terms of the gamma matrices, and the transformation must aect only the 1 direc-
tion, we can use only
1. (For even dimensions, we can identify the
1of dimen-
sional reduction with the one coming from the product of all the other
's, since in
odd dimensions the product of all the
's is proportional to the identity.) We nd
U=exp(
1=2p
2) :
1!
1;
a! p
2
1
a
90 II. SPIN
We perform this transformation directly on the spin operators appearing in the con-
straints, or the inverse transformation on the states. Dimensional reduction, followed
by this transformation, then modies the massless equation of motion as
i@=!i@= m
1! p
2
1(i@=+mp
2)
soi@= =0!(i@=+mp
2) = 0.
The prescription for the vector is
U=exp(1
2ij 1ih 1j):j 1i!ij 1i;h 1j! ih 1j (h 1j 1i=1 )
with the other basis states unchanged. This has the eect of giving each eld a i
for each ( 1)-index. For example, for Maxwell's equations
@[aFbc]!@[aFbc]
@[aFb] 1+imFab!@[aFbc] (redundant)
i(@[aFb] 1 mFab)
@bFab!@bFab+imFa 1
@aF 1a!@bFab+mFa 1
i@aF 1a (redundant)
Note that only the mass-independent equations are redundant. Also, Fa 1appears
explicitly as the potential for Fab, but without gauge invariance. Alternatively, we
can keep the gauge potential:
Fab=@[aAb]!Fab=@[aAb]
Fa 1=@aA 1 imAa!Fab=@[aAb]
iFa 1= i(@aA 1+mAa)
This is known as the \St uckelberg formalism" for a massive vector, which maintains
gauge invariance by having a scalar A 1in addition to the vector: The gauge trans-
formations are now
Aa= @a!Aa= @a
A 1= im!Aa= @a
iA 1= im
Excercise IIB4.1
Consider the general massive eld equations that follow from the general
massless ones by dimensional reduction. One of these is
S 1a@a+wim =0
(before restoring reality). This scalar equation alone gives the complete eld
equations for w=1/2 and 1 (antisymmetric tensors), 0 being trivial.
aShow that for w=1/2 it gives the (massive) Dirac equation.
bExpanding the state over explicit elds, nd the covariant eld equations it
implies for w=1. Show these are sucient to describe spins 0 (vector eld
B. POINCAR E9 1
strength: see excercise IIB2.1) and 1 ( FabandFa 1). Note that S 1aact as
generalized
matrices (the Dirac matrices for spin 1/2, the \Dun-Kemmer
matrices" for w=1), where
Sab= [S 1a;S 1b]
cShow that these covariant eld equations imply the Klein-Gordon equation
for arbitrary antisymmetric tensors. Show that in D=4 all antisymmetric
tensors (coming from 0-5 indices in D=5) are equivalent to either spin 0 orspin 1, or trivial. (Hint: Use
abcd.)
dConsider the reducible representation coming from the direct product of two
Dirac spinors, and represent the wave function itself as a matrix:
Sij = ~Sij + ~Sij
wherei=( 1;a)a n d ~Sijis the usual Dirac-spinor representation. Using the
fact that any 44 (in D=4) matrix can be written as a linear combination of
products of
-matrices (antisymmetric products, since symmetrization yields
anticommutators), nd the irreducible representations of SO(4,1) in , andrelate to part c.
Excercise IIB4.2
Solve the eld equations for massive spins 1/2 and 1 in momentum space bygoing to the rest frame.
The solution to the general massive eld equations can also be found by going
to the rest frame: The combination of that and dimensional reduction is, in terms ofthe massive analog of lightcone components,
p
+=1p
2(p0+p 1)=p
2m; p =1p
2(p0 p 1)=0;pi=0
wherepiare now the other D 1 (spatial) components. This xing of the momentum
is the same as the lightcone frame except that p1has been replaced by p 1,a n d
thuspinow has D 1 components instead of D 2. The solution to the constraints is
thus also the same, except that we are left with an irreducible representation of the
\little group" SO(D 1) as found in the rest frame for the massive particle, vs. one
of SO(D 2) found in the lightcone frame for the massless case.
92 II. SPIN
5. Foldy-Wouthuysen
The other frame we used for the massless analysis, which involved only energy-
independent rotations, can also be applied to the massive case by dimensional reduc-
tion. The result is known as the \Foldy-Wouthuysen transformation", and is useful for
analyzing interacting massive eld equations in the nonrelativistic limit. Replacing
p1!p 1=min our previous result, we have for the free case
U=exp
tan 1j~pj
m
S 1ipi
j~pj
;U H U 1=1
wS 10!
For purposes of generalization to interactions, it was important that in the free trans-
formation (1) we used only the spin part of a rotation, since the orbital part could
introduce explicit xdependence, and (2) we used only rotations, since a Lorentz
boost would introduce p0dependence in the \parameters" of the transformation,
which could generate additional p0(time derivative) terms in the eld equation.
Excercise IIB5.1
Perform this transformation for the Dirac spinor, and then apply the reality-restoring transformation to obtain
H!p
2
0!
We then can use the diagonal representation
0= I
00
I
=p
2. (We can de-
ne this representation, up to phases, by switching
0and
1of the usual
representation.) In general the reality-restoring transformation will be unnec-essary for any spin, since applying the eld equation S
10=wpicks out a
representation of the \little group" SO(D 1).
In the interacting case the result generally can't be obtained in closed form, so it is
derived perturbatively in 1 =m. The goal is again a Hamiltonian diagonal with respect
toS 10, to preserve the separation of positive and negative energies; we then can set
S 10=wto describe just positive energies. We thus choose the transformation to
cancel any terms in Hthat are o-diagonal, which come from odd total numbers of
\ 1" and \0" indices from the spin factors in any term: i.e., odd numbers of S0i
andS 1i(e.g., theS0ipiterm in the original H). For example, for coupling to an
electromagnetic eld, the exponent of Uis generalized by covariantizing derivatives
(minimal coupling @!r =@+iA), but also requires eld-strength ( EandB)t e r m s
to cancel certain ones of those generated from commutators of these derivatives in
the transformation:
ra=@a+iAa) [ra;rb]=iFab
B. POINCAR E9 3
Before performing this transformation explicitly for the rst few orders, we con-
sider some general properties that will allow us to collect similar terms in advance.(Few duplicate terms would appear to the order we consider, but they breed like
rabbits at higher orders.) We start with a eld equation Fthat can be separated into
\even" termsEand \odd" onesO,e a c ho fw h i c hc a nb ee x p a n d e di np o w e r so f1 =m:
F=E+O:E=1X
n= 1m nEn;O=1X
n=0m nOn
Note that the leading ( m+1) term is even; thus we choose only odd generators to
transform away the odd terms in F, perturbatively from this leading term:
F0=eGFe G;G =1X
n=1m nGn
SinceF0is even while Gis odd, we can separate this equation into its even and odd
parts as
F0=cosh(LG)E+sinh(LG)O
0=sinh(LG)E+cosh(LG)O
(withLG=[G;] as in subsection IA3). Since we can perturbatively invert any
Taylor-expandable function of LGthat begins with 1, we can use the second equation
to give a recursion relation for Gn: Separating the leading term of F,
E=mE 1+E; m[G;E 1]=[G;E]+LGcoth(LG)O
which we can expand in 1 =m[after Taylor expanding LGcoth(LG)] to give an expres-
sion for [Gn;E 1]t os o l v ef o r Gn. We can also use the implicit solution for [ G;E]
directly to simplify the expression for F0:
F0=E+tanh(1
2LG)O
For example, to order 1 =m2we have forF0
F0
1=E 1;F0
0=E0;F0
1=E1+1
2[G1;O0]
F0
2=E2+1
2[G2;O0]+1
2[G1;O1]
To this order we therefore need to solve
[G1;E 1]=O0; [G2;E 1]=O1+[G1;E0]
94 II. SPIN
For our applications we will always have
E 1= 1
wS 10
unchanged by interactions. We have oversimplied things a bit in the above deriva-
tion: For general spin we need to consider more than just even and odd terms; weneed to consider all eigenvalues of S
10:
[S 10;Fs]=sFs
and nd the transformation that makes F0commute with it ( s= 0). The procedure
is to rst divide into even and odd values of s, as above, then to divide the remaining
even terms inF0into twice even values of s(multiples of 4) as the new E0and twice
odd as the newO0, which are transformed away with the new twice odd G0,a n ds o
on. This very rapidly removes the lower nonzero values of jsj(1!2!4!:::),
which has a maximum value of 2 w(from the operators that mix the maximum value
S 10=wwith the minimum S 10= w). For example, for the case of most interest,
the Dirac spinor, the only eigenvalues (for operators) are 0 and 1, so the original
even part does commute with S 10, and the procedure need be applied only once.
Furthermore, terms in Fof eigenvalue scan be generated only at order m1 sor
higher; so at any given order the procedure rapidly removes all undesired terms forany spin.
Since the terms we want to cancel are exactly the ones with nonvanishing eigen-
values ofS
10, they can always be written as [ G;S 10]f o rs o m eG, so we can always
nd a transformation to eliminate them:
[S 10;Gsn]=sGsn)Gsn= w
sf[G;E]+LGcoth(LG)Ogsn
(This is just diagonalization of a Hermitian matrix in operator language.) In partic-
ular for the Dirac spinor, since E 1has only1 eigenvalues, it's easy to see that not
only do all even operators commute with it, but all odd operators anticommute with
it. (Consider the diagonal representation of E 1:f 1
00
1; 0
ab
0g=0 . ) W et h e nh a v e
simply
w=1
2) (E 1)2=1) [E 1;E]=fE 1;Og=fE 1;Gg=0
)mG= 1
2f[G;E]+LGcoth(LG)OgE 1
As a nal step, we can apply the usual transformation
U0=eimtS 10=w
B. POINCAR E9 5
which commutes with all but the p0term inE0to have the sole eect of canceling
E 1, eliminating the rest-mass term from the nonrelativistic-style expression for the
energy.
For the minimal electromagnetic coupling described above, we have besides E 1
E0=0;O0=1
wS0ii
w h e r ew eh a v ew r i t t e n a=pa+Aa(instead of a= ira,t os a v es o m e i's). There
are no additional terms in Ffor minimal coupling for spin 1/2, but later we'll need to
include nonminimal eective couplings coming from quantum (eld theoretic) eects.
There are also extra terms for spins 0 and 1 because the eld strength is not the sameas the fundamental eld, so we'll treat only spin 1/2 here, but we'll continue to use
the general notation to illustrate the procedure. Using the above results, we nd to
order 1=m
2forF0
G1=S 1ii;G 2=wS0iiF0i
in agreement with with the free case up to eld strength terms. The diagonalized
Schrodinger equation is then to this order, including the eect of U0,
F0
1=0;F0
0=0;F0
1= 1
2w[1
2fS 1i;S0jgiFij+S 10(i)2]
F0
2= 1
4[fS0i;S0jg(@iF0j) SijfiF0i;jg]
For spin 1/2 we are done, but for other spins we would need a further transformation
(beforeU0)t op i c ko u tt h ep a r to f F0
2that commutes with S 10(by eliminating the
twice odd part); the nal result is
F0
2=1
4[1
2(fS 1i;S 1jg fS0i;S0jg)(@iF0j)+SijfiF0i;jg]
It can also be convenient to translate into notation (as for the massless case, but
with index 1! 1): We then write
F0
1=0;F0
0=0;F0
1= 1
2w[1
2fS+i;S jgiFij+S+ (i)2]
F0
2= 1
4[1
2fS+(i;S j)g(@iF0j) SijfiF0i;jg]
In this notation the eigenvalue of S+ =S 10for any combination of spin operators
can be simply read o as the number of indices minus the number of +.
Excercise IIB5.2
Find the Hamiltonian for spin 1/2 in background electromagnetism, expanded
nonrelativistically to this order, by substituting the appropriate expressions
for the spin operators in terms of
matrices, and applying S+ =won
96 II. SPIN
the right for positive/negative energy. (Ignore the reality-restoring transfor-
mation.)
-matrix algebra can be performed directly with the spin operators:
For the Dirac spinor we have the identities
S(a(bSc)d)=1
2b
(ad
c) acbd)fS+i;S jg=1
2ij 2SijS+
6. Twistors
Besides describing spin 1/2, spinors provide a convenient way to solve the condi-
tionp2= 0 covariantly: Any hermitian matrix with vanishing determinant must have
a zero eigenvalue (consider the diagonalized matrix), and so such a 2 2 matrix can
be simply expressed in terms of its other eigenvector. Absorbing all but the sign of
the nontrivial eigenvalue into the normalization of the eigenvector, we have
p2=0)p.
=pp.
for some spinor p.S i n c ep0is the (canonical) energy, the is the sign of the energy.
This explains why time reversal (actually CT in the usual terminology) is not a lineartransformation. Note that p
is a commuting object, while most spinors are fermionic,
and thus anticommuting (at least in quantum theory). Such commuting spinors are
called \twistors".
Excercise IIB6.1
Show that, in terms of its energy Eand the angular direction ( ;)o fi t s
3-momentum, a massless particle is described by the twistor
p=21=4p
jEj(cos
2e i=2;sin
2ei=2)
One useful way to think of twistors is in terms of the lightcone frame. In spinor
notation, the momentum is
p.
= 1
00
0
If we write an arbitrary massless momentum as a Lorentz transformation from this
lightcone frame, then the twistor is just the part of the SL(2,C) matrix that con-
tributes:
p0.
=p
.
g
g.
.
=
.
.
g
g.
.
=gg.
.
=pp.
For this reason, the twistor formalism can be understood as a Lorentz covariant form
of the lightcone formalism.
The twistor construction thus gives a covariant way of constructing wave functions
satisfying the mass-shell condition (Klein-Gordon equation) for the massless case,
B. POINCAR E9 7
=0 ,w h e r e =@2= p2. We simply Fourier transform, and use the twistor
expression for the momentum, writing the momentum-space wave function in terms
of twistor variables (\Penrose transform"):
(x)=Z
d2pd2p.[exp(ix.
pp.
)+(p;p.)+exp( ix.
pp.
) (p;p.)]
wheredescribe the positive- and negative-energy states, respectively. (The integral
over p.can be performed also, eectively taking the Fourier transform with respect
to that variable only, treating x.
pas the conjugate.)
We can extend the matrix notation of subsection IIA5-6 to twistors:
hpj=p;jpi=p;[pj=p.;jp]=p.
P=jpi[pj; P*=jp]hpj
As a result, we also have for twistors
hpqi= hqpi;[pq]= [qp];hpqihrsi+hqrihpsi+hrpihqsi=0 ;hpqi*=[qp]
These properties do not apply to physical, anticommuting spinors, where h i=
+h i,a n dh i6=0 .
Another natural way to understand twistors is through the conformal group. We
have already seen that the conformal group in D dimensions is SO(D,2). Since thisgroup in four dimensions is the same as SU(2,2), it's simpler to describe its general
representations (and in particular spinors) in SU(2,2) spinor notation. Then the
simplest way to generate representations of this group is to use spinor coordinates:
We therefore write the generators as (see subsection IC1)
G
AB=BA 1
4B
ACC
where we have subtracted out the trace piece to reduce U(2,2) to SU(2,2) and, consis-
tently with the group transformation properties under complex conjugation, we havechosen the complex conjugate of the spinor to also be the canonical conjugate: The
Poisson bracket is dened by
[
A;B]=B
A
To compare with four-dimensional notation, we reduce this four-component spinor by
recognizing it as a particular use of the Dirac spinor. Using the same representation
as in subsection IIA6, we write
A=(p;!.); A=(!;p.); .
AB=0 C..
C0
98 II. SPIN
Now the Poisson brackets are
[!;p]=
;[!.;p.
]=.
.
The group generators themselves reduce to
pp.
;!!.
;p (!);p(.!.
);p!+p.!. 2
which are translations, conformal boosts, SL(2,C) generators and their complex con-
jugates, and dilatations.
Another kind of twistor, related to position space instead of momentum space,
follows from this (D+2)-coordinate description of conformal symmetry for D=4 (see
subsection IA6). In practice, it's more convenient to work with invariances than con-
straints. In this case, we can solve the lightcone constraint on Wick-rotated D=3+3
or 5+1 space, replacing 6-component conformal vector indices with 4-component con-
formal spinor indices, with a position-space twistor:
y2=1
4ABCDyAByCD=0)yAB=zAzB
whereAis an SL(4) (or SU*(4)) index and is an SL(2) (or SU(2)) index, and zA
is real (with either two real or two pseudoreal indices). (Here the SL groups apply
to 3+3 dimensions, the SU groups to 5+1.) Whereas yhad 6 1=5c o m p o n e n t s
due to the constraint, zhas 42 3 = 5 components due to the SL(2) (SU(2)) gauge
invariance of the above relation to y. These coordinates reduce to the usual by an
SL(2) transformation:
A=(;.);zA
=(
;x.)) SL(2) gauge =
wheree=2.
Excercise IIB6.2
Substitute this spinor-notation z(;x)i n t oyz2and compare with the
vector-notation y(e;x) of subsection IA6.
7. Helicity
A sometimes-useful way to treat the transverse spin operators Sijis in terms of
Wabc=1
2P[aJbc]=1
2P[aSbc]
which reduces to Sijin the rest frame, and (like the eld equations) can be written
in terms of just the Poincar e generators. This is the part of Sabwhose commutator
B. POINCAR E9 9
with the eld equations is proportional to the eld equations (i.e., it preserves the
constraints). In D=4 this is the \Pauli-Luba nski (axial) vector"
Wa=1
6Wbcdbcda
We can choose our states to be eigenstates of a component of it: For example, for
massless states W0=P0is called the \helicity". For massive states the helicity is
dened asW0=j~Pj, but is less useful, especially since it is undened (0/0) in the rest
frame. In that case one instead chooses a component in terms of a (momentum-dependent axial) vector s
aassaWa,w h e r esaPa=0a n ds2=1=m2.
Excercise IIB7.1
Show in both the massless and massive cases that Wabcreduces to the little
group generators on shell by going to the appropriate reference frame.
The twistor representation of the conformal group does not give the most general
representation, but it does give all the (free) massless ones. The reason it givesmassless ones is that this representation satisies the constraint (see subsection IIB1)
G
[AB][CD]=G[A[CGB]D] traces =0
which includes p2= 0 as well as all the equations that follow from p2=0b yc o n f o r m a l
transformations. As a consequence, this representation also satises
GACGCB trace =hGAB
wherehis the helicity. This equation may be more recognizable in SO(4,2) notation,
as
1
8ABCDEFGCDGEF=ihGAB
This equation includes, as its lowest mass-dimension part (as dened by dilatations),
the Pauli-Luba nski vector
Wa=1
2bcdaPbJcd=ihPa
(The \i" appears in the last two equations only when we use the antihermitian form of
the generators GABandJab.) Although any massless representation of the conformal
group satises the above conditions (see excercise IIB2.3), the twistor representationsatises the unusual property that helicity is realized as a linear transformation on thecoordinates: For the twistors the implicit denition of helicity can be solved explicitlyto give
h=
1
4fA;Ag=1
2AA+1=1
2(p! p.!.)
100 II. SPIN
which is exactly the U(1) transformation of U(2,2)=SU(2,2)
U(1). (This is similar
to SU(2) in terms of \twistors": See excercise IC1.1.) On functions of pand p.,i t
eectively just counts half the number of p's minus p.'s.
Excercise IIB7.2
These results are pretty clear from symmetry, but we should do some algebra
to check coecients:
aUse the denition of the action of the Lorentz generators on a vector operator
in vector and spinor notations,
[Jab;Vc]=V[ab]c;[J;V
.
]=V
(.
C)
; [J..
;V
.
]=V
(.C.
).
to derive
J
..
= 1
2(CJ..
+C..
J)
bExpressJabandPain terms of the twistors p;p.;!;!.(with normalization
ofJabxed by its action on the twistors themselves), and plug into PJ =ihP
to derive the above expression of hin terms of twistors.
The simple form of the helicity in the twistor formalism is another consequence
of it being a covariantized lightcone formalism. In the lightcone frame, there is still
a residual Lorentz invariance; in particular, a rotation about the spatial direction inwhich the momentum points leaves the momentum invariant. This is another deni-
tion of the helicity, as the part of the angular momentum performing that rotation.
(Only spin contributes, since by denition the momentum is not rotated.) Since theproduct of two Lorentz transformations is another one, this rotation can be inter-
preted as a transformation acting on the Lorentz transformation to the lightcone
frame, i.e., on the twistor, such that the momentum is invariant. This is simply a
phase transformation:
g
0
=ei0
0e i
g
)p0=eip
We can generalize the Penrose transform in a simple way to wave functions car-
rying indices to describe spin:
1:::m.
1:::.
n(x)=Z
d2pd2p.p1pmp.
1p.
n
[exp(ix.
pp.
)+(p;p.)+exp( ix.
pp.
) (p;p.)]
For the integral to give a nonvanishing result, the integrand must be invariant under
the U(1) transformation generated by the helicity operator h: In other words, must
B. POINCAR E 101
have a transformation under h, i.e., a certain helicity, that is exactly the opposite that
of the explicit pfactors that carry the external indices to give a contribution to the
integral, since otherwise integrating over the phase of pwould average it to zero.
(Explicitly, if we derive the helicity by acting on the Penrose transform, this minus
sign comes from integration by parts.) This means that (x) automatically has a
certain helicity, half the number of dotted minus undotted indices:
h=1
2(n m)[w=1
2(m+n)]
as given by the above twistor operator expression acting on . (Alternatively, com-
paring the x-space form of the Pauli-Luba nsky vector, its action plus that of the
twistor-space one must vanish on jip, so the helicity is again minus the twistor-
space helicity operator acting on the prefactor.)
Since, after restricting to the appropriate helicity, the integral over this phase is
trivial, we can also eliminate it by replacing the \volume" integral over the twistor
or its complex conjugate (but not both) with a \surface" (boundary) integral:
Z
d2p!I
pdp
(Alternatively, we can insert a -function in the helicity.) The result is equivalent to
the usual integral over the three independent components of the momentum.
This generalization of the Penrose transform implies that (x) satises some
equations of motion besides p2=0 ,n a m e l y
p.
:::.
:::.
=p
.
:::.
:::.
=0
which are also implied by Sab@b+w@a= 0 (see excercise IIB2.3). Besides Poincar e
invariance, these equations are invariant under the phase transformation
0
:::.
:::.
=ei2h
:::.
:::.
that generalizes duality and chiral transformations. We also see that (anti-)self-
duality and chirality are related to helicity. Another way to understand the twistor
result is to remember its interpretation as a Lorentz transformation from the light
cone: In the light cone frame, where p+.
+is the only nonvanishing component of p.,
the above equations of motion imply the only nonvanishing component of 1:::m.
1:::.
n
is +:::+.
+:::.
+, which can be identied with +(forp+.
+>0) or (forp+.
+<0).
REFERENCES
1E. Wigner, Ann. Math. 40(1939) 149;
V. Bargmann and E.P. Wigner, Proc. Nat. Acad. Sci. US 34(1946) 211:
102 II. SPIN
little group; more general discussion of Poincar e representations and relativistic wave
equations for D=4.
2A.J. Bracken, Lett. Nuo. Cim. 2(1971) 574;
A.J. Bracken and B. Jessup, J. Math. Phys. 23(1982) 1925;
W. Siegel, Nucl. Phys. B263 (1986) 93:
conformal constraints.
3E.C.G. St uckelberg, Helv. Phys. Acta 11(1938) 299.
4P.A.M. Dirac, Rev. Mod. Phys. 21(1949) 392:
lightcone gauge.
5G. Nordstr om, Phys. Z. 15(1914) 504:
dimensional reduction.
6O. Klein, Z. Phys. 37(1926) 895;
V. Fock, Z. Phys. 39(1927) 226:
mass from dimensional reduction.
7R.J. Dun, Phys. Rev. 54(1938) 1114;
N. Kemmer, P r o c .R o y .S o c . A173 (1939) 91.
8M.H.L. Pryce, P r o c .R o y .S o c . A195 (1948) 62;
S. Tani, Soryushiron Kenkyu 1(1949) 15 (in Japanese):
free version of Foldy-Wouthuysen.
9L.L. Foldy and S.A. Wouthuysen, Phys. Rev. 78(1950) 29.
10M. Cini and B. Touschek, Nuo. Cim. 7(1958) 422;
S.K. Bose, A. Gamba, and E.C.G. Sudarshan, Phys. Rev. 113(1959) 1661:
\ultrarelativistic" version of Foldy-Wouthuysen transformation.
11R. Penrose, J. Math. Phys. 8(1967) 345, Int. J. Theor. Phys. 1(1968) 61;
M.A.H. MacCallum and R. Penrose, Phys. Rep. 6(1973) 241:
twistors.
12W. Pauli, unpublished;J.K. Luba nski, Physica IX(1942) 310, 325.
C. SUPERSYMMETRY 103
:::::::::::::::::::: ::::::::::::::::::::
:::::::::::::::::::: C. SUPERSYMMETRY ::::::::::::::::::::
We'll see later that quantum eld theory requires particles with integer spin to be
bosons, and those with half-integer spin to be fermions. This means that any sym-
metry that relates bosonic wave functions/elds to fermionic ones must be generatedby operators with half-integer spin. The simplest (but also the most general, at least
of those that preserve the vacuum) is spin 1/2. Supersymmetry is a major ingredi-
ent in the most promising generalizations of the Standard Model. In this section welook at representations, generalizing the results of the previous sections for Poincar e
symmetry.
1. Algebra
From quantum mechanics we know that for any operator A
h jfA;Aygj i=X
n(h jAjnihnjAyj i+h jAyjnihnjAj i)
=X
n(jhnjAyj ij2+jhnjAj ij2)0
from inserting a complete set of states. In particular,
fA;Ayg=0)A=0
from examining the matrix element for all states j i. This means the anticommuta-
tion relations of the supersymmetry generators must be nontrivial.
We are then led to anticommutation relations of the form, in Dirac (Majorana)
notation,
fq;qg=p=o rfq;qyg=pa
ap
2
0
(We use translations instead of internal symmetry or Lorentz generators because of
dimensional analysis: Bosonic elds dier in dimension from fermionic ones by half
integers.) Note that this implies the positivity of the energy:
trfq;qyg=p
2patr(
a
0)=p
2patr(1
2f
a;
0g)=p01p
2tr I
Similar arguments imply that the supersymmetry generators are constrained, just
as the momentum is constrained by the mass-shell condition. For example, in the
massless case,
fp=q;qp=g=p=p=p== 1
2p2p==0)p=q=0
104 II. SPIN
In four dimensions the commutation relations can be written in terms of irre-
ducible spinors as
fq;q.
g=p.
;fq;qg=fq;qg=0
This generalizes straightforwardly to more than one spinor, carrying a U(N) index:
fqi;qj.
g=j
ip.
2. Supercoordinates
Since the momentum is usually represented as coordinate derivatives, we natu-
rally look for a similar representation for supersymmetry. We therefore introduce
an anticommuting spinor coordinate . Because of the anticommutation relations q
can't be simply @=@ , but the modication is obvious:
q= i@
@+1
2.
@
@x.
; q.= i@
@.+1
2@
@x.
We can also express supersymmetry in terms of its action on the \supercoordinates":
Using the hermitian innitesimal generator q+.q.,
=; .=.; x.
=1
2i(.
+.
)
Note that ( q)y=q.,(q)y= q..
We can also dene \covariant derivatives": derivatives that (anti)commute with
(are invariant under) supersymmetry. These are easily found to be
d=@
@+1
2.
p
.
; d.=@
@.+1
2p.
Besides overall normalization factors of i, leading to the opposite hermiticity condition
(d)y= d., these dier from the q's by the relative sign of the two terms. These
changes combine to preserve
fd;d.
g=p.
as a result of which pis also a covariant derivative as well as being a symmetry
generator (as for the Poincar e group), but now ( d)y=+d..
In classical mechanics, the fact that @=@x commutes with translations is \dual" to
the fact that the innitesimal change dx, or the nite change x x0, is also invariant
under translations. Furthermore, the d'Alembertian =(@=@x )2being Poincar e
invariant is dual to the line element ds2= (dx)2being invariant. This allows the
C. SUPERSYMMETRY 105
construction of the action from.x2. In the supersymmetric case the innitesimal
invariants under the q's (and therefore p)a r e
d;d .;d x.
+1
2i(d).
+1
2i(d.
)
and the corresponding nite ones (by integration) are
0; . 0.;x.
x0.
+1
2i0.
+1
2i.
0
Although these can be used to construct classical mechanics actions, their quantiza-
tion is rather complicated. Just as for particles of one particular spin, direct treatmentof the quantum mechanics has proven much simpler than deriving it by quantization
of a classical system.
Excercise IIC2.1
Check explicitly the invariance of the above innitesimal and nite dierencesunder supersymmetry.
Now that we have a (super)coordinate representation of the supersymmetry gen-
erators, we can examine the wave functions/elds that carry this representation. Such
\superelds" can be Taylor expanded in the 's with a nite number of terms, with
ordinary elds as the coecients. For example, if we expand a real (hermitian) scalar
supereld
(x;;)=(x)+
(x)+. .(x)+:::
and also expand its supersymmetry transformation
= +. .+1
2i.
@
.
+1
2i.@.+:::
we nd the component eld transformations
= +. .; = 1
2i.
@
.
+:::; .= 1
2i@.+:::; :::
which mix the dierent spins.
An alternative, and more convenient, way to dene the expansion is by use of
the covariant derivatives. Using \ j"t om e a n\j=0", we can dene
=j; =(d)j; .=(d.)j; :::
There is some ambiguity at higher orders in because the d's don't anticommute, and
this can be resolved according to whatever is convenient for the particular problem,
avoiding eld redenitions in terms of elds appearing at lower order in :S i n c e
the eld equations must be covariant under supersymmetry (otherwise there is no
106 II. SPIN
advantage to using superelds), they must be written with the covariant derivatives.
Then one denes the component expansions by choosing the same ordering of d's as
appear in the eld equations (where relevant), which gives the component expansion
of the eld equations the simplest form. It also gives a convenient method for deriving
supersymmetry transformations, since the d's anticommute with the q's:
[(d:::d)j]=[d:::d()]j=[d:::d(iq)]j=[ (iq)d:::d]j=[ (d)d:::d]j
w h e r ew eh a v eu s e dt h ef a c tt h a t q= id+-stu, where the -stu is killed by
evaluating at = 0, once it has been pulled in front of all the -derivatives. Covariant
derivatives can also be used for integration, sinceR
d=@=@ =dup to anx-
derivative, which can be dropped when also integratingR
dx.
3. Supergroups
We saw certain relations between the lower-dimensional classical groups that
turned out to be useful for just the cases of physical interest of rotational (SO(D 1)),
Lorentz (SO(D 1,1)), and conformal (SO(D,2)) groups. In particular, the Poincar e
group, though not a classical group, is a certain limit (\contraction") of the groupsSO(D,1) and SO(D 1,2), and a subgroup of the conformal group. Similar remarks
apply to supersymmetry, but because of its relation to spinors, these classical \su-
pergroups" (or \graded" classical groups) exist only for certain lower dimensions, thesame as those where covering groups for the orthogonal groups exist. In higher di-
mensions the supergroups do not correspond to supersymmetry, at least not in any
way that can be represented on physical states.
We'll consider only the graded generalization of the classical groups that appear
in the bosonic case. The basic idea is then to take the group metrics and combine
them in ways that take into account the dierence in symmetry between bosons and
fermions:
Unitary: .
AB
OrthoSymplectic:MAB
Real:.
AB
pseudoreal (*):
.
AB
whereis symmetric and
antisymmetric, as before, while Mis graded symmetric:
ForA=(a;) with bosonic indices aand fermionic ones ,
M[AB)=0:Mab Mba=Ma Ma=M+M=0
C. SUPERSYMMETRY 107
A g a i nw eh a v ei n v e r s em e t r i c s ,e . g . ,
MKIMKJ=I
J
With respect to the usual index-contraction convention (no extra grading signs when
superscript is contracted with subscript immediately following), we should take theordering of indices on as
JI.
T h e r ei sn oa n a l o go ft h e tensor, at least for nite-dimensional groups, since it
would have an innite number of indices when totally symmetric. However, \special"
supergroups can still be dened by generalizing the denition of trace and determinantto supermatrices. One convenient way to do this is by using Gaussian integrals, since
this is a common way that such expressions will arise. As a generalization of the
bosonic and fermionic identities we therefore dene the \superdeterminant"
(sdetM )
1=NZ
dzydz e zyMz
where \N" is a normalization factor dened so sdet I = 1. By explicitly evaluating
the integral, separating out the commuting and anticommuting parts, we nd (see
excercise IB3.3)
sdetAB
CD
=det A
det(D CA 1B)=det(A BD 1C)
det D
The \supertrace" (see also excercise IA2.3c) then can be dened by generalizing the
bosonic identity det(eM)=etrM:
sdet(eM)=estrM
str(MAB)=( 1)AMAA=Maa M=tr A tr D
follows, as in the bosonic case, from lns d e tM =str(M 1M), which is derived by
varying the Gaussian denition.
A useful identity for superdeterminants can be derived by starting with the fol-
lowing identity for the inverse of a matrix for which the range of the indices has been
divided into two pieces:
ab
cd 1
=(a bd 1c) 1(c db 1a) 1
(b ac 1d) 1(d ca 1b) 1
We have assumed all the submatrices are square and invertible; equivalent expressions,
which are more useful in other cases, can be derived easily by multiplying and dividing
by the submatrices: For example,
ab
cd 1
=(a bd 1c) 1 a 1b(d ca 1b) 1
d 1c(a bd 1c) 1(d ca 1b) 1
108 II. SPIN
From either of these we immediately see
AB
CD 1
=~A~B
~C~D
)sdetAB
CD
=detA det ~D=1
det D det ~A
The graded generalizations of the classical groups are then
GL(mjn,C) [SL(mjn,C),SSL(njn,C)]
U: [S]U(m +,m jn) [SSU(n +,n jn++n )]
OSp: OSp(mj2n,C)
R: GL(mjn) [SL(mjn),SSL(njn)]
*: [S]U*(2mj2n) [SSU*(2nj2n)]
U&O S p
R: OSp(m +,m j2n)
*: OSp*(2mj2n)
where \(mjn)" refers to mbosonic and nfermionic indices, or vice versa. In the
matrices of the dening representation, the elements with one bosonic index and onefermionic are anticommuting numbers, while those with both indices of the same kindare commuting. In particular, the commuting parts give the bosonic subgroups:
GL(mjn,C)GL(m,C)
GL(n,C)
SL(mjn,C)GL(m,C)
SL(n,C)
SSL(njn,C)SL(n,C)
SL(n,C)
U(m
+,m jn)U(m +,m )
U(n)
SU(m +,m jn)U(m +,m )
SU(n)
SSU(n +,n jn++n )SU(n +,n )
SU(n ++n )
OSp(mj2n,C)SO(m,C)
Sp(2n,C)
GL(mjn)GL(m)
GL(n)
SL(mjn)GL(m)
SL(n)
SSL(njn)SL(n)
SL(n)
U*(2mj2n)U*(2m)
U*(2n)
SU*(2mj2n)U*(2m)
SU*(2n)
SSU*(2nj2n)SU*(2n)
SU*(2n)
OSp(m +,m j2n)SO(m +,m )
Sp(2n)
OSp*(2mj2n)SO*(2m)
USp(2n)
C. SUPERSYMMETRY 109
When the commuting and anticommuting dimensions are equal, we can impose trace-
lessness conditions on both bosonic parts of the generators separately (\SS": tr A =
tr D = 0). This is related to the fact str(I)=0i ns u c hc a s e s .
4. Superconformal
Since the conformal group is a classical group, its supersymmetric generalization
should be a classical supergroup. Because the fermionic generators must include the
supersymmetry generators, which are spinors, the representation of the conformal
group that appears in the dening representation of the supergroup must be the spinor
representation. However, we have seen that only for n 6 (where covering groups exist)
and n=8 (where the spinor of SO(8) is another of its dening representations) can
the spinor representation of SO(n) be dened by classical group restrictions. This
implies that the superconformal group exists only in D 4a n dD = 6 .
The relevant supergroups can be identied easily by looking at the bosonic sub-
groups:
D= 3 : OSp(Nj4)
4: S U ( 2 , 2jN) (or SSU(2,2j4))
6 : OSp*(8j2N)
(We consider only D >2, since the conformal group is innite-dimensional in D 2.)
These three cases of D=3,4,6 are special for a number of reasons: In particular,these three supergroups can be related to SU(N j4) over the division algebras: the
real numbers, complex numbers, and quaternions, respectively. (Similar remarks
apply to their important classical bosonic subgroups: the conformal, Lorentz, and
rotation groups. Attempts have been made to extend these results to the octonions
for D=10, but with less success, and there seems to be no superconformal group forthat case.) However, just as in the case of the Hilbert space of quantum mechanics,
the complex numbers seems to be the best of these \division algebras", having the
analytic properties the real numbers lack, while avoiding the noncommutativity ofthe quaternions. We'll see later that nontrivial interacting (local, classical) conformal
eld theories exist only in D 4.
For example, for D=4 we nd that the bosonic generators are the conformal group
and the internal symmetry group U(N) (or SU(4) for N=4), while the fermionic gener-ators include supersymmetry (N spinors) and its fraternal twin, \S-supersymmetry".
As supersymmetry is the \square root" of translations, so S-supersymmetry is the
square-root of conformal boosts.
110 II. SPIN
Excercise IIC4.1
For D=4, write the (graded) commutation relations of the superconformal
generators. Decompose them into representations of the Lorentz group, andnd their commutation relations.
5. Supertwistors
We saw that a simple way to nd representations of SO(4,2) was to use the
coordinate representation for SU(2,2): The resulting twistors gave all massless rep-resentations ( p
2= 0 for all helicities). This method generalizes straightforwardly to
the superconformal groups: The generators are
GAB=BA
(For the SU case we should also subtract out the trace, but that generator commutes
with the rest anyway.) The coordinates and their conjugate momenta satisfy
[A;Bg=B
A
Ais then in the dening representation of the supergroup, while the wave function,
which is a function of , contains more general representations.
For D=3, the reality condition sets =,s ot h e's are the graded generalization
of Dirac
matrices. In fact, the anticommuting 's are the
matrices of the SO(N)
subgroup of the OSp(N j4). On the other hand, the commuting 's carry the index of
the dening representation of Sp(4), so they are a spinor of SO(3,2), the 3D conformal
group: They are the bosonic twistor, and can be used in a similar way to the 4Dtwistors discussed earlier.
For D=4, there is a U(1) symmetry acting on under which G
ABis invariant,
generated by ( 1)AGAA, as in the bosonic case: This is the \superhelicity".
For D=6,is pseudoreal. In general, for pseudoreal representations of groups it is
often convenient to introduce a new SU(2) under which the pseudoreal representation
Aand its equivalent complex conjugate representation .
B
.
BAtransform as a doublet
(SU(2) spinor). This is also obvious from construction, since half of the components
are related to the complex conjugate of the other half. We then can write
Ai=(A;i.
B
.
BA)=.
B.
k
.
B.
kAi;
.
B.
kAi=
.
BACki;MAi;Bk=MABCik
GAB=BiAi
C. SUPERSYMMETRY 111
(Thus, OSp*(2mj2n)OSp(4nj4m), and SO*(2m) Sp(4n), USp(2n)SO(4m).) This
means there is now an SU(2) symmetry on , generated by
Gij=A
(iAj)
under which GABis invariant. This is the 6D version of superhelicity. In the D=6
light cone, the manifest part of Lorentz invariance is SO(D 2)=SO(4)=SU(2)
SU(2).
This is one of those SU(2)'s.
We now concentrate on D=4 (although our methods generalize straightforwardly
to D=3 and 6). The simplest way to nd (massless) representations of 4D super-
symmetry is to generalize the Penrose transform. Just as twistors automatically
satisfy the massless eld equations in D=4, supertwistors automatically satisfy their
supersymmetric generalization. The supertwistor is the dening representation of
SU(2,2jN). The SU(2,2) part is the usual twistor, while the SU(N) part is the usual
fermionic creation and annihilation operators for SU(N). Thus, to relate superspace
to supertwistors, we write
p
.
= i@
.
!pp.
qi= i@i+1
2.
i@
.
!aip; qi.= i@i.+1
2i@.!ayip.
This determines the Penrose transform from superspace to supertwistors:
:::.
:::(x;)=Z
d2pd2p.dNaipp.
[ei'+(p;p.;ai)+e i' (p;p.;ai)]
'=(x.
i1
2i.
i)pp.
+iaip
where we have used \chiral superelds" (trivial dependence on ) without loss of
generality. (Instead of treating aias coordinates to be integrated, we can also treat
them as operators; we then make the 's functions of ay, and replace the integration
with vacuum evaluation h0jj0i.) As for ordinary twistors, this result can be related
to the lightcone: For given momentum, we can choose the lightcone frame p=
;
thenq
i=ai, whileq
i= 0 is a result of the supertwistor formalism automatically
incorporating p=q=0 .
Excercise IIC5.1
Find the Penrose transform for D=3. (Warning: The anticommuting part
of the twistor is now like Dirac matrices rather than creation/annihilation
operators.)
112 II. SPIN
Taylor expanding in ai(and thusi, producing terms antisymmetric in i:::jand
symmetric in ::: ), the states then carry the index structure ;i;ij;:::;~i;~, totally
antisymmetric, and terminating with another singlet, where
~=1
N!i1iNi1iN; ~i1=1
(N 1)!i1iNi2iN; :::
From our discussion of helicity in subsection IIB7, we see that the states also decrease
in helicity by 1/2 for each a(i.e., ignoring ,e a c hicomes with a p, simply because
it adds an undotted index). Taking the direct product with any helicity (coming fromthe explicit p
's and p.'s carrying the external Lorentz indices), we see that the states
have helicity h;h 1=2;h 1;:::;h N=2, with multiplicty N
n
for helicity h n=2:
state helicity (Poincar e) multiplicity [SU(N)]
h 1
i h 1
2 N
ij h 1N(N 1)
2.........
i1inh n
2 N
n
...... ...
~ih N
2+1
2N
~ h N
21
This multiplet structure is carried separately by +and by , which are related
by charge (complex) conjugation, one describing the antiparticles of the other, as forordinary twistors. (The existence of both multiplets also follows from CPT invariance,which is required for local actions, to be discussed in subsection IVB1. Here wegeneralized from the Penrose transform, which contained both terms as a consequence
of being the most general solution to S
abpb+wpa= 0, which is CPT invariant.)
Because of the values of the helicities, we can impose a reality condition, identifyingall states with helicity jas the complex conjugates of those with j,o n l yf o r h=
h N=2!h=N=4, whenNis a multiple of 4. We can also get larger representations
by taking the direct product of these smallest representations of supersymmetry with
representations of U(N), in which case the elds will carry those additional SU(N)
indices.
Excercise IIC5.2
For D=4, show that \supergravity", the supersymmetric theory with helicitiesof magnitude 2 and lower, can exist only for N 8. Show that the relevant
C. SUPERSYMMETRY 113
representation for N=8, if real, is the same as the one (complex) for N=7.
Find the analogous statements for \super Yang-M ills", with helicities 1 and
less.
The explicit form of the reality condition is somewhat complicated in terms of
the chiral superelds, because they are really eld strengths of real gauge elds.
(Consider, for example, expressing reality of A
.
in terms of fi nt h ec a s eo fe l e c -
tromagnetism.) However, in terms of the twistor variables, charge conjugation can
be expressed as
C:!*;ai!(ai)y
where the transformation on aiis required because it carries the SU(N) \charge".
Since this violates \chirality" in these variables (dependence on aand notay), it is
accomplished by Fourier transformation:
C:(ai)!CZ
d~ayie~ayiai[(~ai)]*C 1
for some \charge conjugation matrix" C(in case the eld carries an additional index).
REFERENCES
1P. Ramond, Phys. Rev. D3(1971) 86;
A. Neveu and J.H. Schwarz, Nucl. Phys. B31 (1971) 86, Phys. Rev. D4(1971) 1109:
2D supersymmetry in strings.
2Yu.A. Gol'fand and E.P. Likhtman, JETP Lett. 13(1971) 323;
D.V. Volkov and V.P. Akulov, Phys. Lett. 46B (1973) 109;
J. Wess and B. Zumino, Nucl. Phys. B70 (1974) 39:
4D supersymmetry.
3A. Salam and J. Strathdee, Nucl. Phys. B76 (1974) 477;
S. Ferrara, B. Zumino, and J. Wess, Phys. Lett. 51B (1974) 239:
superspace.
4Gates, Grisaru, Ro cek, and Siegel, loc. cit. ;
J. Wess and J. Bagger, Supersymmetry and supergravity , 2nd ed. (Princeton University,
1992);P. West, Introduction to supersymmetry and supergravity , 2nd ed. (World-Scientic,
1990);I.L. Buchbinder and S.M. Kuzenko, Ideas and methods of supersymmetry and super-
gravity, or a walk through superspace (Institute of Physics, 1995).
5Berezin, loc. cit. (IA):
superdeterminant.
6P. Ramond, Physica 15D (1985) 25:
supergroups and their relation to supersymmetry.
7J. Wess and B. Zumino, Nucl. Phys. B70 (1974) 39:
superconformal group in D=4.
8R. Haag, J.T. Lopusza nski, and M. Sohnius, Nucl. Phys. B88 (1975) 257:
superconformal symmetry as the largest symmetry of the S-matrix.
114 II. SPIN
9A. Ferber, Nucl. Phys. B132 (1978) 55:
supertwistors.
10T. Kugo and P. Townsend, Nucl. Phys. B221 (1983) 357;
A. Sudbery, J. Phys. A17 (1984) 939;
K.-W. Chung and A. Sudbery, Phys. Lett. 198B (1987) 161:
spacetime symmetries and division algebras.
A. ACTIONS 115
III. LOCAL
In the previous chapters we considered symmetries acting on coordinates or wave
functions. For the most part, the transformations we considered had constant param-eters: They were \global" transformations. In this chapter we will consider mostlyeld theory. Since elds are functions of spacetime, it will be natural to considertransformations whose parameters are also functions of spacetime, especially thosethat are localized in some small region. Such \local" or \gauge" transformations arefundamental in dening the theories that describe the fundamental interactions.
:::::::::::::::::::::::::::: ::::::::::::::::::::::::::::
::::::::::::::::::::::::::::
A. ACTIONS ::::::::::::::::::::::::::::
A fundamental concept in physics, of as great importance as symmetry, is the
action principle. In quantum physics the dynamics is necessarily formulated in termsof an action (in the path-integral approach), or an equivalent Hamiltonian (in theHeisenberg and Schr odinger approaches). Action principles are also convenient and
powerful for classical physics, allowing all eld equations to be derived from a single
function, and making symmetries simpler to check.
1. General
We begin with some general properties of actions. (For this subsection we'll re-
strict ourselves to bosonic variables; however, in the following subsection we'll ndthat the only modication for fermions is a more careful treatment of signs.) Gen-erally, equations of motion are derived from actions by setting their variation with
respect to their arguments to vanish:
S[]S[+] S[]=0
Here the variables are themselves functions of time; thus, Sis a function of functions,
a \functional". A general principle of mechanics is \locality", that events at one timedirectly aect only those events an innitesimal time away. (In eld theory theseevents can be also only an innitesimal distance away in space.) This means that theaction can be expressed in terms of a Lagrangian:
S[]=Z
dt L[(t)]
whereLat timetis a function of only (t) and a nite number of its derivatives. For
more subtle reasons, this number of time derivatives is restricted to be no more than
116 III. LOCAL
two for any term in L; after integration by parts, each derivative acts on a dierent
factor of. The general form of the action is then
L()= 1
2.
m.
ngmn()+.
mAm()+U()
where \."m e a n s@=@t, and the \metric" g, \vector potential" A, and \scalar po-
tential"Uare not to be varied independently when deriving the equations of motion.
(Specically, U=(m)(@U=@m), etc. Note that our denition of the Lagrangian
diers in sign from the usual.) The equations of motion following from varying an
action that can be written in terms of a Lagrangian are
0=SZ
dt mS
m)S
m=0
where we have eliminated .
mterms by integration by parts (assuming =0a tt h e
boundaries in t), thereby dening the \functional derivative" S=m, and used the
fact that(t) is arbitrary at each value of t. For example,
S= Z
dt1
2.q2) 0=S= Z
dt.q.q=Z
dt(q)..q)S
q=..q=0
Excercise IIIA1.1
Find the equations of motion for mfrom the above general action in terms
of the external elds g,A,a n dU(and their partial derivatives with respect
to).
Sometimes the functional derivative is dened in terms of that of the variable
itself:m(t)
n(t0)=m
n(t t0)
wherem
nis the usual Kronecker delta function, while (t t0) is the \Dirac delta
function". It's not really a function, since it takes only the values 0 or 1, but a
\distribution", meaning it's dened only by integration:
Z
dt0f(t0)(t t0)=f(t)
If we apply this denition of the Dirac to= , we obtain the previous denition
of the functional derivative. (Consider, e.g., S=Rdt f .)
Such actions can be reduced to ones that are only linear in time derivatives
by introducing additional variables. First, separate out the subspace where gis
invertible, with coordinates q(m=(qi; )); the Lagrangian is then written as
L(q; )= 1
2.qi.qjgij(q; )+.qiAi(q; )+.
A(q; )+U(q; )
A. ACTIONS 117
This Lagrangian gives equivalent equations of motion to
L0(q;p; )=[ .qipi+.
A]+[1
2gij(pi+Ai)(pj+Aj)+U]
wheregijis the inverse of gij. (Many other forms are possible by redenitions of
p.) Eliminating the new variables pby their equations of motion gives back L(q; ).
Note that this works only because p's equations of motion are algebraic: For example,
eliminating xfrom the Lagrangian .xp+1
2p2by the equation of motion.x=pis illegal
(it would give the trivial action S=R
dt1
2p2), since it would require solving for the
time dependence of x. On the other hand, pis given explicitly in terms of the other
variables by its equations of motion without inverting time derivatives, so eliminatingit does not lose any of the dynamics. (It is an \auxiliary variable".)
The result is a Hamiltonian form of the Lagrangian:
L
H() =i.
MAM() +H()
in terms of the Hamiltonian H,w h e r e=( q;p; ). It has the \gauge invariance"
AM=@M()
(where@M=@=@M), since that adds only a total derivative term i.
. ClearlyA
will introduce a modication of the Poisson bracket if it is not linear in (e.g., aswhen we make independent nonlinear redenitions of coordinates and momenta on
the usual form of the Lagrangian). To determine this modication we compare the
equation of motion as dened by a Poisson bracket,
.
M= i[M;H]= i[M;N]@NH
with that following from varying the action,
i.
NFNM+@MH=0;FMN=@[MAN]
to nd
[M;N]=(F 1)NM
where \F 1" is the inverse on the maximal subspace where Fis invertible. The
variables in the directions where Fvanishes are \auxiliary", since they appear without
time derivatives: Their equations of motion are not described by the Poisson bracket.In particular, if they appear linearly in Hthey are \Lagrange multipliers", whose
variation imposes algebraic constraints on the rest of .
118 III. LOCAL
Finally, we can make redenitions of the part of describing the invertible sub-
space so thatAis linear:
AM=1
2N
NM)LH() =1
2i.
MN
NM+H()
where
is a constant, hermitian, antisymmetric (and thus imaginary) matrix. For
some purposes it is more convenient to assume this Hamiltonian form of the actionas a starting point. We now have the canonical commutation relations as
[
M;N]=
MN
where
MNis the inverse of
NMon the maximal subspace:
MN
PN=PM
for the projection operator for that subspace.
Excercise IIIA1.2
For electromagnetism, dene ~ =~E+i~B. Show that Maxwell's equations
(in empty space) can be written as two equations in terms of ~ . Interpret
the equation involving the time derivative as a Schr odinger equation for the
wave function ~ , and nd the Hamiltonian operator. Dene the obvious inner
productR
d3x~ *~ : What physical conserved quantity does this represent?
(Note that, unlike electrons, the number of photons is not conserved.)
Note that the requirement of the existence of a Hamiltonian formulation deter-
mines that the kinetic term for a particle in the Lagrangian formulation go as.x2
and notx..x. Although such terms give the same equations of motion, they are not
equivalent quantum mechanically, where boundary terms (dropped when using inte-gration by parts for deriving the equations of motion) contribute. Furthermore, theHamiltonian form of the action
S=Z
dt H dx
ipi
shows that the energy Hrelates to the time in the same way the momentum relates to
the coordinates, except for an interesting minus sign that is explained only by special
relativity.
A. ACTIONS 119
2. Fermions
In nonrelativistic quantum mechanics, spin is usually treated as a quantum eect,
rather than being derived from classical mechanics. Although it is possible to derivespin from classical mechanics, in general it is rather cumbersome, and involves rst
introducing a large number of spins and then constraining away all the undesired ones,
whereas in the quantum mechanics one can just directly introduce some particularrepresentation of the spin angular momentum operators. The one nontrivial exception
is spin 1/2.
We know from quantum mechanics that the spin variables for spin 1/2 are de-
scribed by the Pauli matrices. Since they satisfy anticommutation relations, and
are represented by nite-dimensional matrices, they are interpreted as fermionic. We
have already seen that classical fermions are described by anticommuting numbers,
so we begin by considering general quantization of such objects.
We can now consider actions that depend on both commuting and anticommuting
classical variables,
M=(m; ), where now refers to the bosonic variables and
to the fermionic ones. The Hamiltonian form of the Lagrangian can again be written
as
LH() =1
2i.
MN
NM+H()
When
is invertible, the graded bracket is dened by (see subsection IA2)
[M;Ng=h
MN;
MN
PN=M
P
To describe spin 1/2, we therefore look for particle actions of the form
SH=Z
dt[ .xipi+1
2i.
i i+H(x;p; )]
This corresponds to using
M=(m; )=(i; i)=(xi;pi; i)
mn=
i;j=ijC;
=
ij=ij;
m=
n=0
The fundamental commutation relations are then
[xi;pj]=ihi
j;f i; jg=hij([x;x]=[p;p]=[x; ]=[p; ]=0 )
We recognize ias the Pauli matrices (the Dirac matrices of subsection IC1 for the
special case of SO(3)), i=p
hi. The free Hamiltonian is just
H=p2
2m
120 III. LOCAL
as for spin 0: Spin does not aect the motion of free particles.
A more interesting case is coupling to electromagnetism: Quantum mechanically,
the Hamiltonian can be written in the simple form
H=f i[pi+qAi(x)]g2
mh qA0(x)
in terms of the vector and scalar potentials AiandA0. The classical expression is not
as simple, because the commutation relations must be used to cancel the 1 =hbefore
taking the classical limit. This is an example of \minimal coupling",
H(pi)!H(pi+qAi) qA0
However, this prescription works only if Hfor spin 1/2 is written in the above form:
Using the commutation relations before or after minimal coupling gives dierent re-
sults. The form we have used is justied only by considering the nonrelativistic limit
of the relativistic theory.
Excercise IIIA2.1
Use the multipication rules of the matrices to show that the quantum me-
chanical Hamiltonian for spin 1/2 in an electromagnetic eld can be writtenas a spin-independent piece, identical to the spin-0 Hamiltonian, plus a term
coupling the spin to the magnetic eld.
3. Fields
The eld equations for all eld theories (e.g., electromagnetism) are wave equa-
tions. Wave equations also follow from mechanics upon quantization. Although
classical eld theory and quantum mechanics are not equivalent in their physical in-
terpretation, they are mathematically equivalent in that they have identical wave
equations. This is true not only for the free theories, but also for particles in external
elds, and without direct self-interactions. This is no accident: Classical eld theoryand classical mechanics are two dierent limits of quantum eld theory. They are
both called classical limits, and written as h!0, but since his really 1, this limit
depends on how one inserts h's into the quantum eld theory action.
The wave equation in quantum mechanics is the Schr odinger equation. The cor-
responding eld theory action is then simply the one that gives this wave equation
as the equation of motion, where the wave function is replaced with the eld:
S
ft=Z
d4x *( i@t+H)
A. ACTIONS 121
As usual (cf. electromagnetism), the eld is a function of space and time; thus, we
integrated4x=dt d3xover the three space and one time dimensions. The Hamilto-
nian is some function of coordinates and momenta, with the replacement pi! i@i,
where@i=@=@xiare the space derivatives and @t=@=@t is the time derivative.
The Hamiltonian can contain coupling to other elds. For a general Hamilto-
nian quadratic in momenta, in a notation implied by the corresponding Lagrangian
quadratic in time derivatives,
H=1
2gij( i@i+Ai)( i@j+Aj)+U
wheregij,Ai,a n dUare now interpreted as elds, and thus depend on both xiandt,
as does .I n t h e c a s e gij=ij,w ec a ni d e n t i f y AiandUas the three-vector and scalar
potentials of electromagnetism, and we can add the usual action for electromagnetismto the action for . The action then can be varied also with respect to AandUto
obtain Maxwell's equations with a current in terms of and *. We can also treat
g
ijas a eld, in which case it and parts of AandUare the components of the
gravitational eld.
Field theory actions can be quantized in the same ways as mechanics ones. In
this case, we recognize the *.
term as a special case of the.
term in the
generic Hamiltonian form of the action discussed earlier. Thus, (xi)a n d *(xi)
have replaced xiandpias the variables; xiis now just an index (label) on and *,
just asiwas an index on xiandpi. The eld-theory Hamiltonian is then identied
as
Hft[ ; *] =Z
d3xH;H= *H
In eld theory the Hamiltonian will always be a space integral of a \Hamiltonian
density"H.
We can now dene the two classical limits of quantum eld theory. If we put in
hin the generic way for actions,
Sft!h 1Sft
then we dene the classical limit h!0 as classical eld theory, since in that limit
the classical eld equations are preserved. On the other hand, if we put in h's as
@i!h@i;@t!h@t
which gives the usual hdependence associated with the Schr odinger equation, then
the classical limit h!0 gives classical mechanics. This denes classical mechanics
as the macroscopic limit, the limit of large distances and times.
122 III. LOCAL
A convenient way to implement this limit is to introduce the mechanics action
S=R
dt( .xipi+H) into the eld theory, and then take the limit h!0a f t e rt h e
replacement
S!h 1S
on the mechanics action instead of on the derivatives. The mechanics action can be
introduced when solving the eld equations: The solution to the wave equation canbe expressed in terms of the propagator, which in turn can be written in terms of the
mechanics action or Hamiltonian.
More generally, we can dene actions that are not restricted to be quadratic in
any eld. The Hamiltonian density H(t;x
i) or Lagrangian density L(t;xi),
S[]=Z
dt d3xL[(t;xi)]
should be a function of elds at that point, with only a nite number (usually no
more than two) spacetime derivatives. This is the denition of locality used for gen-
eral quantum systems in subsection IIIA1, but extended from derivatives in time toalso those in space. Although this condition is not always used in nonrelativistic
eld theory (for example, when long-range interactions, such as Coulomb or gravita-
tional, are described without attributing them to elds), it is crucial in relativisticeld theory. For example, global symmetries lead by locality to local (current) conser-
vation laws. Locality is also the reason that spacetime coordinates are so important:
Translation invariance says that the position of the origin is an unphysical, redundant
variable; however, locality is most easily used with this redundancy.
Field equations are derived by the straightforward generalization of the variation
of actions dened in subsection IIIA1: As follows from treating the spatial coordinates
in the same way as discrete indices,
SZ
dt d
3xm(t;xi)S
m(t;xi)
For example,
S= Z
dt d3x1
2.
2)S
=..
4. Relativity
Generalization to relativistic theories is straightforward, except for the fact that
the Klein-Gordon equation is second-order in time derivatives; however, we are fa-miliar with such actions from nonrelativistic quantum mechanics. As usual, we need
to check the sign of the terms in the action: Checking the positivity of the Hamil-
tonian (i.e., the energy), we see from the general relation between the Lagrangian
A. ACTIONS 123
and Hamiltonian (subsection IIIA1) that the terms without time derivatives must be
positive; the time-derivative terms are then determined by Lorentz covariance.
At this point we introduce some normalizations and conventions that will prove
convenient for Fourier transformation and other reasons to be explained later. When-ever D-dimensional integrations are involved (as should be clear from context), we
use Z
dxZd
Dx
(2)D=2;Z
dpZdDp
(2)D=2
(x x0)(2)D=2D(x x0); (p p0)(2)D=2D(p p0)
In particular, this normalization will be used in Green functions and actions. For
example, these implicit 2 's appear in functional variations:
SZ
dx S
)
(x)(x0)=(x x0)
The action for a real scalar is then
S=Z
dx L; L =1
4(@)2+V()
whereV()0, and we now write Lfor the Lagrange density. In particular, V=
1
4m22for the free theory. The free eld equation is then p2+m2= +m2=0 ,
replacing the nonrelativistic i@t+H= 0. For a complex scalar, we replace1
2!
in both terms.
We know from previous considerations (subsection IIB2) that the eld equation
for a free, massless, Dirac spinor is
@ = 0. The generalization to the massive case
(subsection IIB4) is obvious from various considerations, e.g., dimensional analysis;
the action is
S=Z
dx (i@=+mp
2)
in arbitrary dimensions, again using the notation @==
@. In four dimensions, we
can decompose the Dirac spinor into its two Weyl spinors (see subsection IIA6):
L= (i@=+mp
2) = ( .
Li@. L+ .
Ri@. R)+mp
2(
L R+ .
L R.)
For the case of the Majorana spinor, the 4D action reduces to that for a single Weyl
spinor,
S=Z
dx[ i .
@
.
+mp
21
2( + . .)]
Note that in our conventions 0
.
=1p
2(and similarly for the opposite indices, since
a
.
=.
a), so that the time derivative term is always proportional to y( i@0) ,a s
nonrelativistically (previous subsection).
124 III. LOCAL
A scalar eld must be complex to be charged (i.e., a representation of U(1)):
From the gauge transformation
0=ei
we nd the minimal coupling (for q=1 )
S=Z
dx[1
2j(@+iA)j2+1
2m2jj2]
This action is also invariant under charge conjugation
C:!*;A! A
which changes the sign of the charge, since *0=e i*.
Excercise IIIA4.1
Let's consider the semiclassical interpretation of a charged particle as de-
scribed by a complex scalar eld , with Lagrangian
L=1
2(jr j2+m2j j2)
aUse the semiclassical expansion in hdened by
r! h@+iqA; !pe iS=h
Find the Lagrangian in terms of andS(and the background eld A), order-
by-order in h(in this case, just h0and h2).
bTake the semiclassical limit by dropping the h2term inL, to nd
L!1
2[( @S+qA)2+m2]
Vary with respect to Sandto nd the equations of motion. Dening
p @S
show that these eld equations can be interpreted as the mass-shell condition
and current conservation. Show that Acouples to this current by varying L
with respect to A.
The spinor eld also needs doubling for charge. (Actually, the doubling can be
avoided in the massless case; however, problems show up at the quantum level, related
to the fact that there is no charge conjugation transformation without doubling.) The
gauge transformations are similar to the scalar case, and the action again follows from
A. ACTIONS 125
minimal coupling, to an action that has the global invariance ( = constant in the
absence ofA):
0
L=ei
L; 0
R=e i
R
Se=Z
dx[ .
L( i@
.
+A
.
)
L+ .
R( i@
.
A
.
)
R+mp
2(
L R+ .
L R.)]
The current is found from varying with respect to A:
J.
= .
L
L .
R
R
Charge conjugation
C:
L$
R;A! A
(which commutes with Poincar e transformations) changes the sign of the charge and
current.
Excercise IIIA4.2
Show that this action can be rewritten in Dirac notation as
Se=Z
dx (i@= A=+mp
2)
and nd the action of the gauge transformation and charge conjugation on
the Dirac spinor.
As a last example, we consider the action for electromagnetism itself. As before,
we have the gauge invariance and eld strength
A0
.
=A
.
@
.
F
.
;.
=@.
A
.
@
.
A.
=Cf.
.
+C.
.
f;f=1
2@(.
A).
We can write the action for pure electromagnetism as
SA=Z
dx1
2e2ff=Z
dx1
2e2f..
f..
=Z
dx1
8e2FabFab
dropping boundary terms, with the overall sign again determined by positivity of the
Hamiltonian, where eis the electromagnetic coupling constant, i.e., the charge of the
proton. (Other normalizations can be used by rescaling A
.
.) Maxwell's equations
follow from varying the action with a source term added:
S=SA+Z
dx A.
J
.
)1
e2@.
f=J.
Excercise IIIA4.3
By plugging in the appropriate expressions in terms of Aa(and repeatedly
126 III. LOCAL
integrating by parts), show that all of the above expressions for the electro-
magnetism action can be written as
SA= Z
dx1
4e2[AA+(@A)2]
Excercise IIIA4.4
Find all the eld equations for all the elds, found from adding to SAall the
minimally coupled matter actions above.
The energy-momentum tensor for electromagnetism is much simpler in this spinor
notation, and follows (up to normalization) from gauge invariance, dimensional anal-
ysis, Lorentz invariance, and the vanishing of its trace. It has a form similar to that
of the current in electrodynamics:
T
.
.
= 1
e2ff.
.
Note that it is invariant under the duality transformations of subsection IIA7 (as is
the electrodynamic current under chirality).
We have used conventions where eappears multiplying only the action SA,a n d
not in the \covariant derivative"
r=@+iqA
whereqis the charge in units of e: e.g.,q= 1 for the proton, q= 1 for the electron.
Alternatively, we can scale A, as a eld redenition, to produce the opposite situation:
A!eA:SA!Z
dx1
8F2;r!@+iqeA
The former form, which we use unless noted otherwise, has the advantage that the
coupling appears only in the one term SA, while the latter has the advantage that the
kinetic (free) term for Ais normalized the same way as for scalars. The former form
has the further advantage that eappears in the gauge transformations of none of the
elds, making it clear that the group theory does not depend on the value of e.( T h i s
will be more important when generalizing to nonabelian groups in section IIIC.)
Note that the massless part of the kinetic (free) terms in these actions are scale
invariant (in arbitrary dimensions, when the dimension-independent forms are used),when the elds are assigned the scale weights found from conformal arguments in
subsection IIB2.
Excercise IIIA4.5
Using vector notation, minimal coupling, and dimensional analysis, nd the
A. ACTIONS 127
mass dimensions of the electric charge ein arbitrary spacetime dimensions,
and show it is dimensionless only in D=4 .
An interesting distinction between gravity and electromagnetism is that static
bodies always attract gravitationally, whereas electrically they repel if they are likeand attract if they are opposite. This is a direct consequence of the fact that the
graviton has spin 2 while the photon has spin 1: The Lagrangian for a eld of integer
spinscoupled to a current, in an appropriate gauge and the weak-eld approximation,
is
1
4s!a1:::asa1:::as+1
s!a1:::asJa1:::as
where the sign of the rst term is xed by unitarity in quantum eld theory. (Clas-
sically the sign can also be related to positivity of the energy.) From a scalar eld in
the semiclassical approximation (see excercise IIIA4.1 above), starting with
Ja1:::as *$
@a1$
@as
where \A$
@B"m e a n s\A@B (@A)B", we see that the current will be of the form
Ja1:::aspa1pas
for a scalar particle, for some . (The same follows from comparing the expressions
for currents and energy-momentum tensors for particles as in subsection IIIB4 below.
The only way to get vector indices out of a scalar particle, to couple to the vectorindices for the spin of the force eld, is from momentum.) In the static approximation,
only time components contribute: We then can write this Lagrangian as, taking into
account
00= 1,
( 1)s1
4s!0:::00:::0+1
s!0:::0(p0)s
whereE=p0>0f o rap a r t i c l ea n d <0 for an antiparticle. Thus the spin-dependence
of the potential/force between two particles goes as ( E1E2)s. It then follows that
all particles attract by forces mediated by even-spin particles, and a particle and
its antiparticle attract under all forces, while repulsion will occur for odd-spin forces
between two identical particles. (We can substitute \particles of the same sign charge"
for \identical particles", and \particles of opposite sign charge" for \particle and its
antiparticle", where the charge is the coupling constant appropriate for that force.)
Excercise IIIA4.6
Show that the above current is conserved,
@a1Ja1as=0
(and the same for the other indices, by symmetry) if satises the free Klein-
Gordon equation (massless or massive).
128 III. LOCAL
5. Constrained systems
Constraints not only frequently appear in nonrelativistic physics, but are a general
feature of relativistic particles, so we now give a brief description of how they are
incorporated into actions. Consider a general action, with constraints, in Hamiltonianform:
S=Z
dt( .q
mpm+H);H =Hgi(q;p)+iGi(q;p)
(For simplicity, we consider all physical variables to be bosonic for this subsection,
but the method generalizes straightforwardly paying careful attention to signs.) This
action is a functional of qm;pm;i, which are in turn functions of t,w h e r emandirun
over any number of values. We can think of this as describing a nonrelativistic particle
with coordinates qand momenta pin terms of time t, but the form is general enough to
apply to relativistic theories. The.qpterm tells us pis canonically conjugate to q;t h e
rest of the action gives the Hamiltonian, usually quadratic in momenta. The variables
iare \Lagrange multipliers", whose variation in the action implies the constraints
Gi= 0. We then can interpret Hgias the usual (\gauge invariant") Hamiltonian. We
also require that the transformations generated by the constraints close, and that theHamiltonian be invariant:
[G
i;Gj]= ifijkGk;[Gi;Hgi]=0
(More generally, we can allow [ Gi;Hgi]= ifijGj.) This says that the constraints
don't imply any new constraints that we might have missed, and that the \energy"represented by H
giis invariant under these transformations.
We then nd that the action is invariant under the canonical transformations
(q;p)=i[iGi;(q;p)])qm=i@Gi
@pm; pm= i@Gi
@qm
0=d
dt
=@
@t+iH
=i(i)Gi i.
iGi+[jGj;iGi]
)i=.
i+jkfkji
(with(d=dt) dened as in subsection IA1), where @=@t acts on the \explicit" t
dependence (that in everything except qandp): For general expressions, the total
time derivative and total variation are given by commutators as
d
dtA=@
@tA+i[H;A]; A =0A+i[iGi;A]
A. ACTIONS 129
where0acts on everything except qandp. The action then varies under these
transformations as the integral of a total derivative, which vanishes under appropriate
boundary conditions:
SH=Z
dtd
dt[ (qm)pm+iGi]=0
The simplest example is the case with one constraint, which is linear in the vari-
ables: If the constraint is p, the gauge transformation is q=, so we gauge q=0
and use the constraint p= 0. In general, this means that for every degree of freedom
we can gauge away, the conjugate variable can be xed by the constraint. Thus,
for each constraint we eliminate 3 variables: the variable xed by the constraint, its
conjugate, and the Lagrange multiplier that enforced the constraint, which has no
conjugate. (In the Lagrangian form of the action the conjugate may not appear ex-
plicitly, so only 2 variables are eliminated.) As another example, for a nonrelativisticparticle constrained to a sphere, G=(x
i)2 1, we can change to spherical coordinates,
apply the constraint to eliminate the radial coordinate, and use the gauge invariance
to eliminate the radial component of the momentum, leaving an unconstrained the-ory in terms of angles and their conjugates. In most cases in eld theory a similar
procedure can be applied: The result is called a \unitary gauge".
The standard example of a relativistic constrained system is in eld theory |
electromagnetism. Its action can be written in \rst-order (in derivatives) formalism"by introducing an auxiliary eld G
ab:
F2!F2 G2!F2 (G F)2=2GF G2
where in the rst step we added a trivial term for Gand in the second step made
a trivial redenition of G, so elimination of Gby its algebraic equation of motion
returns the original Lagrangian. The Hamiltonian form comes from eliminating only
Gijby its eld equation, since only F0icontains time derivatives:
2GF G2!(Fij)2 4G0iF0i+2 (G0i)2
= 4.
AiG0i+[ 2 (G0i)2+(Fij)2] 4A0@iG0i
which we recognize as the three generic terms for the action in Hamiltonian form,
withG0ias the canonical momenta for Ai,a n dA0as the Lagrange multiplier. The
constraint is Gauss' law, and it generates the usual gauge transformations.
Thusiare also gauge elds for the gauge (time-dependent) transformations i(t).
They allow construction of the gauge-covariant time derivative
r=@t+iiGi;d
dt=r+iHgi)r=i[iGi;r]
130 III. LOCAL
It is convenient to transform the gauge elds away using these gauge transformations,
soH=Hgi. However, with the usual boundary conditionsR1
1dtiis gauge invariant
under the linearized transformations, so the most we could expect is to gauge ito
constants. More precisely, the group element
T
exp
iZ1
1dt i(t)Gi
is gauge invariant, where \ T" is time ordering, meaning we write the exponential of
the integral as the product of exponentials of innitesimal integrals, and order them
with respect to time, later time intervals going to the left of earlier ones. (We treat
Giquantum mechanically or use Poisson brackets when combining the exponentials.)
This is the quantum mechanical version of the time development resulting from thecorresponding term in the classical action. It is also the phase factor coming from
the innite limit of the covariant time translation
e
kr(t)=T
exp
iZt
t kdt0i(t0)Gi
e k@t
as seen from reordering the time derivatives when writing e kr(t)as the product of
exponentials of innitesimal exponents. This allows us to write the explicit gauge
transformation
e i(t)=T
exp
iZt
t0dt0i(t0)Gi
=e (t)r(t)e(t)@t; t=t t0
)r0(t)=ei(t)r(t)e i(t)=e (t)@tr(t)e(t)@t=r(t0)=@t+ii(t0)Gi
(where we dene @tto varytwhile keeping t t0xed). Thus, we can gauge to
its value at a xed time t0. Another way to see this is that varying iin the action
at a xed time gives Gi= 0 at that time, but the remaining eld equations imply.
Gi=0 ,s oGi=0a l w a y s ,a n d iis redundant at other times. This means that if we
carelessly impose i= 0 at all times, we must also impose Gi= 0 at some xed time.
Note that this special gauge transformation itself has a very simple gauge trans-
formation: Transforming the in by an arbitrary nite transformation i(t),
e i0(t)=e ii(t)Gie i(t)eii(t0)Gi
consistent with the transformation law of r0(t) above. Thus, applying the trans-
formation to any gauge-dependent quantity gives a gauge-independent quantity
0(;), which is invariant under the local transformations (t) and transforms only
under the \global" transformations (t0). Thus, xing the gauge (t)=0i se q u i v a l e n t
to working with gauge-invariant quantities.
A. ACTIONS 131
Fixing an invariance of the action is not unique to gauge invariances: Global
invariances also need to be xed, although the procedure is so trivial we seldom
discuss it. For example, even in nonrelativistic systems Galilean invariance needs to
be xed: When analyzing a specic problem, we often choose some object to be at
rest (velocity transformations), choose another to be oriented or moving in a specicdirection (rotations), and choose a specic event to happen at the origin of space and
time (translations). Alternatively, we can work with Galilean invariants, just as in
gauge theories we can work with gauge invariants; however, in practice, for explicitcalculations (as opposed to discussing general properties), it is more convenient to x
the invariance, as this allows simplication of the equations.
REFERENCES
1
M.B. Halpern and W. Siegel, Phys. Rev. D16 (1977) 2486:
Classical mechanics as a limit of quantum eld theory.
2P.A.M. Dirac, P r o c .R o y .S o c . A246 (1958) 326;
L.D. Faddeev, Theo. Math. Phys. 1(1969) 1:
Hamiltonian formalism for constrained systems.
132 III. LOCAL
::::::::::::::::::::::::::: :::::::::::::::::::::::::::
::::::::::::::::::::::::::: B. PARTICLES :::::::::::::::::::::::::::
The simplest relativistic actions are those for the mechanics (as opposed to eld
theory) of particles. These also give the simplest examples of gauge invariance in rela-
tivistic theories. Later we will nd that various properties of the quantum mechanicsof these actions help to explain some features of quantum eld theory.
1. Free
For nonrelativistic mechanics, the fact that the energy is expressed as a function of
the three-momentum is conjugate to the fact that the spatial coordinates are expressedas functions of the time coordinate. In the relativistic generalization, all the spacetime
coordinates are expressed as functions of a parameter : All the points that a particle
occupies in spacetime form a curve, or \worldline", and we can parametrize this curvein an arbitrary way. Such parameters generally can be useful to describe curves: A
circle is better described by x();y()t h a ny(x) (avoiding ambiguities in square roots),
and a cycloid can be described explicitly only this way.
The action for a free, spinless particle then can be written in relativistic Hamil-
tonian form as
S
H=Z
d[ .xmpm+v1
2(p2+m2)]
wherevis a Lagrange multiplier enforcing the constraint p2+m2=0 . T h i sa c -
tion is very similar to nonrelativistic ones, but instead of xi(t);pi(t)w en o wh a v e
xm();pm();v()( w h e r e\ ."n o wm e a n s d=d). The gauge invariance generated by
p2+m2is
x=p; p =0; v =.
A more recognizable form of this invariance can be obtained by noting that any
actionS(A) has invariances of the form
A=ABS
B;AB= BA
which have no physical signicance, since they vanish by the equations of motion. In
this case we can add
x=(.x vp); p =.p; v =0
and set=vto get
x=.x; p =.p; v =.
(v)
B. PARTICLES 133
We then can recognize this as a (innitesimal) coordinate transformation for :
x0(0)=x();p0(0)=p();d 0v0(0)=d v();0= ()
The transformation laws for xandpidentify them as \scalars" with respect to these
\one-dimensional" (worldline) coordinate transformations (but they are vectors with
respect to D-dimensional spacetime). On the other hand, vtransforms as a \density":
The \volume element" d v of the world line transforms as a scalar. This gives us
a way to measure length on the worldline in a way independent of the choice of
parametrization. Because of this geometric interpretation, we are led to constrain
v>0
so that any segment of the worldline will have positive length. Because of this re-
striction,vis not a Lagrange multiplier in the usual sense. This has signicant
physical consequences: p2+m2is treated neither as a constraint nor as the Hamil-
tonian. While in nonrelativistic theories the Schr odinger equation is ( E H) =0
andGi = 0 is imposed on the initial states, in relativistic theories ( p2+m2) =0
is the Schr odinger equation: This is more like H =0 ,s i n c ep2already contains the
necessaryEdependence.
The Lagrangian form of the free particle action follows from eliminating pby its
equation of motion vp=.x:
SL=Z
d1
2(vm2 v 1.x2)
Form6= 0, we can also eliminate vby its equation of motion v 2.x2+m2=0 :
S=mZ
dp
.x2=mZp
dx2=mZ
ds=ms
The action then has the purely geometrical interpretation as the proper time; how-
ever, this last form of the action is awkward to use because of the square root, anddoesn't apply to the massless case. Note that the vequation implies ds=m(d v),
relating the \intrinsic" length of the worldline (as measured with the worldline vol-
ume element) to its \extrinsic" length (as measured by the spacetime metric). As aconsequence, in the massive case we also have the usual relation between momentum
and \velocity"
p
m=mdxm
ds
(Note that p0is the energy, not p0.)
Excercise IIIB1.1
Take the nonrelativistic limit of the Poincar ea l g e b r a :
134 III. LOCAL
aInsert the speed of light cin appropriate places for the structure constants of
the Poincar e group (guided by dimensional analysis) and take the limit c!0
to nd the algebra of the Galilean group.
bDo the same for the representation of the Poincar e group generators in terms
of coordinates and momenta. In particular, take the limit of the Lorentz
boosts to nd the Galilean boosts.
cTake the nonrelativistic limit of the spinless particle action, in the form ms.
(Note that, while the relativistic action is positive, the nonrelativistic one isnegative.)
Excercise IIIB1.2
Consider the following action for a particle with additional fermionic variables
and additional fermionic constraint
p:
S
H=Z
d( .xmpm 1
2i.
m
m+1
2vp2+i
p)
whereis also anticommuting so that each term in the action is bosonic.
Find the algebra of the constraints, and the transformations they generateon the variables appearing in the action. Show that the \Dirac equation"
pj i= 0 implies p
2j i= 0. Find the Lagrangian form of the action as
usual by eliminating pby its equation of motion. (Note 2=0 . )
Excercise IIIB1.3
Consider a \supercoordinate" Xmthat is a function of both a fermionic vari-
ableand the usual :
Xm(;)=xm()+i
m()
where the Taylor expansion in terminates because 2=0 . I d e n t i f y xwith
the usualx,a n d
with its fermionic partner introduced in the previous
problem. In analogy to the way
pwas the square root of the -translation
generator1
2p2, we can dene a square root of @=@ by the \covariant fermionic
derivative"
D=@
@+i@
@)D2=i@
@
We also want to generalize vin the same way as x, to make the action inde-
pendent of coordinate choice for both and. This suggests dening
E=v 1+i
and the gauge invariant action
SL=Z
dd1
2E(D2Xm)DXm
B. PARTICLES 135
Integrate this action over , and show this agrees with the action of the
previous problem after suitable redenitions (including the normalization ofR
d).
The (D+2)-dimensional (conformal) representation of the massless particle (sub-
section IA6) can be derived from the action
S=Z
d1
2( .y2+y2)
whereis a Lagrange multiplier. This action is gauge invariant under
y=.y 1
2.y; =.
+2.+1
2...
If we varyto eliminate it and y as in subsection IA6, the action becomes
S= Z
d1
2e2.x2
which agrees with the previous result, identifying v=e 2, which also guarantees
v>0.
Excercise IIIB1.4
Find the Hamiltonian form of the action for y: The constraints are now y2,
r2,a n dyr, in terms of the conjugate rtoy(see excercise IA6.2). Find
the gauge transformations in the standard way (see subsection IIIA5). Showhow the above Lagrangian form can be obtained from it, including the gauge
transformations.
Using instead the corresponding twistor (subsection IIB6) to satisfy y
2=0 ,t h e
massless, spinless particle now has a single term for its mechanics action:
S=Z
d1
4ABCD.zA.zB
zCzD
Unlike all other relativistic mechanics actions, all variables have been unied into just
z, without the introduction of square roots.
Excercise IIIB1.5
Expressing zin terms of andx.as in subsection IIB6, show this action
reduces to the previous one.
136 III. LOCAL
2. Gauges
Rather than use the equation of motion to eliminate vit's more convenient to use
a gauge choice: The gauge v= 1 is called \ane parametrization" of the worldline.
Note that the gauge transformation of v,v=.
, has no dependence on the coordi-
natesxand momenta p, so that choosing the gauge v= 1 avoids any extraneous x
orpdependence that could arise from the gauge xing. (The appearance of such de-
pendence will be discussed in later chapters.) Since T=R
d v, the intrinsic length,
is gauge invariant, that part of vstill remains when the length is nite, but it can be
incorporated into the limits of integration: The gauge v= 1 is maintained by.
=0 ,
and this constant can be used to gauge one limit of integration to zero, completely
xing the gauge (i.e., the choice of ). We then integrateRT
0,w h e r eT0( s i n c e
originallyv>0), andTis a variable to vary in the action. The gauge-xed action is
then
SH;GF =ZT
0d[ .xmpm+1
2(p2+m2)]
In the massive case, we can instead choose the gauge v=1=m; then the equations
of motion imply that is the proper time. The Hamiltonian p2=2m+ constant then
resembles the nonrelativistic one.
Another useful gauge is the \lightcone gauge"
=x+
p+
which, unlike the Poincar e covariant gauge v=1 , x e scompletely; since the gauge
variation(x+=p+)=,w em u s ts e t = 0 to maintain the gauge. Also, the gauge
transformation is again xandpindependent. In lightcone gauges we always assume
p+6= 0, since we often divide by it. This is usually not too dangerous an assumption,
since we can treat p+= 0 as a limiting case (in D >2).
We saw from our study of constrained systems that, for every degree of freedom we
can gauge away, the conjugate variable can be xed by the constraint that generates
that gauge invariance: In the case where the constraint is p, the gauge transformation
isq=, so we gauge q= 0 and use the constraint p= 0. In lightcone gauges the
constraints are almost linear: The gauge condition is x+=p+and the constraint is
p =:::, so the Lagrange multiplier vis varied to determine p . On the other hand,
varyingp gives
p )v=1
so this gauge is a special case of the gauge v= 1. An important point is that we used
only \auxiliary" equations of motion: those not involving time derivatives. (A slight
B. PARTICLES 137
trick involves the factor of p+: This is a constant by the equations of motion, so we
can ignore.p+terms. However, technically we should not use that equation of motion;
instead, we can redene x !x +:::, which will generate terms to cancel any.p+
terms.) The net result of gauge xing and the auxiliary equation on the action is
SH;GF =Z1
1d[.x p+ .xipi+1
2(pi2+m2)]
wherexa=(x+;x ;xi), etc. In particular, since we have xed one more gauge degree
of freedom (corresponding to constant ), we have also eliminated one more constraint
variable (T, the constant part of v). This is one of the main advantages of lightcone
gauges: They are \unitary", eliminating all unphysical degrees of freedom.
Excercise IIIB2.1
Another obvious gauge is =x0, which works as well as the lightcone gauge
as far as eliminating worldline coordinate invariance is concerned. (The sameis true for=nxfor any constant vector n.) Unfortunately, the same is not
true for the auxiliary equations of motion: After using the gauge condition,
p
0appears without time derivatives, so it and vcan be eliminated by their
equations of motion. Show this gauge is consistent only for p0>0. The
resulting square root is awkward except in the nonrelativistic limit: Take it,
and compare with the usual nonrelativistic mechanics.
3. Coupling
One way to introduce external elds into the mechanics action is by considering
the most general Lagrangian quadratic in derivatives:
SL=Z
d[ 1
2v 1gmn(x).xm.xn+Am(x).xm+v(x)]
I nt h ef r e ec a s ew eh a v ec o n s t a n t e l d s gmn=mn,Am=0 ,a n d=1
2m2.T h ev
dependence has been assigned consistent with worldline coordinate invariance. The
curved-space metric tensor gmndescribes gravity, the D-vector potential Amdescribes
electromagnetism, and is a scalar eld that can be used to introduce mass by
interaction.
Excercise IIIB3.1
Use the method of the problem IIIB1.3 to write the nonrelativistic action
for a spinning particle in terms of a 3-vector (or (D 1)-vector)Xi(;)a n d
the fermionic derivative D. Find the coupling to a magnetic eld, in terms
of the 3-vector potential Ai(X). Integrate the Lagrangian over . Show
138 III. LOCAL
that the quantum mechanical square of i[pi+Ai(x)] is proportional to the
Hamiltonian.
Excercise IIIB3.2
Derive the Lorentz force law by varying the Lagrangian form of the action forthe relativistic particle, in an external electromagnetic eld (but
at metric),with respect to x.
This action also has very simple transformation properties under D-dimensional
gauge transformations on the external elds:
g
mn=p@pgmn+gp(m@n)p; Am=p@pAm+Ap@mp @m; =p@p
)SL[x]+SL[x]=SL[x+] (xf)+(xi)
where we have integrated the actionRf
idand setx(i)=xi,x(f)=xf.T h e s e
transformations have a very natural interpretation in the quantum theory, where
Z
Dx e iS=hxfjxii
Then thetransformation of Ais canceled by the U(1) (phase) transformation
0(x)=ei(x) (x)
in the inner product
h fj ii=Z
dxfdxih fjxfihxfjxiihxij ii=Z
dxfdxi f*(xf)hxfjxii i(xi)
while thetransformation associated with gmnis canceled by the D-dimensional
coordinate transformation
0(x)= (x+)
4. Conservation
There are two types of conservation laws generally found in physics: In mechanics
we usually have global conservation laws, of the form.
Q= 0, associated with a
symmetry of the Hamiltonian Hgenerated by a conserved quantity Q:
0=H=i[Q;H ]= .
Q
On the other hand, in eld theory we have local conservation laws, since the action
for a eld is written as an integralRdDxof a Lagrangian density that depends only
B. PARTICLES 139
on elds at x, and a nite number of their derivatives. The local conservation law
implies a global one, since
@mJm=0) 0=ZdDx
(2)D=2@mJmd
dtZdD 1x
(2)D=2J0=.
Q=0
where we have integrated over a volume whose boundaries in space are at inn-
ity (whereJvanishes), and whose boundaries in time are innitesimally separated.
Equivalently, the global symmetry is a special case of the local one.
A simple way to derive the local conservation laws is by coupling gauge elds: We
couple the electromagnetic eld Amto arbitrary charged matter elds and demand
gauge invariance of the matter part of the action, the matter-free part of the action
being separately invariant. We then have
0=SM=Z
dx
(Am)SM
Am+()SM
using just the denition of the functional derivative =. Applying the matter eld
equationsSM== 0, integration by parts, and the gauge transformation Am=
@m, we nd
0=Z
dx
@mSM
Am
)Jm=SM
Am;@mJm=0
Similar remarks apply to gravity, but only if we evaluate the \current", in this case the
energy-momentum tensor, in
at space gmn=mn, since gravity is self-interacting.
We then nd
Tmn= 2SM
gmn
gmn=mn;@mTmn=0
where the normalization factor of 2 will be found later for consistency with the
particle. In this case the corresponding \charge" is the D-momentum:
Pm=ZdD 1x
(2)D=2T0m
In particular we see that the condition for the energy in any region of space to be
nonnegative is
T000
To apply this to the action for the particle in external elds, we must rst dis-
tinguish the particle coordinates X() from coordinates xfor all of spacetime: The
particle exists only at x=X()f o rs o m e , but the elds exist at all x.I n t h i s
notation we can write the mechanics action as
SL=Z
dx
gmn(x)Z
d (x X)1
2v 1.
Xm.
Xn
+Am(x)Z
d (x X).
Xm+(x)Z
d (x X)v
140 III. LOCAL
usingR
dx (x X()) = 1. We then have
Jm(x)=Z
d (x X).
Xm
Tmn=Z
d (x X)v 1.
Xm.
Xn
Note thatT000( s i n c ev>0). Integrating to nd the charge and momentum:
Q=Z
d (x0 X0).
X0=Z
dX0(.
X0)(x0 X0)=(p0)
Pm=Z
d (x0 X0)v 1.
X0.
Xm=Z
dX0(.
X0)(x0 X0)v 1.
Xm=(p0)pm
w h e r ew eh a v eu s e d p=v 1.
X(for the free particle), where pis the momentum
conjugate to X, not to be confused with P. The factor of (p0)((u)=u=jujis the
sign ofu) comes from the Jacobian from changing integration variables from toX0.
The result is that our naive expectations for the momentum and charge of the
particle can dier from the correct result by a sign. In particular p0,w h i c hs e m i -
classically is identied with the angular frequency of the corresponding wave, can
be either positive or negative, while the true energy P0=jp0jis always positive, as
physically required. (Otherwise all states could decay into lower-energy ones: There
would be no lowest-energy state, the \vacuum".) When p0is negative, the charge Q
anddX0=dare also negative. In the massive case, we also have dX0=dsnegative.
This means that as the proper time sincreases,X0decreases. Since the proper time
is the time as measured in the rest frame of the particle, this means that the particle
is traveling backward in time: Its clock changes in the direction opposite to that of
the coordinate system xm. Particles traveling backward in time are called \antiparti-
cles", and have charges opposite to their corresponding particles. They have positive
true energy, but the \energy" p0conjugate to the time is negative.
Excercise IIIB4.1
Compare these expressions for the current and energy-momentum tensor to
those from the semiclassical expansion in excercise IIIA4.1. (Include the in-
verse metric to dene the square of @mS+qAmthere.)
B. PARTICLES 141
5. Pair creation
Free particles travel in straight lines. Nonrelativistically, external elds can alter
the motion of a particle to the extent of changing the signs of spatial components of
the momentum. Relativistically, we might then expect that interactions could also
change the sign of the energy, or at least the canonical energy p0. As an extreme case,
consider a worldline that is a closed loop: We can pick as an angular coordinate
around the loop. As increases,X0will either increase or decrease. For example, a
circle in the x0-x1plane will be viewed by the particle as repeating its history after
some nite , moving forward with respect to time x0until reaching a latest time tf,
and then backward until some earliest time ti. On the other hand, from the point of
view of an observer at rest with respect to the xmcoordinate system, there are no
particles until x0=ti, at which time both a particle and an antiparticle appear at
the same position in space, move away from each other, and then come back togetherand disappear. This process is known as \pair creation and annihilation".
t
tf
ix
x0
1
t
Whether such a process can actually occur is determined by solving the equations
of motion. A simple example is a particle in the presence of only a static electriceld, produced by the time component A
0of the potential. We consider the case of a
piecewise constant potential, vanishing outside a certain region and constant inside.
Then the electric eld vanishes except at the boundaries, so the particle travels in
straight lines except at the boundaries. For simplicity we reduce the problem to two
dimensions:
A0= Vf o r 0x1L; 0otherwise
for some constant V. The action is, in Hamiltonian form,
SH=Z
df .xmpm+v1
2[(p+A)2+m2]g
and the equations of motion are
.pm= v(p+A)n@mAn)p0=E
(p+A)2= m2)p1=p
(E+A0)2 m2
v 1.x=p+A)v 1.x1=p1;v 1.x0=E+A0
142 III. LOCAL
whereEis a constant (the canonical energy at x1=1)a n dt h ee q u a t i o n.p1=:::
is redundant because of gauge invariance. We assume E> 0, so initially we have a
particle and not an antiparticle.
We look only at the cases where the worldline begins at x0=x1= 1 (lower
left) and continues toward the right till it reaches x0=x1=+1(upper right), so
thatp1=v 1.x1>0 everywhere (no re
ection). However, the worldline might bend
backward in time (.x0<0) inside the potential: To the outside viewer, this looks
like pair creation at the right edge before the rst particle reaches the left edge; the
antiparticle then annihilates the original particle when it reaches the left edge, whilethe new particle continues on to the right. From the particle's point of view, it has
simply traveled backward in time so that it exits the right of the potential before it
enters the left, but it is the same particle that travels out the right as came in theleft. The velocity of the particle outside and inside the potential is
dx
1
dx0=8
>>><
>>>:p
E2 m2
Eoutside
p
(E V)2 m2
E Vinside
From the sign of the velocity we then see that we have normal transmission (no
antiparticles) for E>m +VandE>m , and pair creation/annihilation when
V m>E>m)V> 2m
The true \kinetic" energy of the antiparticle (which appears only inside the potential)
is then (E V)>m.
Excercise IIIB5.1
This solution might seem to violate causality. However, in mechanics as well
as eld theory, causality is related to boundary conditions at innite times.
Describe another solution to the equations of motion that would be inter-
preted by an outside observer as pair creation without any initial particles :
What happens ultimately to the particle and antiparticle? What are the al-lowed values of their kinetic energies (maximum and minimum)? Since many
such pairs can be created by the potential alone, it can be accidental (and not
acausal) that an external particle meets up with such an antiparticle. Notethat the generator of the potential, to maintain its value, continuously loses
energy (and charge) by emitting these particles.
REFERENCES
1
A. Barducci, R. Casalbuoni, and L. Lusanna, Nuo. Cim. 35A (1976) 377;
L. Brink, S. Deser, B. Zumino, P. DiVecchia, and P. Howe, Phys. Lett. 64B (1976) 435;
B. PARTICLES 143
P.A. Collins and R.W. Tucker, Nucl. Phys. B121 (1977) 307:
world-line metric; classical mechanics for relativistic spinors.
2R. Marnelius, Phys. Rev. D20 (1979) 2091;
W. Siegel, Int. J. Mod. Phys. A 3(1988) 2713:
conformal invariance in classical mechanics actions.
3W. Siegel, hep-th/9412011, Phys. Rev. D52 (1995) 1042:
twistor formulation of classical mechanics.
4R.P. Feynman, Phys. Rev. 74(1948) 939:
classical pair creation.
144 III. LOCAL
::::::::::::::::::::::::: :::::::::::::::::::::::::
::::::::::::::::::::::::: C. YANG-MILLS :::::::::::::::::::::::::
The concept of a \covariant derivative" allows the straightforward generalization
of electromagnetism to a self-interacting theory, once U(1) has been generalized to a
nonabelian group. Yang-Mills theory is an essential part of the Standard Model.
1. Nonabelian
The group U(1) of electromagnetism is Abelian: Group elements commute, which
makes group multiplication equivalent to multiplication of real numbers, or addition
if we write U=eiG. The linearity of this addition is directly related to the linearity
of the eld equations for electromagnetism without matter. On the other hand, the
nonlinearity of nonabelian groups causes the corresponding particles to interact with
themselves: Photons are neutral, but \gluons" have charge and \gravitons" haveweight.
In coupling electromagnetism to the particle, the relation of the canonical mo-
mentum to the velocity is modied: Classically, the covariant momentum is dx=d =
p+qAfor a particle of charge q(e.g.,q= 1 for the proton). Quantum mechanically,
the net eect is that the wave equation is modied by the replacement
@!r =@+iqA
which accounts for all dependence on A(\minimal coupling"). This \covariant deriva-
tive" has a fundamental role in the formulation of gauge theories, including gravity.Its main purpose is to preserve gauge invariance of the action that gives the wave
equation, which would otherwise be spoiled by derivatives acting on the coordinate-
dependent gauge parameters: In electromagnetism,
0=eiq ;A0=A @) (r )0=eiq(r )
or more simply
r0=eiqre iq
(More generally, qis some Hermitian matrix when is a reducible representation of
U(1).)
Yang-Mills theory then can be obtained as a straightforward generalization of elec-
tromagnetism, the only dierence being that the gauge transformation, and thereforethe covariant derivative, now depends on the generators of some nonabelian group.
We begin with the hermitian generators
[G
i;Gj]= ifijkGk;Giy=Gi
C. YANG-MILLS 145
and exponentiate linear combinations of them to obtain the unitary group elements
g=ei; =iGi;i*=i)gy=g 1
We then can dene representations of the group (see subsection IB1)
0=ei ; y0= ye i;(Gi )A=(Gi)AB B
For compact groups charge is quantized: For example, for SU(2) the spin (or, for
internal symmetry, \isospin") is integral or half-integral. On the other hand, with
Abelian groups the charge can take continuous values: For example, in principle the
proton might decay into a particle of charge and another of charge 1 .T h e
experimental fact that charge is quantized suggests already semiclassically that all
interactions should be descibed by (semi)simple groups.
Ifis coordinate dependent (a local, or \gauge" transformation), the ordinary
partial derivative spoils gauge covariance, so we introduce the covariant derivative
ra=@a+iAa;Aa=AaiGi
Thus, the covariant derivative acts on matter in a way similar to the innitesimal
gauge transformation,
A=iiGiAB B;ra A=@a A+iAaiGiAB B
Gauge covariance is preserved by demanding it have a covariant transformation law
r0=eire i)A= [r;]= @ i[A;]
The gauge covariance of the eld strength follows from dening it in a manifestly
covariant way:
[ra;rb]=iFab)F0=eiFe i;Fab=FabiGi=@[aAb]+i[Aa;Ab]
)Fabi=@[aAb]i+AajAbkfjki
The Jacobi identity for the covariant derivative is the Bianchi identity for the eld
strength:
0=[r[a;[rb;rc]]] =i[r[a;Fbc]]
(If we choose instead to use antihermitian generators, all the explicit i's go away;
however, with hermitian generators the i's will cancel with those from the derivatives
when we Fourier transform for purposes of quantization.) Since the adjoint represen-
tation can be treated as either matrices or vectors (see subsection IB2), the covariant
146 III. LOCAL
derivative on it can be written as either a commutator or multiplication: For example,
we may write either or [ r;F]o rrF, depending on the context.
Actions then can be constructed in a manifestly covariant way: For matter, we
take a Lagrangian LM;0(@; ) that is invariant under global (constant) group trans-
formations, and couple to Yang-Mills as
LM;0(@; )!LM;A=LM;0(r; )
(This is the analog of minimal coupling in electrodynamics.) The representation we
use forGiinra=@a+iAi
aGiis determined by how represents the group. (For
an Abelian group factor U(1), Gis just the charge q, in multiples of the gfor that
factor.) For example, the Lagrangian for a massless scalar is simply
L0=1
2(ra)y(ra)
(normalized for a complex representation).
For the part of the action describing Yang-Mills itself we take (in analogy to the
U(1) case)
LA(Ai
a)=1
8g2
AFiabFj
abij
whereijis the Cartan metric (see subsection IB2). This way of writing the action
is independent of our choice of normalization of the structure constants, and so gives
one unambiguous denition for the normalization of the coupling constant g.( I t i s
invariant under any simultaneous redenition of the elds and the generators that
leaves the covariant derivative invariant.) Generally, for simple groups we can choose
to (ortho)normalize the generators Giwith the condition
ij=cAij
for some constant cA; for groups that are products of simple groups (semisimple),
we might choose dierent normalization factors (but, of course, also dierent g's) for
each simple group. For Abelian groups (U(1) factors) ij= 0, but then the gauge
eld has no self-interactions, so the normalization of the coupling constant is dened
only by matter terms in the action, and we can replace ijwithijin the above.
Usually it will prove more convenient to use matrix notation: Choosing some
convenient representation RofGi(not necessarily the adjoint), we write
LA(Ai
a)=1
8g2
RtrRFabFab
The normalization of the trace is determined by R, and thus so is the normalization
convention for the coupling constant; a change in the representation used in the action
C. YANG-MILLS 147
can also be absorbed by a redenition of the coupling. For example, comparing the
dening and adjoint representations of SU(N) (see subsection IB2),
LA=1
8g2
DtrDFabFab=1
8g2
AtrAFabFab)g2
A=2Ng2
D
In general, we specify our normalization of the structure constants by xing cRfor
someR, and our normalization of the coupling constant by specifying the choice of
representation used in the trace (or use explicit adjoint indices). As a rule, we nd
the most convenient choices of normalization are
cD=1;g =gD
(see subsection IB5).
Excercise IIIC1.1
Write the action for SU(N) Yang-Mills coupled to a massless (2-component)
spinor in the dening representation. Make all (internal and Lorentz) indicesexplicit (no \ tr", etc.), and use dening (N-component) indices on the Yang-
Mills eld.
We have chosen a normalization where the Yang-Mills coupling constant gappears
only as an overall factor multiplying the F
2term (and similarly for the electromagnetic
coupling, as discussed in previous chapters). An alternative is to rescale A!gAand
F!gFeverywhere; then r=@+igAandF=@A+ig[A;A], and theF2term has
no extra factor. This allows the Yang-Mills coupling to be treated similarly to othercouplings, which are usually not written multiplying kinetic terms (unless analogies to
Yang-Mills are being drawn), since (almost) only for Yang-Mills is there a nonlinear
symmetry relating kinetic and interaction terms.
Current conservation works a bit dierently in the nonabelian case: Applying
the same argument as in subsection IIIB4, but taking into account the modied
(innitesimal) gauge transformation law, we nd
J
m=SM
Am;rmJm=0
Since@mJm6= 0, there is no corresponding covariant conserved charge.
Excercise IIIC1.2
Let's look at the eld equations:
aUsing properties of the trace, show the entire covariant derivative can be
integrated by parts as
Z
dx tr (A[r;B]) = Z
dx tr ([r;A]B);Z
dx yr= Z
dx(r )y
148 III. LOCAL
for matricesA;Band column vectors ;.
bShow
Fab=r[aAb]
cUsing the denition of the current as for electromagnetism (subsection IIIB4),
derive the eld equations with arbitrary matter,
1
g21
2rbFba=Ja
dShow that gauge invariance of the action SAimplies
ra(rbFba)=0
Also show this is true directly, using the Jacobi identity, but not the eld
equations. (Hint: Write the covariant derivatives as commutators.)
eExpand the left-hand side in the eld, as
1
g21
2rbFba=1
g21
2@b@[bAa] ja
wherejcontains the quadratic and higher-order terms. Show the noncovari-
antcurrent
Ja=Ja+ja
is conserved. The jterm can be considered the gluon contribution to the
current: Unlike photons, gluons are charged. Although the current is gaugedependent, and thus physically meaningless, the corresponding charge can
be gauge independent under situations where the boundary conditions are
suitable.
2. Lightcone
Since gauge parameters are always of the same form as the gauge eld, but with
one less vector index, an obvious type of gauge choice (at least from the point of viewof counting components) is to require the gauge eld to vanish when one vector index
is xed to a certain value. Explicitly, in terms of the covariant derivative we set
nr=n@)nA=0
for some constant vector n
a. We then can distinguish three types of \axial gauges":
(1) \Arnowitt-Fickler", or spacelike ( n2>0), (2) \lightcone", or lightlike ( n2=0 ) ,
and \temporal", or timelike ( n2<0). By appropriate choice of reference frame, and
C. YANG-MILLS 149
with the usual notation, we can write these gauge conditions as r1=@1,r+=@+,
andr0=@0.
One way to apply this gauge in the action is to keep the same set of elds, but
have explicit ndependence. A much simpler choice is to use a gauge choice such as
A0= 0 simply to eliminate A0explicitly from the action. For example, for Yang-Mills
we nd
A0=0)F0i=.
Ai)1
8(Fab)2= 1
4(.
Ai)2+1
8(Fij)2
where \.
" here refers to the time derivative. Canonical quantization is simple in
this gauge, because we have the canonical time-derivative term. However, the gaugecondition can't be imposed everywhere, as seen for the corresponding gauge for the
one-dimensional metric in subsection IIIB2, and in our general discussion in subsec-
tion IIIA5: Here we can generalize the time-ordered integral for the temporal gauge
to an integral path-ordered with respect to a straight-line path in the ndirection:
e
knr(x)=e i(x;x kn)e kn@;e i(x;x kn)=P
exp
iZx
x kndx0A(x0)
Applying this gauge transformation to nr, as in subsection IIIA5, xes nAto a
constant with respect to n@; the eect on all of ris:
r0(x)=ei(x;x kn)r(x)e i(x;x kn),r0(x+kn)=eknr(x)r(x)e knr(x)
For example, for the temporal gauge, if we choose \ x" to be on the initial hypersurface
x0=t0,t h e nw ec a nc h o o s e k=t t0so thatr0is evaluated at arbitrary time t:
r0
a(x)=[ (ekr0(x)ra(x)e kr0(x))jx0=t0]jk=x0
By Taylor expanding in k, this gives an explicit expression for Aaat all times in terms
ofAa,a n dFaband its covariant time-derivatives, evaluated at some initial time, but
with simply A0(t;xi)=A0(t0;xi). Thus, we still need to impose the A0eld equation
[ri;.
Ai] = 0 as a constraint at some initial time.
Excercise IIIC2.1
Show explicitly that the eld equations for Aifollowing from the action for
Yang-Mills in the gauge A0= 0 (see excercise IIIC1.2 for J=0 )i m p l yt h a t
thetime derivative of the constraint [ ri;.
Ai]=0v a n i s h e s .
In the case of the lightcone gauge we can carry this analysis one step further.
In subsection IIB3 we saw that lightcone formalisms are described by massless eldswith (D 2)-dimensional (\transverse") indices. In the present analysis, gauge xing
alone gives us, again for the example of pure Yang-M ills,
A
+=0)F+i=@+Ai;F+ =@+A ;F i=@ Ai [ri;A ]
150 III. LOCAL
)1
8(Fab)2= 1
4(@+A )2 1
2(@+Ai)(@ Ai [ri;A ]) +1
8(Fij)2
In the lightcone formalism @ ( @+) is to be treated as a time derivative, while
@+can be freely inverted (i.e., modes propagate to innity in the x+direction, but
boundary conditions set them to vanish in the x direction). Thus, we can treat A
as an auxiliary eld. The solution to its eld equation is
A =1
@+2[ri;@+Ai]
which can be substituted directly into the action:
1
8(Fab)2=1
2Ai@+@ Ai+1
8(Fij)2 1
4[ri;@+Ai]1
@+2[rj;@+Aj]
= 1
4AiAi+i1
2[Ai;Aj]@iAj+i1
2(@iAi)1
@+[Aj;@+Aj]
1
8[Ai;Aj]2+1
4[Ai;@+Ai]1
@+2[Aj;@+Aj]
We can save a couple of steps in this derivation by noting that elimination of any
auxiliary eld, appearing quadratically (as in going from Hamiltonian to Lagrangian
formalisms), has the eect
L=1
2ax2+bx+c! 1
2ax2j@L=@x =0+Ljx=0
In this case, the quadratic term is ( F +)2, and we have
1
8(Fab)2=1
8(Fij)2 1
2F+iF i 1
4(F+ )2!1
8(Fij)2 1
2(@+Ai)(@ Ai)+1
4(F+ )2
where the last term is evaluated at
0=[ra;F+a]= @+F+ +[ri;F+i])F+ =1
@+[ri;F+i]
)L=1
8(Fij)2+1
2Ai@+@ Ai 1
4[ri;@+Ai]1
@+2[rj;@+Aj]
as above.
In this case, canonical quantization is even simpler, since interpreting @ as the
time derivative makes the action look like that for a nonrelativistic eld theory, witha kinetic term linear in time derivatives (as well as interactions without them). The
free part of the eld equation is also simpler, since the kinetic operator is now just
. (This is true in general in lightcone formalisms from the analysis of free theories
in chapter XII.) In general, lightcone gauges are the simplest for analyzing physical
degrees of freedom (within perturbation theory), since the maximum number of de-
grees of freedom is eliminated, and thus kinetic operators look like those of scalars.
C. YANG-MILLS 151
On the other hand, interaction terms are more complicated because of the nonlo-
cal Coulomb-like terms involving 1 =@+: The inverse of a derivative is an integral.
(However, in practice we often work in momentum space, where 1 =p+is local, but
Fourier transformation itself introduces multiple integrals.) This makes lightconegauges useful for discussing unitarity (they are \unitary gauges"), but inconvenientfor explicit calculations. However, in subsection VIB6 we'll nd a slight modication
of the lightcone that makes it the most convenient method for certain calculations.
(In the literature, \lightcone gauge" is sometimes used to refer to an axial gaugewhereA
+is set to vanish but A is not eliminated, and D-vector notation is still
used, so unitarity is not manifest. Here we always eliminate both components andexplicitly use ( D 2)-vectors, which has distinct technical advantages.)
Although spin 1/2 has no gauge invariance, the second step of the lightcone
formalism, eliminating auxiliary elds, can also be applied there: For example, for amassless spinor in D=4, identifying @
.
=@ as the lightcone \time" derivative, we
vary .
(or ) as the auxiliary eld:
iL= .
@.
+ .
@ .
.
@ .
.
@.
) =1
@.
@ .
)L= .
1
2
i@.
This tells us that a 4D massless spinor, like a 4D massless vector (or a complex scalar)
has only 1 complex (2 real) degree of freedom, describing a particle of helicity +1/2and its antiparticle of helicity 1/2 (1 for the vector, 0 for the scalar), in agreement
with our general discussion of helicity in subsection IIB7. On the other hand, in the
massive case we can always go to a rest frame, so the analysis is in terms of spin(SU(2) for D=4) rather than helicity. For a massive Weyl spinor we can perform thesame analysis as above, with the modications
L!L+im
p
2( + .
.
))L= .
1
2( m2)
i@.
where we have dropped some terms that vanish upon using integration by parts and
the antisymmetry of the fermions. So now we have the two states of an SU(2) spinor,but these are identied with their antiparticles. This diers from the vector: While
for the spinor we have 2 states of a given energy for both the massless and massive
cases, for a vector we have 2 for the massless but 3 for the massive, since for SU(2)spin s has 2s+1 states.
152 III. LOCAL
Excercise IIIC2.2
Show that integration by parts for 1 =@gives just a sign change, just as for @.
In general dimensions, massless particles are representations of the \little group"
SO(D 2) (the helicity SO(2) in D=4), as described in subsection IIB3. Massive
particles represent the little group SO(D 1), corresponding to dimensional reduction
from an extra dimension, as described in subsection IIB4.
3. Plane waves
The simplest nontrivial solutions to nonabelian eld equations are the general-
izations of the plane wave solutions of the free theory. We begin with general, free,massless theories, as analyzed in subsection IIB3. In the lightcone frame only p
+is
nonvanishing. In position space this means the eld strength depends only on x .
This describes a wave traveling at the speed of light in the positive x1direction, with
no other spatial dependence (i.e., a plane wave). We allow arbitrary dependence on
x , corresponding to a superposition of waves with parallel momenta (but dierent
values ofp+). While its dependence on only x solves the Klein-Gordon equation,
Maxwell's equations are solved by giving the eld strength as many upper + indices
as possible, and no upper 's.
Generalizing to interactions, we notice that the Yang-Mills eld equations and
Bianchi identities dier from Maxwell's equations only by the covariantization of the
derivatives (at least for pure Yang-Mills). Because Maxwell's equations were satised
by just restricting the index structure, we can do the same for the covariant derivativesby assuming that only r
+is novanishing on the eld strengths. In other words, we
can solve the eld equations and Bianchi identities by choosing the only nontrivial
components of the gauge elds to be those in r+.
The nal step is to solve the relation between covariant derivative and eld
strength. This is simple because the index structure we found implies the only non-
trvial commutators are
[@i;r+]=iFi+; [@ ;r+]=0
In particular, this implies that the gauge elds have no x+dependence, and only a
very simple dependence on xi. We nd directly
A+=xiFi+(x )
whereFi+(x ) is unrestricted (other than the explicit index structure and coordinate
dependence). Of course, this result can also be used in the free theory, although it
diers from the usual lightcone gauge.
C. YANG-MILLS 153
4. Self-duality
The simplest and most important solutions to the eld equations are those that
are invariant under the \duality" symmetry that relates electric and magnetic charge:
[ra;rb]=1
2abcd[rc;rd]
Applying the self-duality condition twice, we nd
abefefcd=+c
[ad
b]
which requires an even number of time dimensions. For example, since the action
is usually Wick rotated anyway for perturbative purposes, we might assume thatwe should do the same for classical solutions that are not considered as \small"
uctations about the usual vacuum. (Such a Euclidean denition of eld theory
has been considered for a mathematically rigorous formalism, called \constructive
quantum eld theory", since the Gaussian path integrals for scalars and vectors arethen well-dened and convergent. However, other spins, such as for fermions orgravity, are a problem in this approach.) Alternatively, we can replace
abcdwith
iabcdand complexify our elds. The self-duality condition, when combined with the
Bianchi identities, implies the eld equations: For Yang-Mills,
r[aFbc]=0) 0=1
2abcdraFbc=1
4abcdrabcefFef=raFad
Since the self-duality condition is only rst-order in derivatives, it's easier to solve
than the usual eld equations.
Plane wave solutions provide a simple example of self-duality, since the eld
strengths can easily be written as the sum of self-dual and anti-self-dual parts: In
Minkowski space we dene the self-dual part as helicity +1 ( f..
), and anti-self-dual
as 1(f). For example, for a wave traveling in the \1" direction, the F+2iF+3
components give the two self-dualities for Yang-Mills, describing helicities 1 (the
two circular polarizations).
Before further analyzing solutions to the self-duality condition, we consider ac-
tions that use self-dual elds directly. This will allow us to describe not only theorieswhose only solutions are self-dual, but also more standard theories as perturbationsabout self-duality, and even massive theories. The most unusual feature of this ap-
proach is that complex elds are used without their complex conjugates, since this
is implied in D=3+1 by self-duality. (Alternatively, we can Wick rotate to 2+2 di-mensions, where all Lorentz representations are real.) There are two stages to this
154 III. LOCAL
approach: (1) Use a rst-order formalism where the auxiliary eld is self-dual. The
usual rst-order actions for spin 1/2 (Weyl or Dirac) already can be interpreted inthis way, where \self-duality" means \chirality". (2) For the massive theory, elimi-nate the non-self-dual eld (as an auxiliary eld, as allowed by the mass term), sothat the dynamics is described by the self-dual eld, which was formerly considered
as auxiliary. The massless theory then can be treated as a limiting case.
The simplest (and perhaps most useful) example is massive spin 1/2 coupled in
a real representation to Yang-Mills elds:
L=
Tir. .+1
2p
2m( T + T. .)
where the transposition (\T") refers to the Yang-Mills group index (with respect to
which the spinors are column vectors). Note that must be a real representation of
this group ( AT= A) for the mass term to be gauge invariant (unless the mass term
includes scalars: see the following chapter). Even though and are complex conju-
gates, they can be treated independently as far as eld equations are concerned, sincethey are just dierent linear combinations of their real and imaginary parts. (Com-plex conjugation can be treated as just a symmetry, related to unitarity.) Noticingthat the quadratic term for has no derivatives, we can treat it as an auxiliary eld,
and integrate it out (i.e., eliminate it by its equation of motion, which gives an explicit
local expression for it):
L!
p
2
m[1
4 T( m2) +1
2 Tif ]
w h e r ew eh a v eu s e dt h ei d e n t i t y
r.
r.
=1
2fr.
;r.
g+1
2[r.
;r.
]= 1
2
if
whose simplicity followed from being a real representation of the Yang-Mills group.
(Of course, we could have eliminated instead, but not both.) For convenience we
also scale by a constant
!2 1=4pm
to nd the nal result
L! 1
4 T( m2) 1
2 Tif
Now the massless limit can be taken easily. This action resembles that of a scalar, plus
a \magnetic-moment coupling", which couples the \(anti-)self-dual" (chiral) spinor
to only the (anti-)self-dual part fof the Yang-Mills eld strength.
C. YANG-MILLS 155
For the same reason, the kinetic operator can be written in terms of just the
self-dual part Sof the spin operator:
L= 1
4 T( m2 ifS)
This operator is of the same form found by squaring the Dirac operator:
2r=2= 2(
r)2= (f
a;
bg+[
a;
b])rarb= iFabSba
except for the self-duality. The simple form of this result again depends on the
reality (parity invariance) of the Yang-Mills representation; although this squaring
trick can be applied for complex representations (parity violating), the coupling does
not simplify. This is related to the fact that real representations are required for our
derivation of the self-dual form.
In the special case where the real representation is the direct sum of a complex
one +with its complex conjugate (as for quarks in the Standard Model, or
electrons in electrodynamics), we can rewrite the Lagrangian as
Lc= 1
2 T
+( m2) T
+if
The method can also be generalized to the case of scalar couplings, but the action
becomes nonpolynomial.
For spin 1, we start with the massless case. We can write the Lagrangian for
Yang-Mills as
L=tr(Gf 1
2g2G2
)
whereGis a (anti-)self-dual auxiliary eld. Although this action is complex, elim-
inatingGby its algebraic eld equation gives the usual Yang-Mills action up to a
total derivative term ( abcdFabFcd), which can be dropped for purposes of perturbation
theory. For g=0 ,t h i si sa na c t i o nw h e r e Gacts as a Lagrange multiplier, enforcing
the self-duality of the Yang-Mills eld strength.
If we simply add a mass term
Lm=1
4(m
g)2A2
thenAcan be eliminated by its eld equation, giving a nonpolynomial action of the
form
L+Lm! 1
2(@G)[(m
g)2+G] 1(@G) 1
2g2G2
Just as the spin-1/2 action contained only a 2-component spinor describing the 2
polarizations of spin 1/2, this action contains only the 3-component G, describing
the 3 polarizations of (massive) spin 1.
156 III. LOCAL
Excercise IIIC4.1
Find the Abelian part of this action. Show the free eld equation is
( m2)G= 0 (without gauge xing).
5. Twistors
In four dimensions with an even number of time dimensions, the \Lorentz" group
factorizes (into SU(2)2for D=4+0 and SL(2)2for D=2+2). This makes self-duality
especially simple in spinor notation: For Yang-Mills (cf. electromagnetism in subsec-tion IIA7),
[r
0;r
0]=iC
f00(f=0 )
where we have written primes instead of dots to emphasize that the two kinds of
indices transform independently (instead of as complex conjugates, as in D=3+1). For
purposes of analyzing self-duality within perturbation theory, we can use a lightconemethod that breaks only one of the two SL(2)'s (or SU(2)'s), by separating out its
indices into theand components:
[r
0;r0]=0)r0=@0
where we have chosen a lightcone gauge: The vanishing of all eld strengths for the
covariant derivative r0says that it is pure gauge (as seen by ignoring all but the
x 0coordinates). We now solve
[r[0;r ]0]=0)r 0=@ 0+i@0
i.e.,r 0 @ 0has vanishing curl, and is therefore a gradient. We therefore have
A0=0;A 0=@0;f00= i@0@0
These can also be written in terms of an arbitrary constant twistor (=
above)
as
A0=@
0( i
);f00=@
0@0(i
)
The nal self-duality condition [ r 0;r 0] = 0 then gives the equation of motion
1
2+(@0)(@
0)=0
Excercise IIIC5.1
Show that the sign convention for Wick rotation of the Levi-Civita tensor
consistent with the above equations is
Fab=1
2abcdFcd;F00=Cf00)
C. YANG-MILLS 157
00
00=CC
C00C0
0 CC
C00C
00
Excercise IIIC5.2
Look at the action Gffor self-dual Yang-Mills in the lightcone gauge,
using the results above. Show that this action is equivalent to the lightcone
action for ordinary Yang-Mills (subsection IIIC2), with some terms in the
interaction dropped.
At least for 4D Yang-Mills, advantage can be taken of the conformal invariance
of the classical interacting theory by using a formalism where this invariance is man-
ifest. We saw in subsection IA6 that classical mechanics could be made manifestly
conformal by use of extra coordinates. Covariant derivatives can be dened in termsof projective lightcone coordinates, but the twistor coordinates z
A(see subsections
IIB6 and IIIB1) are more useful. The self-dual covariant derivatives then satisfy
[rA;rB]=iCfAB
in direct analogy to 4D spinor notation. This equation also can be solved by the
lightcone method used above, but now this method breaks only the internal SL(2)symmetry, leaving SL(4) conformal symmetry manifest. More general self-dual eld
strengths in this twistor space are also of the form f
A:::B, totally symmetric in the
indices. We also need to impose the constraint on the eld strength
zAfAB=0
(and similarly for the more general case) to restrict the range of indices to the usual
4D spinor indices (in which the eld strengths are totally symmetric). Self-dualityimplies the Bianchi identity
r
[AfB]C=0
which also generalizes to the other eld strengths, and is the equivalent of the usual
rst-order dierential equations (Dirac, Maxwell, etc.) satised by 4D eld strengths.
As usual, it in turn implies the interacting Klein-Gordon equation, which in theYang-Mills case is
1
2r[ArB]fCD= i[fC[A;fB]D]
The Bianchi together with the zindex constraint imply the constraint on coordinate
dependence
(zArA+
)fBC=0
which eliminates dependence on all but the usual 4D coordinates. These four equa-
tions are generically satisied by self-dual eld strengths. The self-duality itself of the
158 III. LOCAL
eld strengths is a consequence of their total symmetry in their indices, and the fact
that they are all lower (SL(4)) indices. (The zindex constraint then reduces them to
SL(2) Weyl indices all of the same chirality.)
Excercise IIIC5.3
Derive the last three equations from the previous two (self-duality and zf=0).
Excercise IIIC5.4
Show that non-self-dual Yang-Mills is conformally invariant in D=4 by extend-ing the (4+2)-dimensional formalism of subsections IA6 and IIIB1 (especiallyexcercise IIIB1.4): Show the eld strength
F
ABC= i1
2y[A[rB;rC]]
satises the gauge covariances
AA= [rA;] yA^
and Bianchi identities
y[AFBCD ]=r[AFBCD ]=0
The duality transformation
FABC!1
6ABCDEFFDEF
then suggests the eld equations
yAFABC=rAFABC=0
in addition to the usual constraint
y2FABC=0
By reducing to D=4 coordinates with the aid of the above yFconditions, show
Freduces to the usual eld strength, and the remaining equations reduce to
the usual gauge transformation, Bianchi identity, duality, and eld equation.
C. YANG-MILLS 159
6. Instantons
Another interesting class of self-dual solutions to Yang-Mills theory are \instan-
tons", so called because the eld strength is maximum at points in spacetime, unlikethe plane waves, whose wavefronts propagate from and toward timelike innity. Aparticular subset of these can be expressed in a very simple form by the 't Hooftansatz in terms of a scalar eld: In twistor notation, choosing the Yang-Mills gaugegroup GL(2) (in 2+2 dimensions, or SU(2)
GL(1) for 4+0),
iA
A=
@Aln)iAA= @Aln
so the GL(1) piece is pure gauge, and has been included just for convenience. Note
that this ansatz ties the SL(2) twistor index with the SL(2) gauge group indices ( ;),
but in this notation the index that carries the spacetime (conformal) symmetry is free.Imposing the self-duality condition on the eld strength, and separating out the terms
symmetric and antisymmetric in AB, we nd
if
AB= 1
2@(A@B) 1
1@A@B=0
The \eld equation" for is just the twistor version of the (free) Klein-Gordon
equation, and its solution is the projective lightcone version of 4D point sources (seesubsection IA6): Since for any two 6D lightlike vectors yandy
0
y=e(x;1;1
2x2))yy0= 1
2ee0(x x0)2
we have the solution
=k+1X
i=1(yyi) 1;y2
i=0
withyg i v e ni nt e r m so f zas before, and yiare constant null vectors. \ k"i st h en u m b e r
of instantons. (The one term for k= 0 is pure gauge.) The usual singularities in
the Klein-Gordon equation at y=yiare killed by the extra factor of 1in the eld
equation.
Excercise IIIC6.1
Let's check the Klein-Gordon equation for y6=yidirectly in twistor space.
We will need the identity
y2
i=0)yiA[ByiCD]=0
noX
i!
160 III. LOCAL
in the product
yyi=1
2yAByiAB
Prove this identity in two ways:
aShow it follows from the denition
y2
i=1
4ABCDyiAByiCD
bShow it follows from plugging in the solution to the lightlike condition,
yiAB=ziAziB
cNow use the identity to show the above solution satises its eld equation by
evaluating the zderivatives.
We can rewrite this in the usual 4D coordinates by transforming from zAto
andx0aszA=(
;x0) (see subsection IIB6):
dzAAA=dx0A0+[ (d
) 1
]zAAA=dx0A0 i 1
d
where in the rst step we have used the expression for zin terms of andx,a n di n
the second we used the result that
(zA@A+
)=0
We now recognize that the gauge transformation that gets rid of all but the \ x
components" of A(whose existence is guaranteed by the condition zAfAB=0 )u s e s
itself as the gauge parameter:
dzAAA= i 1
d
+ 1
(dzAA0
A)
The net result is that Acan be reduced to an ordinary 4-dimensional expression by
just setting =in the original expression. Then
iA0=
@0ln; =X
i1
ei(x xi)2
withyin terms of xand a scale factor (worldline metric) eas in subsections IA6
and IIIB1 (and dropping an overall factor that doesn't contribute to A). Note that,
unlike the expression in twistor space, where conformal invariance is manifest, here
Lorentz invariance is tied to the Yang-Mills symmetry.
Excercise IIIC6.2
Show in 4D coordinates that the gauge-invariant quantity tr(f2) is nite at
C. YANG-MILLS 161
the pointsx=xi,w h e r eAis singular. (This means that the gauge choice is
singular, not physical quantities.)
Another important property of instantons is that they give nite contributions to
the action. In vector notation, we have
Fab=1
2abcdFcd)S=1
8g2trZd4x
(2)2FabFab=1
16g2trZd4x
(2)2abcdFabFcd
The last expression can be reduced to a boundary term, since
1
8tr F [abFcd]=1
6@[aBbcd]
in terms of the \Chern-Simons form"
Babc=tr(1
2A[a@bAc]+i1
3A[aAbAc])
Excercise IIIC6.3
Although the Chern-Simons form is not manifestly invariant, its variation is,
up to a total derivative:
aShow that its general variation is
Babc=1
2tr[(A[a)Fbc] @[aAbAc]]
bShow the gauge transformation of Bis
Babc= 1
2@[abc];ab=1
2tr(@[aAb])
If we assume boundary conditions such that Fdrops o rapidly at innity, then
Amust drop o to pure gauge at innity:
iAm!g 1@mg
Since instantons always deal with an SU(2) subgroup of the gauge group, we'll assume
now for simplicity that the whole group is itself SU(2). Then the action can be givena group theory interpretation directly, since the integral over the surface at innity
is an integral over the 3-sphere, which covers the group space of SO(3), and thus half
the group space of SU(2). Explicitly,
S=
1
82g2I
d3m1
6mnpqBnpq
!1
82g2I
d3m1
6mnpqtr(g 1@ng)(g 1@pg)(g 1@qg)
162 III. LOCAL
=1
82g2I
d3x1
6ijktr(g 1@ig)(g 1@jg)(g 1@kg)
where in the last step we have switched to coordinates for the 3-sphere, using the factR
d4xmnpqfmnpq is independent of coordinate choice. In fact, in the case where gis
a one-to-one map between the 3-sphere and the group SO(3), this last expression isjust the denition of the invariant volume of the SO(3) group space. In that case, the
integral gives just the volume of the 3-sphere (2
2). In general, the map gwill cover
the SU(2) group space an integer number qof times, and thus cover the SO(3) group
space 2qtimes, so the result will be
S=jqj
2g2
where we have used the fact that self-dual solutions have q>0 while anti-self-dual
haveq<0.
(Anti-)self-dual solutions give relative minima of the action with respect to more
general eld congurations:
0trZ
1
2(Fab1
2abcdFcd)2=trZ
(F21
2abcdFabFcd)
)Sjqj
2g2
qis an integer, and thus can't be changed by continuous variations: It is a topological
property of nite-action congurations. Thus the self-dual solutions give absolute
minima for a given topology. (All these solutions will be given implicitly by twistor
construction in the following subsection. Note that our normalization for the structureconstants of SU(2) diers from the usual, since we use eectively tr(G
iGj)=ij
instead of the more common tr(GiGj)=1
2ij, which would normalize the structure
constants as in SO(3): fijk=ijk. The net eect is that our g2contains a relative
extra factor of 1/2, in addition to the eective extra factors coming from our dierent
normalization of the action.)
Excercise IIIC6.4
Explicitly evaluate the integral for the instanton number qfor the solutions
of the 't Hooft ansatz. Show that the asymptotic form can be expressed in
terms of ( g)=x0.(detg6= 1 because of the GL(1) piece.) Note that
there are boundary contributions not only at x=1but also around the
singular points x=xi, which are of the same form but opposite sign. (Since
the singular parts of Aare pure gauge, they cancel in F.)
At the quantum level, instantons are important mostly because they are an ex-
ample of elds that don't fall o rapidly at innity, and thus contribute toRFF.
C. YANG-MILLS 163
However, once the restriction on boundary conditions is relaxed, there can be many
such eld congurations. The instantons are then distinguished by the fact that theyare the minimal action solutions for a given topology; this makes them important fordescribing low-energy behavior.
7. ADHM
Much more general solutions of this form can be constructed using twistor meth-
ods. (In fact, they can be shown to be the most general self-dual solutions that fallo fast enough at innity in all directions.) The rst step of the Atiyah-Drinfel'd-Hitchin-Manin (ADHM) construction is to introduce a scalar square matrix in a larger
group space
U
II0=(uI;vIi)(I0=(;i))
The indexis the usual two-valued twistor index, for SU(2) in Euclidean space or
SL(2) in 2+2 dimensions. The other indices are
(H) I(G) i
SO(N) SO(N+4k) GL(2k)
SU(N) (SL(N)) SU(N+2k) (SL(N+2k)) GL(k,C)
USp(2N) (Sp(2N)) USp(2N+2k) (Sp(2N+2k)) GL(k)
The index is for the dening representation of the Yang-Mills group H, which is
any of the compact classical groups for Euclidean space, but is its real Wick rotationfor 2+2 dimensions. The index Iis for the dening representation of the group G,
a larger version of H, where k is the instanton number. Finally, the index iis for a
general linear group. We also have the matrix
U
I
I0=(uI
;vI
i0)
For the SO and (U)Sp cases both Umatrices are real, for the SU case they are complex
conjugates of each other, and for the SL case they are real and independent. We next
relate the two U's by
uI
uI=
;uI
vIi=vI
i0uI=0;vI
i0vIi=Cgii0
so they are almost inverses of each other, except that the \metric" gis not constrained
to be a Kronecker . We then write the gauge eld as a generalization of pure gauge:
iAA=uI
@AuI
164 III. LOCAL
(This is similar to the method used for nonlinear models of coset spaces G/H as
discussed in subsection IVA3 below, except for g.)
Self-duality then follows from requiring a certain coordinate dependence of the
U's: This is xed by giving the explicit dependence of the v's as
vIi=bIiAzA
;vI
i0=bI
i0AzA
where theb's are constants. The orthonormality conditions on the U's then implies
the constraint on the b's
bI
i0(AbIiB)=0
as well as determining the u's in terms of the b's (with much messier dependence than
thev's), and thus A. Note that the zdependence of ucan be written in terms of just
x, as follows from rewriting the uvorthogonality as (after multiplying by z)
uI
bIiAyAB=uIbI
i0AyAB=0
and noting scale invariance. Then the xcomponents of Acan also be written in
terms of just x. We then can check the self-duality condition by calculating f:T h e
orthonormality condition on the U's can be written as
J
I=uIuJ
+vIigii0vJ
i0
wheregii0is the inverse of gii0. Then schematically we have
iF=@iA+iAiA
=(@u)(@u) (@u)uu(@u)
=(@u)vgv(@u)
=u(@v)g(@v)u
=ubgbu
or more explicitly
ifAB= (uI
bIi(A)gii0(uJbJ
i0B))
where self-duality is FA;B=CfAB. We can also directly show zAfAB=0 .
C. YANG-MILLS 165
8. Monopoles
Instantons are essentially 0-dimensional objects, localized near a point in 4-
dimensional spacetime (or many points for multi-instanton solutions). Another type
of solution is 1-dimensional; this represents a particle (with a 1D worldline). Unlikethe plane-wave solutions, which represent the massless particles already described
explicitly by elds in the action, we now look for time-independent solutions, which
describe massive (since they have a rest frame), bound-state particles.
Looking at time-independent solutions is similar to the dimensional reduction
that we considered in subsection IIB4 to introduce masses into free theories, only
(1) this mass vanishes, and (2) we reduce the time dimension, not a spatial one. In
our case, the dimensional reduction of a 4-vector (the Yang-Mills potential) gives a3-vector and a scalar, both in the adjoint representation of the group. Let's consider
the reduction in Euclidean space, so the scalar kinetic term comes out with the right
sign. Then the 4D Yang-Mills action reduces as
1
8F2
ab!1
8F2
ij+1
4[ri;]2
where we have labeled the scalar A0=and by dimensional reduction @0!0. Note
that this is the same action that would have been obtained by starting out with Yang-
Mills coupled to an adjoint scalar in four dimensions, either Minkowski or Euclidean,
and choosing the gauge A0= 0. Thus, time-independent solutions to Euclidean
Yang-Mills theory are also time-independent solutions to Minkowskian Yang-Mills
coupled to an adjoint scalar (although not the most general, since the gauge A0=0
is not generally possible globally, especially when we assume time independence of
even gauge-dependent quantities). In particular, this means that time-independent
solutions to self-dual Yang-Mills are also solutions of Minkowskian Yang-Mills cou-pled to an adjoint scalar. This allows us to use the rst-order dierential equations
and topological properties of self-dual Yang-Mills theory to nd physical bo und-state
particles in this vector-scalar theory.
Dimensionally reducing the (Euclidean) self-duality condition, we have
[r
i;]=1
2ijkFjk
As for instantons, the simplest solutions are for SU(2). As for the 't Hooft ansatz,
we look for a solution that is covariant under the combined SU(2) of the gauge group
and 3D rotations: In SO(3) vector notation for both kinds of indices (using the SO(3)
normalization of the structure constants [ iGi;iGj]=ijkiGk),
i=xi'(r); (Ai)j=ijkxkA(r)
166 III. LOCAL
(We know to use an tensor inAbecause of covariance under parity.) The self-
duality equation then reduces to two nonlinear rst-order dierential equations (the
coecients of ijandxixj=r2):
'